This page was machine-translated and may differ from the original. View original
Providing Proactive Services Tailored to User Situations
40 Students Collected 3,500 Total User Response Data Points Over One Week
Professor Lee Eui-jin's research team at the School of Computing at KAIST announced on the 28th that they have discovered important user context factors enabling AI assistants in smart speakers to determine optimal speech timing at the right moment.

40 Students Collected 3,500 Total User Response Data Points Over One Week
Professor Lee Eui-jin's research team at the School of Computing at KAIST announced on the 28th that they have discovered important user context factors enabling AI assistants in smart speakers to determine optimal speech timing at the right moment.

▲ (From left) Cha Na-rae, first author at our university, Professor Kim A-uk (Kangwon National University), Professor Lee Eui-jin at our university [Photo=KAIST]
While existing smart speaker AI assistants or those in commercial circulation only provide services upon user request, the smart speaker developed by Professor Lee's team provides proactive services tailored to the user's situation.
The intelligent voice assistant is being developed to accurately understand the user's situation and proactively help with schedule and health management. However, if the assistant speaks without regard to timing at any moment, it could become an obstacle rather than helpful.
Professor Lee Eui-jin's research team conducted collaborative research to find the optimal timing for smart speakers to proactively provide voice services. As a result, the team identified important user context factors that determine optimal speech timing in smart home environments.
Reasoning about optimal speech timing is an essential technology for AI assistants to autonomously decide and control the initiation, cessation, or resumption of voice services. Stakeholders anticipate that the important context factors discovered by the research team will enhance accuracy in inferring optimal speech timing.
The intelligent voice assistant is being developed to accurately understand the user's situation and proactively help with schedule and health management. However, if the assistant speaks without regard to timing at any moment, it could become an obstacle rather than helpful.
Professor Lee Eui-jin's research team conducted collaborative research to find the optimal timing for smart speakers to proactively provide voice services. As a result, the team identified important user context factors that determine optimal speech timing in smart home environments.
Reasoning about optimal speech timing is an essential technology for AI assistants to autonomously decide and control the initiation, cessation, or resumption of voice services. Stakeholders anticipate that the important context factors discovered by the research team will enhance accuracy in inferring optimal speech timing.

▲ Proactive Dialogue Management Based on Multimodal Sensor Data [Photo=KAIST]
To identify the optimal timing for AI assistants in smart speakers to proactively initiate conversations, the research team designed an experimental smart speaker. The smart speaker asked "Is now a good time to chat?" periodically when the user's movement was detected or after a certain period of time had passed. Participants answered "yes" or "no" to whether it was a good time to chat and explained what they were doing. The research team then installed smart speakers in the dorm rooms of 40 students (double occupancy) residing in the campus dormitory and collected a total of 3,500 user response data points over one week.
Data analysis revealed that 47% of total participant responses indicated that chatting was inappropriate. The research team created 19 indoor activity categories to identify key situational factors determining good timing for conversation. Through this, the team identified three major context factors determining appropriate timing: personal factors, movement factors, and social factors.
Personal factors consist of four categories: 'activity concentration level', 'degree of urgency and busyness', 'mental and physical state', and 'possibility of listening or speaking for multitasking'. For example, it was difficult to chat with the speaker when concentrating on studying or blow-drying hair. Movement factors consist of three categories: 'leaving home', 'returning home', and 'activity transition'.
In particular, when users moved, the distance at which conversation with the speaker was possible had a significant impact on determining the optimal timing. Leaving home is the movement of going outside the speaker's conversation range, while returning home is the movement of entering within the range. When returning home within range, most responses were classified as good timing for conversation.
Generally, smart speakers are installed in shared living spaces such as living rooms. Half of the collected user responses were gathered when roommates were present together. The research team discovered that not only phone conversations but also the presence of someone else affects the optimal timing for conversation with smart speakers. The reason is to minimize conflict from conversation with the smart speaker when a roommate is sleeping or concentrating on an activity.
Na-rae Cha, who participated in this research as the first author, stated, "This research will serve as an important foundation for future smart speaker development" and added, "In the future, smart speakers will be able to proactively detect the optimal timing to initiate, pause, or resume conversations by utilizing context information detected from sensor data, thereby providing intelligent voice services."
Meanwhile, this research was published in the September issue of 'Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies', a top-tier international journal in the field of ubiquitous computing.
Data analysis revealed that 47% of total participant responses indicated that chatting was inappropriate. The research team created 19 indoor activity categories to identify key situational factors determining good timing for conversation. Through this, the team identified three major context factors determining appropriate timing: personal factors, movement factors, and social factors.
Personal factors consist of four categories: 'activity concentration level', 'degree of urgency and busyness', 'mental and physical state', and 'possibility of listening or speaking for multitasking'. For example, it was difficult to chat with the speaker when concentrating on studying or blow-drying hair. Movement factors consist of three categories: 'leaving home', 'returning home', and 'activity transition'.
In particular, when users moved, the distance at which conversation with the speaker was possible had a significant impact on determining the optimal timing. Leaving home is the movement of going outside the speaker's conversation range, while returning home is the movement of entering within the range. When returning home within range, most responses were classified as good timing for conversation.
Generally, smart speakers are installed in shared living spaces such as living rooms. Half of the collected user responses were gathered when roommates were present together. The research team discovered that not only phone conversations but also the presence of someone else affects the optimal timing for conversation with smart speakers. The reason is to minimize conflict from conversation with the smart speaker when a roommate is sleeping or concentrating on an activity.
Na-rae Cha, who participated in this research as the first author, stated, "This research will serve as an important foundation for future smart speaker development" and added, "In the future, smart speakers will be able to proactively detect the optimal timing to initiate, pause, or resume conversations by utilizing context information detected from sensor data, thereby providing intelligent voice services."
Meanwhile, this research was published in the September issue of 'Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies', a top-tier international journal in the field of ubiquitous computing.
To request a correction, reply or follow-up report on this article, see how to file a request. Previously published statements are collected in corrections & replies.
김동우 Reporter













