IP Library › Granted Patent US 12,731,586
Granted Patent B2
US 12,731,586 · App. 17/846,401 · Granted Sep 8, 2026

Emotional detection and processing for barge-in speech

Inventors: Raymond Brueckner (Ulm, DE); Daniel Mario Kindermann (Aachen, DE); Markus Funk (Ulm, DE)
Assignee: Cerence Operating Company
G10L15/222B60Q9/00G10L15/26G10L25/18G10L25/63G10L25/78G10L2015/223G10L2015/227
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,586
App. No.
17/846,401
Filed
Jun 22, 2022
Granted
Sep 8, 2026
Kind
B2
Art Unit
2658
USPC
704/251
Abstract

A method for managing an interaction between a user and a driver interaction system in a vehicle, the method comprising presenting a first audio output to a user from an output device of the driver interaction system, and, while presenting the first audio output to the user, receiving sensed input at the driver interaction system, processing the sensed input including determining an emotional content of the driver, and controlling the interaction based at least in part on the emotional content of the sensed input.

Claims (38)

1 . A method for managing an interaction between a user and a driver-interaction system in a vehicle, the method comprising:

presenting a first audio output to a user from an output device of the driver-interaction system, and

while presenting the first audio output to the user, receiving sensed input at the driver-interaction system, processing the sensed input, wherein processing the sensed input comprises determining whether the sensed input includes speech of the driver defining a barge-in event, and determining an emotional content of the driver from the speech of the driver, and controlling the interaction based at least in part on the emotional content responsive to the sensed input being determined to include speech of the driver defining the barge-in event.

2 . The method of claim 1 , wherein controlling the interaction includes aborting presentation of the first audio output according to the processing, determining a dialog state according to the processing, and presenting a subsequent audio output based on the determined dialog state.

3 . The method of claim 1 , wherein the sensed input comprises spoken input and wherein processing the sensed input comprises determining an amplitude of the spoken input.

4 . The method of claim 1 , wherein processing the sensed input comprises determining the presence of speech in the sensed input based on frequency content of the sensed input.

5 . The method of claim 1 , wherein determining the emotional content of the sensed input includes classifying features of the sensed input using an emotion detector.

6 . The method of claim 1 , wherein determining the emotional content of the sensed input comprises classifying features of the sensed input into discrete emotion categories.

7 . The method of claim 1 , wherein determining emotion content of the sensed input comprises classifying the features of the sensed output into emotions from a discrete set of emotions and assigning scores to each of the emotions.

8 . The method of claim 1 , wherein determining emotion content of the sensed input comprises assigning scores to each emotion in a discrete set of emotions.

9 . The method of claim 1 , wherein the sensed input is a barge-in event and wherein the barge-in event occurs when the driver begins speaking during presentation of the first audio output by the driver-interaction system, wherein the first audio output comprises a first verbal output.

10 . The method of claim 1 , wherein processing the sensed input comprises determining the presence of speech in the sensed input based on periodicity of the sensed input.

11 . The method of claim 1 , wherein processing the sensed input comprises determining the presence of speech in the sensed input based on energy of the sensed input.

12 . The method of claim 1 , wherein determining the emotional content of the sensed input includes processing the sensed input to determine a dimensional representation of the emotional content of the sensed input and wherein the dimensional representation of the emotional content includes a scalar representation of the emotional content in a substantially continuous range of scalar values corresponding to a range of emotions.

13 . The method of claim 1 , wherein the sensed input comprises spoken input and wherein processing the sensed input comprises determining a pitch of the spoken input.

14 . The method of claim 1 , wherein the sensed input comprises spoken input and wherein processing the sensed input comprises processing spectral features of the spoken input.

15 . The method of claim 1 , wherein the sensed input includes one or more of camera input, physiological sensor input, radar sensor input, proximity sensor input, location information, and temperature input.

16 . The method of claim 1 , wherein controlling the interaction includes aborting presentation of the first audio output in response to having determined that the emotional content of the spoken input indicates a negative emotion toward the first audio output.

17 . The method of claim 1 , wherein controlling the interaction includes aborting presentation of the first audio output according to the processing, wherein aborting the presentation is carried out in response to determining that the emotional content of the spoken input indicates a lack of understanding of the first audio input.

18 . The method of claim 1 , wherein controlling the interaction includes continuing presentation of the first audio output based on a determination that the emotional content of the spoken input indicates a positive emotion toward the first audio output.

19 . The method of claim 1 , wherein controlling the interaction includes aborting presentation of the first audio output in response to having detected a negative emotion towards the first audio output in said sensed input, said sensed input being a barge-in event that occurred when the driver began speaking during the first audio output.

20 . The method of claim 1 , wherein the driver-interaction system constantly senses for sensed input.

21 . The method of claim 1 , wherein controlling the interaction comprises determining that presentation of the first audio output is to continue notwithstanding having received the sensed input.

22 . The method of claim 1 , wherein the sensed input is a first sensed input, wherein the emotional content is a first emotional content, wherein the method further comprises,

based on the first emotional content, aborting the presentation of the first audio output,

after having aborted presentation of the first audio output, presenting a second audio output to the user, and

while presenting the second audio output to the user, receiving a second sensed input at the driver-interaction system and processing the second sensed input,

wherein processing the second sensed input comprises determining a second emotional content of the driver based on the second sensed input and, based on the second emotional content, continuing presentation of the second audio output.

23 . A driver-interaction system for interacting with a driver in a vehicle, the driver-interaction system comprising:

a microphone that senses sensed signals from the driver, wherein the sensed signals includes speech;

an output device through which the driver-interaction system presents audio output to the driver;

a speech detector, wherein the speech detector processes the sensed signals and generates output speech signals corresponding to speech from the driver;

a speech recognizer, wherein the speech recognizer processes the sensed signals to generate a transcript of the speech signals;

an emotion detector, wherein the emotion detector processes the sensed signals to generate a classified emotion of the driver; and

a barge-in detector, wherein, while the driver-interaction system is presenting audio output to the driver, the barge-in detector carries out processing steps to determine that a barge-in event has occurred involving speech of the driver and to determine whether or not to interrupt the audio output being presented in response to the barge-in event based on the classified emotion of the driver determined from the speech of the driver defining the barge-in event, wherein the processing steps carried out by the barge-in detector rely on one or more of: the sensed signals, the output speech signals, the transcript, and the classified emotion.

24 . The driver-interaction system of claim 23 , further comprising an interaction-control module that is configured for controlling an interaction with the driver by determining a dialog state based on processing carried out by one or more of the barge-in detector, the driver-sensing components, the speech detector, and the speech recognizer and presenting a subsequent audio output to the driver based on the determined dialog state.

25 . The driver-interaction system of claim 23 , wherein the barge-in event occurs when the driver begins speaking during presentation of the verbal output by the driver-interaction system.

26 . A non-transitory computer-readable medium having instructions stored thereon, the instructions, when executed by a processor of a driver-interaction system, cause the driver-interaction system to, as the driver-interaction system outputs a first audio output to a driver, execute a first action, a second action, and a third action, wherein: the first action is that of receiving sensed input; the second action is that of processing the sensed input including determining whether the sensed input includes speech of the driver defining a barge-in event, and determining an emotional content of the driver from the speech of the driver, and the third action is that of controlling an interaction with the driver based at least in part on the emotional content of the sensed input responsive to the sensed input being determined to include speech of the driver defining the barge-in event.

Assignments (3)
RELEASE (REEL 067417 / FRAME 0303) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0422 →
SECURITY AGREEMENT Recorded Apr 15, 2024
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 067417/0303 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 19, 2022
From: BRUECKNER, RAYMOND; KINDERMANN, DANIEL MARIO; FUNK, MARKUS
To: CERENCE OPERATING COMPANY
Reel/Frame 060548/0912 →
Continuity (1)
Related Publication 20230419965A1 · Dec 28, 2023
References Cited (14)
US 6711536B2 · Rees · 2004 [cited by examiner]
US 10943604B1 · Bone · 2021 [cited by examiner]
US 20110295607A1 · Krishnan · 2011 [cited by examiner]
US 20140229175A1 · Fischer · 2014 [cited by examiner]
US 20150254955A1 · Fields et al. · 2015 [cited by applicant]
US 20190109878A1 · Boyadjiev · 2019 [cited by examiner]
US 20210295833A1 · Rastrow · 2021 [cited by examiner]
US 20220147510A1 · Singh · 2022 [cited by examiner]
EP 3057091A1 · 2016 [cited by applicant]
WO WO0175555A2 · 2001 [cited by examiner]
Kamaruddin et al., “Driver Behavior Analysis through Speech Emotion Understanding” Jun. 21-24, 2010 (Year: 2010). [cited by examiner]
Automobile Driver Profiling—Infering Emotions Through Actions for safe driving (Year: 2016). [cited by examiner]
Kleinschmidt et al., “Assessment of Speech Dialog Systems using Multi-Modal Cognitive Load Analysis and Driving Performance Metrics”; IEEE Xplore (Year: 2009). [cited by examiner]
Michael Braun et al., “Affective Automotive User Interfaces—Reviewing the State of Emotion Regulation in the Car,” arxiv.org, Cornell University Library, 201 Olin Library, Cornell University, Ithaca, NY 14853, pp. 1-25 … [cited by applicant]