IP Library › Granted Patent US 11,605,384
Granted Patent B1
US 11,605,384 · App. 17/390,118 · Granted Mar 14, 2023

Duplex communications for conversational AI by dynamically responsive interrupting content

Inventors: Steven Dalton (Cary, NC); Siddha Ganju (Santa Clara, CA); Ruthie Lyle (Durham, NC)
Assignee: NVIDIA Corporation
G10L15/222G10L15/08G10L2015/225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,605,384
App. No.
17/390,118
Granted
Mar 14, 2023
Kind
B1
Abstract

Systems and methods of presenting interrupting content during human speech are disclosed. The proposed systems offer improved duplex communications in conversational AI platforms. In some embodiments, the system receives speech data and evaluates the data using linguistic models. If the linguistic models detect indications of linguistic irregularities such as mispronunciation, a smart feedback assistant can determine that the system should interrupt the speaker in near-real-time and provide feedback regarding their pronunciation. In addition, conversational irregularities may also be detected, causing the smart feedback assistant to interrupt with presentation of moderating guidance. In some cases, emotion models may also be utilized to detect emotional states based on the speaker's voice in order to offer near-immediate feedback. Users can also customize the manner and occasions in which they are interrupted.

Claims (54)

1. A computer-implemented method of presenting interrupting content during speech, the method comprising:

receiving, at a first time and by an application accessed via a computing device, first audio data of a first user speaking;

detecting, via the application, at least a first indicator of a first type of speech irregularity in the first audio data;

determining, by the application and based on the first indicator, that a triggering event has occurred; and

causing, via the application, first interrupting content to be presented by the computing device, wherein the first interrupting content:

includes feedback about the detected speech irregularity, and

is presented at a subsequent second time regardless of whether the first user is still speaking.

2. The method of claim 1 , wherein the first type of speech irregularity is one of a mispronunciation, lexico-grammatical inaccuracy, prosodic error, semantic error, speech disfluency, and incorrect phraseology.

3. The method of claim 1 , further comprising:

receiving, via the application, second audio data of the first user speaking;

detecting, by the application, a second indicator of a content irregularity in the second audio data; and

causing, via the application, second interrupting content to be presented that includes feedback identifying misinformation identified in the second audio data.

4. The method of claim 3 , wherein the second interrupting content also includes feedback correcting the misinformation.

5. The method of claim 1 , wherein the feedback includes guidance about correcting the detected speech irregularity.

6. The method of claim 1 , wherein the first interrupting content is presented as audio output that interrupts the first user while the first user is speaking.

7. The method of claim 1 , wherein the first interrupting content is presented as visual output that interrupts a second user while the first user is speaking.

8. The method of claim 1 , wherein the second time is less than ten seconds after the first time.

9. A computer-implemented method of presenting interrupting content during a conversation including at least a first participant and a second participant, the method comprising:

receiving, at a first time and by an application accessed via a computing device, first audio data of at least the first participant speaking;

detecting, by the application, a first indicator of a first type of conversational irregularity in the first audio data;

determining, by the application and based on the first indicator, that a triggering event has occurred; and

causing, via the application, first interrupting content to be presented, wherein the first interrupting content:

includes moderating guidance associated with the detected conversational irregularity, and

is presented at a subsequent second time while one or both of the first participant and second participant are speaking.

10. The method of claim 9 , wherein the first type of conversational irregularity is one of an instance in which the first participant cut off the second participant mid-speech, the first participant and second participant are speaking over one another, the first participant is raising their voice, and the first participant is repeating what the second participant said earlier in the conversation.

11. The method of claim 9 , wherein the moderating guidance includes a suggestion that the second participant be allowed to continue speaking.

12. The method of claim 9 , wherein the moderating guidance includes a recognition that the second participant previously raised an idea that is now being raised in first audio data.

13. The method of claim 9 , further comprising:

receiving, by the application and prior to the first time, a first input describing an agenda for the conversation; and

wherein the moderating guidance includes a reminder that content in the first audio data is off-topic with respect to the agenda.

14. The method of claim 9 , further comprising:

receiving, by the application and prior to the first time, a first data input corresponding to a selection of a duration;

storing the first data input in a preferences module for the first participant;

determining, by the application, the first audio data includes substantially continuous speech made by the first participant that exceeds the selected duration; and

causing, via the application, second interrupting content to be presented, wherein the second interrupting content includes a notification that the first user has exceeded the selected duration.

15. The method of claim 9 , further comprising wherein the first interrupting content is presented as either audio output that interrupts the first participant while the first participant is speaking or visual output that interrupts the first participant while either the first participant or second participant is speaking.

16. A computer-implemented method of determining whether a speech irregularity has occurred, the method comprising:

receiving, at a first time and by an application accessed via a computing device, first audio data of a first user speaking;

classifying, via a speech-based machine learning model for the application, one or more speech characteristics associated with the first user based on the first audio data;

receiving, at a second time and by the application, second audio data of the first user speaking;

determining, via the application, that a speech irregularity has occurred in the second audio data, the determination based at least in part by a comparison of the second audio data with the classified speech characteristics; and

causing, via the application, first interrupting content to be presented by the computing device in response to the determination that a speech irregularity has occurred, wherein the first interrupting content:

includes feedback about the detected speech irregularity, and

is presented at a subsequent third time regardless of whether the first user is still speaking.

17. The method of claim 16 , wherein the speech-based machine learning model resides on a first computer device associated with the first user.

18. The method of claim 16 , further comprising generating, based on the speech characteristics, a user profile for the first user, the user profile including an identification of a language and an accent of the first user.

19. The method of claim 18 , further comprising:

receiving from the first user, at a third time and by the application, a first selection of a first type of speech irregularity that should elicit presentation of interrupting content from the application;

storing, by the application, the first selection in the user profile; and

referring to the user profile when determining whether a speech irregularity has occurred.

20. The method of claim 18 , further comprising:

receiving from the first user, at a third time and by the application, a first selection of a first type of speech irregularity that should be excluded from triggering a presentation of interrupting content from the application;

storing, by the application, the first selection in the user profile; and

referring to the user profile when determining whether a speech irregularity has occurred.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 2, 2021
From: DALTON, STEVEN; GANJU, SIDDHA; LYLE, RUTHIE
To: NVIDIA CORPORATION
Reel/Frame 057371/0170 →
Cited By (7)
US 12,230,262 US 12,425,382 US 12,444,409 US 12,676,166 US 12,694,068 US 12,705,999 US 12,712,970