IP Library Granted Patent US 12670913
Granted Patent B1
US 12670913 · App. 16/846,194 · Granted Jun 30, 2026

Machine-learning techniques for dialog processing

Inventors: Anshul Jain (San Jose, CA); Sumit Kumar Bhattacharya (Pune, IN); Ratin Kumar (Cupertino, CA); Jason Conrad Roche (Santa Clara, CA); Shubhadeep Das (Pune, IN); Bangqi Wang (Sunnyvale, CA); Rajath Bellipady Shetty (Mountain View, CA)
Assignee: NVIDIA Corporation
G10L17/18G06V40/165G06V40/172G06V40/173G10L15/25
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12670913
App. No.
16/846,194
Granted
Jun 30, 2026
Kind
B1
Abstract

Apparatuses, systems, and techniques to identify speakers based on content of speech. In at least one embodiment, one or more speakers are identified based on content of speech.

Claims (80)

1 . One or more processors, comprising:

circuitry to:

use one or more neural networks to identify at least two or more related conversation agent queries among a plurality of conversation agent queries based at least in part on grouping conversation agent queries into one or more conversations represented by one or more dialog states;

determine, using the one or more neural networks, contextual data corresponding to one or more of the plurality of conversation agent queries;

select content from the two or more related conversation agent queries based, at least in part, on the determined contextual data; and

generate a response to at least one of the plurality of conversation agent queries based, at least in part, on comparison of the contextual data to the one or more dialog states.

2 . The one or more processors of claim 1 , wherein the circuitry is to use the one or more neural networks to select the content based, at least in part, on a confidence score indicating a correlation between the content and the content of at least one of the plurality of conversation agent queries.

3 . The one or more processors of claim 1 , wherein the circuitry is to use the one or more neural networks to select the content based, at least in part, on inferring one or more purposes of at least one of the plurality of conversation agent queries.

4 . The one or more processors of claim 1 , wherein the circuitry is to use the one or more neural networks to select the content based, at least in part, on a determination, by the one or more neural networks, of whether the content is meaningful with respect to at least one of the plurality of conversation agent queries.

5 . The one or more processors of claim 1 , wherein the circuitry is to use the one or more neural networks to select the content based, at least in part, on a comparison between the content and at least one of the plurality of conversation agent queries.

6 . The one or more processors of claim 1 , wherein the circuitry is to use the one or more neural networks to select the content based, at least in part, on conversation histories of users asking at least one of the plurality of conversation agent queries.

7 . A method, comprising:

identifying, using one or more neural networks, two or more related conversation agent queries among a plurality of conversation agent queries based at least in part on grouping conversation agent queries into one or more conversations represented by one or more dialog states;

determining, using the one or more neural networks, contextual data corresponding to one or more of the plurality of conversation agent queries;

selecting content from the two or more related conversation agent queries based, at least in part, on the determined contextual data; and

generating a response to at least one of the plurality of conversation agent queries based, at least in part, on comparison of the contextual data to the one or more dialog states.

8 . The method of claim 7 , further comprising:

selecting the content based, at least in part, on a confidence score indicating a correlation between the content and the dialog state.

9 . The method of claim 7 , further comprising:

selecting the content based, at least in part, on using the one or more neural networks to determine one or more purposes of at least one of the plurality of conversation agent queries.

10 . The method of claim 7 , further comprising:

selecting the content based, at least in part, on a determination, by the one or more neural networks, of whether the content is meaningful with respect to at least one of the plurality of conversation agent queries.

11 . The method of claim 7 , further comprising:

selecting the content based, at least in part, on conversation histories of users asking at least one of the plurality of conversation agent queries.

12 . The method of claim 7 , further comprising:

selecting the content based, at least in part, on parsing text data using the one or more neural networks to determine the contextual data to be used to update conversation histories of users asking at least one of the plurality of conversation agent queries.

13 . The method of claim 7 , further comprising:

selecting the content from the two or more related conversation agent queries based, at least in part, on one or more groupings of conversation agent queries into one or more conversations.

14 . A non-transitory computer-readable storage medium having stored thereon an application programming interface (API), which if performed by one or more processors, causes the one or more processors to at least:

identify, using one or more neural networks, two or more related conversation agent queries among a plurality of conversation agent queries based at least in part on grouping conversation agent queries into one or more conversations represented by one or more dialog states;

determine, using the one or more neural networks, contextual data corresponding to one or more of the plurality of conversation agent queries;

select content from the two or more related conversation agent queries based, at least in part, on the determined contextual data; and

generate a response to at least one of the plurality of conversation agent queries based, at least in part, on comparison of the contextual data to the one or more dialog states.

15 . The non-transitory computer-readable storage medium of claim 14 , wherein the one or more processors are to select the content from the two or more related conversation agent queries by at least:

determining an identifier from a video component of speech corresponding to at least one of the plurality of conversation agent queries to update the dialog state associated with one or more speakers of the at least one of the plurality of conversation agent queries.

16 . The non-transitory computer-readable storage medium of claim 14 , wherein the one or more processors are to select the content from the two or more related conversation agent queries by at least:

determining the contextual data from an audio component of speech corresponding to at least one of the plurality of conversation agent queries to update the dialog state associated with one or more speakers of the at least one of the plurality of conversation agent queries.

17 . The non-transitory computer-readable storage medium of claim 14 , wherein the one or more processors are to select the content from the two or more related conversation agent queries based, at least in part, on using the one or more neural networks to perform gaze tracking, facial landmarks estimation, and lip activity detection on one or more faces of speakers of at least one of the plurality of conversation agent queries.

18 . The non-transitory computer-readable storage medium of claim 14 , wherein the one or more processors are further to:

determine, using the one or more neural networks, whether one or more faces of speakers of at least one of the plurality of conversation agent queries are registered users.

19 . The non-transitory computer-readable storage medium of claim 14 , wherein the dialog state comprises a data representation of a conversation history corresponding to a speaker.

20 . The non-transitory computer-readable storage medium of claim 14 , wherein the one or more processors are to select the content from the two or more related conversation agent queries by at least:

determining, using the one or more neural networks, a confidence score based, at least in part, on a comparison between the selected content and the plurality of conversation agent queries.

21 . A system, comprising:

one or more processors configured to:

use one or more neural networks to identify at least two or more related conversation agent queries among a plurality of conversation agent queries based at least in part on grouping conversation agent queries into one or more conversations represented by one or more dialog states;

determine, using the one or more neural networks, contextual data corresponding to one or more of the plurality of conversation agent queries;

select content from the two or more related conversation agent queries based, at least in part, on the determined contextual data; and

generate a response to at least one of the plurality of conversation agent queries based, at least in part, on comparison of the contextual data to the one or more dialog states.

22 . The system of claim 21 , wherein the one or more processors are to use the one or more neural networks to select the content from the two or more related conversation agent queries based, at least in part, on a confidence score indicating a correlation between the selected content and the content of at least one other of the plurality of conversation agent queries.

23 . The system of claim 21 , wherein the one or more processors are to use the one or more neural networks to select the content from the two or more related conversation agent queries based, at least in part, on one or more purposes, of at least one of the plurality of conversation agent queries, determined by the one or more neural networks.

24 . The system of claim 21 , wherein the one or more processors are to use the one or more neural networks to select the content from the two or more related conversation agent queries based, at least in part, on a determination by the one or more neural networks of whether the content is meaningful with respect to at least one of the plurality of conversation agent queries.

25 . The system of claim 22 , wherein the one or more processors are to use the one or more neural networks to select the content based, at least in part, on a comparison between the selected content and the plurality of conversation agent queries comprising one or more voice queries.

26 . The system of claim 21 , wherein the one or more processors are to use the one or more neural networks to select the content from the two or more related conversation agent queries based, at least in part, on conversation histories of users asking at least one of the plurality of conversation agent queries.

27 . A human-interaction device, comprising: one or more processors comprising one or more circuits;

one or more storage devices;

one or more neural networks; and

one or more cameras;

wherein the one or more circuits are to:

identify, using the one or more neural networks, at least two or more related conversation agent queries among a plurality of conversation agent queries, based at least in part on grouping conversation agent queries into one or more conversations represented by one or more dialog states;

determine, using the one or more neural networks, contextual data corresponding to one or more of the plurality of conversation agent queries;

select content from the two or more related conversation agent queries based, at least in part, on the determined contextual data; and

generate a response to at least one of the plurality of conversation agent queries based, at least in part, on comparison of the contextual data to the one or more dialog states.

28 . The human-interaction device of claim 27 , wherein the one or more circuits are to select the content from the two or more related conversation agent queries by at least:

using the one or more cameras to obtain video and audio components of speech corresponding to at least one of the plurality of conversation agent queries.

29 . The human-interaction device of claim 27 , wherein the one or more circuits are further to:

generate the response to at least one of the plurality of conversation agent queries by updating the dialog state corresponding to one or more speakers of at least one of the plurality of conversation agent queries based, at least in part, on a comparison between the selected content and the at least one of the plurality of conversation agent queries.

30 . The human-interaction device of claim 27 , wherein:

the human-interaction device further comprises a display; and

the one or more circuits are to determine an identifier of a video component of speech corresponding to speakers of at least one of the plurality of conversation agent queries based, at least in part, on detecting that the speakers are looking at the display.

31 . The human-interaction device of claim 27 , wherein the human-interaction device comprises legs or wheels that can be programmatically controlled by the human-interaction device.

32 . The human-interaction device of claim 27 , where the one or more circuits are further to:

determine, based at least in part on an updated dialog state correlating to one or more speakers of at least one of the plurality of conversation agent queries, a location that is being requested by the one or more speakers; and

use legs or wheels to navigate to the location.

33 . The human-interaction device of claim 27 , wherein the one or more circuits are to use the one or more neural networks to determine a contextual intent and one or more subject categories of words in at least one of the plurality of conversation agent queries.

34 . The one or more processors of claim 1 , wherein at least one of the conversation agent queries are portions of speech that are less than all speech spoken by speakers of at least one of the plurality of conversation agent queries.

35 . The one or more processors of claim 1 , wherein the circuitry is to use the one or more neural networks to select content by at least identifying at least one of the plurality of conversation agent queries.

36 . The one or more processors of claim 1 , wherein the circuitry is to use the one or more neural networks to select the content based, at least in part, on one or more rules associating at least one of the plurality of conversation agent queries with one or more speakers.

37 . The one or more processors of claim 1 , wherein the circuitry is to use the one or more neural networks to select the content based, at least in part, on one or more contexts of at least one of the plurality of conversation agent queries corresponding to historical conversation data of one or more speakers.

38 . The one or more processors of claim 1 , wherein the circuitry is to use one or more neural networks to select the content from the two or more related conversation agent queries, spoken by two or more speakers and a conversational artificial intelligence (AI) system, to be used to generate the response to one of a plurality of voice queries spoken by two or more speakers.