IP Library Granted Patent US 12711946
Granted Patent B2
US 12711946 · App. 18/447,846 · Granted Aug 18, 2026

Speech model personalization via ambient context harvesting

Inventors: Gabriel Amores (Seville, ES); Guillermo Perez (Seville, ES); Moshe Wasserblat (Maccabim, IL); Michael Deisher (Hillsboro, OR); Loic Dufrensne de Virel (Hillsboro, OR)
Assignee: Intel Corporation
G10L15/063G10L15/065G10L15/075G10L15/16G10L15/183G10L2015/0631G10L2015/0633G10L2015/0635G10L2015/0636G10L2015/0638
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12711946
App. No.
18/447,846
Granted
Aug 18, 2026
Kind
B2
Abstract

An apparatus for speech model with personalization via ambient context harvesting, is described herein. The apparatus includes a microphone, context harvesting module, confidence module, and training module. The context harvesting module is to determine a context associated with the captured audio signals. A confidence module is to determine a confidence of the context as applied to the audio signals. A training module is to train a neural network in response to the confidence being above a predetermined threshold.

Claims (57)

1 . An apparatus comprising:

a microphone to capture audio;

interface circuitry;

machine readable instructions stored in a non-transitory memory; and

programmable circuitry to at least one of execute or instantiate the machine readable instructions to:

detect speech based on the audio collected by the microphone;

identify situational data associated with the audio;

determine a context associated with the audio based on the situational data;

determine a confidence score indicative of a likelihood that the speech is associated with the context;

recognize a dialog pattern based on the speech and the situational data, including comparing the speech to reference dialog data, wherein the dialog pattern includes a sequence of interactions in the reference dialog data;

assign a similarity metric to the speech based on the comparison;

classify the speech based on the dialog pattern and the similarity metric; and

based on the confidence score satisfying a confidence threshold, add the speech and the similarity metric to a database.

2 . The apparatus of claim 1 , wherein the situational data includes one or more of a location or a time of day associated with collection of the audio.

3 . The apparatus of claim 1 , wherein the situational data includes the image data representative of an environment in which the audio was collected.

4 . The apparatus of claim 1 , wherein the situational data includes ambient noise in an environment in which the audio was collected.

5 . The apparatus of claim 1 , wherein the programmable circuitry is to:

recognize a speaker associated with the speech; and

identify the situational data based on the speaker.

6 . The apparatus of claim 1 , wherein the database is a database of structured interactions including reference dialog data, and wherein the programmable circuitry is to recognize the dialogue pattern based on a comparison of the speech to the reference dialog data.

7 . The apparatus of claim 6 , wherein the programmable circuitry is to identify the dialog pattern as a sequence of interactions in the reference dialog data.

8 . The apparatus of claim 6 , wherein the programmable circuitry is to assign a similarity metric to the speech based on the comparison.

9 . The apparatus of claim 8 , wherein the programmable circuitry is to include data associated with the classified speech and the similarity metric in training data to train a neural network for speech recognition.

10 . At least one non-transitory memory comprising instructions to cause programmable circuitry to at least:

detect speech from a user based on an audio collected by a microphone;

identify a location associated with collection of the audio;

determine a context associated with the audio based on the location;

determine a confidence score indicative of a likelihood that the speech is associated with the context;

recognize a structured interaction involving the user based on the speech and the location, including comparing the speech to reference dialog data, wherein the structured interaction includes a sequence of interactions in the reference dialog data;

assign a similarity metric to the speech based on the comparison;

classify the speech based on the structured interaction and the similarity metric; and

add the speech and the similarity metric to a database, based on the confidence score satisfying a confidence threshold.

11 . The at least one memory of claim 10 , wherein the instructions cause the programmable circuitry to identify the location based on location data generated by a mobile device.

12 . The at least one memory of claim 10 , wherein the instructions cause the programmable circuitry to identify image data representative of an environment in which the audio was collected.

13 . The at least one memory of claim 10 , wherein the instructions cause the programmable circuitry to:

identify the user associated with the speech; and

recognize the structured interaction based on the identification of the user.

14 . The at least one memory of claim 10 , wherein the instructions cause the programmable circuitry to

access textual data associated with the location, the textual data not associated with the speech; and

generate training data to train a neural network, the training data including the classified speech and the textual data.

15 . An apparatus comprising:

a microphone to capture audio signals;

interface circuitry;

machine readable instructions stored in a non-transitory memory; and

programmable circuitry to at least one of execute or instantiate the machine readable instructions to:

identify a speaker associated with speech;

identify situational data associated with the speech;

determine a context associated with the audio based on the situational data;

determine a confidence score indicative of a likelihood that the speech is associated with the context;

recognize a dialog pattern based on the speech, the identity of the speaker, and the situational data, including comparing the speech to reference dialog data, wherein the dialog pattern includes a sequence of interactions in the reference dialog data; update training data based on the dialog pattern and the confidence score satisfying a confidence threshold, the training data to train a neural network model.

16 . The apparatus of claim 15 , wherein the speaker is a first speaker and the programmable circuitry is to:

recognize the dialogue pattern is an interaction between the first speaker and the second speaker; and

associate at least one of the dialogue pattern or the speech with the interaction.

17 . The apparatus of claim 15 , wherein the situational data includes one or more of a location or a time of day associated with the collection of the speech.

18 . The apparatus of claim 15 , wherein the programmable circuitry is to identify the speaker based on the situational data.

19 . The apparatus of claim 15 , wherein the processor circuitry is to recognize the dialogue pattern based on a comparison of the speech to reference dialog data.

20 . The apparatus of claim 19 , wherein the reference dialog data is associated with the speaker.