IP Library › Patent Application 16516986
Patent Application
App. No. 16/516,986

USER-SPECIFIC ACOUSTIC MODELS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/516,986
Abstract

Systems and processes for providing user-specific acoustic models are provided. In accordance with one example, a method includes, at an electronic device having one or more processors, receiving a plurality of speech inputs, each of the speech inputs associated with a same user of the electronic device; providing each of the plurality of speech inputs to a user-independent acoustic model, the user-independent acoustic model providing a plurality of speech results based on the plurality of speech inputs; initiating a user-specific acoustic model on the electronic device; and adjusting the user-specific acoustic model based on the plurality of speech inputs and the plurality of speech results.

Claims (46)

1 . A computer-implemented method, comprising:

at a first electronic device with one or more processors and memory:

receiving a plurality of speech inputs, each of the speech inputs associated with a same user of the first electronic device;

providing each of the plurality of speech inputs to a user-independent acoustic model, the user-independent acoustic model providing a plurality of speech results based on the plurality of speech inputs;

initiating a user-specific acoustic model on the first electronic device; and

adjusting the user-specific acoustic model by providing the plurality of speech inputs and the plurality of speech results to the user-specific acoustic model.

2 . The method of claim 1 , further comprising:

providing the user-specific acoustic model to another electronic device.

3 . The method of claim 2 , wherein providing the user-specific acoustic model to another electronic device comprises:

determining whether the user-specific acoustic model has been trained on a threshold number of speech inputs;

in accordance with a determination that the user-specific acoustic model has been trained on a threshold number of speech inputs, providing the user-specific acoustic model to the another electronic device; and

in accordance with a determination that the user-specific acoustic model has not been trained on a threshold number of speech inputs:

adjusting the user-specific model based on a second plurality of speech inputs and a second plurality of speech results; and

providing the user-specific acoustic model to the another electronic device.

4 . The method of claim 2 , wherein, at the another electronic device:

receiving the user-specific acoustic model;

receiving a second speech input; and

identifying, with the user-specific acoustic model, a speaker of the second speech input.

5 . The method of claim 4 , wherein identifying, with the user-specific acoustic model, a speaker of the second speech input comprises:

providing the second speech input to the user-specific acoustic model to provide a first speech result and a first accuracy score corresponding to the first speech result;

providing the second speech input to another user-specific acoustic model to provide a second speech result and a second accuracy score corresponding to the second speech result; and

identifying the speaker of the second speech input based on the first accuracy score and the second accuracy score.

6 . The method of claim 4 , wherein receiving a plurality of speech inputs, each of the speech inputs associated with a same user of the first electronic device comprises:

receiving one or more speech inputs of the plurality of speech inputs from the another electronic device.

7 . The method of claim 1 , wherein receiving a plurality of speech inputs comprises:

receiving one or more speech inputs of the plurality of speech inputs at the first electronic device.

8 . The method of claim 7 , wherein receiving one or more speech inputs of the plurality of speech inputs at the first electronic device comprises:

obtaining the one or more speech inputs of the plurality of speech inputs from a user utterance corresponding to a phone call.

9 . The method of claim 7 , wherein receiving one or more speech inputs of the plurality of speech inputs at the first electronic device comprises:

obtaining the one or more speech inputs of the plurality of speech inputs from a user utterance corresponding to a request for a digital assistant.

10 . The method of claim 1 , wherein the user-independent acoustic model is based on a dataset, and

wherein the user-specific acoustic model is initiated using the dataset.

11 . The method of claim 1 , wherein the user-independent acoustic model has a first number of parameters and the user-specific acoustic model has a second number of parameters, wherein the first number is greater than the second number

12 . The method of claim 1 , wherein the user-independent acoustic model is an ensemble of two or more acoustic models.

13 . A system, comprising:

a first electronic device, comprising:

one or more first processors; a first memory; and one or more first programs, wherein the one or more first programs are stored in the first memory and configured to be executed by the one or more first processors, the one or more first programs including instructions for:

receiving a plurality of speech inputs, each of the speech inputs associated with a same user of the first electronic device;

providing each of the plurality of speech inputs to a user-independent acoustic model, the user-independent acoustic model providing a plurality of speech results based on the plurality of speech inputs;

initiating a user-specific acoustic model on the first electronic device; and

adjusting the user-specific acoustic model by providing the plurality of speech inputs and the plurality of speech results to the user-specific acoustic model.

14 . A plurality of non-transitory computer-readable storage media storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a first electronic device, cause the first electronic device to:

receiving a plurality of speech inputs, each of the speech inputs associated with a same user of the first electronic device;

providing each of the plurality of speech inputs to a user-independent acoustic model, the user-independent acoustic model providing a plurality of speech results based on the plurality of speech inputs;

initiating a user-specific acoustic model on the first electronic device; and

adjusting the user-specific acoustic model by providing the plurality of speech inputs and the plurality of speech results to the user-specific acoustic model.