IP Library › Granted Patent US 9,972,304
Granted Patent B2
US 9,972,304 · App. 15/266,949 · Granted May 15, 2018

Privacy preserving distributed evaluation framework for embedded personalized systems

Inventors: Matthias Paulik (San Jose, CA); Henry G. Mason (San Francisco, CA); Matthew S. Seigel (Cupertino, CA)
Assignee: Apple Inc.
G10L15/01G10L15/063G10L15/07G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,972,304
App. No.
15/266,949
Filed
Sep 15, 2016
Granted
May 15, 2018
Kind
B2
Examiner
AZAD, ABUL K
Art Unit
2657
USPC
704/243
Abstract

Systems and processes for evaluating embedded personalized systems are provided. In one example process, instructions that define an experiment associated with a personalized speech recognition system can be received. The instructions can define one or more experimental parameters. In accordance with the received instructions, a second personalized speech recognition system can be generated based on the personalized speech recognition system and the one or more experimental parameters. Additionally, the plurality of user speech samples can be processed using the second personalized speech recognition system to generate a plurality of speech recognition results and a plurality of accuracy scores corresponding to the plurality of speech recognition results. Second instructions can be received based on the plurality of accuracy scores. In accordance with the second instructions, the second speech recognition system can be activated.

Claims (68)

1. An electronic device for evaluating personalized embedded systems, the device comprising:

one or more processors; and

memory storing a plurality of user speech samples, a personalized speech recognition system, and instructions, the instructions, when executed by the one or more processors, cause the one or more processors to:

receive second instructions that define an experiment associated with the personalized speech recognition system, wherein the second instructions define one or more experimental parameters, and wherein the one or more experimental parameters include one or more weighting parameters for interpolating between a general speech recognition model and a personalized speech recognition model of the personalized speech recognition system;

in accordance with the received second instructions:

generate a second personalized speech recognition system based on the personalized speech recognition system and the one or more experimental parameters; and

process the plurality of user speech samples using the second personalized speech recognition system to generate a plurality of speech recognition results and a plurality of accuracy scores corresponding to the plurality of speech recognition results;

provide, to one or more remote devices, the plurality of accuracy scores to evaluate, wherein third instructions for applying the second personalized speech recognition system are generated based on the plurality of accuracy scores;

receive the third instructions;

in accordance with the third instructions, activate the second personalized speech recognition system to perform speech recognition;

receive user speech input;

process the user speech input using the activated second personalized speech recognition system to generate a speech recognition result; and

output a response to the user speech input based on the speech recognition result.

2. The device of claim 1 , wherein the one or more experimental parameters further include one or more machine learning hyperparameters.

3. The device of claim 1 , wherein the plurality of speech recognition results are not transmitted to a remote electronic device.

4. The device of claim 1 , wherein the plurality of accuracy scores are confidence scores generated without comparing the plurality of speech recognition results to a plurality of reference text.

5. The device of claim 1 , wherein:

the memory stores a plurality of verified text, the plurality of verified text generated based on user input received at the electronic device; and

the plurality of accuracy scores are generated by comparing the plurality of speech recognition results to the plurality of verified text.

6. The device of claim 1 , wherein the instructions further cause the one or more processors to:

transmit the plurality of user speech samples to a remote electronic device; and

receive, from the remote electronic device, a plurality of second verified text corresponding to the plurality of user speech samples, wherein the plurality of accuracy scores are generated by comparing the plurality of speech recognition results to the plurality of second verified text.

7. The device of claim 6 , wherein the plurality of second verified text is generated by processing the plurality of user speech samples using a large vocabulary automatic speech recognition system of the remote electronic device.

8. The device of claim 6 , wherein the plurality of second verified text is generated based on second user input received at the remote electronic device.

9. The device of claim 1 , wherein the plurality of speech recognition results are based on one or more speech recognition models of the personalized speech recognition system and one or more second speech recognition models generated based on the one or more experimental parameters, and wherein the plurality of accuracy scores are derived from the one or more speech recognition models of the personalized speech recognition system.

10. The device of claim 9 , wherein the plurality of accuracy scores are not derived from the one or more second speech recognition models.

11. The device of claim 1 , wherein the instructions further cause the one or more processors to:

prior to receiving the second instructions, receive a plurality of speech inputs at the electronic device, wherein the plurality of speech inputs are associated with a user, and wherein the user speech samples are derived from the plurality of speech inputs.

12. The device of claim 1 , wherein the instructions further cause the one or more processors to:

process the plurality of user speech samples using the personalized speech recognition system to generate a plurality of reference speech recognition results and a plurality of reference accuracy scores corresponding to the plurality of reference speech recognition results, wherein the third instructions are based on the plurality of reference accuracy scores.

13. The device of claim 1 , wherein the plurality of speech recognition results are generated by combining a plurality of second speech recognition results generated from the second personalized speech recognition system with a plurality of third speech recognition results generated from a remote speech recognition system.

14. The device of claim 1 , wherein the third instructions are based on a plurality of sets of accuracy scores obtained from a plurality of remote electronic devices in accordance with experimental instructions defining the one or more experimental parameters.

15. The device of claim 1 , wherein the speech recognition result is generated based on a word-level combination of a second speech recognition result generated from the second personalized speech recognition system with a third speech recognition result generated from a remote speech recognition system.

16. A method for evaluating personalized embedded systems implemented on a device, comprising:

at an electronic device having one or more processors and memory storing a plurality of user speech samples and a personalized speech recognition system:

receiving instructions that define an experiment associated with the personalized speech recognition system, wherein the instructions define one or more experimental parameters, and wherein the one or more experimental parameters include one or more weighting parameters for interpolating between a general speech recognition model and a personalized speech recognition model of the personalized speech recognition system;

in accordance with the received instructions:

generating a second personalized speech recognition system based on the personalized speech recognition system and the one or more experimental parameters; and

processing the plurality of user speech samples using the second personalized speech recognition system to generate a plurality of speech recognition results and a plurality of accuracy scores corresponding to the plurality of speech recognition results;

providing, to one or more remote devices, the plurality of accuracy scores to evaluate, wherein second instructions for applying the second personalized speech recognition system are generated based on the plurality of accuracy scores;

receiving the second instructions ;

in accordance with the second instructions, activating the second personalized speech recognition system to perform speech recognition;

receiving user speech input;

processing the user speech input using the activated second personalized speech recognition system to generate a speech recognition result; and

outputting a response to the user speech input based on the speech recognition result.

17. The method of claim 16 , wherein the one or more experimental parameters further include one or more machine learning hyperparameters.

18. The method of claim 16 , wherein the plurality of speech recognition results are not transmitted to a remote electronic device.

19. The method of claim 16 , wherein the plurality of accuracy scores are confidence scores generated without comparing the plurality of speech recognition results to a plurality of reference text.

20. The method of claim 16 , further comprising:

transmitting the plurality of user speech samples to a remote electronic device; and

receiving, from the remote electronic device, a plurality of second verified text corresponding to the plurality of user speech samples, wherein the plurality of accuracy scores are generated by comparing the plurality of speech recognition results to the plurality of second verified text.

21. A non-transitory computer readable storage medium having instructions stored thereon, the instructions, when executed by one or more processors, cause the one or more processors to:

receive second instructions that define an experiment associated with a personalized speech recognition system, wherein the second instructions define one or more experimental parameters, and wherein the one or more experimental parameters include one or more weighting parameters for interpolating between a general speech recognition model and a personalized speech recognition model of the personalized speech recognition system;

in accordance with the received second instructions:

generate a second personalized speech recognition system based on the personalized speech recognition system and the one or more experimental parameters; and

process a plurality of user speech samples using the second personalized speech recognition system to generate a plurality of speech recognition results and a plurality of accuracy scores corresponding to the plurality of speech recognition results;

provide, to one or more remote devices, the plurality of accuracy scores to evaluate, wherein third instructions for applying the second personalized speech recognition system are generated based on the plurality of accuracy scores;

receive the third instructions ;

in accordance with the third instructions, activate the second personalized speech recognition system to perform speech recognition;

receive user speech input;

process the user speech input using the activated second personalized speech recognition system to generate a speech recognition result; and

output a response to the user speech input based on the speech recognition result.

22. The computer readable storage medium of claim 21 , wherein the one or more experimental parameters further include one or more machine learning hyperparameters.

23. The computer readable storage medium of claim 21 , wherein the plurality of speech recognition results are not transmitted to a remote electronic device.

24. The computer readable storage medium of claim 21 , wherein the plurality of accuracy scores are confidence scores generated without comparing the plurality of speech recognition results to a plurality of reference text.

25. The computer readable storage medium of claim 21 , wherein the instructions further cause the one or more processors to:

transmit the plurality of user speech samples to a remote electronic device; and

receive, from the remote electronic device, a plurality of second verified text corresponding to the plurality of user speech samples, wherein the plurality of accuracy scores are generated by comparing the plurality of speech recognition results to the plurality of second verified text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 22, 2016
From: PAULIK, MATTHIAS; MASON, HENRY G.; SEIGEL, MATTHEW S.
To: APPLE INC.
Reel/Frame 040405/0533 →
Continuity (2)
Provisional Application 62345401 · Jun 3, 2016
Related Publication 20170352346A1 · Dec 7, 2017