IP Library Granted Patent US 10,192,552
Granted Patent B2
US 10,192,552 · App. 15/266,932 · Granted Jan 29, 2019

Digital assistant providing whispered speech

Inventors: Tuomo J. Raitio (Sunnyvale, CA); Melvyn J. Hunt (Cheltenham, GB); Hywel B. Richards (Rhiwbina, GB); Madhusudan Chinthakunta (Saratoga, CA)
Assignee: Apple Inc.
G10L15/22G10L13/033G10L25/18G10L25/24G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,192,552
App. No.
15/266,932
Granted
Jan 29, 2019
Kind
B2
Abstract

Systems and processes for detecting and/or providing a whispered speech response are provided. In one example process, speech is received from a user, and based on the speech input, determined that a whispered speech response is to be provided. Upon determining that a whispered speech response is to be provided, the whispered speech response is generated and provided to the user.

Claims (202)

1. An electronic device, comprising:

one or more processors;

memory; and

one or more programs stored in memory, the one or more programs including instructions for:

receiving a speech input from a user;

determining, based on the speech input, that a whispered speech response is to be provided;

upon determining that a whispered speech response is to be provided, generating the whispered speech response, wherein generating the whispered speech response comprises:

generating text based on the speech input;

performing natural language processing of the text;

generating an intermediate speech based on a result of the natural language processing;

obtaining a residual signal based on a linear prediction analysis of the intermediate speech;

modifying the residual signal; and

obtaining the whispered speech response based on a linear prediction synthesis of the modified residual signal; and

providing the whispered speech response to the user.

2. The electronic device of claim 1 , wherein the speech input comprises at least one of an informational request or a request to perform a task.

3. The electronic device of claim 2 , wherein the whispered speech response comprises at least one of a response to the informational request or a response associated with performing the task.

4. The electronic device of claim 1 , wherein determining that the whispered speech response is to be provided comprises at least one of:

determining whether the speech input includes a whispered speech input; and

determining whether context data indicates that the whispered speech response is expected.

5. The electronic device of claim 4 , wherein the whispered speech input is associated with a first spectrum having one or more first spectrum characteristics associated with a whispered speech.

6. The electronic device of claim 5 , wherein the one or more first spectrum characteristics comprise at least one of:

a first amplitude, wherein the first amplitude is less than a second amplitude below a threshold frequency, the second amplitude being associated with the non-whispered speech;

a first energy, wherein the first energy is less than a second energy below the threshold frequency, the second energy being associated with the non-whispered speech;

a first volume, wherein the first volume is less than a second volume by a threshold volume percentage, the second volume being associated with the non-whispered speech; and

a first slope of the first spectrum, wherein the first slope of the first spectrum is shifted by a threshold slope percentage with respect to a second slope of the second spectrum, the second slope of the second spectrum being associated with the non-whispered speech.

7. The electronic device of claim 5 , wherein determining whether the speech input includes a whispered speech input comprises:

determining whether the speech input includes a whispered speech input using one or more features of the speech input, wherein the one or more features represent one or more spectrum characteristics associated with a spectrum of the speech input.

8. The electronic device of claim 7 , wherein determining whether the speech input includes a whispered speech input using the one or more features comprises:

obtaining the spectrum of the speech input;

determining the one or more spectrum characteristics associated with the spectrum of the speech input; and

determining a first feature and a second feature based on the one or more spectrum characteristics associated with the spectrum of the speech input.

9. The electronic device of claim 8 ,

wherein the first feature is a first mel-frequency cepstrum coefficient (MFCC0) representing an energy or an amplitude associated with the spectrum of the speech input; and

wherein the second feature is a second mel-frequency cepstrum coefficient (MFCC1) representing a slope associated with the spectrum of the speech input.

10. The electronic device of claim 8 , wherein the one or more programs include further instructions for:

obtaining a whisper score based on the first feature to the second feature; and

determining whether the whisper score satisfies a score threshold.

11. The electronic device of claim 4 , wherein determining whether the context data indicates that the whispered speech response is expected comprises:

obtaining the context data provided by at least one of the electronic device or one or more additional devices communicatively connected to the electronic device; and

determining whether the context data satisfy one or more conditions for providing the whispered speech response.

12. The electronic device of claim 1 , wherein the intermediate speech has substantially the same content as the whispered speech response.

13. The electronic device of claim 1 , wherein obtaining the residual signal based on a linear prediction analysis of the intermediate speech comprises:

obtaining a plurality of speech frames using the intermediate speech; and

performing the linear prediction analysis of the plurality of speech frames.

14. The electronic device of claim 13 , wherein performing the linear prediction analysis of the plurality of speech frames comprises:

pre-emphasizing the plurality of speech frames;

estimating a plurality of linear prediction coefficients; and

inverse filtering the pre-emphasized speech frames to obtain the residual signal.

15. The electronic device of claim 14 , wherein estimating the plurality of linear prediction coefficients comprises:

performing a windowing on the pre-emphasized plurality of speech frames.

16. The electronic device of claim 1 , wherein modifying the residual signal comprises:

receiving a white noise sequence;

estimating energy of the white noise sequence and the residual signal;

correlating the energy of the white noise sequence and the energy of the residual signal; and

compensating the correlated white noise sequence.

17. The electronic device of claim 16 , wherein compensating the correlated white noise sequence comprises performing at least one of differentiating, high-pass filtering, or band-pass filtering with respect to the correlated white noise sequence.

18. The electronic device of claim 1 , wherein obtaining the whispered speech response based on the linear prediction synthesis of the modified residual signal comprises:

obtaining a plurality of linear prediction coefficients;

modifying the linear prediction coefficients; and

performing a linear prediction synthesis of the modified residual signal using the modified linear prediction coefficients.

19. The electronic device of claim 18 , wherein modifying the linear prediction coefficients comprises:

converting the plurality of linear prediction coefficients to line spectral frequencies;

modifying the line spectral frequencies; and

generating modified linear prediction coefficients based on the modified line spectral frequencies.

20. The electronic device of claim 18 , wherein performing a linear prediction synthesis of the modified residual signal using the modified linear prediction coefficients comprises:

generating a plurality of whispered frames using a synthesis filter and the modified linear prediction coefficients; and

generating the whispered speech response using the plurality of whispered frames.

21. The electronic device of claim 1 , wherein the one or more programs include further instructions for, prior to determining that a whispered speech response is to be provided:

determining whether providing a whispered speech response is disabled; and

in accordance with a determination that providing the whispered speech response is disabled,

generating a non-whispered speech response, and

providing the non-whispered speech response to the user in lieu of the whispered speech response.

22. The electronic device of claim 1 , wherein generating the intermediate speech based on the result of the natural language processing comprises:

identifying a user intent based on the result of the natural language processing; and

generating the intermediate speech according to the user intent.

23. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device, cause the electronic device to:

receive a speech input from a user;

determine, based on the speech input, that a whispered speech response is to be provided;

upon determining that a whispered speech response is to be provided, generate the whispered speech response, wherein generating the whispered speech response comprises:

generating text based on the speech input;

performing natural language processing of the text;

generating an intermediate speech based on a result of the natural language processing;

obtaining a residual signal based on a linear prediction analysis of the intermediate speech;

modifying the residual signal; and

obtaining the whispered speech response based on a linear prediction synthesis of the modified residual signal; and

provide the whispered speech response to the user.

24. The non-transitory computer-readable storage medium of claim 23 , wherein determining that the whispered speech response is to be provided comprises at least one of:

determining whether the speech input includes a whispered speech input; and

determining whether context data indicates that the whispered speech response is expected.

25. The non-transitory computer-readable storage medium of claim 24 , wherein the whispered speech input is associated with a first spectrum having one or more first spectnim characteristics associated with a whispered speech.

26. The non-transitory computer-readable storage medium of claim 25 , wherein the one or more first spectrum characteristics comprise at least one of:

a first amplitude, wherein the first amplitude is less than a second amplitude below a threshold frequency, the second amplitude being associated with the non-whispered speech;

a first energy, wherein the first energy is less than a second energy below the threshold frequency, the second energy being associated with the non-whispered speech;

a first volume, wherein the first volume is less than a second volume by a threshold volume percentage, the second volume being associated with the non-whispered speech; and

a first slope of the first spectrum, wherein the first slope of the first spectrum is shifted by a threshold slope percentage with respect to a second slope of the second spectrum, the second slope of the second spectrum being associated with the non-whispered speech.

27. The non-transitory computer-readable storage medium of claim 25 , wherein determining whether the speech input includes a whispered speech input comprises:

determining whether the speech input includes a whispered speech input using one or more features of the speech input, wherein the one or more features represent one or more spectrum characteristics associated with a spectrum of the speech input.

28. The non-transitory computer-readable storage medium of claim 27 , wherein determining whether the speech input includes a whispered speech input using the one or more features comprises:

obtaining the spectrum of the speech input;

determining the one or more spectrum characteristics associated with the spectrum of the speech input; and

determining a first feature and a second feature based on the one or more spectrum characteristics associated with the spectrum of the speech input.

29. The non-transitory computer-readable storage medium of claim 28 , wherein the one or more programs comprise further instructions, which when executed by one or more processors of the electronic device, cause the electronic device to:

obtain a whisper score based on the first feature to the second feature; and

determine whether the whisper score satisfies a score threshold.

30. The non-transitory computer-readable storage medium of claim 24 , wherein determining whether the context data indicates that the whispered speech response is expected comprises:

obtaining the context data provided by at least one of the electronic device or one or more additional devices communicatively connected to the electronic device; and

determining whether the context data satisfy one or more conditions for providing the whispered speech response.

31. The non-transitory computer-readable storage medium of claim 23 , wherein obtaining the residual signal based on a linear prediction analysis of the intermediate speech comprises:

obtaining a plurality of speech frames using the intermediate speech; and

performing the linear prediction analysis of the plurality of speech frames.

32. The non-transitory computer-readable storage medium of claim 31 , wherein performing the linear prediction analysis of the plurality of speech frames comprises:

pre-emphasizing the plurality of speech frames;

estimating a plurality of linear prediction coefficients; and

inverse filtering the pre-emphasized speech frames to obtain the residual signal.

33. The non-transitory computer-readable storage medium of claim 23 , wherein modifying the residual signal comprises:

receiving a white noise sequence;

estimating energy of the white noise sequence and the residual signal;

correlating the energy of the white noise sequence and the energy of the residual signal; and

compensating the correlated white noise sequence.

34. The non-transitory computer-readable storage medium of claim 23 , wherein obtaining the whispered speech response based on the linear prediction synthesis of the modified residual signal comprises:

obtaining a plurality of linear prediction coefficients;

modifying the linear prediction coefficients; and

performing a linear prediction synthesis of the modified residual signal using the modified linear prediction coefficients.

35. The non-transitory computer-readable storage medium of claim 34 , wherein modifying the linear prediction coefficients comprises:

converting the plurality of linear prediction coefficients to line spectral frequencies;

modifying the line spectral frequencies; and

generating modified linear prediction coefficients based on the modified line spectral frequencies.

36. The non-transitory computer-readable storage medium of claim 34 , wherein performing a linear prediction synthesis of the modified residual signal using the modified linear prediction coefficients comprises:

generating a plurality of whispered frames using a synthesis filter and the modified linear prediction coefficients; and

generating the whispered speech response using the plurality of whispered frames.

37. The non-transitory computer-readable storage medium of claim 23 , wherein the one or more programs comprise further instructions, which when executed by one or more processors of the electronic device, cause the electronic device to, prior to determining that a whispered speech response is to be provided:

determine whether providing a whispered speech response is disabled; and

in accordance with a determination that providing the whispered speech response is disabled,

generate a non-whispered speech response, and

provide the non-whispered speech response to the user in lieu of the whispered speech response.

38. The non-transitory computer-readable storage medium of claim 23 , wherein generating the intermediate speech based on the result of the natural language processing comprises:

identifying a user intent based on the result of the natural language processing; and

generating the intermediate speech according to the user intent.

39. A method for operating a digital assistant, comprising:

at a user device with one or more processors and memory:

receiving a speech input from a user;

determining, based on the speech input, that a whispered speech response is to be provided;

upon determining that a whispered speech response is to be provided, generating the whispered speech response, wherein generating the whispered speech response comprises:

generating text based on the speech input;

performing natural language processing of the text;

generating an intermediate speech based on a result of the natural language processing;

obtaining a residual signal based on a linear prediction analysis of the intermediate speech;

modifying the residual signal; and

obtaining the whispered speech response based on a linear prediction synthesis of the modified residual signal; and

providing the whispered speech response to the user.

40. The method of claim 39 , wherein determining that the whispered speech response is to be provided comprises at least one of:

determining whether the speech input includes a whispered speech input; and

determining whether context data indicates that the whispered speech response is expected.

41. The method of claim 40 , wherein the whispered speech input is associated with a first spectrum having one or more first spectrum characteristics associated with a whispered speech.

42. The method of claim 41 , wherein the one or more first spectrum characteristics comprise at least one of:

a first amplitude, wherein the first amplitude is less than a second amplitude below a threshold frequency, the second amplitude being associated with the non-whispered speech;

a first energy, wherein the first energy is less than a second energy below the threshold frequency, the second energy being associated with the non-whispered speech;

a first volume, wherein the first volume is less than a second volume by a threshold volume percentage, the second volume being associated with the non-whispered speech; and

a first slope of the first spectrum, wherein the first slope of the first spectrum is shifted by a threshold slope percentage with respect to a second slope of the second spectrum, the second slope of the second spectrum being associated with the non-whispered speech.

43. The method of claim 41 , wherein determining whether the speech input includes a whispered speech input comprises:

determining whether the speech input includes a whispered speech input using one or more features of the speech input, wherein the one or more features represent one or more spectrum characteristics associated with a spectrum of the speech input.

44. The method of claim 43 , wherein determining whether the speech input includes a whispered speech input using the one or more features comprises:

obtaining the spectrum of the speech input;

determining the one or more spectrum characteristics associated with the spectrum of the speech input; and

determining a first feature and a second feature based on the one or more spectrum characteristics associated with the spectrum of the speech input.

45. The method of claim 44 , further comprising:

obtaining a whisper score based on the first feature to the second feature; and

determining whether the whisper score satisfies a score threshold.

46. The method of claim 40 , wherein determining whether the context data indicates that the whispered speech response is expected comprises:

obtaining the context data provided by at least one of the electronic device or one or more additional devices communicatively connected to the electronic device; and

determining whether the context data satisfy one or more conditions for providing the whispered speech response.

47. The method of claim 39 , wherein obtaining the residual signal based on a linear prediction analysis of the intermediate speech comprises:

obtaining a plurality of speech frames using the intermediate speech; and

performing the linear prediction analysis of the plurality of speech frames.

48. The method of claim 47 , wherein performing the linear prediction analysis of the plurality of speech frames comprises:

pre-emphasizing the plurality of speech frames;

estimating a plurality of linear prediction coefficients; and

inverse filtering the pre-emphasized speech frames to obtain the residual signal.

49. The method of claim 39 , wherein modifying the residual signal comprises:

receiving a white noise sequence;

estimating energy of the white noise sequence and the residual signal;

correlating the energy of the white noise sequence and the energy of the residual signal; and

compensating the correlated white noise sequence.

50. The method of claim 39 , wherein obtaining the whispered speech response based on the linear prediction synthesis of the modified residual signal comprises:

obtaining a plurality of linear prediction coefficients;

modifying the linear prediction coefficients; and

performing a linear prediction synthesis of the modified residual signal using the modified linear prediction coefficients.

51. The method of claim 50 , wherein modifying the linear prediction coefficients comprises:

converting the plurality of linear prediction coefficients to line spectral frequencies;

modifying the line spectral frequencies; and

generating modified linear prediction coefficients based on the modified line spectral frequencies.

52. The method of claim 50 , wherein performing a linear prediction synthesis of the modified residual signal using the modified linear prediction coefficients comprises:

generating a plurality of whispered frames using a synthesis filter and the modified linear prediction coefficients; and

generating the whispered speech response using the plurality of whispered frames.

53. The method of claim 39 , further comprising, prior to determining that a whispered speech response is to be provided:

determining whether providing a whispered speech response is disabled; and

in accordance with a determination that providing the whispered speech response is disabled,

generating a non-whispered speech response, and

providing the non-whispered speech response to the user in lieu of the whispered speech response.

54. The method of claim 39 , wherein generating the intermediate speech based on the result of the natural language processing comprises:

identifying a user intent based on the result of the natural language processing; and

generating the intermediate speech according to the user intent.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2016
From: RAITIO, TUOMO J.; HUNT, MELVYN J.; RICHARDS, HYWEL B.; CHINTHAKUNTA, MADHUSUDAN
To: APPLE INC.
Reel/Frame 040122/0676 →
Continuity (2)
Provisional Application 62348705 · Jun 10, 2016
Related Publication 20170358301A1 · Dec 14, 2017
Cited By (39)
US 12,197,712 US 12,197,817 US 12,200,297 US 12,204,932 US 12,211,502 US 12,216,894 US 12,219,314 US 12,223,282 US 12,236,952 US 12,249,333 US 12,254,887 US 12,260,234 US 12,277,954 US 12,284,256 US 12,293,203 US 12,293,763 US 12,293,764 US 12,301,635 US 12,333,404 US 12,361,943 US 12,367,879 US 12,379,894 US 12,380,281 US 12,380,876 US 12,386,434 US 12,386,491 US 12,424,218 US 12,431,128 US 12,477,470 US 12,556,890 US 12,567,415 US 12,608,171 US 12,613,621 US 12,613,730 US 12,619,452 US 12,620,179 US 12,640,151 US 12,670,639 US 12,675,839