IP Library Granted Patent US 10,074,360
Granted Patent B2
US 10,074,360 · App. 14/834,239 · Granted Sep 11, 2018

Providing an indication of the suitability of speech recognition

Inventor: Yoon Kim (Cupertino, CA)
Assignee: Apple Inc.
G10L15/01G10L15/22G10L25/60H04R29/008
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,074,360
App. No.
14/834,239
Granted
Sep 11, 2018
Kind
B2
Abstract

This relates to providing an indication of the suitability of an acoustic environment for performing speech recognition. One process can include receiving an audio input and determining a speech recognition suitability based on the audio input. The speech recognition suitability can include a numerical, textual, graphical, or other representation of the suitability of an acoustic environment for performing speech recognition. The process can further include displaying a visual representation of the speech recognition suitability to indicate the likelihood that a spoken user input will be interpreted correctly. This allows a user to determine whether to proceed with the performance of a speech recognition process, or to move to a different location having a better acoustic environment before performing the speech recognition process. In some examples, the user device can disable operation of a speech recognition process in response to determining that the speech recognition suitability is below a threshold suitability.

Claims (92)

1. A method for operating a virtual assistant, the method comprising:

at an electronic device:

receiving an audio input from an acoustic environment;

determining a speech recognition suitability value based on the audio input, wherein the speech recognition suitability value represents a suitability of the acoustic environment of the electronic device for speech recognition;

in accordance with a determination of the speech recognition suitability value, displaying a visual representation of the speech recognition suitability value;

determining whether the speech recognition suitability value satisfies a predetermined criterion; and

in accordance with a determination that the speech recognition suitability value does not satisfy the predetermined criterion, disabling, by the electronic device, speech recognition functionality on the electronic device.

2. The method of claim 1 , wherein determining the speech recognition suitability value based on the audio input comprises:

determining one or more characteristics of the acoustic environment based on the audio input; and

determining the speech recognition suitability based on the one or more characteristics of the acoustic environment.

3. The method of claim 2 , wherein the one or more characteristics of the acoustic environment comprises a signal to noise ratio for a first frequency band of the acoustic environment.

4. The method of claim 3 , wherein the one or more characteristics of the acoustic environment comprises a type of noise detected in the first frequency band.

5. The method of claim 2 , wherein the one or more characteristics of the acoustic environment comprises a signal to noise ratio for a second frequency band of the acoustic environment.

6. The method of claim 5 , wherein the one or more characteristics of the acoustic environment comprises a type of noise detected in the second frequency band.

7. The method of claim 2 , wherein the one or more characteristics of the acoustic environment comprises a number of transient noises detected in a buffer comprising previously recorded audio of the acoustic environment.

8. The method of claim 2 , wherein determining the speech recognition suitability value comprises:

determining a speech recognition suitability vector based on the audio input, wherein the speech recognition suitability vector comprises one or more elements that represent the one or more characteristics of the acoustic environment; and

using a neural network to determine the speech recognition suitability value based on the speech recognition suitability vector.

9. The method of claim 1 , wherein the visual representation comprises one or more bars, and wherein a value of the speech recognition suitability value is represented by a number of the one or more bars.

10. The method of claim 1 , wherein the visual representation comprises an icon, and wherein the speech recognition suitability value is represented by a color of the icon.

11. The method of claim 10 , wherein the icon comprises an image of a microphone.

12. The method of claim 10 , wherein displaying the visual representation of the speech recognition suitability value comprises:

determining whether the speech recognition suitability value is less than a threshold value;

in accordance with a determination that the speech recognition suitability value is less than the threshold value, displaying the icon in a grayed out state; and

in accordance with a determination that the speech recognition suitability value is not less than the threshold value, displaying the icon in a non-grayed out state.

13. The method of claim 1 , wherein the method further comprises:

receiving a user selection of the visual representation of the speech recognition suitability value;

in accordance with a determination that the speech recognition suitability value is not less than a threshold value, performing speech recognition on an audio input received subsequent to receiving the user selection; and

in accordance with a determination that the speech recognition suitability value is less than the threshold value, forgoing the performance of speech recognition on the audio input received subsequent to receiving the user selection.

14. The method of claim 12 , wherein the method further comprises:

in accordance with a determination that the speech recognition suitability value is less than the threshold value, outputting a message indicating a low suitability of the acoustic environment of the electronic device for speech recognition.

15. The method of claim 1 , wherein the visual representation comprises a textual representation of the speech recognition suitability value.

16. The method of claim 1 , wherein determining the speech recognition suitability value comprises periodically determining the speech recognition suitability value, and wherein displaying the visual representation of the speech recognition suitability value comprises updating the display of the visual representation of the speech recognition suitability value in accordance with the periodically determined speech recognition suitability value.

17. The method of claim 1 , wherein the speech recognition suitability value comprises a numerical value.

18. The method of claim 1 , wherein the audio input does not include speech from a user of the electronic device.

19. A non-transitory computer-readable storage medium for operating a virtual assistant, the computer-readable storage medium comprising instructions for:

receiving an audio input from an acoustic environment;

determining a speech recognition suitability value based on the audio input, wherein the speech recognition suitability value represents a suitability of the acoustic environment of the electronic device for speech recognition;

in accordance with a determination of the speech recognition suitability value, displaying a visual representation of the speech recognition suitability value;

determining whether the speech recognition suitability value satisfies a predetermined criterion; and

in accordance with a determination that the speech recognition suitability value does not satisfy the predetermined criterion, disabling, by the electronic device, speech recognition functionality on the electronic device.

20. The storage medium of claim 19 , wherein the visual representation comprises an icon, and wherein the speech recognition suitability value is represented by a color of the icon.

21. The storage medium of claim 20 , wherein displaying the visual representation of the speech recognition suitability value comprises:

determining whether the speech recognition suitability value is less than a threshold value;

in accordance with a determination that the speech recognition suitability value is less than the threshold value, displaying the icon in a grayed out state; and

in accordance with a determination that the speech recognition suitability value is not less than the threshold value, displaying the icon in a non-grayed out state.

22. The storage medium of claim 21 , further comprising:

in accordance with a determination that the speech recognition suitability value is less than the threshold value, outputting a message indicating a low suitability of the acoustic environment of the electronic device for speech recognition.

23. The storage medium of claim 20 , wherein the icon comprises an image of a microphone.

24. The storage medium of claim 19 , wherein determining the speech recognition suitability value comprises periodically determining the speech recognition suitability value, and wherein displaying the visual representation of the speech recognition suitability value comprises updating the display of the visual representation of the speech recognition suitability value in accordance with the periodically determined speech recognition suitability value.

25. The storage medium of claim 19 , wherein the speech recognition suitability value is determined based on a signal to noise ratio for a first frequency band of the acoustic environment.

26. The storage medium of claim 19 , wherein determining the speech recognition suitability value based on the audio input comprises:

determining one or more characteristics of the acoustic environment based on the audio input; and

determining the speech recognition suitability based on the one or more characteristics of the acoustic environment.

27. The storage medium of claim 26 , further comprising:

determining a speech recognition suitability vector based on the audio input, wherein the speech recognition suitability vector comprises one or more elements that represent the one or more characteristics; and

using a neural network to determine the speech recognition suitability value based on the speech recognition suitability vector.

28. The storage medium of claim 26 , wherein the one or more characteristics of the acoustic environment comprises a type of noise detected in a first frequency band.

29. The storage medium of claim 26 , wherein the one or more characteristics of the acoustic environment comprises a number of transient noises detected in a buffer comprising previously recorded audio of the acoustic environment.

30. The storage medium of claim 19 , wherein the visual representation comprises one or more bars, and wherein a value of the speech recognition suitability value is represented by a number of the one or more bars.

31. The storage medium of claim 19 , wherein the visual representation comprises a textual representation of the speech recognition suitability value.

32. The storage medium of claim 19 , wherein the speech recognition suitability value comprises a numerical value.

33. A system for operating a virtual assistant, the system comprising:

one or more processors;

memory; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:

receiving an audio input from an acoustic environment;

determining a speech recognition suitability value based on the audio input, wherein the speech recognition suitability value represents a suitability of acoustic environment of the electronic device for speech recognition;

in accordance with a determination of the speech recognition suitability value, displaying a visual representation of the speech recognition suitability value;

determining whether the speech recognition suitability value satisfies a predetermined criterion; and

in accordance with a determination that the speech recognition suitability value does not satisfy the predetermined criterion, disabling, by the electronic device, speech recognition functionality on the electronic device.

34. The system of claim 33 , wherein determining the speech recognition suitability value based on the audio input comprises:

determining one or more characteristics of the acoustic environment based on the audio input; and

determining the speech recognition suitability based on the one or more characteristics of the acoustic environment.

35. The system of claim 34 , further comprising:

determining a speech recognition suitability vector based on the audio input, wherein the speech recognition suitability vector comprises one or more elements that represent the one or more characteristics; and

using a neural network to determine the speech recognition suitability value based on the speech recognition suitability vector.

36. The system of claim 34 , wherein the one or more characteristics of the acoustic environment comprises a type of noise detected in a first frequency band.

37. The system of claim 34 , wherein the one or more characteristics of the acoustic environment comprises a number of transient noises detected in a buffer comprising previously recorded audio of the acoustic environment.

38. The system of claim 33 , wherein the visual representation comprises an icon, and wherein the speech recognition suitability value is represented by a color of the icon.

39. The system of claim 38 , wherein displaying the visual representation of the speech recognition suitability value comprises:

determining whether the speech recognition suitability value is less than a threshold value;

in accordance with a determination that the speech recognition suitability value is less than the threshold value, displaying the icon in a grayed out state; and

in accordance with a determination that the speech recognition suitability value is not less than the threshold value, displaying the icon in a non-grayed out state.

40. The system of claim 39 , further comprising:

in accordance with a determination that the speech recognition suitability value is less than the threshold value, outputting a message indicating a low suitability of the acoustic environment of the electronic device for speech recognition.

41. The system of claim 38 , wherein the icon comprises an image of a microphone.

42. The system of claim 33 , wherein determining the speech recognition suitability value comprises periodically determining the speech recognition suitability value, and wherein displaying the visual representation of the speech recognition suitability value comprises updating the display of the visual representation of the speech recognition suitability value in accordance with the periodically determined speech recognition suitability value.

43. The system of claim 33 , wherein the speech recognition suitability value is determined based on a signal to noise ratio for a first frequency band of the acoustic environment.

44. The system of claim 33 , wherein the visual representation comprises one or more bars, and wherein a value of the speech recognition suitability value is represented by a number of the one or more bars.

45. The system of claim 33 , wherein the visual representation comprises a textual representation of the speech recognition suitability value.

46. The system of claim 33 , wherein the speech recognition suitability value comprises a numerical value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 24, 2015
From: KIM, YOON
To: APPLE INC.
Reel/Frame 036406/0303 →
Continuity (2)
Provisional Application 62057979 · Sep 30, 2014
Related Publication 20160093291A1 · Mar 31, 2016
Cited By (36)
US 12,197,712 US 12,197,817 US 12,200,297 US 12,204,932 US 12,211,502 US 12,216,894 US 12,219,314 US 12,223,282 US 12,236,952 US 12,254,887 US 12,260,234 US 12,277,954 US 12,293,203 US 12,293,763 US 12,301,635 US 12,315,495 US 12,333,404 US 12,361,934 US 12,361,943 US 12,367,879 US 12,380,876 US 12,386,434 US 12,386,491 US 12,431,128 US 12,477,470 US 12,556,890 US 12,567,415 US 12,608,171 US 12,613,621 US 12,613,730 US 12,619,452 US 12,620,179 US 12,640,151 US 12,670,639 US 12,675,839 US 12,688,425