IP Library Granted Patent US 11,609,739
Granted Patent B2
US 11,609,739 · App. 17/056,126 · Granted Mar 21, 2023

Providing audio information with a digital assistant

Inventors: Rahul Nair (Daly City, CA); Golnaz Abdollahian (San Francisco, CA); Avi Bar-Zeev (Oakland, CA); Niranjan Manjunath (San Jose, CA)
Assignee: Apple Inc.
G06F3/167G06F3/011G06F3/165G10L15/22G10L2015/227H04R1/406
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,609,739
App. No.
17/056,126
Granted
Mar 21, 2023
Kind
B2
Abstract

In an exemplary technique for providing audio information, an input is received, and audio information responsive to the received input is provided using a speaker. While providing the audio information, an external sound is detected. If it is determined that the external sound is a communication of a first type, then the provision of the audio information is stopped. If it is determined that the external sound is a communication of a second type, then the provision of the audio information continues.

Claims (92)

1. An electronic device, comprising:

one or more processors; and

memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for:

providing, using a speaker, audio information responsive to received input;

while providing the audio information, detecting an external sound;

in accordance with a determination that the external sound is a communication of a first type, stopping the provision of the audio information;

in accordance with a determination that the external sound is a communication of a second type, continuing the provision of the audio information;

after stopping the provision of the audio information:

detecting one or more visual characteristics associated with the communication of the first type; and

detecting the communication of the first type has stopped;

in response to detecting the communication of the first type has stopped, determining whether the one or more visual characteristics indicate that further communication of the first type is expected;

in accordance with a determination that further communication of the first type is not expected, providing resumed audio information; and

in accordance with a determination that further communication of the first type is expected, continuing to stop the provision of the audio information.

2. The electronic device of claim 1 , wherein the one or more visual characteristics include eye gaze, facial expression, hand gesture, or a combination thereof.

3. The electronic device of claim 1 , wherein stopping the provision of the audio information includes fading out the audio information.

4. The electronic device of claim 1 , the one or more programs further including instructions for:

after stopping the provision of the audio information and in accordance with a determination that the communication of the first type has stopped, providing resumed audio information.

5. The electronic device of claim 4 , wherein the audio information is divided into predefined segments, and the resumed audio information begins with the segment where the audio information was stopped.

6. The electronic device of claim 5 , wherein the resumed audio information includes a rephrased version of a previously provided segment of the audio information.

7. The electronic device of claim 1 , wherein the received input includes a triggering command.

8. The electronic device of claim 1 , wherein the communication of the first type includes a directly-vocalized lexical utterance.

9. The electronic device of claim 8 , wherein the directly-vocalized lexical utterance excludes silencing commands.

10. The electronic device of claim 8 , the one or more programs further including instructions for:

determining the external sound is a directly-vocalized lexical utterance by determining a location corresponding to a source of the external sound.

11. The electronic device of claim 10 , wherein the location is determined with a directional microphone array.

12. The electronic device of claim 1 , wherein the communication of the second type includes conversational sounds.

13. The electronic device of claim 1 , wherein the communication of the second type includes compressed audio.

14. The electronic device of claim 1 , wherein the communication of the second type includes a lexical utterance reproduced by an electronic device.

15. The electronic device of claim 14 , the one or more programs further including instructions for:

determining the external sound is a lexical utterance reproduced by an electronic device by determining a location corresponding to a source of the external sound.

16. The electronic device of claim 15 , wherein the location is determined with a directional microphone array.

17. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device, the one or more programs including instructions for:

providing, using a speaker, audio information responsive to received input;

while providing the audio information, detecting an external sound;

in accordance with a determination that the external sound is a communication of a first type, stopping the provision of the audio information;

in accordance with a determination that the external sound is a communication of a second type, continuing the provision of the audio information;

after stopping the provision of the audio information:

detecting one or more visual characteristics associated with the communication of the first type; and

detecting the communication of the first type has stopped;

in response to detecting the communication of the first type has stopped, determining whether the one or more visual characteristics indicate that further communication of the first type is expected;

in accordance with a determination that further communication of the first type is not expected, providing resumed audio information; and

in accordance with a determination that further communication of the first type is expected, continuing to stop the provision of the audio information.

18. A method, comprising:

providing, using a speaker, audio information responsive to received input;

while providing the audio information, detecting an external sound;

in accordance with a determination that the external sound is a communication of a first type, stopping the provision of the audio information;

in accordance with a determination that the external sound is a communication of a second type, continuing the provision of the audio information;

after stopping the provision of the audio information:

detecting one or more visual characteristics associated with the communication of the first type; and

detecting the communication of the first type has stopped;

in response to detecting the communication of the first type has stopped, determining whether the one or more visual characteristics indicate that further communication of the first type is expected;

in accordance with a determination that further communication of the first type is not expected, providing resumed audio information; and

in accordance with a determination that further communication of the first type is expected, continuing to stop the provision of the audio information.

19. The electronic device of claim 1 , wherein the one or more visual characteristics include a gaze direction of a user associated with the communication of the first type.

20. The non-transitory computer-readable storage medium of claim 17 , wherein the one or more visual characteristics include eye gaze, facial expression, hand gesture, or a combination thereof.

21. The non-transitory computer-readable storage medium of claim 17 , wherein stopping the provision of the audio information includes fading out the audio information.

22. The non-transitory computer-readable storage medium of claim 17 , the one or more programs further including instructions for:

after stopping the provision of the audio information and in accordance with a determination that the communication of the first type has stopped, providing resumed audio information.

23. The non-transitory computer-readable storage medium of claim 22 , wherein the audio information is divided into predefined segments, and the resumed audio information begins with the segment where the audio information was stopped.

24. The non-transitory computer-readable storage medium of claim 23 , wherein the resumed audio information includes a rephrased version of a previously provided segment of the audio information.

25. The non-transitory computer-readable storage medium of claim 17 , wherein the received input includes a triggering command.

26. The non-transitory computer-readable storage medium of claim 17 , wherein the communication of the first type includes a directly-vocalized lexical utterance.

27. The non-transitory computer-readable storage medium of claim 26 , wherein the directly-vocalized lexical utterance excludes silencing commands.

28. The non-transitory computer-readable storage medium of claim 26 , the one or more programs further including instructions for:

determining the external sound is a directly-vocalized lexical utterance by determining a location corresponding to a source of the external sound.

29. The non-transitory computer-readable storage medium of claim 28 , wherein the location is determined with a directional microphone array.

30. The non-transitory computer-readable storage medium of claim 17 , wherein the communication of the second type includes conversational sounds.

31. The non-transitory computer-readable storage medium of claim 17 , wherein the communication of the second type includes compressed audio.

32. The non-transitory computer-readable storage medium of claim 17 , wherein the communication of the second type includes a lexical utterance reproduced by an electronic device.

33. The non-transitory computer-readable storage medium of claim 32 , the one or more programs further including instructions for:

determining the external sound is a lexical utterance reproduced by an electronic device by determining a location corresponding to a source of the external sound.

34. The non-transitory computer-readable storage medium of claim 33 , wherein the location is determined with a directional microphone array.

35. The non-transitory computer-readable storage medium of claim 17 , wherein the one or more visual characteristics include a gaze direction of a user associated with the communication of the first type.

36. The method of claim 18 , wherein the one or more visual characteristics include eye gaze, facial expression, hand gesture, or a combination thereof.

37. The method of claim 18 , wherein stopping the provision of the audio information includes fading out the audio information.

38. The method of claim 18 , further comprising:

after stopping the provision of the audio information and in accordance with a determination that the communication of the first type has stopped, providing resumed audio information.

39. The method of claim 38 , wherein the audio information is divided into predefined segments, and the resumed audio information begins with the segment where the audio information was stopped.

40. The method of claim 39 , wherein the resumed audio information includes a rephrased version of a previously provided segment of the audio information.

41. The method of claim 18 , wherein the received input includes a triggering command.

42. The method of claim 18 , wherein the communication of the first type includes a directly-vocalized lexical utterance.

43. The method of claim 42 , wherein the directly-vocalized lexical utterance excludes silencing commands.

44. The method of claim 42 , further comprising:

determining the external sound is a directly-vocalized lexical utterance by determining a location corresponding to a source of the external sound.

45. The method of claim 44 , wherein the location is determined with a directional microphone array.

46. The method of claim 18 , wherein the communication of the second type includes conversational sounds.

47. The method of claim 18 , wherein the communication of the second type includes compressed audio.

48. The method of claim 18 , wherein the communication of the second type includes a lexical utterance reproduced by an electronic device.

49. The method of claim 48 , further comprising:

determining the external sound is a lexical utterance reproduced by an electronic device by determining a location corresponding to a source of the external sound.

50. The method of claim 49 , wherein the location is determined with a directional microphone array.

51. The method of claim 18 , wherein the one or more visual characteristics include a gaze direction of a user associated with the communication of the first type.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 12, 2021
From: NAIR, RAHUL; ABDOLLAHIAN, GOLNAZ; BAR-ZEEV, AVI; MANJUNATH, NIRANJAN
To: APPLE INC.
Reel/Frame 054894/0796 →
Continuity (2)
Provisional Application 62679644 · Jun 1, 2018
Related Publication 20210224031A1 · Jul 22, 2021