IP Library › Granted Patent US 10,395,659
Granted Patent B2
US 10,395,659 · App. 15/679,098 · Granted Aug 27, 2019

Providing an auditory-based interface of a digital assistant

Inventors: Aimee Piercy (Mountain View, CA); Cyrus Daniel Irani (Los Altos, CA); Yoon Kim (Cupertino, CA); David Chance Graham (Campbell, CA); Patrick L. Coffman (San Francisco, CA)
Assignee: Apple Inc.
G10L17/22G06F3/167G06F16/2457G06F16/9537G10L13/033G10L15/22G06N20/00G10L13/027G10L15/30G10L21/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,395,659
App. No.
15/679,098
Granted
Aug 27, 2019
Kind
B2
Abstract

Systems and processes for operating an intelligent automated assistant are provided. In accordance with one example, a method includes, at an electronic device with one or more processors and memory, receiving a natural-language speech input indicative of a request to the digital assistant; obtaining, by the digital assistant, context information; determining, by the digital assistant, a text-to-speech mode from a plurality of text-to-speech modes based on the obtained context information; and providing, by the digital assistant, an audio output with the determined text-to-speech mode, where the audio output is indicative of a speech response to the user request.

Claims (97)

1. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device, cause the electronic device to:

receive a natural-language speech input indicative of a request to a digital assistant;

obtain a representation of user intent based on the natural-language speech input;

determine one or more parameters for a task based on the representation of user intent;

obtain, by the digital assistant, context information including the determined one or more parameters;

determine, by the digital assistant, a text-to-speech mode from a plurality of text-to-speech modes based on the obtained context information; and

provide, by the digital assistant, an audio output with the determined text-to-speech mode, wherein the audio output is indicative of a speech response to the user request.

2. The non-transitory computer-readable storage medium of claim 1 , wherein the context information comprises a current time at the electronic device.

3. The non-transitory computer-readable storage medium of claim 1 , wherein the context information comprises a location of the electronic device.

4. The non-transitory computer-readable storage medium of claim 1 , wherein the context information comprises information corresponding to an environment of the electronic device.

5. The non-transitory computer-readable storage medium of claim 1 , further comprising instructions, which when executed by one or more processors of the electronic device, cause the electronic device to:

obtain a text string by performing a speech-to-text analysis of the natural-language speech input; and

wherein obtaining the representation of user intent includes interpreting the text string to obtain the representation of user intent, wherein the interpreting is based on a natural-language processing of the text string.

6. The non-transitory computer-readable storage medium of claim further comprising instructions, which when executed by one or more processors of the electronic device, cause the electronic device to:

perform the task in accordance with the one or more parameters; and

obtain one or more results based on the performance of the task.

7. The non-transitory computer-readable storage medium of claim 6 , wherein the context information comprises information corresponding to the one or more results.

8. The non-transitory computer-readable storage medium of claim 1 , wherein the context information comprises information corresponding to one or more previous user interactions with the digital assistant.

9. The non-transitory computer-readable storage medium of claim 1 , wherein at least one text-to-speech mode of the plurality of text-to-speech modes corresponds to an emotional mood.

10. The non-transitory computer-readable storage medium of claim 1 , wherein at least one text-to-speech mode of the plurality of text-to-speech modes corresponds to a tone.

11. The non-transitory computer-readable storage medium of claim 1 , wherein at least one text-to-speech mode of the plurality of text-to-speech modes corresponds to a voice profile.

12. The non-transitory computer-readable storage medium of claim 1 ,

wherein the context information is a first set of context information,

wherein the text-to-speech mode is a first text-to-speech mode,

wherein the audio output is a first audio output, and wherein the one or more programs further comprise instructions, which when executed by one or more processors of the electronic device, cause the electronic device to:

obtain, by the digital assistant, a second set of context information;

determine, by the digital assistant, a second text-to-speech mode from the plurality of text-to-speech modes based on the second set of context information; and

provide, by the digital assistant, a second audio output with the second text-to-speech mode.

13. The non-transitory computer-readable storage medium of claim 1 , further comprising instructions, which when executed by one or more processors of the electronic device, cause the electronic device to:

determine by the digital assistant, the audio output based on the context information.

14. The non-transitory computer-readable storage medium of claim 1 , wherein the electronic device is a computer, a set-top box, a speaker, a smart watch, a phone, or any combination thereof.

15. An electronic device, comprising:

one or more processors;

a memory; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:

receiving a natural-language speech input indicative of a request to a digital assistant;

obtaining a representation of user intent based on the natural-language speech input;

determining one or more parameters for a task based on the representation of user intent;

obtaining, by the digital assistant, context information including the determined one or more parameters;

determining, by the digital assistant, a text-to-speech mode from a plurality of text-to-speech modes based on the obtained context information; and

providing, by the digital assistant, an audio output with the determined text-to-speech mode, wherein the audio output is indicative of a speech response to the user request.

16. The electronic device of claim 15 , wherein the context information comprises a current time at the electronic device.

17. The electronic device of claim 15 , wherein the context information comprises a location of the electronic device.

18. The electronic device of claim 15 , wherein the context information comprises information corresponding to an environment of the electronic device.

19. The electronic device of claim 15 , the one or more programs further including instructions for:

obtaining a text string by performing a speech-to-text analysis of the natural-language speech input; and

wherein obtaining the representation of user intent includes interpreting the text string to obtain the representation of user intent, wherein the interpreting is based on a natural-language processing of the text string.

20. The electronic device of claim 15 , the one or more programs further including instructions for:

performing the task in accordance with the one or more parameters; and

obtaining one or more results based on the performance of the task.

21. The electronic device of claim 20 , wherein the context on comprises information corresponding to the one or more results.

22. The electronic device of claim 15 , wherein the context information comprises information corresponding to one or more previous user interactions with the digital assistant.

23. The electronic device of claim 15 , wherein at least one text-to-speech mode of the plurality of text-to-speech modes corresponds to an emotional mood.

24. The electronic device of claim 15 , wherein at least one text-to-speech mode of the plurality of text-to-speech modes corresponds to a tone.

25. The electronic device of claim 15 , wherein at least one text-to-speech mode of the plurality of text-to-speech modes corresponds to a voice profile.

26. The electronic device of claim 15 ,

wherein the context information is a first set of context information,

wherein the text-to-speech mode is a first text-to-speech mode,

wherein the audio output is a first audio output, and the one or more programs further including instructions for:

obtaining, by the digital assistant, a second set of context information;

determining, by the digital assistant, a second text-to-speech mode from the plurality of text-to-speech modes based on the second set of context information; and

providing, by the digital assistant, a second audio output with the second text-to-speech mode.

27. The electronic device of claim 15 , the one or more programs further including instructions for:

determining, by the digital assistant, the audio output based on the context information.

28. The electronic device of claim 15 , wherein the electronic device is a computer, a set-top box, a speaker, a smart watch, a phone, or any combination thereof.

29. A method for operating a digital assistant, comprising:

at an electronic device with one or more processors and memory:

receiving a natural-language speech input indicative of a request to the digital assistant,

obtaining a representation of user intent based on the natural-language speech input;

determining one or more parameters for a task based on the representation of user intent;

obtaining, by the digital assistant, context information including the determined one or more parameters;

determining, by the digital assistant, a text-to-speech mode from a plurality of text-to-speech modes based on the obtained context information; and

providing, by the digital assistant, an audio output with the determined text-to-speech mode, wherein the audio output is indicative of a speech response to the user request.

30. The method of claim 29 , wherein the context information comprises a current time at the electronic device.

31. The method of claim 29 , wherein the context information comprises a location of the electronic device.

32. The method of claim 29 ; wherein the context information comprises information corresponding to an environment of the electronic device.

33. The method of claim 29 , further comprising:

obtaining a text string by performing a speech-to-text analysis of the natural-language speech input; and

wherein obtaining the representation of user intent includes interpreting the text string to obtain the representation of user intent, wherein the interpreting is based on a natural-language processing of the text string.

34. The method of claim 29 , further comprising:

performing the task in accordance with the one or more parameters; and

obtaining one or more results based on the performance of the task.

35. The method of claim 34 , wherein the context information comprises information corresponding to the one or more results.

36. The method of claim 29 , wherein the context information comprises information corresponding to one or more previous user interactions with the digital assistant.

37. The method of claim 29 , wherein at least one text-to-speech mode of the plurality of text-to-speech modes corresponds to an emotional mood.

38. The method of claim 29 , wherein at least one text-to-speech mode of the plurality of text-to-speech modes corresponds to a tone.

39. The method of claim 29 , wherein at least one text-to-speech mode of the plurality of text-to-speech modes corresponds to a voice profile.

40. The method of claim 29 ,

wherein the context information is a first set of context information,

wherein the text-to-speech mode is a first text-to-speech mode,

wherein the audio output is a first audio output, and the method further comprising:

obtaining, by the digital assistant, a second set of context information;

determining, by the digital assistant, a second text-to-speech mode from the plurality of text-to-speech modes based on the second set of context information; and

providing, by the digital assistant, a second audio output with the second text-to-speech mode.

41. The method of claim 29 , further comprising:

determining, by the digital assistant, the audio output based on the context information.

42. The method of claim 29 , wherein the electronic device is a computer, a set-top box, a speaker, a smart watch, a phone, or any combination thereof.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 17, 2018
From: PIERCY, AIMEE; IRANI, CYRUS DANIEL; KIM, YOON; GRAHAM, DAVID CHANCE; COFFMAN, PATRICK L.
To: APPLE INC.
Reel/Frame 044644/0810 →
Continuity (2)
Provisional Application 62507056 · May 16, 2017
Related Publication 20180336904A1 · Nov 22, 2018
Cited By (5)
US 12,393,967 US 12,469,033 US 12,547,963 US 12,561,721 US 12,586,582