IP Library Granted Patent US 12,190,873
Granted Patent B2
US 12,190,873 · App. 17/952,005 · Granted Jan 7, 2025

Determining whether speech input is intended for a digital assistant

Inventors: Ahmed S. Hussen Abdelaziz (San Ramon, CA); Saurabh Adya (San Jose, CA); Alexander W. Churchill (London, GB); Pranay Dighe (Berkeley, CA); Sachin S. Kajarekar (Sunnyvale, CA); Chaitanya Mannemala (San Ramon, CA); Erik Marchi (Zurich, CH); Seyedmahdad Mirsamadi (Santa Clara, CA); Ognjen Rudovic (Seattle, WA); Ahmed H. Tewfik (Los Altos, CA); Barry-John Theobald (San Jose, CA); Srikanth Vishnubhotla (Santa Clara, CA)
Assignee: Apple Inc.
G10L15/197G06T7/70G06V40/161G10L15/16G10L15/22G10L25/78G06T2207/30201G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,873
App. No.
17/952,005
Granted
Jan 7, 2025
Kind
B2
Abstract

An example process includes: receiving a speech input representing a user utterance; determining, based on a textual representation of the speech input, a first score corresponding to a type of the user utterance; determining, based on the textual representation of the speech input, a second score representing a correspondence between the user utterance and a domain recognized by a digital assistant; determining, based on the first score and the second score, whether the speech input is intended for the digital assistant; in accordance with a determination that the speech input is intended for the digital assistant: initiating, by the digital assistant, a task based on the speech input; and providing an output indicative of the initiated task.

Claims (157)

1. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device, cause the electronic device to:

receive a first speech input representing a first user utterance;

initiate, by a digital assistant operating on the electronic device, a first task based on the first speech input;

provide a first output indicative of the initiated first task; and

after providing the first output:

receive a second speech input following the first speech input, the second speech input representing a second user utterance;

determine, based on a textual representation of the second speech input, a first score representing a correspondence between the second user utterance and a domain recognized by the digital assistant;

determine, based on the textual representation of the second speech input, a second score representing contextual continuity between the first user utterance and the second user utterance;

determine, based on the first score and the second score, whether the second speech input is intended for the digital assistant; and

in accordance with a determination that the second speech input is intended for the digital assistant:

initiate, by the digital assistant, a second task based on the second speech input; and

provide a second output indicative of the initiated second task.

2. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

in accordance with a determination that the second speech input is not intended for the digital assistant:

forgo initiating the second task.

3. The non-transitory computer-readable storage medium of claim 1 , wherein the first speech input and the second speech input are received within a same digital assistant session.

4. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

determine, based on detecting a spoken trigger for initiating a digital assistant session, that the first speech input is intended for the digital assistant, wherein:

initiating the first task is performed in accordance with a determination that the first speech input is intended for the digital assistant; and

determining whether the second speech input is intended for the digital assistant is performed without detecting the spoken trigger.

5. The non-transitory computer-readable storage medium of claim 1 , wherein determining whether the second speech input is intended for the digital assistant is performed without detecting a selection of a displayed affordance and without detecting a selection of a button of the electronic device.

6. The non-transitory computer-readable storage medium of claim 1 , wherein determining the first score includes determining whether the second user utterance corresponds to a plurality of utterances recognized by the digital assistant.

7. The non-transitory computer-readable storage medium of claim 1 , wherein determining the first score includes determining whether the second user utterance corresponds to a vocabulary associated with the domain recognized by the digital assistant.

8. The non-transitory computer-readable storage medium of claim 1 , wherein determining the first score includes determining the first score using a binary classification neural network.

9. The non-transitory computer-readable storage medium of claim 1 , wherein determining the second score includes determining the second score based on a textual representation of the first user utterance.

10. The non-transitory computer-readable storage medium of claim 1 , wherein determining the second score includes determining the second score using a second binary classification neural network.

11. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

determine, based on the textual representation of the second speech input, a third score corresponding to a type of the second user utterance, wherein determining whether the second speech input is intended for the digital assistant is further based on the third score.

12. The non-transitory computer-readable storage medium of claim 11 , wherein determining the third score includes:

determining respective probabilities that the second user utterance corresponds to each of a plurality of user utterance types;

selecting a subset of the respective probabilities, wherein each probability of the subset of the respective probabilities corresponds to a respective predetermined user utterance type of the plurality of user utterance types; and

determining the third score based on the subset of the respective probabilities.

13. The non-transitory computer-readable storage medium of claim 12 , wherein the plurality of user utterance types include:

a first type of command;

a second type of command;

a first type of question;

a second type of question;

a first type of answer;

a second type of answer;

an opinion; and

a statement.

14. The non-transitory computer-readable storage medium of claim 13 , wherein the respective predetermined user utterance types include:

the first type of command; and

the first type of question.

15. The non-transitory computer-readable storage medium of claim 11 , wherein determining whether the second speech input is intended for the digital assistant includes:

weighting the first score, the second score, and the third score to obtain a final score indicating whether the second speech input is intended for the digital assistant;

comparing the final score to a threshold;

in accordance with a determination that the final score is above the threshold:

determining that the second speech input is intended for the digital assistant; and

in accordance with a determination that the final score is below the threshold:

determining that the second speech input is not intended for the digital assistant.

16. An electronic device, comprising:

one or more processors;

a memory; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:

receiving a first speech input representing a first user utterance;

initiating, by a digital assistant operating on the electronic device, a first task based on the first speech input;

providing a first output indicative of the initiated first task; and

after providing the first output:

receiving a second speech input following the first speech input, the second speech input representing a second user utterance;

determining, based on a textual representation of the second speech input, a first score representing a correspondence between the second user utterance and a domain recognized by the digital assistant;

determining, based on the textual representation of the second speech input, a second score representing contextual continuity between the first user utterance and the second user utterance;

determining, based on the first score and the second score, whether the second speech input is intended for the digital assistant; and

in accordance with a determination that the second speech input is intended for the digital assistant:

initiating, by the digital assistant, a second task based on the second speech input; and

providing a second output indicative of the initiated second task.

17. The electronic device of claim 16 , the one or more programs further including instructions for:

in accordance with a determination that the second speech input is not intended for the digital assistant:

forgoing initiating the second task.

18. The electronic device of claim 16 , wherein the first speech input and the second speech input are received within a same digital assistant session.

19. The electronic device of claim 16 , the one or more programs further including instructions for:

determining, based on detecting a spoken trigger for initiating a digital assistant session, that the first speech input is intended for the digital assistant, wherein:

initiating the first task is performed in accordance with a determination that the first speech input is intended for the digital assistant; and

determining whether the second speech input is intended for the digital assistant is performed without detecting the spoken trigger.

20. The electronic device of claim 16 , wherein determining whether the second speech input is intended for the digital assistant is performed without detecting a selection of a displayed affordance and without detecting a selection of a button of the electronic device.

21. The electronic device of claim 16 , wherein determining the first score includes determining whether the second user utterance corresponds to a plurality of utterances recognized by the digital assistant.

22. The electronic device of claim 16 , wherein determining the first score includes determining whether the second user utterance corresponds to a vocabulary associated with the domain recognized by the digital assistant.

23. The electronic device of claim 16 , wherein determining the first score includes determining the first score using a binary classification neural network.

24. The electronic device of claim 16 , wherein determining the second score includes determining the second score based on a textual representation of the first user utterance.

25. The electronic device of claim 16 , wherein determining the second score includes determining the second score using a second binary classification neural network.

26. The electronic device of claim 16 , the one or more programs further including instructions for:

determining, based on the textual representation of the second speech input, a third score corresponding to a type of the second user utterance, wherein determining whether the second speech input is intended for the digital assistant is further based on the third score.

27. The electronic device of claim 26 , wherein determining the third score includes:

determining respective probabilities that the second user utterance corresponds to each of a plurality of user utterance types;

selecting a subset of the respective probabilities, wherein each probability of the subset of the respective probabilities corresponds to a respective predetermined user utterance type of the plurality of user utterance types; and

determining the third score based on the subset of the respective probabilities.

28. The electronic device of claim 27 , wherein the plurality of user utterance types include:

a first type of command;

a second type of command;

a first type of question;

a second type of question;

a first type of answer;

a second type of answer;

an opinion; and

a statement.

29. The electronic device of claim 28 , wherein the respective predetermined user utterance types include:

the first type of command; and

the first type of question.

30. The electronic device of claim 26 , wherein determining whether the second speech input is intended for the digital assistant includes:

weighting the first score, the second score, and the third score to obtain a final score indicating whether the second speech input is intended for the digital assistant;

comparing the final score to a threshold;

in accordance with a determination that the final score is above the threshold:

determining that the second speech input is intended for the digital assistant; and

in accordance with a determination that the final score is below the threshold:

determining that the second speech input is not intended for the digital assistant.

31. A method, comprising:

at an electronic device with one or more processors and memory:

receiving a first speech input representing a first user utterance;

initiating, by a digital assistant operating on the electronic device, a first task based on the first speech input;

providing a first output indicative of the initiated first task; and

after providing the first output:

receiving a second speech input following the first speech input, the second speech input representing a second user utterance;

determining, based on a textual representation of the second speech input, a first score representing a correspondence between the second user utterance and a domain recognized by the digital assistant;

determining, based on the textual representation of the second speech input, a second score representing contextual continuity between the first user utterance and the second user utterance;

determining, based on the first score and the second score, whether the second speech input is intended for the digital assistant; and

in accordance with a determination that the second speech input is intended for the digital assistant:

initiating, by the digital assistant, a second task based on the second speech input, and

providing a second output indicative of the initiated second task.

32. The method of claim 31 , further comprising:

in accordance with a determination that the second speech input is not intended for the digital assistant:

forgoing initiating the second task.

33. The method of claim 31 , wherein the first speech input and the second speech input are received within a same digital assistant session.

34. The method of claim 31 , further comprising:

determining, based on detecting a spoken trigger for initiating a digital assistant session, that the first speech input is intended for the digital assistant, wherein:

initiating the first task is performed in accordance with a determination that the first speech input is intended for the digital assistant; and

determining whether the second speech input is intended for the digital assistant is performed without detecting the spoken trigger.

35. The method of claim 31 , wherein determining whether the second speech input is intended for the digital assistant is performed without detecting a selection of a displayed affordance and without detecting a selection of a button of the electronic device.

36. The method of claim 31 , wherein determining the first score includes determining whether the second user utterance corresponds to a plurality of utterances recognized by the digital assistant.

37. The method of claim 31 , wherein determining the first score includes determining whether the second user utterance corresponds to a vocabulary associated with the domain recognized by the digital assistant.

38. The method of claim 31 , wherein determining the first score includes determining the first score using a binary classification neural network.

39. The method of claim 31 , wherein determining the second score includes determining the second score based on a textual representation of the first user utterance.

40. The method of claim 31 , wherein determining the second score includes determining the second score using a second binary classification neural network.

41. The method of claim 31 , further comprising:

determining, based on the textual representation of the second speech input, a third score corresponding to a type of the second user utterance, wherein determining whether the second speech input is intended for the digital assistant is further based on the third score.

42. The method of claim 41 , wherein determining the third score includes:

determining respective probabilities that the second user utterance corresponds to each of a plurality of user utterance types;

selecting a subset of the respective probabilities, wherein each probability of the subset of the respective probabilities corresponds to a respective predetermined user utterance type of the plurality of user utterance types; and

determining the third score based on the subset of the respective probabilities.

43. The method of claim 42 , wherein the plurality of user utterance types include:

a first type of command;

a second type of command;

a first type of question;

a second type of question;

a first type of answer;

a second type of answer;

an opinion; and

a statement.

44. The method of claim 43 , wherein the respective predetermined user utterance types include:

the first type of command; and

the first type of question.

45. The method of claim 41 , wherein determining whether the second speech input is intended for the digital assistant includes:

weighting the first score, the second score, and the third score to obtain a final score indicating whether the second speech input is intended for the digital assistant;

comparing the final score to a threshold;

in accordance with a determination that the final score is above the threshold:

determining that the second speech input is intended for the digital assistant; and

in accordance with a determination that the final score is below the threshold:

determining that the second speech input is not intended for the digital assistant.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2023
From: MARCHI, ERIK; ADYA, SAURABH; DIGHE, PRANAY; RUDOVIC, OGNJEN; THEOBALD, BARRY-JOHN; MIRSAMADI, SEYEDMAHDAD; HUSSEN ABDELAZIZ, AHMED S.; KAJAREKAR, SACHIN S.; VISHNUBHOTLA, SRIKANTH; MANNEMALA, CHAITANYA; TEWFIK, AHMED H.; CHURCHILL, ALEXANDER W.
To: APPLE INC.
Reel/Frame 065782/0246 →
Continuity (2)
Provisional Application 63341893 · May 13, 2022
Related Publication 20230368783A1 · Nov 16, 2023
References Cited (37)
US 7620549B2 · Di Cristo et al. · 2009 [cited by applicant]
US 8731942B2 · Cheyer et al. · 2014 [cited by applicant]
US 8898568B2 · Bull et al. · 2014 [cited by applicant]
US 9262612B2 · Cheyer · 2016 [cited by applicant]
US 9548050B2 · Gruber et al. · 2017 [cited by applicant]
US 9668121B2 · Naik et al. · 2017 [cited by applicant]
US 9721566B2 · Newendorp et al. · 2017 [cited by applicant]
US 9858927B2 · Williams · 2018 [cited by examiner]
US 9986419B2 · Naik et al. · 2018 [cited by applicant]
US 10049663B2 · Orr et al. · 2018 [cited by applicant]
US 10083688B2 · Piernot · 2018 [cited by examiner]
US 10102359B2 · Cheyer · 2018 [cited by applicant]
US 10170123B2 · Orr et al. · 2019 [cited by applicant]
US 10249300B2 · Booker et al. · 2019 [cited by applicant]
US 10269345B2 · Sanchez et al. · 2019 [cited by applicant]
US 10311871B2 · Newendorp et al. · 2019 [cited by applicant]
US 10410637B2 · Paulik et al. · 2019 [cited by applicant]
US 10748546B2 · Kim · 2020 [cited by examiner]
US 11133008B2 · Piernot · 2021 [cited by examiner]
US 11195524B2 · Mukherjee · 2021 [cited by examiner]
US 11423898B2 · Shum · 2022 [cited by examiner]
US 20090112572A1 · Thorn · 2009 [cited by examiner]
US 20090304198A1 · Herre · 2009 [cited by examiner]
US 20100332003A1 · Yaguez · 2010 [cited by applicant]
US 20150370531A1 · Faaborg · 2015 [cited by applicant]
US 20170358304A1 · Castillo Sanchez · 2017 [cited by examiner]
US 20180090143A1 · Saddler · 2018 [cited by examiner]
US 20180233139A1 · Finkelstein · 2018 [cited by examiner]
US 20190387352A1 · Jot et al. · 2019 [cited by applicant]
US 20200105260A1 · Piernot · 2020 [cited by examiner]
US 20200279556A1 · Gruber et al. · 2020 [cited by applicant]
US 20200380980A1 · Shum · 2020 [cited by examiner]
US 20220093101A1 · Krishnan · 2022 [cited by examiner]
US 20230186921A1 · Paulik · 2023 [cited by examiner]
US 20230368812A1 · Marchi · 2023 [cited by examiner]
WO 2013173504A1 · 2013 [cited by applicant]
WO 2015184186A1 · 2015 [cited by applicant]
Cited By (1)
US 12,640,142