IP Library › Granted Patent US 12,640,151
Granted Patent B2
US 12,640,151 · App. 18/368,333 · Granted May 26, 2026

Voice control with contextual keywords

Inventors: Devang K. Naik (San Jose, CA); Madhu Chinthakunta (Saratoga, CA); Paul R. Dixon (Zurich, CH); Kumari Nishu (Santa Clara, CA); Harry J. Saddler (Berkeley, CA)
Assignee: Apple Inc.
G10L15/22G10L15/02G10L15/1815
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,640,151
App. No.
18/368,333
Granted
May 26, 2026
Kind
B2
Abstract

Systems and processes for operating an intelligent automated assistant are provided. In some embodiments, contextual data is obtained and used to select a set of keywords (e.g., words or phrases) for voice control of an electronic device. When a speech input is received by the electronic device, a determination is made whether the speech input includes any of the selected keywords. If the speech input does include a selected keyword, an action is performed in response.

Claims (97)

1 . An electronic device, comprising:

one or more processors;

a memory; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:

obtaining first contextual information;

selecting, based on the first contextual information, a first set of one or more keywords, wherein selecting the first set of one or more keywords includes:

determining, based on the first contextual information, a first action that is available for performance; and

in response to determining the first action, selecting a first keyword for the first set of one or more keywords, wherein the first keyword corresponds to the first action;

receiving a speech input;

determining whether the speech input includes at least one keyword of the first set of one or more keywords; and

in response to determining that the speech input includes at least one keyword of the first set of one or more keywords, performing a first an action that corresponds to the at least one keyword.

2 . The electronic device of claim 1 , wherein obtaining the first contextual information includes receiving at least a first portion of the first contextual information from one or more sensors.

3 . The electronic device of claim 1 , wherein the first contextual information includes current device context.

4 . The electronic device of claim 3 , wherein the current device context includes user interface information.

5 . The electronic device of claim 3 , wherein the current device context includes application information.

6 . The electronic device of claim 1 , wherein the first contextual information includes user data.

7 . The electronic device of claim 1 , wherein selecting the first set of one or more keywords includes:

selecting a respective keyword of the first set of one or more keywords from a plurality of keywords.

8 . The electronic device of claim 7 , the one or more programs including instructions for:

prior to selecting the first set of one or more keywords:

receiving a natural-language speech input directed to a digital assistant, wherein the natural-language speech input includes a request to perform a third action; and

in accordance with a determination that the natural-language speech input satisfies keyword criteria, adding a candidate keyword to the plurality of keywords, wherein the candidate keyword corresponds to at least a portion of the natural-language speech input.

9 . The electronic device of claim 7 , wherein at least one keyword of the plurality of keywords is obtained from an application.

10 . The electronic device of claim 1 , wherein determining, based on the first contextual information, the first action that is available for performance includes:

determining whether the first action satisfies attention criteria.

11 . The electronic device of claim 1 , the one or more programs including instructions for:

obtaining second contextual information;

selecting, based on the second contextual information, a second set of one or more keywords different from the first set of one or more keywords;

receiving a second speech input;

determining whether the second speech input includes at least one keyword of the second set of one or more keywords; and

in accordance with a determination that the speech input includes at least one keyword of the second set of one or more keywords, performing a second action.

12 . The electronic device of claim 11 , wherein selecting the second set of one or more keywords in accordance with a determination that the second contextual information differs from the first contextual information.

13 . The electronic device of claim 1 , the one or more programs including instructions for:

providing an output indicating at least a first portion of the first set of one or more keywords.

14 . The electronic device of claim 1 , the one or more programs including instructions for:

in accordance with a determination that the speech input includes a trigger word in addition to the at least one keyword of the first set of one or more keywords, providing an output indicating that performance of the action can be caused with the at least one keyword and without the trigger word.

15 . The electronic device of claim 1 , wherein determining whether the speech input includes at least one keyword of the first set of one or more keywords includes:

extracting a set of acoustic features from the speech input.

16 . The electronic device of claim 15 , wherein determining whether the speech input includes at least one keyword of the first set of one or more keywords includes:

determining, based on the set of acoustic features, a confidence score representing a likelihood that the speech input includes the at least one keyword; and

determining whether the confidence score satisfies confidence criteria.

17 . The electronic device of claim 1 , wherein determining whether the speech input includes at least one keyword of the first set of one or more keywords is performed using a first language model.

18 . The electronic device of claim 1 , the one or more programs including instructions for:

in accordance with a determination that the speech input does not include at least one keyword of the first set of one or more keywords, forgoing performance of the action.

19 . The electronic device of claim 1 , wherein performing the action includes:

determining, using a second language model, at least one word included in the speech input.

20 . The electronic device of claim 19 , wherein the at least one keyword included in the speech input corresponds to a first portion of the speech input; and

wherein determining, using the second language model, the at least one word included in the speech input is performed in accordance with a determination that the speech input includes a second portion not corresponding to at least one keyword of the first set of one or more keywords.

21 . A method, comprising:

at an electronic device with one or more processors and memory:

obtaining first contextual information;

selecting, based on the first contextual information, a first set of one or more keywords, wherein selecting the first set of one or more keywords includes:

determining, based on the first contextual information, a first action that is available for performance; and

in response to determining the first action, selecting a first keyword for the first set of one or more keywords, wherein the first keyword corresponds to the first action;

receiving a speech input;

determining whether the speech input includes at least one keyword of the first set of one or more keywords; and

in response to determining that the speech input includes at least one keyword of the first set of one or more keywords, performing a first an action that corresponds to the at least one keyword.

22 . The method of claim 21 , wherein obtaining the first contextual information includes receiving at least a first portion of the first contextual information from one or more sensors.

23 . The method of claim 21 , wherein the first contextual information includes current device context.

24 . The method of claim 23 , wherein the current device context includes user interface information.

25 . The method of claim 23 , wherein the current device context includes application information.

26 . The method of claim 21 , wherein the first contextual information includes user data.

27 . The method of claim 21 , wherein selecting the first set of one or more keywords includes:

selecting a respective keyword of the first set of one or more keywords from a plurality of keywords.

28 . The method of claim 27 , further comprising:

prior to selecting the first set of one or more keywords:

receiving a natural-language speech input directed to a digital assistant, wherein the natural-language speech input includes a request to perform a third action; and

in accordance with a determination that the natural-language speech input satisfies keyword criteria, adding a candidate keyword to the plurality of keywords, wherein the candidate keyword corresponds to at least a portion of the natural-language speech input.

29 . The method of claim 27 , wherein at least one keyword of the plurality of keywords is obtained from an application.

30 . The method of claim 21 , wherein determining, based on the first contextual information, the first action that is available for performance includes:

determining whether the first action satisfies attention criteria.

31 . The method of claim 21 , further comprising:

in accordance with a determination that the speech input does not include at least one keyword of the first set of one or more keywords, forgoing performance of the action.

32 . A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a first electronic device, cause the first electronic device to:

obtain first contextual information;

select, based on the first contextual information, a first set of one or more keywords, wherein selecting the first set of one or more keywords includes:

determining, based on the first contextual information, a first action that is available for performance; and

in response to determining the first action, selecting a first keyword for the first set of one or more keywords, wherein the first keyword corresponds to the first action;

receive a speech input;

determine whether the speech input includes at least one keyword of the first set of one or more keywords; and

in response to determining that the speech input includes at least one keyword of the first set of one or more keywords, perform an action that corresponds to the at least one keyword.

33 . The non-transitory computer-readable storage medium of claim 32 , wherein obtaining the first contextual information includes receiving at least a first portion of the first contextual information from one or more sensors.

34 . The non-transitory computer-readable storage medium of claim 32 , wherein the first contextual information includes current device context.

35 . The non-transitory computer-readable storage medium of claim 34 , wherein the current device context includes user interface information.

36 . The non-transitory computer-readable storage medium of claim 34 , wherein the current device context includes application information.

37 . The non-transitory computer-readable storage medium of claim 32 , wherein the first contextual information includes user data.

38 . The non-transitory computer-readable storage medium of claim 32 , wherein selecting the first set of one or more keywords includes:

selecting a respective keyword of the first set of one or more keywords from a plurality of keywords.

39 . The non-transitory computer-readable storage medium of claim 38 , the one or more programs further including instructions for:

prior to selecting the first set of one or more keywords:

receiving a natural-language speech input directed to a digital assistant, wherein the natural-language speech input includes a request to perform a third action; and

in accordance with a determination that the natural-language speech input satisfies keyword criteria, adding a candidate keyword to the plurality of keywords, wherein the candidate keyword corresponds to at least a portion of the natural-language speech input.

40 . The non-transitory computer-readable storage medium of claim 38 , wherein at least one keyword of the plurality of keywords is obtained from an application.

41 . The non-transitory computer-readable storage medium of claim 32 , wherein determining, based on the first contextual information, the first action that is available for performance includes:

determining whether the first action satisfies attention criteria.

42 . The non-transitory computer-readable storage medium of claim 32 , the one or more programs further including instructions for:

in accordance with a determination that the speech input does not include at least one keyword of the first set of one or more keywords, forgoing performance of the action.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2023
From: NAIK, DEVANG K.; SADDLER, HARRY J.; CHINTHAKUNTA, MADHU; DIXON, PAUL R.; NISHU, KUMARI
To: APPLE INC.
Reel/Frame 065640/0451 →
Continuity (3)
Provisional Application 63434425 · Dec 21, 2022
Provisional Application 63407531 · Sep 16, 2022
Related Publication 20240096321A1 · Mar 21, 2024
References Cited (90)
US 7596269B2 · King et al. · 2009 [cited by applicant]
US 7899666B2 · Varone · 2011 [cited by applicant]
US 7899720B1 · Nesladek · 2011 [cited by examiner]
US 8798995B1 · Edara · 2014 [cited by applicant]
US 9628955B1 · Fleizach et al. · 2017 [cited by applicant]
US 9633004B2 · Giuli et al. · 2017 [cited by applicant]
US 9633660B2 · Haughay · 2017 [cited by applicant]
US 9633674B2 · Sinha · 2017 [cited by applicant]
US 9668121B2 · Naik et al. · 2017 [cited by applicant]
US 9679570B1 · Edara · 2017 [cited by applicant]
US 9697822B1 · Naik et al. · 2017 [cited by applicant]
US 9721566B2 · Newendorp et al. · 2017 [cited by applicant]
US 9818400B2 · Paulik et al. · 2017 [cited by applicant]
US 9858925B2 · Gruber et al. · 2018 [cited by applicant]
US 9886953B2 · Lemay et al. · 2018 [cited by applicant]
US 9922642B2 · Pitschel et al. · 2018 [cited by applicant]
US 9966065B2 · Gruber et al. · 2018 [cited by applicant]
US 9966068B2 · Cash et al. · 2018 [cited by applicant]
US 9972304B2 · Paulik et al. · 2018 [cited by applicant]
US 9986419B2 · Naik et al. · 2018 [cited by applicant]
US 10049663B2 · Orr et al. · 2018 [cited by applicant]
US 10049668B2 · Huang et al. · 2018 [cited by applicant]
US 10074360B2 · Kim · 2018 [cited by applicant]
US 10078487B2 · Gruber et al. · 2018 [cited by applicant]
US 10083688B2 · Piernot et al. · 2018 [cited by applicant]
US 10083690B2 · Giuli et al. · 2018 [cited by applicant]
US 10089072B2 · Piersol et al. · 2018 [cited by applicant]
US 10102359B2 · Cheyer · 2018 [cited by applicant]
US 10169329B2 · Futrell et al. · 2019 [cited by applicant]
US 10170123B2 · Orr et al. · 2019 [cited by applicant]
US 10176167B2 · Evermann · 2019 [cited by applicant]
US 10185542B2 · Carson et al. · 2019 [cited by applicant]
US 10186254B2 · Williams et al. · 2019 [cited by applicant]
US 10192552B2 · Raitio et al. · 2019 [cited by applicant]
US 10199051B2 · Binder et al. · 2019 [cited by applicant]
US 10223066B2 · Martel et al. · 2019 [cited by applicant]
US 10241644B2 · Gruber et al. · 2019 [cited by applicant]
US 10249300B2 · Booker et al. · 2019 [cited by applicant]
US 10269345B2 · Castillo Sanchez et al. · 2019 [cited by applicant]
US 10276170B2 · Gruber et al. · 2019 [cited by applicant]
US 10296160B2 · Shah et al. · 2019 [cited by applicant]
US 10297253B2 · Walker, II et al. · 2019 [cited by applicant]
US 10311871B2 · Newendorp et al. · 2019 [cited by applicant]
US 10417037B2 · Gruber et al. · 2019 [cited by applicant]
US 10475446B2 · Gruber et al. · 2019 [cited by applicant]
US 10497365B2 · Gruber et al. · 2019 [cited by applicant]
US 10540976B2 · Van Os et al. · 2020 [cited by applicant]
US 10568032B2 · Freeman et al. · 2020 [cited by applicant]
US 10659851B2 · Lister et al. · 2020 [cited by applicant]
US 10671428B2 · Zeitlin · 2020 [cited by applicant]
US 10706841B2 · Gruber et al. · 2020 [cited by applicant]
US 10748529B1 · Milden · 2020 [cited by applicant]
US 10791176B2 · Phipps et al. · 2020 [cited by applicant]
US 10978090B2 · Binder et al. · 2021 [cited by applicant]
US 11038934B1 · Hansen et al. · 2021 [cited by applicant]
US 11080336B2 · Van Dusen · 2021 [cited by applicant]
US 11133008B2 · Piernot et al. · 2021 [cited by applicant]
US 11151899B2 · Pitschel et al. · 2021 [cited by applicant]
US 11217255B2 · Kim et al. · 2022 [cited by applicant]
US 20040117191A1 · Seshadri · 2004 [cited by applicant]
US 20080133235A1 · Simoneau et al. · 2008 [cited by applicant]
US 20130006572A1 · Hayner · 2013 [cited by applicant]
US 20140163751A1 · Davis et al. · 2014 [cited by applicant]
US 20140337036A1 · Haiut et al. · 2014 [cited by applicant]
US 20150199965A1 · Leak et al. · 2015 [cited by applicant]
US 20150302847A1 · Yun et al. · 2015 [cited by applicant]
US 20160314792A1 · Alvarez et al. · 2016 [cited by applicant]
US 20170316779A1 · Mohapatra et al. · 2017 [cited by applicant]
US 20210125609A1 · Dusan et al. · 2021 [cited by applicant]
US 20220101827A1 · Chang · 2022 [cited by examiner]
US 20220172727A1 · Sharifi et al. · 2022 [cited by applicant]
WO 2012160567A1 · 2012 [cited by applicant]
WO 2015036817A1 · 2015 [cited by applicant]
Bell, Jason, “Machine Learning Hands-On for Developers and Technical Professionals”, Wiley, Nov. 3, 2014, 82 pages. [cited by applicant]
Bocchieri et al., “Use of Geographical Meta-Data in ASR Language and Acoustic Models”, IEEE International Conference on Acoustics Speech and Signal Processing, 2010, pp. 5118-5121. [cited by applicant]
Cambria et al., “Jumping NLP curves: A review of natural language processing research.”, IEEE Computational Intelligence magazine, 2014, vol. 9, May 2014, pp. 48-57. [cited by applicant]
Coulouris et al., “Distributed Systems: Concepts and Design (Fifth Edition)”, Addison-Wesley, May 7, 2011, 391 pages. [cited by applicant]
Jefford et al., “Professional BizTalk Server 2006”, Wrox, May 7, 2007, 398 pages. [cited by applicant]
Michalevsky et al., “Gyrophone: Recognizing Speech from Gyroscope Signals”, Proceedings of the 23rd USENIX Security Symposium, Aug. 20-22, 2014, pp. 1053-1067. [cited by applicant]
Nakamura, Satoshi, “Overcoming the Language Barrier with Speech Translation Technology”, Science & Technology Trends, Quarterly Review No. 31, Apr. 2009, pp. 36-49. [cited by applicant]
“Nuance Dragon Naturally Speaking, Version 13 End-User Workbook, Nuance Communications, Inc.”, Online Available at:https://www.nuance.com/content/dam/nuance/en_us/collateral/dragon/guide/gd-dragon-naturally-speaking-13-… [cited by applicant]
Phoenix Solutions, Inc., “Declaration of Christopher Schmandt Regarding the MIT Galaxy System”, West Interactive Corp., a Delaware Corporation, Document 40, Jul. 2, 2010, 162 pages. [cited by applicant]
Qian et al., “Single-channel Multi-talker Speech Recognition With Permutation Invariant Training”, Speech Communication, Issue 104, 2018, pp. 1-11. [cited by applicant]
Rowland et al., “Designing Connected Products: UX for the Consumer Internet of Things”, O'Reilly, May 31, 2015, 452 pages. [cited by applicant]
Samsung, “SGH-a885 Series—Portable Quad-Band Mobile Phone-User Manual”, Retrieved from the Internet: URL: http://web.archive.org/web/20100106113758/http://www.comparecellular.com/images/phones/userguide1896.pdf, Jan. 1,… [cited by applicant]
Seroter et al., “SOA Patterns with BizTalk Server 2013 and Microsoft Azure”, Second Edition, Packt Publishing, Jun. 30, 2015, 454 pages. [cited by applicant]
Stent et al., “Geo-Centric Language Models for Local Business Voice Search”, AT&T Labs—Research, 2009, pp. 389-396. [cited by applicant]
Zhao et al., “SeeingVR: A Set of Tools to Make Virtual Reality More Accessible to People with Low Vision”, In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (CHI '19). ACM, Article 111, Gla… [cited by applicant]
Zheng et al., “Intent Detection and Semantic Parsing for Navigation Dialogue Language Processing”, 2017 IEEE 20th International Conference on Intelligent Transportation Systems (ITSC), 2017, 6 pages. [cited by applicant]
Zhong et al., “JustSpeak: Enabling Universal Voice Control on Android”, W4A'14, Proceedings of the 11th Web for All Conference, No. 36, Apr. 7-9, 2014, 8 pages. [cited by applicant]