IP Library › Granted Patent US 11,783,815
Granted Patent B2
US 11,783,815 · App. 17/732,243 · Granted Oct 10, 2023

Multimodality in digital assistant systems

Inventors: Pierre P. Greborio (Sunnyvale, CA); Didier Rene Guzzoni (Montsur-Rolle, CH); Philippe P. Piernot (Palo Alto, CA)
Assignee: Apple Inc.
G10L15/1815G10L15/1822G10L15/22G10L15/30G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,783,815
App. No.
17/732,243
Granted
Oct 10, 2023
Kind
B2
Abstract

Systems and processes for operating an intelligent automated assistant are provided. An example process for determining user intent includes receiving a natural language input and detecting an event. The process further includes, determining, at a first time, based on the natural language input, a first value for a first node of a parsing structure; and determining, at a second time, based on the detected data event, a second value for a second node of the parsing structure. The process further includes in accordance with a determination that the first time and the second time are within the predetermined time: determining, using the parsing structure, the first value, and the second value, a user intent associated with the natural language input; initiating a task based on the determined intent; and providing an output indicative of the task.

Claims (89)

1. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device, cause the electronic device to:

receive a natural language input;

detect an object with a camera of the electronic device;

determine, based on the natural language input, a first value for a first node of a plurality of nodes of a parsing structure, the first node corresponding to a natural language data stream;

determine, based on the object and without using data from the natural language data stream, a second value for a second node of the plurality of nodes of the parsing structure; and

determine, using the parsing structure, the first value, and the second value, a user intent associated with the natural language input.

2. The non-transitory computer-readable storage medium of claim 1 , wherein detecting the object with the camera is an event separate from the natural language input.

3. The non-transitory computer-readable storage medium of claim 1 , wherein the first value is determined at a first time, the second value is determined at a second time, and the user intent is determined in accordance with a determination that the first time and the second time are within a predetermined time.

4. The non-transitory computer-readable storage medium of claim 1 , wherein the second value is determined without processing the natural language input.

5. The non-transitory computer-readable storage medium of claim 1 , wherein the object represents a point of interest.

6. The non-transitory computer-readable storage medium of claim 5 , wherein the point of interest is determined using an image recognition application.

7. The non-transitory computer-readable storage medium of claim 1 , wherein data representing the object is provided to the parsing structure from an application associated with the camera.

8. The non-transitory computer-readable storage medium of claim 1 , wherein:

the first node and the second node depend from a group node of the plurality of nodes; and

determining the user intent associated with the natural language input comprises determining, using the group node, one or more values for a property of the user intent based on at least one of the first value and the second value.

9. The non-transitory computer-readable storage medium of claim 8 , wherein the group node determines the one or more values for the property in accordance with a rule associated with the group node being satisfied.

10. The non-transitory computer-readable storage medium of claim 9 , wherein the rule associated with the group node is satisfied when the first value and the second value are each determined within a predetermined time.

11. The non-transitory computer-readable storage medium of claim 9 , wherein the rule associated with the group node is satisfied when the first value is determined before the second value is determined.

12. The non-transitory computer-readable storage medium of claim 1 , wherein

determining the user intent comprises:

determining a candidate intent based on the first value and the second value;

determining a confidence score associated with the candidate intent; and

selecting, using a second group node of the plurality of nodes, the candidate intent as the user intent based on the confidence score.

13. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions, which when executed by one or more processors of the electronic device, cause the electronic device to:

in accordance with a determination that detection of the object satisfies a predetermined criterion;

initiate a session of a digital assistant; and

wherein the natural language input is received in accordance with initiating the session of the digital assistant.

14. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions, which when executed by one or more processors of the electronic device, cause the electronic device to:

in accordance with a determination that the first value is not determined within a second predetermined time after determining the second value, provide a request for user input, wherein the natural language input is received responsive to the request for user input.

15. A method for processing natural language requests, the method comprising:

at an electronic device with one or more processors and memory:

receiving a natural language input;

detecting an object with a camera of the electronic device;

determining, based on the natural language input, a first value for a first node of a plurality of nodes of a parsing structure, the first node corresponding to a natural language data stream;

determining, based on the object and without using data from the natural language data stream, a second value for a second node of the plurality of nodes of the parsing structure; and

determining, using the parsing structure, the first value, and the second value, a user intent associated with the natural language input.

16. The method of claim 15 , wherein detecting the object with the camera is an event separate from the natural language input.

17. The method of claim 15 , wherein the first value is determined at a first time, the second value is determined at a second time, and the user intent is determined in accordance with a determination that the first time and the second time are within a predetermined time.

18. The method of claim 15 , wherein the second value is determined without processing the natural language input.

19. The method of claim 15 , wherein the object represents a point of interest.

20. The method of claim 19 , wherein the point of interest is determined using an image recognition application.

21. The method of claim 15 , wherein data representing the object is provided to the parsing structure from an application associated with the camera.

22. The method of claim 15 , wherein:

the first node and the second node depend from a group node of the plurality of nodes; and

determining the user intent associated with the natural language input comprises determining, using the group node, one or more values for a property of the user intent based on at least one of the first value and the second value.

23. The method of claim 22 , wherein the group node determines the one or more values for the property in accordance with a rule associated with the group node being satisfied.

24. The method of claim 23 , wherein the rule associated with the group node is satisfied when the first value and the second value are each determined within a predetermined time.

25. The method of claim 23 , wherein the rule associated with the group node is satisfied when the first value is determined before the second value is determined.

26. The method of claim 15 , wherein determining the user intent comprises:

determining a candidate intent based on the first value and the second value;

determining a confidence score associated with the candidate intent; and

selecting, using a second group node of the plurality of nodes, the candidate intent as the user intent based on the confidence score.

27. The method of claim 15 , further comprising:

in accordance with a determination that detection of the object satisfies a predetermined criterion;

initiating a session of a digital assistant; and

wherein the natural language input is received in accordance with initiating the session of the digital assistant.

28. The method of claim 15 , further comprising:

in accordance with a determination that the first value is not determined within a second predetermined time after determining the second value, providing a request for user input, wherein the natural language input is received responsive to the request for user input.

29. An electronic device, comprising:

one or more processors;

memory; and

one or more programs stored in the memory, the one or more programs including instructions for:

receiving a natural language input;

detecting an object with a camera of the electronic device;

determining, based on the natural language input, a first value for a first node of a plurality of nodes of a parsing structure, the first node corresponding to a natural language data stream;

determining, based on the object and without using data from the natural language data stream, a second value for a second node of the plurality of nodes of the parsing structure; and

determining, using the parsing structure, the first value, and the second value, a user intent associated with the natural language input.

30. The electronic device of claim 29 , wherein detecting the object with the camera is an event separate from the natural language input.

31. The electronic device of claim 29 , wherein the first value is determined at a first time, the second value is determined at a second time, and the user intent is determined in accordance with a determination that the first time and the second time are within a predetermined time.

32. The electronic device of claim 29 , wherein the second value is determined without processing the natural language input.

33. The electronic device of claim 29 , wherein the object represents a point of interest.

34. The electronic device of claim 33 , wherein the point of interest is determined using an image recognition application.

35. The electronic device of claim 29 , wherein data representing the object is provided to the parsing structure from an application associated with the camera.

36. The electronic device of claim 29 , wherein:

the first node and the second node depend from a group node of the plurality of nodes; and

determining the user intent associated with the natural language input comprises determining, using the group node, one or more values for a property of the user intent based on at least one of the first value and the second value.

37. The electronic device of claim 36 , wherein the group node determines the one or more values for the property in accordance with a rule associated with the group node being satisfied.

38. The electronic device of claim 37 , wherein the rule associated with the group node is satisfied when the first value and the second value are each determined within a predetermined time.

39. The electronic device of claim 37 , wherein the rule associated with the group node is satisfied when the first value is determined before the second value is determined.

40. The electronic device of claim 29 , wherein determining the user intent comprises:

determining a candidate intent based on the first value and the second value;

determining a confidence score associated with the candidate intent; and

selecting, using a second group node of the plurality of nodes, the candidate intent as the user intent based on the confidence score.

41. The electronic device of claim 29 , the one or more programs further including instructions for:

in accordance with a determination that detection of the object satisfies a predetermined criterion:

initiating a session of a digital assistant; and

wherein the natural language input is received in accordance with initiating the session of the digital assistant.

42. The electronic device of claim 29 , the one or more programs further including instructions for:

in accordance with a determination that the first value is not determined within a second predetermined time after determining the second value, providing a request for user input, wherein the natural language input is received responsive to the request for user input.

Continuity (3)
Continuation 16441461 · Jun 14, 2019
Provisional Application 62820164 · Mar 18, 2019
Related Publication 20220262354A1 · Aug 18, 2022
Cited By (1)
US 12,518,109