IP Library Granted Patent US 11,837,232
Granted Patent B2
US 11,837,232 · App. 18/115,721 · Granted Dec 5, 2023

Digital assistant interaction in a video communication session environment

Inventors: Niranjan Manjunath (Sunnyvale, CA); Willem Mattelaer (San Jose, CA); Jessica Peck (Morgan Hill, CA); Lily Shuting Zhang (Seattle, WA)
Assignee: Apple Inc.
G10L15/22G10L15/083G10L15/1815G10L15/26G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,837,232
App. No.
18/115,721
Granted
Dec 5, 2023
Kind
B2
Abstract

This relates to an intelligent automated assistant in a video communication session environment. An example method includes, during a video communication session between at least two user devices, and at a first user device: receiving a first user voice input; in accordance with a determination that the first user voice input represents a communal digital assistant request, transmitting a request to provide context information associated with the first user voice input to the first user device; receiving context information associated with the first user voice input; obtaining a first digital assistant response based at least on a portion of the context information received from the second user device and at least a portion of context information associated with the first user voice input that is stored on the first user device; providing the first digital assistant response to the second user device; and outputting the first digital assistant response.

Claims (100)

1. A method, comprising:

during a video communication session between at least two user devices, and at a first user device of the at least two user devices;

receiving a first user voice input;

determining whether the first user voice input represents a communal digital assistant request;

in accordance with a determination that the first user voice input represents a communal digital assistant request, transmitting a request to a second user device of the at least two user devices for the second user device to provide context information associated with the first user voice input to the first user device;

receiving the context information associated with the first user voice input from the second user device;

obtaining a first digital assistant response based at least on a portion of the context information received from the second user device and at least a portion of context information associated with the first user voice input that is stored on the first user device;

providing the first digital assistant response to the second user device; and

outputting the first digital assistant response.

2. The method of claim 1 , wherein the context information received from the second user device includes at least one of user-specific calendar data, user-specific music preference data, user-specific restaurant preference data, and a current location of the second user device.

3. The method of claim 1 , wherein the first digital assistant response comprises at least one of:

a natural-language expression corresponding to a task performed by a digital assistant of the first user device based on the first user voice input, the at least a portion of the context information received from the second user device, and the at least a portion of the context information associated with the first user voice input that is stored on the first user device; and

data retrieved by the digital assistant of the first user device based on the first user voice input, the at least a portion of the context information received from the second user device, and the at least a portion of the context information associated with the first user voice input that is stored on the first user device.

4. The method of claim 1 , wherein obtaining the first digital assistant response includes:

performing one or more tasks based on the first user voice input, the at least a portion of the context information received from the second user device, and the at least a portion of the context information associated with the first user voice input that is stored on the first user device; and

determining the first digital assistant response based on results of the performance of the one or more tasks.

5. The method of claim 1 , wherein obtaining the first digital assistant response includes comparing the at least a portion of the context information received from the second user device to the at least a portion of the context information associated with the first user voice input that is stored on the first user device.

6. The method of claim 1 , wherein the context information received from the second user device was stored by a software application stored on the second user device, wherein obtaining the first digital assistant response includes processing the context information received from the second user device based on one or more protocols corresponding to the software application.

7. The method of claim 1 , wherein the first user device provides the first digital assistant response to the second user device as text data.

8. The method of claim 1 , wherein the first user device provides the first digital assistant response to the second user device as audio data via an audio stream.

9. The method of claim 1 , further comprising:

prior to transmitting the request to the second user device, transmitting the first user voice input to the second user device.

10. The method of claim 1 , further comprising:

prior to outputting the first digital assistant response, providing, to the second user device, the at least a portion of the context information associated with the first user voice input that is stored on the first user device and additional context information that is not used to obtain the first digital assistant response.

11. The method of claim 1 , wherein providing the first digital assistant response to the second user device causes the second user device to output the first digital assistant response.

12. The method of claim 1 , further comprising:

in response to receiving the first user voice input and prior to determining whether the first user voice input represents a communal digital assistant request, determining whether the first user voice input includes a digital assistant trigger; and

in accordance with a determination that the first user voice input includes the digital assistant trigger, determining whether the first user device currently holds a digital assistant invocation permission, wherein the first user device determines whether the first user voice input represents a communal digital assistant request in accordance with a determination that the first user device currently holds the digital assistant invocation permission.

13. The method of claim 12 , further comprising:

in accordance with a determination that the first user device does not currently hold the digital assistant invocation permission, transmitting a request to the second user device for the second user device to provide the digital assistant invocation permission to the first user device,

wherein the first user device determines whether the first user voice input represents a communal digital assistant request in response to receiving the digital assistant invocation permission from the second user device, and

wherein the first user device forgoes determining whether the first user voice input represents a communal digital assistant request in response to not receiving the digital assistant invocation permission from the second user device within a predetermined period of time.

14. A first user device, comprising:

a display;

one or more processors;

a memory; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, wherein the one or more programs include instructions for:

during a video communication session between the first user device and at least a second user device:

receiving a first user voice input;

determining whether the first user voice input represents a communal digital assistant request;

in accordance with a determination that the first user voice input represents a communal digital assistant request, transmitting a request to a second user device of the at least two user devices for the second user device to provide context information associated with the first user voice input to the first user device;

receiving the context information associated with the first user voice input from the second user device;

obtaining a first digital assistant response based at least on a portion of the context information received from the second user device and at least a portion of context information associated with the first user voice input that is stored on the first user device;

providing the first digital assistant response to the second user device; and

outputting the first digital assistant response.

15. The first user device of claim 14 , wherein the context information received from the second user device includes at least one of user-specific calendar data, user-specific music preference data, user-specific restaurant preference data, and a current location of the second user device.

16. The first user device of claim 14 , wherein the first digital assistant response comprises at least one of:

a natural-language expression corresponding to a task performed by a digital assistant of the first user device based on the first user voice input, the at least a portion of the context information received from the second user device, and the at least a portion of the context information associated with the first user voice input that is stored on the first user device; and

data retrieved by the digital assistant of the first user device based on the first user voice input, the at least a portion of the context information received from the second user device, and the at least a portion of the context information associated with the first user voice input that is stored on the first user device.

17. The first user device of claim 14 , wherein obtaining the first digital assistant response includes:

performing one or more tasks based on the first user voice input, the at least a portion of the context information received from the second user device, and the at least a portion of the context information associated with the first user voice input that is stored on the first user device; and

determining the first digital assistant response based on results of the performance of the one or more tasks.

18. The first user device of claim 14 , wherein obtaining the first digital assistant response includes comparing the at least a portion of the context information received from the second user device to the at least a portion of the context information associated with the first user voice input that is stored on the first user device.

19. The first user device of claim 14 , wherein the context information received from the second user device was stored by a software application stored on the second user device, wherein obtaining the first digital assistant response includes processing the context information received from the second user device based on one or more protocols corresponding to the software application.

20. The first user device of claim 14 , wherein the first user device provides the first digital assistant response to the second user device as text data.

21. The first user device of claim 14 , wherein the first user device provides the first digital assistant response to the second user device as audio data via an audio stream.

22. The first user device of claim 14 , wherein the one or more programs further include instructions for:

prior to transmitting the request to the second user device, transmitting the first user voice input to the second user device.

23. The first user device of claim 14 , wherein the one or more programs further include instructions for:

prior to outputting the first digital assistant response, providing, to the second user device, the at least a portion of the context information associated with the first user voice input that is stored on the first user device and additional context information that is not used to obtain the first digital assistant response.

24. The first user device of claim 14 , wherein providing the first digital assistant response to the second user device causes the second user device to output the first digital assistant response.

25. The first user device of claim 14 , wherein the one or more programs further include instructions for:

in response to receiving the first user voice input and prior to determining whether the first user voice input represents a communal digital assistant request, determining whether the first user voice input includes a digital assistant trigger; and

in accordance with a determination that the first user voice input includes the digital assistant trigger, determining whether the first user device currently holds a digital assistant invocation permission, wherein the first user device determines whether the first user voice input represents a communal digital assistant request in accordance with a determination that the first user device currently holds the digital assistant invocation permission.

26. The first user device of claim 25 , wherein the one or more programs further include instructions for:

in accordance with a determination that the first user device does not currently hold the digital assistant invocation permission, transmitting a request to the second user device for the second user device to provide the digital assistant invocation permission to the first user device,

wherein the first user device determines whether the first user voice input represents a communal digital assistant request in response to receiving the digital assistant invocation permission from the second user device, and

wherein the first user device forgoes determining whether the first user voice input represents a communal digital assistant request in response to not receiving the digital assistant.

27. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a first user device with a display, cause the first user device to:

during a video communication session between the first user device and at least a second user device:

receive a first user voice input;

determine whether the first user voice input represents a communal digital assistant request;

in accordance with a determination that the first user voice input represents a communal digital assistant request, transmit a request to a second user device of the at least two user devices for the second user device to provide context information associated with the first user voice input to the first user device;

receive the context information associated with the first user voice input from the second user device;

obtain a first digital assistant response based at least on a portion of the context information received from the second user device and at least a portion of context information associated with the first user voice input that is stored on the first user device;

provide the first digital assistant response to the second user device; and

output the first digital assistant response.

28. The non-transitory computer-readable storage medium of claim 27 , wherein the context information received from the second user device includes at least one of user-specific calendar data, user-specific music preference data, user-specific restaurant preference data, and a current location of the second user device.

29. The non-transitory computer-readable storage medium of claim 27 , wherein the first digital assistant response comprises at least one of:

a natural-language expression corresponding to a task performed by a digital assistant of the first user device based on the first user voice input, the at least a portion of the context information received from the second user device, and the at least a portion of the context information associated with the first user voice input that is stored on the first user device; and

data retrieved by the digital assistant of the first user device based on the first user voice input, the at least a portion of the context information received from the second user device, and the at least a portion of the context information associated with the first user voice input that is stored on the first user device.

30. The non-transitory computer-readable storage medium of claim 27 , wherein obtaining the first digital assistant response includes:

performing one or more tasks based on the first user voice input, the at least a portion of the context information received from the second user device, and the at least a portion of the context information associated with the first user voice input that is stored on the first user device; and

determining the first digital assistant response based on results of the performance of the one or more tasks.

31. The non-transitory computer-readable storage medium of claim 27 , wherein obtaining the first digital assistant response includes comparing the at least a portion of the context information received from the second user device to the at least a portion of the context information associated with the first user voice input that is stored on the first user device.

32. The non-transitory computer-readable storage medium of claim 27 , wherein the context information received from the second user device was stored by a software application stored on the second user device, wherein obtaining the first digital assistant response includes processing the context information received from the second user device based on one or more protocols corresponding to the software application.

33. The non-transitory computer-readable storage medium of claim 27 , wherein the first user device provides the first digital assistant response to the second user device as text data.

34. The non-transitory computer-readable storage medium of claim 27 , wherein the first user device provides the first digital assistant response to the second user device as audio data via an audio stream.

35. The non-transitory computer-readable storage medium of claim 27 , wherein the one or more programs further include instructions, which when executed by the one or more processors, cause the first user device to:

prior to transmitting the request to the second user device, transmit the first user voice input to the second user device.

36. The non-transitory computer-readable storage medium of claim 27 , wherein the one or more programs further include instructions, which when executed by the one or more processors, cause the first user device to:

prior to outputting the first digital assistant response, provide, to the second user device, the at least a portion of the context information associated with the first user voice input that is stored on the first user device and additional context information that is not used to obtain the first digital assistant response.

37. The non-transitory computer-readable storage medium of claim 27 , wherein providing the first digital assistant response to the second user device causes the second user device to output the first digital assistant response.

38. The non-transitory computer-readable storage medium of claim 27 , wherein the one or more programs further include instructions, which when executed by the one or more processors, cause the first user device to:

in response to receiving the first user voice input and prior to determining whether the first user voice input represents a communal digital assistant request, determine whether the first user voice input includes a digital assistant trigger; and

in accordance with a determination that the first user voice input includes the digital assistant trigger, determine whether the first user device currently holds a digital assistant invocation permission, wherein the first user device determines whether the first user voice input represents a communal digital assistant request in accordance with a determination that the first user device currently holds the digital assistant invocation permission.

39. The non-transitory computer-readable storage medium of claim 38 , wherein the one or more programs further include instructions, which when executed by the one or more processors, cause the first user device to:

in accordance with a determination that the first user device does not currently hold the digital assistant invocation permission, transmit a request to the second user device for the second user device to provide the digital assistant invocation permission to the first user device,

wherein the first user device determines whether the first user voice input represents a communal digital assistant request in response to receiving the digital assistant invocation permission from the second user device, and

wherein the first user device forgoes determining whether the first user voice input represents a communal digital assistant request in response to not receiving the digital assistant.

Continuity (4)
Division 17158703 · Jan 26, 2021
Provisional Application 63016083 · Apr 27, 2020
Provisional Application 62975643 · Feb 12, 2020
Related Publication 20230215435A1 · Jul 6, 2023
Cited By (5)
US 12,230,264 US 12,333,404 US 12,386,434 US 12,477,470 US 12,633,289