IP Library › Granted Patent US 10,984,798
Granted Patent B2
US 10,984,798 · App. 16/896,426 · Granted Apr 20, 2021

Voice interaction at a primary device to access call functionality of a companion device

Inventors: Karl Ferdinand Schramm (Monte Sereno, CA); Justin Binder (Oakland, CA); Benjamin S. Phipps (San Francisco, CA); Po Keng Sung (San Francisco, CA)
Assignee: Apple Inc.
G10L15/22G10L15/1815G10L15/30H04M3/42212G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,984,798
App. No.
16/896,426
Filed
Jun 9, 2020
Granted
Apr 20, 2021
Kind
B2
Art Unit
2674
USPC
704/257
Abstract

The present disclosure generally relates to using voice interaction to access call functionality of a companion device. In an example process, a user utterance is received. Based on the user utterance and contextual information, the process causes a server to determine a user intent corresponding to the user utterance. The contextual information is based on a signal received from the companion device. In accordance with the user intent corresponding to an actionable intent of answering the incoming call, a command is received. Based on the command, instructions are provided to the companion device, which cause the companion device to answer the incoming call and provide audio data of the answered incoming call. Audio is outputted according to the audio data of the answered incoming call.

Claims (82)

1. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device with a speaker, cause the electronic device to:

receive, from a companion device, audio data of an active call, wherein the audio data represents an utterance received by the companion device from a far-end party of the active call via an active call connection;

determine, based on the audio data, whether at least a portion of the utterance corresponds to a predefined digital assistant trigger;

in accordance with a determination that at least a portion of the utterance corresponds to the predefined digital assistant trigger, cause a server to determine a user intent corresponding to the utterance;

receive, from the server, one or more commands representing one or more tasks to satisfy the user intent;

determine, based on results of performing the one or more tasks, a digital assistant response; and

transmit the digital assistant response to the companion device, wherein the companion device transmits the digital assistant response to the far-end party via the active call connection.

2. The non-transitory computer-readable storage medium of claim 1 , wherein the far-end party is a second electronic device having a processor, memory, and a speaker, and

wherein transmitting the digital assistant response to the far-end party via the active call connection causes the speaker of the far-end party to output the digital assistant response as a first audio signal during the active call.

3. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs include further instructions, which when executed by the one or more processors, cause the electronic device to:

after transmitting the digital assistant response to the companion device, output the digital assistant response as a second audio signal via the speaker of the electronic device.

4. The non-transitory computer-readable storage medium of claim 3 , wherein the far-end party and the electronic device concurrently output the first audio signal and the second audio signal, respectively.

5. The non-transitory computer-readable storage medium of claim 1 , wherein causing the server to determine the user intent corresponding to the utterance includes transmitting the audio data representing the utterance to the server, and

wherein the server determines the user intent corresponding to the utterance based on results of performing natural language processing of the audio data.

6. The non-transitory computer-readable storage medium of claim 5 , wherein causing the server to determine the user intent corresponding to the utterance further includes transmitting contextual information to the server with the audio data representing the utterance, and

wherein the server determines the user intent further based on the contextual information.

7. The non-transitory computer-readable storage medium of claim 6 , wherein the contextual information indicates that the companion device is engaged in an active call.

8. The non-transitory computer-readable storage medium of claim 6 , wherein the contextual information indicates that a wireless communication connection is established between the electronic device and the companion device and that the companion device is a registered device of the electronic device.

9. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more commands specify an actionable intent node corresponding to the user intent.

10. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs include further instructions, which when executed by the one or more processors, cause the electronic device to:

prior to receiving the audio data of the active call, exchange authentication information with the companion device; and

in accordance with a determination, based on the authentication information, that the companion device is a registered device of the electronic device, establish a wireless communication connection with the companion device,

wherein the electronic device receives the audio data of the active call from the companion device via the established wireless communication connection, and

wherein the electronic device transmits the digital assistant response to the companion device via the established wireless communication connection.

11. The non-transitory computer-readable storage medium of claim 1 , wherein the active call is an active phone call, and wherein the active call connection is an active phone call connection.

12. The non-transitory computer-readable storage medium of claim 1 , wherein the active call is an active video call, and wherein the active call connection is an active video call connection.

13. A method for using voice interaction to access functionality of a digital assistant of an electronic device during an active call at a companion device, the method comprising:

at the electronic device having a processor, memory, and a speaker:

receiving, from the companion device, audio data of the active call, wherein the audio data represents an utterance received by the companion device from a far-end party of the active call via an active call connection;

determining, based on the audio data, whether at least a portion of the utterance corresponds to a predefined digital assistant trigger;

in accordance with a determination that at least a portion of the utterance corresponds to the predefined digital assistant trigger, causing a server to determine a user intent corresponding to the utterance;

receiving, from the server, one or more commands representing one or more tasks to satisfy the user intent;

determining, based on results of performing the one or more tasks, a digital assistant response; and

transmitting the digital assistant response to the companion device, wherein the companion device transmits the digital assistant response to the far-end party via the active call connection.

14. An electronic device, comprising:

a speaker;

one or more processors; and

memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for:

receiving, from a companion device, audio data of an active call, wherein the audio data represents an utterance received by the companion device from a far-end party of the active call via an active call connection;

determining, based on the audio data, whether at least a portion of the utterance corresponds to a predefined digital assistant trigger;

in accordance with a determination that at least a portion of the utterance corresponds to the predefined digital assistant trigger, causing a server to determine a user intent corresponding to the utterance;

receiving, from the server, one or more commands representing one or more tasks to satisfy the user intent;

determining, based on results of performing the one or more tasks, a digital assistant response; and

transmitting the digital assistant response to the companion device, wherein the companion device transmits the digital assistant response to the far-end party via the active call connection.

15. The method of claim 13 , wherein the far-end party is a second electronic device having a processor, memory, and a speaker, and

wherein transmitting the digital assistant response to the far-end party via the active call connection causes the speaker of the far-end party to output the digital assistant response as a first audio signal during the active call.

16. The method of claim 13 , further comprising:

after transmitting the digital assistant response to the companion device, outputting the digital assistant response as a second audio signal via the speaker of the electronic device.

17. The method of claim 16 , wherein the far-end party and the electronic device concurrently output the first audio signal and the second audio signal, respectively.

18. The method of claim 13 , wherein causing the server to determine the user intent corresponding to the utterance includes transmitting the audio data representing the utterance to the server, and

wherein the server determines the user intent corresponding to the utterance based on results of performing natural language processing of the audio data.

19. The method of claim 18 , wherein causing the server to determine the user intent corresponding to the utterance further includes transmitting contextual information to the server with the audio data representing the utterance, and

wherein the server determines the user intent further based on the contextual information.

20. The method of claim 19 , wherein the contextual information indicates that the companion device is engaged in an active call.

21. The method of claim 19 , wherein the contextual information indicates that a wireless communication connection is established between the electronic device and the companion device and that the companion device is a registered device of the electronic device.

22. The method of claim 13 , wherein the one or more commands specify an actionable intent node corresponding to the user intent.

23. The method of claim 13 , further comprising:

prior to receiving the audio data of the active call, exchanging authentication information with the companion device; and

in accordance with a determination, based on the authentication information, that the companion device is a registered device of the electronic device, establishing a wireless communication connection with the companion device,

wherein the electronic device receives the audio data of the active call from the companion device via the established wireless communication connection, and

wherein the electronic device transmits the digital assistant response to the companion device via the established wireless communication connection.

24. The method of claim 13 , wherein the active call is an active phone call, and wherein the active call connection is an active phone call connection.

25. The method of claim 13 , wherein the active call is an active video call, and wherein the active call connection is an active video call connection.

26. The electronic device of claim 14 , wherein the far-end party is a second electronic device having a processor, memory, and a speaker, and

wherein transmitting the digital assistant response to the far-end party via the active call connection causes the speaker of the far-end party to output the digital assistant response as a first audio signal during the active call.

27. The electronic device of claim 14 , the one or more programs further including instructions for:

after transmitting the digital assistant response to the companion device, outputting the digital assistant response as a second audio signal via the speaker of the electronic device.

28. The electronic device of claim 27 , wherein the far-end party and the electronic device concurrently output the first audio signal and the second audio signal, respectively.

29. The electronic device of claim 14 , wherein causing the server to determine the user intent corresponding to the utterance includes transmitting the audio data representing the utterance to the server, and

wherein the server determines the user intent corresponding to the utterance based on results of performing natural language processing of the audio data.

30. The electronic device of claim 29 , wherein causing the server to determine the user intent corresponding to the utterance further includes transmitting contextual information to the server with the audio data representing the utterance, and

wherein the server determines the user intent further based on the contextual information.

31. The electronic device of claim 30 , wherein the contextual information indicates that the companion device is engaged in an active call.

32. The electronic device of claim 30 , wherein the contextual information indicates that a wireless communication connection is established between the electronic device and the companion device and that the companion device is a registered device of the electronic device.

33. The electronic device of claim 14 , wherein the one or more commands specify an actionable intent node corresponding to the user intent.

34. The electronic device of claim 14 , the one or more programs further including instructions for:

prior to receiving the audio data of the active call, exchanging authentication information with the companion device; and

in accordance with a determination, based on the authentication information, that the companion device is a registered device of the electronic device, establishing a wireless communication connection with the companion device,

wherein the electronic device receives the audio data of the active call from the companion device via the established wireless communication connection, and

wherein the electronic device transmits the digital assistant response to the companion device via the established wireless communication connection.

35. The electronic device of claim 14 , wherein the active call is an active phone call, and wherein the active call connection is an active phone call connection.

36. The electronic device of claim 14 , wherein the active call is an active video call, and wherein the active call connection is an active video call connection.

Continuity (4)
Continuation 16504782 · Jul 8, 2019
Continuation 16113119 · Aug 27, 2018
Provisional Application 62679177 · Jun 1, 2018
Related Publication 20200302932A1 · Sep 24, 2020
Cited By (1)
US 12,664,969