IP Library Granted Patent US 12,633,289
Granted Patent B2
US 12,633,289 · App. 18/589,236 · Granted May 19, 2026

Digital assistant interactions in a voice communication session between an electronic device and a remote electronic device

Inventors: Miles Munro (Sunnyvale, CA); Jonathan H. Russell (Incline Village, NV); Felicia W. Edwards (Loveland, CO); Keith C. Strickling (San Francisco, CA)
Assignee: Apple Inc.
G10L15/22G10L15/1815G10L15/30G10L21/0208G10L21/034G06F3/1454G10L2015/223G10L2021/02082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,633,289
App. No.
18/589,236
Granted
May 19, 2026
Kind
B2
Abstract

An example process includes: while an electronic device is engaged in a voice communication session with at least one remote device: receiving a request to invoke a digital assistant operating on the electronic device; receiving a natural language input; in accordance with a determination that a set of one or more criteria is satisfied, where the set of one or more criteria includes a criterion that is satisfied when the natural language input is received from a near-end user of the voice communication session: providing a response to the near-end user, where the response is generated by the digital assistant based on the natural language input; and forgoing causing the response to be provided to the at least one remote device; and in accordance with a determination that the criterion is not satisfied: forgoing generating, by the digital assistant, the response.

Claims (160)

1 . A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device with a microphone and a speaker, cause the electronic device to:

while the electronic device is engaged in a voice communication session with at least one remote device:

receive a request to invoke a digital assistant operating on the electronic device;

receive a natural language input;

in accordance with a determination that a set of one or more criteria is satisfied, wherein the set of one or more criteria includes a first criterion that is satisfied when the electronic device is engaged in the voice communication session with the at least one remote device and a second criterion that is satisfied when the natural language input is received from a near-end user of the voice communication session:

while an audio channel that is configured to transmit audio data to the at least one remote device is open, provide, via the speaker, a response to the near-end user, wherein the response is generated by the digital assistant based on the natural language input;

detect, via the microphone, the response; and

forgo causing the response to be provided to the at least one remote device, including applying echo cancellation to the detected response to prevent transmission of the detected response via the audio channel; and

in accordance with a determination that the second criterion is not satisfied:

forgo generating, by the digital assistant, the response.

2 . The non-transitory computer-readable storage medium of claim 1 , wherein the voice communication session is between the near-end user and at least one respective far-end user of the at least one remote device.

3 . The non-transitory computer-readable storage medium of claim 1 , wherein the natural language input is received from the near-end user, and wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

transmit the natural language input to the at least one remote device.

4 . The non-transitory computer-readable storage medium of claim 1 , wherein the second criterion is not satisfied when the natural language input is received from a far-end user of the voice communication session.

5 . The non-transitory computer-readable storage medium of claim 4 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

detect, via the microphone, the natural language input received from the far- end user, wherein forgoing generating, by the digital assistant, the response includes:

applying echo cancellation to the detected natural language input received from the far-end user to prevent the digital assistant from receiving the detected natural language input received from the far-end user.

6 . The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

while providing the response, receive, from the near end user, a speech input; and

transmit the speech input to the at least one remote device.

7 . The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

in accordance with a determination that the electronic device is engaged in the voice communication session, select, from a plurality of digital assistant response modes, a first digital assistant response mode, wherein:

the first digital assistant response mode specifies that the response is to be displayed; and

the response is provided according to the first digital assistant response mode.

8 . The non-transitory computer-readable storage medium of claim 1 , wherein:

the natural language input requests to remove a first participant from the voice communication session; and

the response indicates that the digital assistant is unable to remove the first participant.

9 . The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

in accordance with a determination that the set of one or more criteria is satisfied:

initiate, by the digital assistant, a task based on the natural language input, wherein the response indicates the initiated task.

10 . The non-transitory computer-readable storage medium of claim 9 , wherein:

the natural language input requests to add a second participant to the voice communication session; and

initiating the task includes adding the second participant to the voice communication session.

11 . The non-transitory computer-readable storage medium of claim 10 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

before adding the second participant to the voice communication session:

in accordance with a determination, based on a conversation history between the second participant and a current participant of the voice communication session, to provide a first output requesting to confirm to add the second participant to the voice communication session:

provide the first output; and

in accordance with a determination, based on the conversation history, to not provide the first output:

forgo providing the first output.

12 . The non-transitory computer-readable storage medium of claim 10 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

before adding the second participant to the voice communication session:

in accordance with a determination that the natural language input is ambiguous regarding the second participant or a third participant:

in accordance with a determination that the second participant has a first conversation history with at least one current participant of the voice communication session and that the third participant does not have a second conversation history with at least one current participant of the voice communication session:

provide a second output requesting to confirm to add the second participant to the voice communication session, wherein the second output includes an identity of the second participant and does not include an identity of the third participant; and

in accordance with a determination that the second participant has the first conversation history with at least one current participant of the voice communication session and that the third participant has the second conversation history with at least one current participant of the voice communication session:

provide a third output requesting disambiguation between the second participant and the third participant, wherein the third output includes the identity of the second participant and includes the identity of the third participant.

13 . The non-transitory computer-readable storage medium of claim 9 , wherein:

the natural language input requests to mute the near-end user; and

initiating the task includes muting the near-end user.

14 . The non-transitory computer-readable storage medium of claim 9 , wherein:

the natural language input requests to change a volume of the voice communication session; and

initiating the task includes changing the volume of the voice communication session.

15 . The non-transitory computer-readable storage medium of claim 14 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

while the electronic device is engaged in the voice communication session, provide media playback at the electronic device, wherein initiating the task further includes changing a volume of the media playback.

16 . The non-transitory computer-readable storage medium of claim 9 , wherein:

when the natural language input is received, the voice communication session does not provide video call functionality;

the natural language input requests to change the voice communication session to provide the video call functionality; and

initiating the task includes changing the voice communication session to provide the video call functionality.

17 . The non-transitory computer-readable storage medium of claim 9 , wherein:

the natural language input requests to share content displayed on the electronic device with the at least one remote device; and

initiating the task includes sharing the content displayed on the electronic device with the at least one remote device.

18 . The non-transitory computer-readable storage medium of claim 17 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

before initiating the task:

provide an output requesting confirmation, from the near-end user, to share the content; and

receive, from the near-end user, the confirmation to share the content.

19 . The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

while the electronic device is engaged in the voice communication session, provide playback of media that is concurrently consumed by each participant of the voice communication session; and

in response to receiving the natural language input, lower, for the near-end user, a volume of the media without causing the at least one remote device to adjust playback of the media.

20 . The non-transitory computer-readable storage medium of claim 1 , wherein:

the natural language input corresponds to a user request to record or transmit audio input; and

the response indicates that the digital assistant is unable to satisfy the user request.

21 . The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

while the electronic device is engaged in the voice communication session with the at least one remote device:

after providing the response, forgo operating the digital assistant in a listening state until a second request to invoke the digital assistant is received.

22 . The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

while the electronic device is not engaged in a voice communication session with the at least one remote device:

receive a second natural language input;

provide a second response to the near-end user, wherein the second response is generated by the digital assistant based on the second natural language input; and

in accordance with a determination that the second natural language input corresponds to a predetermined type of task, dismiss a session of the digital assistant a first predetermined duration after providing the second response; and

while the electronic device is engaged in the voice communication session with the at least one remote device:

in accordance with a determination that the natural language input corresponds to the predetermined type of task, dismiss a second session of the digital assistant a second predetermined duration after providing the response, wherein the second predetermined duration is less than the first predetermined duration.

23 . An electronic device, comprising:

a microphone;

a speaker;

one or more processors;

a memory; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:

while the electronic device is engaged in a voice communication session with at least one remote device:

receiving a request to invoke a digital assistant operating on the electronic device;

receiving a natural language input;

in accordance with a determination that a set of one or more criteria is satisfied, wherein the set of one or more criteria includes a first criterion that is satisfied when the electronic device is engaged in the voice communication session with the at least one remote device and a second criterion that is satisfied when the natural language input is received from a near-end user of the voice communication session:

while an audio channel that is configured to transmit audio data to the at least one remote device is open, providing, via the speaker, a response to the near-end user, wherein the response is generated by the digital assistant based on the natural language input;

detecting, via the microphone, the response; and

forgoing causing the response to be provided to the at least one remote device, including applying echo cancellation to the detected response to prevent transmission of the detected response via the audio channel; and

in accordance with a determination that the second criterion is not satisfied:

forgoing generating, by the digital assistant, the response.

24 . The electronic device of claim 23 , wherein the voice communication session is between the near-end user and at least one respective far-end user of the at least one remote device.

25 . The electronic device of claim 23 , wherein the natural language input is received from the near-end user, and wherein the one or more programs further include instructions for:

transmitting the natural language input to the at least one remote device.

26 . The electronic device of claim 23 , wherein the second criterion is not satisfied when the natural language input is received from a far-end user of the voice communication session.

27 . The electronic device of claim 26 , wherein the one or more programs further include instructions for:

detecting, via the microphone, the natural language input received from the far-end user, wherein forgoing generating, by the digital assistant, the response includes:

applying echo cancellation to the detected natural language input received from the far-end user to prevent the digital assistant from receiving the detected natural language input received from the far-end user.

28 . The electronic device of claim 23 , wherein the one or more programs further include instructions for:

while providing the response, receiving, from the near end user, a speech input; and

transmitting the speech input to the at least one remote device.

29 . The electronic device of claim 23 , wherein the one or more programs further include instructions for:

in accordance with a determination that the electronic device is engaged in the voice communication session, selecting, from a plurality of digital assistant response modes, a first digital assistant response mode, wherein:

the first digital assistant response mode specifies that the response is to be displayed; and p 1 the response is provided according to the first digital assistant response mode.

30 . The electronic device of claim 23 , wherein the one or more programs further include instructions for:

in accordance with a determination that the set of one or more criteria is satisfied:

initiating, by the digital assistant, a task based on the natural language input, wherein the response indicates the initiated task.

31 . The electronic device of claim 23 , wherein:

the natural language input corresponds to a user request to record or transmit audio input; and

the response indicates that the digital assistant is unable to satisfy the user request.

32 . The electronic device of claim 23 , wherein the one or more programs further include instructions for:

while the electronic device is not engaged in a voice communication session with the at least one remote device:

receiving a second natural language input;

providing a second response to the near-end user, wherein the second response is generated by the digital assistant based on the second natural language input; and

in accordance with a determination that the second natural language input corresponds to a predetermined type of task, dismissing a session of the digital assistant a first predetermined duration after providing the second response; and

while the electronic device is engaged in the voice communication session with the at least one remote device:

in accordance with a determination that the natural language input corresponds to the predetermined type of task, dismissing a second session of the digital assistant a second predetermined duration after providing the response, wherein the second predetermined duration is less than the first predetermined duration.

33 . A method, comprising:

at an electronic device with one or more processors, memory, a microphone, and a speaker:

while the electronic device is engaged in a voice communication session with at least one remote device:

receiving a request to invoke a digital assistant operating on the electronic device;

receiving a natural language input;

in accordance with a determination that a set of one or more criteria is satisfied, wherein the set of one or more criteria includes a first criterion that is satisfied when the electronic device is engaged in the voice communication session with the at least one remote device and a second criterion that is satisfied when the natural language input is received from a near-end user of the voice communication session:

while an audio channel that is configured to transmit audio data to the at least one remote device is open, providing, via the speaker, a response to the near-end user, wherein the response is generated by the digital assistant based on the natural language input;

detecting, via the microphone, the response; and

forgoing causing the response to be provided to the at least one remote device, including applying echo cancellation to the detected response to prevent transmission of the detected response via the audio channel; and

in accordance with a determination that the second criterion is not satisfied:

forgoing generating, by the digital assistant, the response.

34 . The method of claim 33 , wherein the voice communication session is between the near- end user and at least one respective far-end user of the at least one remote device.

35 . The method of claim 33 , wherein the natural language input is received from the near- end user, the method further comprising:

transmitting the natural language input to the at least one remote device.

36 . The method of claim 33 , wherein the second criterion is not satisfied when the natural language input is received from a far-end user of the voice communication session.

37 . The method of claim 36 , further comprising:

detecting, via the microphone, the natural language input received from the far-end user, wherein forgoing generating, by the digital assistant, the response includes:

applying echo cancellation to the detected natural language input received from the far-end user to prevent the digital assistant from receiving the detected natural language input received from the far-end user.

38 . The method of claim 33 , further comprising:

while providing the response, receiving, from the near end user, a speech input; and

transmitting the speech input to the at least one remote device.

39 . The method of claim 33 , further comprising:

in accordance with a determination that the electronic device is engaged in the voice communication session, selecting, from a plurality of digital assistant response modes, a first digital assistant response mode, wherein:

the first digital assistant response mode specifies that the response is to be displayed; and

the response is provided according to the first digital assistant response mode.

40 . The method of claim 33 , further comprising:

in accordance with a determination that the set of one or more criteria is satisfied:

initiating, by the digital assistant, a task based on the natural language input, wherein the response indicates the initiated task.

41 . The method of claim 33 , wherein:

the natural language input corresponds to a user request to record or transmit audio input; and

the response indicates that the digital assistant is unable to satisfy the user request.

42 . The method of claim 33 , further comprising:

while the electronic device is not engaged in a voice communication session with the at least one remote device:

receiving a second natural language input;

providing a second response to the near-end user, wherein the second response is generated by the digital assistant based on the second natural language input; and

in accordance with a determination that the second natural language input corresponds to a predetermined type of task, dismissing a session of the digital assistant a first predetermined duration after providing the second response; and

while the electronic device is engaged in the voice communication session with the at least one remote device:

in accordance with a determination that the natural language input corresponds to the predetermined type of task, dismissing a second session of the digital assistant a second predetermined duration after providing the response, wherein the second predetermined duration is less than the first predetermined duration.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 12, 2024
From: MUNRO, MILES; RUSSELL, JONATHAN H.; EDWARDS, FELICIA W.; STRICKLING, KEITH C.
To: APPLE INC.
Reel/Frame 066728/0826 →
Continuity (2)
Provisional Application 63464354 · May 5, 2023
Related Publication 20240371373A1 · Nov 7, 2024
References Cited (59)
US 8554559B1 · Aleksic et al. · 2013 [cited by applicant]
US 8909693B2 · Frissora et al. · 2014 [cited by applicant]
US 9578173B2 · Sanghavi et al. · 2017 [cited by applicant]
US 9967381B1 · Kashimba et al. · 2018 [cited by applicant]
US 10135965B2 · Woolsey et al. · 2018 [cited by applicant]
US 10356243B2 · Sanghavi et al. · 2019 [cited by applicant]
US 10403272B1 · Fanty et al. · 2019 [cited by applicant]
US 10671428B2 · Zeitlin · 2020 [cited by applicant]
US 10691473B2 · Karashchuk et al. · 2020 [cited by applicant]
US 10847142B2 · Newendorp et al. · 2020 [cited by applicant]
US 10942703B2 · Martel et al. · 2021 [cited by applicant]
US 11769497B2 · Manjunath et al. · 2023 [cited by applicant]
US 11837232B2 · Manjunath et al. · 2023 [cited by applicant]
US 20110047246A1 · Frissora et al. · 2011 [cited by applicant]
US 20110202594A1 · Ricci · 2011 [cited by applicant]
US 20140297288A1 · Yu et al. · 2014 [cited by applicant]
US 20150088514A1 · Typrin · 2015 [cited by applicant]
US 20150309691A1 · Seo et al. · 2015 [cited by applicant]
US 20150347630A1 · Li · 2015 [cited by applicant]
US 20150373183A1 · Woolsey et al. · 2015 [cited by applicant]
US 20160316349A1 · Lee et al. · 2016 [cited by applicant]
US 20160335532A1 · Sanghavi et al. · 2016 [cited by applicant]
US 20160360039A1 · Sanghavi et al. · 2016 [cited by applicant]
US 20160373571A1 · Woolsey et al. · 2016 [cited by applicant]
US 20170132019A1 · Karashchuk et al. · 2017 [cited by applicant]
US 20170346949A1 · Sanghavi et al. · 2017 [cited by applicant]
US 20180146089A1 · Rauenbuehler et al. · 2018 [cited by applicant]
US 20180152558A1 · Chan et al. · 2018 [cited by applicant]
US 20180242219A1 · Deluca et al. · 2018 [cited by applicant]
US 20190005024A1 · Somech et al. · 2019 [cited by applicant]
US 20190116264A1 · Sanghavi et al. · 2019 [cited by applicant]
US 20190222684A1 · Li et al. · 2019 [cited by applicant]
US 20190378511A1 · Vuskovic · 2019 [cited by examiner]
US 20200169637A1 · Sanghavi et al. · 2020 [cited by applicant]
US 20200249985A1 · Zeitlin · 2020 [cited by applicant]
US 20200272485A1 · Karashchuk et al. · 2020 [cited by applicant]
US 20200286472A1 · Newendorp et al. · 2020 [cited by applicant]
US 20200342039A1 · Bakir · 2020 [cited by examiner]
US 20210035567A1 · Newendorp et al. · 2021 [cited by applicant]
US 20210249009A1 · Manjunath et al. · 2021 [cited by applicant]
US 20210314440A1 · Matias et al. · 2021 [cited by applicant]
US 20220329691A1 · Chinthakunta et al. · 2022 [cited by applicant]
US 20220343066A1 · Kwong et al. · 2022 [cited by applicant]
US 20230013615A1 · Sanghavi et al. · 2023 [cited by applicant]
US 20230017115A1 · Sanghavi et al. · 2023 [cited by applicant]
US 20230026764A1 · Karashchuk et al. · 2023 [cited by applicant]
US 20230058929A1 · Lasko et al. · 2023 [cited by applicant]
US 20230179704A1 · Chinthakunta et al. · 2023 [cited by applicant]
US 20230215435A1 · Manjunath et al. · 2023 [cited by applicant]
US 20230386464A1 · Manjunath et al. · 2023 [cited by applicant]
CN 106465074A · 2017 [cited by applicant]
CN 107852436A · 2018 [cited by applicant]
CN 112153223A · 2020 [cited by applicant]
EP 4281855A1 · 2023 [cited by applicant]
WO 2015047932A1 · 2015 [cited by applicant]
WO 2016187149A1 · 2016 [cited by applicant]
International Preliminary Report on Patentability received for PCT Patent Application No. PCT/US2024/026528, mailed on Nov. 20, 2025, 11 pages. [cited by applicant]
International Search Report and Written Opinion received for PCT Patent Application No. PCT/US2024/026528, mailed on Aug. 26, 2024, 17 pages. [cited by applicant]
Invitation to Pay Additional Fees and Partial International Search Report received for PCT Patent Application No. PCT/US2024/026528, mailed on Jul. 3, 2024, 12 pages. [cited by applicant]