IP Library › Granted Patent US 12,626,704
Granted Patent B2
US 12,626,704 · App. 18/132,793 · Granted May 12, 2026

Detecting visual attention during user speech

Inventors: Maxwell C. Horton (Santa Monica, CA); Stephen A. Berardi (Seattle, WA); Yanzi Jin (Bothell, WA); Sophie Lebrecht (Seattle, WA); Richard P. Muffoletto (Bainbridge Island, WA); Daniel Tormoen (Seattle, WA)
Assignee: Apple Inc.
G10L15/25G06V20/40G06V40/20G10L15/22G10L25/57G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,704
App. No.
18/132,793
Granted
May 12, 2026
Kind
B2
Abstract

An example process includes: concurrently receiving an audio stream and a video stream; determining, based on a first portion of the audio stream received within a predetermined duration before a current time and a first portion of the video stream received within the predetermined duration before the current time, whether a visual attention of a user is directed to an electronic device while the user is speaking; and in accordance with a determination that the visual attention of the user is directed to the electronic device while the user is speaking: identifying a second portion of the audio stream to include user speech intended for the electronic device; initiating, by a digital assistant operating on the electronic device, a task based the second portion of the audio stream; and providing an output indicative of the initiated task.

Claims (130)

1 . A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device, cause the electronic device to:

concurrently receive an audio stream and a video stream;

determine, based on a first portion of the audio stream received within a predetermined duration before a current time and a first portion of the video stream received within the predetermined duration before the current time, whether a visual attention of a user is directed to the electronic device while the user is speaking, wherein the first portion of the video stream includes a plurality of video frames, and wherein determining whether the visual attention of the user is directed to the electronic device while the user is speaking includes:

determining a first confidence score for a first video frame of the plurality of video frames, wherein the first video frame corresponds to a first time, wherein the first confidence score indicates, for the first video frame, whether the visual attention of the user is directed to the electronic device while the user is speaking, and wherein determining the first confidence score includes:

determining an initial first confidence score based on the first video frame; and

adjusting, based on processing a second video frame of the plurality of video frames, the initial first confidence score to obtain the first confidence score, wherein the second video frame corresponds to a second time after the first time; and

in accordance with a determination that the visual attention of the user is directed to the electronic device while the user is speaking:

identify a second portion of the audio stream to include user speech intended for the electronic device;

initiate, by a digital assistant operating on the electronic device, a task based the second portion of the audio stream; and

provide an output indicative of the initiated task.

2 . The non-transitory computer-readable storage medium of claim 1 , wherein determining whether the visual attention of the user is directed to the electronic device while the user is speaking includes:

determining whether the visual attention of the user is directed to a display of the electronic device.

3 . The non-transitory computer-readable storage medium of claim 1 , wherein determining whether the visual attention of the user is directed to the electronic device while the user is speaking includes:

determining whether the visual attention of the user is directed to an affordance displayed by the electronic device.

4 . The non-transitory computer-readable storage medium of claim 1 , wherein determining the initial first confidence score includes processing a third video frame of the plurality of video frames, wherein the third video frame corresponds to a third time before the first time.

5 . The non-transitory computer-readable storage medium of claim 1 , wherein:

determining the first confidence score for the first video frame of the plurality of video frames includes determining a respective confidence score for each video frame of the plurality of video frames to obtain a plurality of respective confidence scores;

the plurality of respective confidence scores include a fourth confidence score for a fourth video frame of the plurality of video frames; and

identifying the second portion of the audio stream to include user speech intended for the electronic device includes:

in accordance with a determination that the fourth confidence score exceeds a threshold, determining that a fourth time corresponding to the fourth video frame is the start time of the second portion of the audio stream.

6 . The non-transitory computer-readable storage medium of claim 1 , wherein:

the plurality of video frames include a fifth video frame and a sixth video frame consecutive to the fifth video frame; and

identifying the second portion of the audio stream to include user speech intended for the electronic device includes:

in accordance with a determination that a fifth confidence score for the fifth video frame is above a second threshold and that a sixth confidence score for the sixth video frame is below the second threshold:

determining that a sixth time corresponding to the sixth video frame is the end time of the second portion of the audio stream.

7 . The non-transitory computer-readable storage medium of claim 1 , wherein determining whether the visual attention of the user is directed to the electronic device while the user is speaking includes:

determining, using a machine learned model, whether the visual attention of the user is directed to the electronic device while the user is speaking, including:

processing a representation of the first portion of the audio stream and the first portion of the video stream using parameters of the machine learned model representing a correlation between mouth movement of the user and speech input.

8 . The non-transitory computer-readable storage medium of claim 1 , wherein determining that the visual attention of the user is directed to the electronic device while the user is speaking includes:

determining that the user faces the electronic device while the user is speaking.

9 . The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

in accordance with a determination that the visual attention of the user is not directed to the electronic device while the user is speaking:

forgo identifying the second portion of the audio stream to include user speech intended for the electronic device.

10 . The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

in accordance with a determination that the visual attention of the user is directed to the electronic device while the user is not speaking:

forgo identifying the second portion of the audio stream to include user speech intended for the electronic device.

11 . The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

in accordance with a determination that the visual attention of the user is not directed to the electronic device while the user is not speaking:

forgo identifying the second portion of the audio stream to include user speech intended for the electronic device.

12 . The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

in accordance with a determination that the user is not visible in the first portion of the video stream:

forgo identifying the second portion of the audio stream to include user speech intended for the electronic device.

13 . An electronic device, comprising:

one or more processors;

a memory; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:

concurrently receiving an audio stream and a video stream;

determining, based on a first portion of the audio stream received within a predetermined duration before a current time and a first portion of the video stream received within the predetermined duration before the current time, whether a visual attention of a user is directed to the electronic device while the user is speaking, wherein the first portion of the video stream includes a plurality of video frames, and wherein determining whether the visual attention of the user is directed to the electronic device while the user is speaking includes:

determining a first confidence score for a first video frame of the plurality of video frames, wherein the first video frame corresponds to a first time, wherein the first confidence score indicates, for the first video frame, whether the visual attention of the user is directed to the electronic device while the user is speaking, and wherein determining the first confidence score includes:

determining an initial first confidence score based on the first video frame; and

adjusting, based on processing a second video frame of the plurality of video frames, the initial first confidence score to obtain the first confidence score, wherein the second video frame corresponds to a second time after the first time; and

in accordance with a determination that the visual attention of the user is directed to the electronic device while the user is speaking:

identifying a second portion of the audio stream to include user speech intended for the electronic device;

initiating, by a digital assistant operating on the electronic device, a task based the second portion of the audio stream; and

providing an output indicative of the initiated task.

14 . The electronic device of claim 13 , wherein determining whether the visual attention of the user is directed to the electronic device while the user is speaking includes:

determining whether the visual attention of the user is directed to a display of the electronic device.

15 . The electronic device of claim 13 , wherein determining whether the visual attention of the user is directed to the electronic device while the user is speaking includes:

determining whether the visual attention of the user is directed to an affordance displayed by the electronic device.

16 . The electronic device of claim 13 , wherein determining the initial first confidence score includes processing a third video frame of the plurality of video frames, wherein the third video frame corresponds to a third time before the first time.

17 . The electronic device of claim 13 , wherein:

determining the first confidence score for the first video frame of the plurality of video frames includes determining a respective confidence score for each video frame of the plurality of video frames to obtain a plurality of respective confidence scores;

the plurality of respective confidence scores include a fourth confidence score for a fourth video frame of the plurality of video frames; and

identifying the second portion of the audio stream to include user speech intended for the electronic device includes:

in accordance with a determination that the fourth confidence score exceeds a threshold, determining that a fourth time corresponding to the fourth video frame is the start time of the second portion of the audio stream.

18 . The electronic device of claim 13 , wherein:

the plurality of video frames include a fifth video frame and a sixth video frame consecutive to the fifth video frame; and

identifying the second portion of the audio stream to include user speech intended for the electronic device includes:

in accordance with a determination that a fifth confidence score for the fifth video frame is above a second threshold and that a sixth confidence score for the sixth video frame is below the second threshold:

determining that a sixth time corresponding to the sixth video frame is the end time of the second portion of the audio stream.

19 . The electronic device of claim 13 , wherein determining whether the visual attention of the user is directed to the electronic device while the user is speaking includes:

determining, using a machine learned model, whether the visual attention of the user is directed to the electronic device while the user is speaking, including:

processing a representation of the first portion of the audio stream and the first portion of the video stream using parameters of the machine learned model representing a correlation between mouth movement of the user and speech input.

20 . The electronic device of claim 13 , wherein determining that the visual attention of the user is directed to the electronic device while the user is speaking includes:

determining that the user faces the electronic device while the user is speaking.

21 . The electronic device of claim 13 , the one or more programs further including instructions for:

in accordance with a determination that the visual attention of the user is not directed to the electronic device while the user is speaking:

forgoing identifying the second portion of the audio stream to include user speech intended for the electronic device.

22 . The electronic device of claim 13 , the one or more programs further including instructions for:

in accordance with a determination that the visual attention of the user is directed to the electronic device while the user is not speaking:

forgoing identifying the second portion of the audio stream to include user speech intended for the electronic device.

23 . The electronic device of claim 13 , the one or more programs further including instructions for:

in accordance with a determination that the visual attention of the user is not directed to the electronic device while the user is not speaking:

forgoing identifying the second portion of the audio stream to include user speech intended for the electronic device.

24 . The electronic device of claim 13 , the one or more programs further including instructions for:

in accordance with a determination that the user is not visible in the first portion of the video stream:

forgoing identifying the second portion of the audio stream to include user speech intended for the electronic device.

25 . A method, comprising:

at an electronic device with one or more processors and memory:

concurrently receiving an audio stream and a video stream;

determining, based on a first portion of the audio stream received within a predetermined duration before a current time and a first portion of the video stream received within the predetermined duration before the current time, whether a visual attention of a user is directed to the electronic device while the user is speaking, wherein the first portion of the video stream includes a plurality of video frames, and wherein determining whether the visual attention of the user is directed to the electronic device while the user is speaking includes:

determining a first confidence score for a first video frame of the plurality of video frames, wherein the first video frame corresponds to a first time, wherein the first confidence score indicates, for the first video frame, whether the visual attention of the user is directed to the electronic device while the user is speaking, and wherein determining the first confidence score includes:

determining an initial first confidence score based on the first video frame; and

adjusting, based on processing a second video frame of the plurality of video frames, the initial first confidence score to obtain the first confidence score, wherein the second video frame corresponds to a second time after the first time; and

in accordance with a determination that the visual attention of the user is directed to the electronic device while the user is speaking:

identifying a second portion of the audio stream to include user speech intended for the electronic device;

initiating, by a digital assistant operating on the electronic device, a task based the second portion of the audio stream; and

providing an output indicative of the initiated task.

26 . The method of claim 25 , wherein determining whether the visual attention of the user is directed to the electronic device while the user is speaking includes:

determining whether the visual attention of the user is directed to a display of the electronic device.

27 . The method of claim 25 , wherein determining whether the visual attention of the user is directed to the electronic device while the user is speaking includes:

determining whether the visual attention of the user is directed to an affordance displayed by the electronic device.

28 . The method of claim 25 , wherein determining the initial first confidence score includes processing a third video frame of the plurality of video frames, wherein the third video frame corresponds to a third time before the first time.

29 . The method of claim 25 , wherein:

determining the first confidence score for the first video frame of the plurality of video frames includes determining a respective confidence score for each video frame of the plurality of video frames to obtain a plurality of respective confidence scores;

the plurality of respective confidence scores include a fourth confidence score for a fourth video frame of the plurality of video frames; and

identifying the second portion of the audio stream to include user speech intended for the electronic device includes:

in accordance with a determination that the fourth confidence score exceeds a threshold, determining that a fourth time corresponding to the fourth video frame is the start time of the second portion of the audio stream.

30 . The method of claim 25 , wherein:

the plurality of video frames include a fifth video frame and a sixth video frame consecutive to the fifth video frame; and

identifying the second portion of the audio stream to include user speech intended for the electronic device includes:

in accordance with a determination that a fifth confidence score for the fifth video frame is above a second threshold and that a sixth confidence score for the sixth video frame is below the second threshold:

determining that a sixth time corresponding to the sixth video frame is the end time of the second portion of the audio stream.

31 . The method of claim 25 , wherein determining whether the visual attention of the user is directed to the electronic device while the user is speaking includes:

determining, using a machine learned model, whether the visual attention of the user is directed to the electronic device while the user is speaking, including:

processing a representation of the first portion of the audio stream and the first portion of the video stream using parameters of the machine learned model representing a correlation between mouth movement of the user and speech input.

32 . The method of claim 25 , wherein determining that the visual attention of the user is directed to the electronic device while the user is speaking includes:

determining that the user faces the electronic device while the user is speaking.

33 . The method of claim 25 , further comprising:

in accordance with a determination that the visual attention of the user is not directed to the electronic device while the user is speaking:

forgoing identifying the second portion of the audio stream to include user speech intended for the electronic device.

34 . The method of claim 25 , further comprising:

in accordance with a determination that the visual attention of the user is directed to the electronic device while the user is not speaking:

forgoing identifying the second portion of the audio stream to include user speech intended for the electronic device.

35 . The method of claim 25 , further comprising:

in accordance with a determination that the visual attention of the user is not directed to the electronic device while the user is not speaking:

forgoing identifying the second portion of the audio stream to include user speech intended for the electronic device.

36 . The method of claim 25 , further comprising:

in accordance with a determination that the user is not visible in the first portion of the video stream:

forgoing identifying the second portion of the audio stream to include user speech intended for the electronic device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2023
From: HORTON, MAXWELL C.; BERARDI, STEPHEN A.; JIN, YANZI; LEBRECHT, SOPHIE; MUFFOLETTO, RICHARD P.; TORMOEN, DANIEL
To: APPLE INC.
Reel/Frame 063415/0257 →
Continuity (3)
Provisional Application 63456639 · Apr 3, 2023
Provisional Application 63346693 · May 27, 2022
Related Publication 20230386469A1 · Nov 30, 2023
References Cited (100)
US 7617094B2 · Aoki et al. · 2009 [cited by applicant]
US 8219406B2 · Yu et al. · 2012 [cited by applicant]
US 8571851B1 · Tickner et al. · 2013 [cited by applicant]
US 9098467B1 · Blanksteen et al. · 2015 [cited by applicant]
US 9250703B2 · Hernandez-Abrego et al. · 2016 [cited by applicant]
US 9274598B2 · Beymer et al. · 2016 [cited by applicant]
US 9990921B2 · Vanblon et al. · 2018 [cited by applicant]
US 10176808B1 · Lovitt et al. · 2019 [cited by applicant]
US 10228904B2 · Raux · 2019 [cited by applicant]
US 10242501B1 · Pusch et al. · 2019 [cited by applicant]
US 10732708B1 · Roche et al. · 2020 [cited by applicant]
US 11133008B2 · Piernot et al. · 2021 [cited by applicant]
US 20100079508A1 · Hodge et al. · 2010 [cited by applicant]
US 20120259638A1 · Kalinli · 2012 [cited by applicant]
US 20120295708A1 · Hernandez-Abrego et al. · 2012 [cited by applicant]
US 20130144616A1 · Bangalore · 2013 [cited by applicant]
US 20130304479A1 · Teller et al. · 2013 [cited by applicant]
US 20130342672A1 · Gray et al. · 2013 [cited by applicant]
US 20140122086A1 · Kapur et al. · 2014 [cited by applicant]
US 20140350924A1 · Zurek et al. · 2014 [cited by applicant]
US 20150033130A1 · Scheessele · 2015 [cited by applicant]
US 20150130716A1 · Sridharan et al. · 2015 [cited by applicant]
US 20150235434A1 · Miller et al. · 2015 [cited by applicant]
US 20150346845A1 · Di Censo et al. · 2015 [cited by applicant]
US 20150348548A1 · Piernot et al. · 2015 [cited by applicant]
US 20160026242A1 · Burns et al. · 2016 [cited by applicant]
US 20160091967A1 · Prokofieva et al. · 2016 [cited by applicant]
US 20160116980A1 · George-Svahn et al. · 2016 [cited by applicant]
US 20160132290A1 · Raux · 2016 [cited by applicant]
US 20160217794A1 · Imoto et al. · 2016 [cited by applicant]
US 20160284350A1 · Yun et al. · 2016 [cited by applicant]
US 20160320838A1 · Teller et al. · 2016 [cited by applicant]
US 20170169818A1 · Vanblon et al. · 2017 [cited by applicant]
US 20170235361A1 · Rigazio et al. · 2017 [cited by applicant]
US 20170242478A1 · Ma · 2017 [cited by applicant]
US 20170262051A1 · Tall et al. · 2017 [cited by applicant]
US 20180012596A1 · Piernot · 2018 [cited by examiner]
US 20180046851A1 · Kienzle et al. · 2018 [cited by applicant]
US 20180225131A1 · Tommy et al. · 2018 [cited by applicant]
US 20180233142A1 · Koishida et al. · 2018 [cited by applicant]
US 20190138268A1 · Andersen et al. · 2019 [cited by applicant]
US 20190139541A1 · Andersen et al. · 2019 [cited by applicant]
US 20190187787A1 · White et al. · 2019 [cited by applicant]
US 20190362557A1 · Lacey et al. · 2019 [cited by applicant]
US 20190391726A1 · Iskandar et al. · 2019 [cited by applicant]
US 20200098362A1 · Piernot et al. · 2020 [cited by applicant]
US 20200103963A1 · Kelly et al. · 2020 [cited by applicant]
US 20200193997A1 · Piernot et al. · 2020 [cited by applicant]
US 20200227034A1 · Summa et al. · 2020 [cited by applicant]
US 20200285327A1 · Hindi et al. · 2020 [cited by applicant]
US 20210089124A1 · Manjunath et al. · 2021 [cited by applicant]
US 20210216134A1 · Fukunaga et al. · 2021 [cited by applicant]
US 20210249009A1 · Manjunath et al. · 2021 [cited by applicant]
US 20210256980A1 · George-Svahn et al. · 2021 [cited by applicant]
US 20210271333A1 · Hindi et al. · 2021 [cited by applicant]
US 20210390955A1 · Piernot et al. · 2021 [cited by applicant]
US 20220093093A1 · Krishnan et al. · 2022 [cited by applicant]
US 20220093101A1 · Krishnan et al. · 2022 [cited by applicant]
US 20220114327A1 · Faaborg et al. · 2022 [cited by applicant]
US 20220155857A1 · Lee et al. · 2022 [cited by applicant]
US 20220293124A1 · Weinberg · 2022 [cited by examiner]
US 20220300094A1 · Hindi et al. · 2022 [cited by applicant]
US 20230035941A1 · Herman et al. · 2023 [cited by applicant]
US 20230081605A1 · O'mara et al. · 2023 [cited by applicant]
US 20240135959A1 · Kulkarni et al. · 2024 [cited by applicant]
US 20240406150A1 · Capelis et al. · 2024 [cited by applicant]
CN 104094580A · 2014 [cited by applicant]
CN 105320726A · 2016 [cited by applicant]
EP 3389045A1 · 2018 [cited by applicant]
JP 2012220959A · 2012 [cited by applicant]
JP 2015514254A · 2015 [cited by applicant]
JP 2017211608A · 2017 [cited by applicant]
JP 2018180523A · 2018 [cited by applicant]
JP 2021521497A · 2021 [cited by applicant]
JP 2021144259A · 2021 [cited by applicant]
KR 1020140007282A · 2014 [cited by applicant]
KR 1020150138109A · 2015 [cited by applicant]
TW 201610982A · 2016 [cited by applicant]
WO 2012158407A1 · 2012 [cited by applicant]
WO 2014159581A1 · 2014 [cited by applicant]
WO 2016049439A1 · 2016 [cited by applicant]
WO 2019190646A2 · 2019 [cited by applicant]
WO 2019212569A1 · 2019 [cited by applicant]
WO WO2019212569 · 2019 [cited by examiner]
WO 2021076164A1 · 2021 [cited by applicant]
WO 2023229989A1 · 2023 [cited by applicant]
Invitation to Pay Additional Fees and Partial International Search Report received for PCT Patent Application No. PCT/US2023/023090, mailed on Aug. 11, 2023, 11 pages. [cited by applicant]
International Search Report and Written Opinion received for PCT Patent Application No. PCT/US2023/023090, mailed on Oct. 2, 2023, 20 pages. [cited by applicant]
International Preliminary Report on Patentability received for PCT Patent Application No. PCT/US2023/023090, mailed on Dec. 12, 2024, 14 pages. [cited by applicant]
Office Action received for European Patent Application No. 23732278.9, mailed on Aug. 19, 2025, 5 pages. [cited by applicant]
Applicant-Initiated Interview Summary received for U.S. Appl. No. 17/500,518, mailed on Aug. 22, 2024, 2 pages. [cited by applicant]
Applicant-Initiated Interview Summary received for U.S. Appl. No. 17/500,518, mailed on Mar. 12, 2024, 5 pages. [cited by applicant]
Final Office Action received for U.S. Appl. No. 17/500,518, mailed on Jul. 17, 2024, 30 pages. [cited by applicant]
Intention to Grant received for European Patent Application No. 23732278.9, mailed on Nov. 17, 2025, 9 pages. [cited by applicant]
International Preliminary Report on Patentability received for PCT Patent Application No. PCT/US2022/036123, mailed on Jan. 25, 2024, 11 pages. [cited by applicant]
International Search Report and Written Opinion received for PCT Patent Application No. PCT/US2022/036123, mailed on Nov. 10, 2022, 14 pages. [cited by applicant]
Non-Final Office Action received for U.S. Appl. No. 17/500,518, mailed on Jan. 29, 2024, 23 pages. [cited by applicant]
Notice of Allowance received for U.S. Appl. No. 17/500,518, mailed on Nov. 7, 2024, 40 pages. [cited by applicant]
Office Action received for Japanese Patent Application No. 2024-569639, mailed on Jan. 19, 2026, 14 pages (6 pages of English Translation and 8 pages of Official Copy). [cited by applicant]
Invitation to Pay Search Fees received for European Patent Application No. 23732278.9, mailed on Feb. 11, 2026, 5 pages. [cited by applicant]