IP Library Granted Patent US 12,327,573
Granted Patent B2
US 12,327,573 · App. 16/850,965 · Granted Jun 10, 2025

Identifying input for speech recognition engine

Inventors: Anthony Robert Sheeder (Fort Lauderdale, FL); Tushar Arora (Hollywood, FL)
Assignee: Magic Leap, Inc.
G10L25/87G10L15/22G10L2015/227
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,327,573
App. No.
16/850,965
Granted
Jun 10, 2025
Kind
B2
Abstract

A method of presenting a signal to a speech recognition engine is disclosed. According to an example of the method, an audio signal is received from a user. A portion of the audio signal is identified, the portion having a first time and a second time. A pause in the portion of the audio signal, the pause comprising the second time, is identified. It is determined whether the pause indicates the completion of an utterance of the audio signal. In accordance with a determination that the pause indicates the completion of the utterance, the portion of the audio signal is presented as input to the speech recognition engine. In accordance with a determination that the pause does not indicate the completion of the utterance, the portion of the audio signal is not presented as input to the speech recognition engine.

Claims (136)

1. A method comprising:

receiving, via a microphone and a first processor of a head-wearable device, an audio signal, wherein the audio signal comprises voice activity;

dividing, via the first processor, the audio signal into a plurality of audio signal segments;

receiving, via the first processor and one or more sensors of the head-wearable device, non-verbal sensor data corresponding to a user of the head-wearable device, wherein:

the one or more sensors comprise a camera,

the receiving the non-verbal sensor data comprises receiving, via the camera, information associated with an environment of the user concurrently with the receiving the audio signal,

the information indicates a position and an orientation of the user, and

the information is determined based on one or more of Simultaneous Localization and Mapping (SLAM) and visual odometry performed via the first processor and further via visual data from the camera;

classifying, via a first classifier, audio data corresponding to the audio signal;

for each of the plurality of the audio signal segments:

determining, via the first processor, one or more of an amplitude and frequency components of a respective audio signal segment;

receiving, via a second processor, the respective audio signal segment of the plurality of the audio signal segments and data associated with one or more of the amplitude and the frequency components of the respective audio signal segment;

determining, via the second processor and based on one or more of the amplitude and the frequency components of the respective audio signal segment, whether the respective audio signal segment comprises a pause in the voice activity;

responsive to determining that the respective audio signal segment comprises a pause in the voice activity, determining, based on the classifying of the audio data, whether the pause in the voice activity corresponds to an end point of the voice activity; and

responsive to determining that the pause in the voice activity corresponds to an end point of the voice activity, presenting a response to the user based on the voice activity, wherein:

determining whether the respective audio signal segment comprises the pause in the voice activity comprises:

determining, based on the information determined based on the one or more of SLAM and visual odometry, head poses of the user associated with the respective audio signal segment, and

determining whether the head poses comprise a head pose change, wherein the respective audio signal segment comprises the pause in accordance with a determination that the head poses comprise the head pose change, and

determining whether the pause in the voice activity corresponds to an end point of the voice activity comprises:

determining a probability of interest based on the classifying of the audio data and based further on applying the non-verbal sensor data as input to a machine learning process, the probability determined based on relative distances between the non-verbal sensor data and its neighbors in an N-dimensional space,

determining whether the probability of interest exceeds a threshold,

in accordance with a determination that the probability of interest exceeds the threshold, determining that the pause in the voice activity corresponds to an end point of the voice activity, and

in accordance with a determination that the probability of interest does not exceed the threshold, prompting the user to speak, and

presenting the response to the user comprises presenting the response via a transmissive display of the head-wearable device, concurrently with presenting to the user a view of an external environment via the transmissive display.

2. The method of claim 1 further comprising:

in accordance with the determination that the probability of interest does not exceed the threshold:

determining that the pause in the voice activity does not correspond to an end point of the voice activity, and

forgoing presenting the response to the user based on the voice activity.

3. The method of claim 1 , wherein determining whether the respective audio signal segment comprises a pause in the voice activity further comprises determining whether the amplitude of the respective audio signal segment falls below a second threshold for a predetermined period of time.

4. The method of claim 1 further comprising:

in accordance with the determination that the probability of interest does not exceed the threshold:

determining that the pause in the voice activity does not correspond to an end point of the voice activity, and

determining whether the audio signal comprises a second pause corresponding to an end point of the voice activity.

5. The method of claim 1 , wherein determining whether the respective audio signal segment comprises a pause in the voice activity further comprises determining whether the respective audio signal segment comprises one or more verbal cues corresponding to an end point of the voice activity.

6. The method of claim 5 , wherein the one or more verbal cues comprise a characteristic of the user's prosody.

7. The method of claim 5 , wherein the one or more verbal cues comprise a terminating phrase.

8. The method of claim 1 , wherein the non-verbal sensor data is indicative of the user's gaze.

9. The method of claim 1 , wherein the non-verbal sensor data is indicative of the user's facial expression.

10. The method of claim 1 , wherein the non-verbal sensor data is indicative of the user's heart rate.

11. The method of claim 1 , wherein determining whether the pause in the voice activity corresponds to an end point of the voice activity comprises identifying one or more interstitial sounds.

12. The method of claim 1 , wherein the machine learning process comprises one or more of an artificial neural network, a genetic algorithm, a nearest neighbor interpolation, and a support vector machine.

13. The method of claim 1 , further comprising: further in accordance with a determination that the probability of interest does not exceed the threshold, comparing a confidence value to a second threshold, the confidence value indicative of whether the voice activity is complete;

wherein prompting the user to speak is performed further in accordance with a determination that the confidence value does not exceed the second threshold.

14. The system of claim 1 , wherein the probability of interest is determined via the second processor.

15. The method of claim 1 , wherein the first classifier is configured to classify the audio data into a first class indicating a pause, a second class indicating a completed utterance, and a third class indicating a presence of an interstitial sound.

16. A system comprising:

a microphone of a head-wearable device;

one or more sensors of the head-wearable device;

a transmissive display of the head-wearable device; and

one or more processors, the one or more processor comprising a first processor of the head-wearable device and a second processor, configured to execute a method comprising:

receiving, via the microphone and the first processor of the head-wearable device, an audio signal, wherein the audio signal comprises voice activity;

dividing, via the first processor, the audio signal into a plurality of audio signal segments;

receiving, via the first processor and the one or more sensors of the head-wearable device, non-verbal sensor data corresponding to a user of the head-wearable device, wherein:

the one or more sensors comprise a camera,

the receiving the non-verbal sensor data comprises receiving, via the camera, information associated with an environment of the user concurrently with the receiving the audio signal,

the information indicates a position and an orientation of the user, and

the information is determined based on one or more of Simultaneous Localization and Mapping (SLAM) and visual odometry performed via the first processor and further via visual data from the camera;

classifying, via a first classifier, audio data corresponding to the audio signal;

for each of the plurality of the audio signal segments:

determining, via the first processor, one or more of an amplitude and frequency components of a respective audio signal segment;

receiving, via a second processor, the respective audio signal segment of the plurality of the audio signal segments and data associated with one or more of the amplitude and the frequency components of the respective audio signal segment;

determining, via the second processor and based on one or more of the amplitude and the frequency components of the respective audio signal segment, whether the respective audio signal segment comprises a pause in the voice activity;

responsive to determining that the respective audio signal segment comprises a pause in the voice activity, determining, based on the classifying of the audio data, whether the pause in the voice activity corresponds to an end point of the voice activity; and

responsive to determining that the pause in the voice activity corresponds to an end point of the voice activity, presenting a response to the user based on the voice activity, wherein:

determining whether the respective audio signal segment comprises the pause in the voice activity comprises:

determining, based on the information determined based on the one or more of SLAM and visual odometry, head poses of the user associated with the respective audio signal segment, and

determining whether the head poses comprise a head pose change, wherein the respective audio signal segment comprises the pause in accordance with a determination that the head poses comprise the head pose change, and

determining whether the pause in the voice activity corresponds to an end point of the voice activity comprises:

determining a probability of interest based on the classifying of the audio data and based further on applying the non-verbal sensor data as input to a machine learning process, the probability determined based on relative distances between the non-verbal sensor data and its neighbors in an N-dimensional space,

determining whether the probability of interest exceeds a threshold,

in accordance with a determination that the probability of interest exceeds the threshold, determining that the pause in the voice activity corresponds to an end point of the voice activity, and

in accordance with a determination that the probability of interest does not exceed the threshold, prompting the user to speak, and

presenting the response to the user comprises presenting the response via the transmissive display, concurrently with presenting to the user a view of an external environment via the transmissive display.

17. The system of claim 16 further comprising:

in accordance with the determination that the probability of interest does not exceed the threshold:

determining that the pause in the voice activity does not correspond to an end point of the voice activity, and

forgoing presenting the response to the user based on the voice activity.

18. The system of claim 16 , wherein determining whether the respective audio signal segment comprises a pause in the voice activity further comprises determining whether the amplitude of the respective audio signal segment falls below a second threshold for a predetermined period of time.

19. The system of claim 16 further comprising:

in accordance with the determination that the probability of interest does not exceed the threshold:

determining that the pause in the voice activity does not correspond to an end point of the voice activity, and

determining whether the audio signal comprises a second pause corresponding to an end point of the voice activity.

20. The system of claim 16 , wherein determining whether the respective audio signal segment comprises a pause in the voice activity further comprises determining whether the respective audio signal segment comprises one or more verbal cues corresponding to an end point of the voice activity.

21. The system of claim 20 , wherein the one or more verbal cues comprise a characteristic of the user's prosody.

22. The system of claim 20 , wherein the one or more verbal cues comprise a terminating phrase.

23. The system of claim 16 , wherein the non-verbal sensor data is indicative of the user's gaze.

24. The system of claim 16 , wherein the non-verbal sensor data is indicative of the user's facial expression.

25. The system of claim 16 , wherein the non-verbal sensor data is indicative of the user's heartrate.

26. The system of claim 16 , wherein determining whether the pause in the voice activity corresponds to an end point of the voice activity comprises identifying one or more interstitial sounds.

27. The system of claim 16 , wherein the machine learning process comprises one or more of an artificial neural network, a genetic algorithm, a nearest neighbor interpolation, and a support vector machine.

28. The system of claim 16 , wherein the method further comprises:

further in accordance with a determination that the probability of interest does not exceed the threshold, comparing a confidence value to a second threshold, the confidence value indicative of whether the voice activity is complete;

wherein prompting the user to speak is performed further in accordance with a determination that the confidence value does not exceed the second threshold.

29. A non-transitory computer-readable medium storing one or more instructions, which, when executed by one or more processors of an electronic device, cause the device to perform a method comprising:

receiving, via a microphone and a first processor of a head-wearable device, an audio signal, wherein the audio signal comprises voice activity;

dividing, via the first processor, the audio signal into a plurality of audio signal segments;

receiving, via the first processor and one or more sensors of the head-wearable device, non-verbal sensor data corresponding to a user of the head-wearable device, wherein:

the one or more sensors comprise a camera,

the receiving the non-verbal sensor data comprises receiving, via the camera, information associated with an environment of the user concurrently with the receiving the audio signal;

the information indicates a position and an orientation of the user, and

the information is determined based on one or more of Simultaneous Localization and Mapping (SLAM) and visual odometry performed via the first processor and further via visual data from the camera;

classifying, via a first classifier, audio data corresponding to the audio signal;

for each of the plurality of the audio signal segments:

determining, via the first processor, one or more of an amplitude and frequency components of a respective audio signal segment;

receiving, via a second processor, the respective audio signal segment of the plurality of the audio signal segments and data associated with one or more of the amplitude and the frequency components of the respective audio signal segment;

determining, via the second processor and based on one or more of the amplitude and the frequency components of the respective audio signal segment, whether the respective audio signal segment comprises a pause in the voice activity;

responsive to determining that the respective audio signal segment comprises a pause in the voice activity, determining, based on the classifying of the audio data, whether the pause in the voice activity corresponds to an end point of the voice activity; and

responsive to determining that the pause in the voice activity corresponds to an end point of the voice activity, presenting a response to the user based on the voice activity, wherein:

determining whether the respective audio signal segment comprises the pause in the voice activity comprises:

determining, based on the information determined based on the one or more of SLAM and visual odometry, head poses of the user associated with the respective audio signal segment, and

determining whether the head poses comprise a head pose change, wherein the respective audio signal segment comprises the pause in accordance with a determination that the head poses comprise the head pose change, and

determining whether the pause in the voice activity corresponds to an end point of the voice activity comprises:

determining a probability of interest based on the classifying of the audio data and based further on applying the non-verbal sensor data as input to a machine learning process, the probability determined based on relative distances between the non-verbal sensor data and its neighbors in an N-dimensional space,

determining whether the probability of interest exceeds a threshold,

in accordance with a determination that the probability of interest exceeds the threshold, determining that the pause in the voice activity corresponds to an end point of the voice activity, and

in accordance with a determination that the probability of interest does not exceed the threshold, prompting the user to speak, and

presenting the response to the user comprises presenting the response via a transmissive display of the head-wearable device, concurrently with presenting to the user a view of an external environment via the transmissive display.

30. The non-transitory computer-readable medium of claim 29 , the method further comprising:

in accordance with the determination that the probability of interest does not exceed the threshold:

determining that the pause in the voice activity does not correspond to an end point of the voice activity, and

forgoing presenting the response to the user based on the voice activity.

31. The non-transitory computer-readable medium of claim 29 , wherein determining whether the respective audio signal segment comprises a pause in the voice activity further comprises determining whether the amplitude of the respective audio signal segment falls below a second threshold for a predetermined period of time.

32. The non-transitory computer-readable medium of claim 29 , the method further comprising:

in accordance with the determination that the probability of interest does not exceed the threshold:

determining that the pause in the voice activity does not correspond to an end point of the voice activity, and

determining whether the audio signal comprises a second pause corresponding to an end point of the voice activity.

33. The non-transitory computer-readable medium of claim 29 , wherein determining whether the respective audio signal segment comprises a pause in the voice activity further comprises determining whether the respective audio signal segment comprises one or more verbal cues corresponding to an end point of the voice activity.

34. The non-transitory computer-readable medium of claim 33 , wherein the one or more verbal cues comprise a characteristic of the user's prosody.

35. The non-transitory computer-readable medium of claim 33 , wherein the one or more verbal cues comprise a terminating phrase.

36. The non-transitory computer-readable medium of claim 29 , wherein the non-verbal sensor data is indicative of the user's gaze.

37. The non-transitory computer-readable medium of claim 29 , wherein the non-verbal sensor data is indicative of the user's facial expression.

38. The non-transitory computer-readable medium of claim 29 , wherein the non-verbal sensor data is indicative of the user's heartrate.

39. The non-transitory computer-readable medium of claim 29 , wherein determining whether the pause in the voice activity corresponds to an end point of the voice activity comprises identifying one or more interstitial sounds.

40. The non-transitory computer-readable medium of claim 29 , wherein the machine learning process comprises one or more of an artificial neural network, a genetic algorithm, a nearest neighbor interpolation, and a support vector machine.

41. The non-transitory computer-readable medium of claim 29 , wherein the method further comprises: further in accordance with a determination that the probability of interest does not exceed the threshold, comparing a confidence value to a second threshold, the confidence value indicative of whether the voice activity is complete;

wherein prompting the user to speak is performed further in accordance with a determination that the confidence value does not exceed the second threshold.

Assignments (3)
SECURITY INTEREST Recorded Oct 31, 2025
From: MAGIC LEAP, INC.; MENTOR ACQUISITION ONE, LLC; MOLECULAR IMPRINTS, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 073439/0168 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 9, 2023
From: SHEEDER, ANTHONY ROBERT; ARORA, TUSHAR
To: MAGIC LEAP, INC.
Reel/Frame 062641/0092 →
SECURITY INTEREST Recorded May 24, 2022
From: MOLECULAR IMPRINTS, INC.; MENTOR ACQUISITION ONE, LLC; MAGIC LEAP, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 060338/0665 →
Continuity (2)
Provisional Application 62836593 · Apr 19, 2019
Related Publication 20200335128A1 · Oct 22, 2020
References Cited (251)
US 4158750A · Sakoe · 1979 [cited by applicant]
US 4852988A · Velez · 1989 [cited by applicant]
US 6433760B1 · Vaissie · 2002 [cited by applicant]
US 6491391B1 · Blum et al. · 2002 [cited by applicant]
US 6496799B1 · Pickering · 2002 [cited by applicant]
US 6738482B1 · Jaber · 2004 [cited by applicant]
US 6820056B1 · Harif · 2004 [cited by applicant]
US 6847336B1 · Lemelson · 2005 [cited by applicant]
US 6943754B2 · Aughey · 2005 [cited by applicant]
US 6977776B2 · Volkenandt et al. · 2005 [cited by applicant]
US 7346654B1 · Weiss · 2008 [cited by applicant]
US 7347551B2 · Fergason et al. · 2008 [cited by applicant]
US 7488294B2 · Torch · 2009 [cited by applicant]
US 7587319B2 · Catchpole · 2009 [cited by applicant]
US 7979277B2 · Larri et al. · 2011 [cited by applicant]
US 8154588B2 · Burns · 2012 [cited by applicant]
US 8235529B1 · Raffle · 2012 [cited by applicant]
US 8611015B2 · Wheeler · 2013 [cited by applicant]
US 8638498B2 · Bohn et al. · 2014 [cited by applicant]
US 8696113B2 · Lewis · 2014 [cited by applicant]
US 8929589B2 · Publicover et al. · 2015 [cited by applicant]
US 9010929B2 · Lewis · 2015 [cited by applicant]
US 9274338B2 · Robbins et al. · 2016 [cited by applicant]
US 9292973B2 · Bar-zeev et al. · 2016 [cited by applicant]
US 9294860B1 · Carlson · 2016 [cited by applicant]
US 9323325B2 · Perez et al. · 2016 [cited by applicant]
US 9715875B2 · Piernot · 2017 [cited by applicant]
US 9720505B2 · Gribetz et al. · 2017 [cited by applicant]
US 10013053B2 · Cederlund et al. · 2018 [cited by applicant]
US 10025379B2 · Drake et al. · 2018 [cited by applicant]
US 10062377B2 · Larri et al. · 2018 [cited by applicant]
US 10134425B1 · Johnson, Jr. · 2018 [cited by applicant]
US 10289205B1 · Sumter · 2019 [cited by examiner]
US 10839789B2 · Larri et al. · 2020 [cited by applicant]
US 10971140B2 · Catchpole · 2021 [cited by applicant]
US 11151997B2 · Sugiyama et al. · 2021 [cited by applicant]
US 11328740B2 · Lee et al. · 2022 [cited by applicant]
US 11587563B2 · Sheeder et al. · 2023 [cited by applicant]
US 11790935B2 · Lee et al. · 2023 [cited by applicant]
US 11854550B2 · Sheeder · 2023 [cited by applicant]
US 11854566B2 · Leider · 2023 [cited by applicant]
US 11917384B2 · Roach · 2024 [cited by applicant]
US 12094489B2 · Lee · 2024 [cited by applicant]
US 12238496B2 · Roach et al. · 2025 [cited by applicant]
US 12243531B2 · Sheeder et al. · 2025 [cited by applicant]
US 20010055985A1 · Matt et al. · 2001 [cited by applicant]
US 20030030597A1 · Geist · 2003 [cited by applicant]
US 20050033571A1 · Huang · 2005 [cited by applicant]
US 20050069852A1 · Janakiraman · 2005 [cited by applicant]
US 20060023158A1 · Howell et al. · 2006 [cited by applicant]
US 20060072767A1 · Zhang et al. · 2006 [cited by applicant]
US 20060098827A1 · Paddock et al. · 2006 [cited by applicant]
US 20060178876A1 · Sato et al. · 2006 [cited by applicant]
US 20070225982A1 · Washio · 2007 [cited by applicant]
US 20080124690A1 · Redlich · 2008 [cited by examiner]
US 20080201138A1 · Visser et al. · 2008 [cited by applicant]
US 20090180626A1 · Nakano · 2009 [cited by applicant]
US 20100245585A1 · Fisher et al. · 2010 [cited by applicant]
US 20100323652A1 · Visser et al. · 2010 [cited by applicant]
US 20110211056A1 · Publicover et al. · 2011 [cited by applicant]
US 20110213664A1 · Osterhout · 2011 [cited by applicant]
US 20110238407A1 · Kent · 2011 [cited by applicant]
US 20110288860A1 · Schevciw et al. · 2011 [cited by applicant]
US 20120021806A1 · Maltz · 2012 [cited by applicant]
US 20120130713A1 · Shin et al. · 2012 [cited by applicant]
US 20120209601A1 · Jing · 2012 [cited by applicant]
US 20130077147A1 · Efimov · 2013 [cited by applicant]
US 20130204607A1 · Baker, IV · 2013 [cited by examiner]
US 20130226589A1 · Largey · 2013 [cited by applicant]
US 20130236040A1 · Crawford · 2013 [cited by applicant]
US 20130339028A1 · Rosner et al. · 2013 [cited by applicant]
US 20140016793A1 · Gardner · 2014 [cited by applicant]
US 20140194702A1 · Tran · 2014 [cited by applicant]
US 20140195918A1 · Friedlander · 2014 [cited by applicant]
US 20140200887A1 · Nakadal et al. · 2014 [cited by applicant]
US 20140222430A1 · Rao · 2014 [cited by applicant]
US 20140270202A1 · Ivanov et al. · 2014 [cited by applicant]
US 20140270244A1 · Fan · 2014 [cited by applicant]
US 20140310595A1 · Acharya et al. · 2014 [cited by applicant]
US 20140337023A1 · Mcculloch et al. · 2014 [cited by applicant]
US 20140379336A1 · Bhatnagar · 2014 [cited by applicant]
US 20150006181A1 · Fan et al. · 2015 [cited by applicant]
US 20150168731A1 · Robbins · 2015 [cited by applicant]
US 20150310857A1 · Habets et al. · 2015 [cited by applicant]
US 20150348572A1 · Thornburg et al. · 2015 [cited by applicant]
US 20160019910A1 · Faubel et al. · 2016 [cited by applicant]
US 20160066113A1 · Elkhatib et al. · 2016 [cited by applicant]
US 20160112817A1 · Fan et al. · 2016 [cited by applicant]
US 20160142830A1 · Hu · 2016 [cited by applicant]
US 20160165340A1 · Benattar · 2016 [cited by applicant]
US 20160180837A1 · Gustavsson · 2016 [cited by applicant]
US 20160216130A1 · Abramson et al. · 2016 [cited by applicant]
US 20160217781A1 · Zhong et al. · 2016 [cited by applicant]
US 20160284350A1 · Yun et al. · 2016 [cited by applicant]
US 20160358598A1 · Williams · 2016 [cited by examiner]
US 20160379629A1 · Hofer · 2016 [cited by applicant]
US 20160379632A1 · Hoffmeister · 2016 [cited by applicant]
US 20160379638A1 · Basye · 2016 [cited by examiner]
US 20170078819A1 · Habets · 2017 [cited by applicant]
US 20170091169A1 · Bellegarda · 2017 [cited by applicant]
US 20170092276A1 · Sun et al. · 2017 [cited by applicant]
US 20170110116A1 · Tadpatrikar · 2017 [cited by applicant]
US 20170148429A1 · Hayakawa · 2017 [cited by applicant]
US 20170270919A1 · Parthasarathi · 2017 [cited by applicant]
US 20170280239A1 · Sekiya · 2017 [cited by applicant]
US 20170316780A1 · Lovitt · 2017 [cited by applicant]
US 20170330555A1 · Kawano · 2017 [cited by applicant]
US 20170332187A1 · Lin · 2017 [cited by applicant]
US 20180011534A1 · Poulos et al. · 2018 [cited by applicant]
US 20180053284A1 · Rodriguez et al. · 2018 [cited by applicant]
US 20180077095A1 · Deyle et al. · 2018 [cited by applicant]
US 20180129469A1 · Vennström et al. · 2018 [cited by applicant]
US 20180227665A1 · Elko et al. · 2018 [cited by applicant]
US 20180316939A1 · Todd · 2018 [cited by applicant]
US 20180336902A1 · Cartwright · 2018 [cited by examiner]
US 20180349946A1 · Nguyen · 2018 [cited by examiner]
US 20180358021A1 · Mistica · 2018 [cited by examiner]
US 20180366114A1 · Anbazhagan · 2018 [cited by applicant]
US 20190129944A1 · Kawano · 2019 [cited by applicant]
US 20190362704A1 · Nicolis et al. · 2019 [cited by applicant]
US 20190362741A1 · Li et al. · 2019 [cited by applicant]
US 20190373362A1 · Ansai et al. · 2019 [cited by applicant]
US 20190392641A1 · Taylor · 2019 [cited by applicant]
US 20200027455A1 · Sugiyama · 2020 [cited by examiner]
US 20200064921A1 · Kang · 2020 [cited by examiner]
US 20200194028A1 · Lipman · 2020 [cited by applicant]
US 20200213729A1 · Soto · 2020 [cited by applicant]
US 20200279552A1 · Piersol et al. · 2020 [cited by applicant]
US 20200279561A1 · Sheeder · 2020 [cited by applicant]
US 20200286465A1 · Wang et al. · 2020 [cited by applicant]
US 20200296521A1 · Wexler · 2020 [cited by examiner]
US 20210043223A1 · Lee et al. · 2021 [cited by applicant]
US 20210056966A1 · Bilac · 2021 [cited by examiner]
US 20210125609A1 · Dusan et al. · 2021 [cited by applicant]
US 20210192413A1 · Shirazipour · 2021 [cited by examiner]
US 20210264931A1 · Leider · 2021 [cited by applicant]
US 20210306751A1 · Roach et al. · 2021 [cited by applicant]
US 20220230658A1 · Lee et al. · 2022 [cited by applicant]
US 20230135768A1 · Sheeder et al. · 2023 [cited by applicant]
US 20230386461A1 · Leider · 2023 [cited by applicant]
US 20230410835A1 · Lee · 2023 [cited by applicant]
US 20240087565A1 · Sheeder · 2024 [cited by applicant]
US 20240087587A1 · Leider · 2024 [cited by applicant]
US 20240163612A1 · Roach · 2024 [cited by applicant]
US 20240420718A1 · Audfray · 2024 [cited by applicant]
US 20250006219A1 · Lee · 2025 [cited by applicant]
CA 2316473A1 · 2001 [cited by applicant]
CA 2362895A1 · 2002 [cited by applicant]
CA 2388766A1 · 2003 [cited by applicant]
CN 105529033A · 2016 [cited by applicant]
EP 2950307A1 · 2015 [cited by applicant]
EP 3211918A1 · 2017 [cited by applicant]
JP S52144205A · 1977 [cited by applicant]
JP 06075588A · 1994 [cited by applicant]
JP 2000148184A · 2000 [cited by applicant]
JP 2002135173A · 2002 [cited by applicant]
JP 2005196134A · 2005 [cited by applicant]
JP 2008242067A · 2008 [cited by applicant]
JP 2010273305A · 2010 [cited by applicant]
JP 2013183358A · 2013 [cited by applicant]
JP 2014137405A · 2014 [cited by applicant]
JP 2014178339A · 2014 [cited by applicant]
JP 2016004270A · 2016 [cited by applicant]
JP 2017211596A · 2017 [cited by applicant]
JP 2018523156A · 2018 [cited by applicant]
JP 2018179954A · 2018 [cited by applicant]
WO 2014113891A1 · 2014 [cited by applicant]
WO 2014159581A1 · 2014 [cited by applicant]
WO 2015169618A1 · 2015 [cited by applicant]
WO 2016063587A1 · 2016 [cited by applicant]
WO 2016151956A1 · 2016 [cited by applicant]
WO 2016153712A1 · 2016 [cited by applicant]
WO 2017003903A1 · 2017 [cited by applicant]
WO 2017017591A1 · 2017 [cited by applicant]
WO 2017191711A1 · 2017 [cited by applicant]
WO 2018163648A1 · 2018 [cited by applicant]
WO 2019224292A1 · 2019 [cited by applicant]
WO 2020180719A1 · 2020 [cited by applicant]
WO 2022072752A1 · 2022 [cited by applicant]
WO 2023064875A1 · 2023 [cited by applicant]
Backstrom, T. (Oct. 2015). “Voice Activity Detection Speech Processing,” Aalto University, vol. 58, No. 10; Publication [online], retrieved Apr. 19, 2020, retrieved from the Internet: URL: https://mycourses.aalto.fi/plu… [cited by applicant]
International Search Report and Written Opinion mailed May 18, 2020, for PCT Application No. PCT/US20/20469, filed Feb. 28, 2020, twenty pages. [cited by applicant]
Non-Final Office Action mailed Jun. 24, 2021, for U.S. Appl. No. 16/805,337, filed Feb. 28, 2020, fourteen pages. [cited by applicant]
Bilac, et al. (Nov. 15, 2017). Gaze and Filled Pause Detection for Smooth Human-Robot Conversations. www.angelicalim.com, retrieved on Jun. 17, 2020, Retrieved from the internet URL: http://www.angelicalim.com/papers/hu… [cited by applicant]
International Search Report and Written Opinion mailed Jul. 2, 2020, for PCT Application No. PCT/US2020/028570, filed Apr. 16, 2020, nineteen pages. [cited by applicant]
Kitayama, et al. (Sep. 30, 2003). “Speech Starter: Noise-Robust Endpoint Detection by Using Filled Pauses.” Eurospeech 2003, retrieved on Jun. 17, 2020, retrieved from the internet URL: http://clteseerx.ist.psu.edu/view… [cited by applicant]
Harma, A. et al. (Jun. 2004). “Augmented Reality Audio for Mobile and Wearable Appliances,” J. Audio Eng. Soc., vol. 52, No. 6, retrieved on Aug. 20, 2019, Retrieved from the Internet: URL:https://pdfs.semanticscholar.o… [cited by applicant]
International Preliminary Report and Patentability mailed Dec. 22, 2020, for PCT Application No. PCT/US2019/038546, 13 pages. [cited by applicant]
International Search Report and Written Opinion mailed Sep. 17, 2019, for PCT Application No. PCT/US2019/038546, sixteen pages. [cited by applicant]
Tonges, R. (Dec. 2015). “An augmented Acoustics Demonstrator with Realtime stereo up-mixing and Binaural Auralization,” Technische University Berlin, Audio Communication Group, retrieved on Aug. 22, 2019, Retrieved from… [cited by applicant]
European Search Report dated Nov. 12, 2021, for EP Application No. 19822754.8, ten pages. [cited by applicant]
International Search Report and Written Opinion mailed Jan. 24, 2022, for PCT Application No. PCT/US2021/53046, filed Sep. 30, 2021, 15 pages. [cited by applicant]
Non-Final Office Action mailed Mar. 17, 2022, for U.S. Appl. No. 16/805,337, filed Feb. 28, 2020, sixteen pages. [cited by applicant]
Notice of Allowance mailed Mar. 3, 2022, for U.S. Appl. No. 16/987,267, filed Aug. 6, 2020, nine pages. [cited by applicant]
Final Office Action mailed Oct. 6, 2021, for U.S. Appl. No. 16/805,337, filed Feb. 28, 2020, fourteen pages. [cited by applicant]
International Preliminary Report and Written Opinion mailed Oct. 28, 2021, for PCT Application No. PCT/US2020/028570, filed Apr. 16, 2020, 17 pages. [cited by applicant]
International Preliminary Report and Written Opinion mailed Sep. 16, 2021, for PCT Application No. PCT/US2020/020469, filed Feb. 28, 2020, nine pages. [cited by applicant]
Non-Final Office Action mailed Nov. 17, 2021, for U.S. Appl. No. 16/987,267, filed Aug. 6, 2020, 21 pages. [cited by applicant]
Non-Final Office Action mailed Aug. 10, 2022, for U.S. Appl. No. 17/214,446, filed Mar. 26, 2021, fifteen pages. [cited by applicant]
Final Office Action mailed Aug. 5, 2022, for U.S. Appl. No. 16/805,337, filed Feb. 28, 2020, eighteen pages. [cited by applicant]
Final Office Action mailed Jan. 11, 2023, for U.S. Appl. No. 17/214,446, filed Mar. 26, 2021, sixteen pages. [cited by applicant]
European Search Report dated Oct. 6, 2022, for EP Application No. 20766540.7, nine pages. [cited by applicant]
Jacob, R. “Eye Tracking in Advanced Interface Design”, Virtual Environments and Advanced Interface Design, Oxford University Press, Inc. (Jun. 1995). [cited by applicant]
Rolland, J. et al., “High-resolution inset head-mounted display”, Optical Society of America, vol. 37, No. 19, Applied Optics, (Jul. 1, 1998). [cited by applicant]
Tanriverdi, V. et al. (Apr. 2000). “Interacting With Eye Movements In Virtual Environments,” Department of Electrical Engineering and Computer Science, Tufts University, Medford, MA 02155, USA, Proceedings of the SIGCHI… [cited by applicant]
Yoshida, A. et al., “Design and Applications of a High Resolution Insert Head Mounted Display”, (Jun. 1994). [cited by applicant]
Notice of Allowance mailed Nov. 30, 2022, for U.S. Appl. No. 16/805,337, filed Feb. 28, 2020, six pages. [cited by applicant]
European Search Report dated Nov. 21, 2022, for EP Application No. 20791183.5 nine pages. [cited by applicant]
Liu, Baiyang, et al.: (Sep. 6, 2015). “Accurate Endpointing with Expected Pause Duration,” Interspeech 2015, pp. 2912-2916, retrieved from: https://scholar.google.com/scholar?q=BAIYANG,+Liu+et+al.: +(September+6,+2015).… [cited by applicant]
Shannon, Matt et al. (Aug. 20-24, 2017). “Improved End-of-Query Detection for Streaming Speech Recognition”, Interspeech 2017, Stockholm, Sweden, pp. 1909-1913. [cited by applicant]
European Office Action dated Jun. 1, 2023, for EP Application No. 19822754.8, six pages. [cited by applicant]
Chinese Office Action dated Jun. 2, 2023, for ON Application No. 2020-571488, with English translation, 9 pages. [cited by applicant]
International Preliminary Report and Written Opinion mailed Apr. 13, 2023, for PCT Application No. PCT/US2021/53046, filed Sep. 30, 2021, nine pages. [cited by applicant]
Non-Final Office Action mailed Apr. 12, 2023, for U.S. Appl. No. 17/214,446, filed Mar. 26, 2021, seventeen pages. [cited by applicant]
Non-Final Office Action mailed Apr. 13, 2023, for U.S. Appl. No. 17/714,708, filed Apr. 6, 2022, sixteen pages. [cited by applicant]
Non-Final Office Action malled Apr. 27, 2023, for U.S. Appl. No. 17/254,832, filed Dec. 21, 2020, fourteen pages. [cited by applicant]
Non-Final Office Action mailed Jun. 23, 2023, for U.S. Appl. No. 18/148,221, filed Dec. 29, 2022, thirteen pages. [cited by applicant]
Notice of Allowance mailed Jul. 31, 2023, for U.S. Appl. No. 17/714,708, filed Apr. 6, 2022, eight pages. [cited by applicant]
Final Office Action mailed Aug. 4, 2023, for U.S. Appl. No. 17/254,832, filed Dec. 21. 2020, seventeen pages. [cited by applicant]
Chinese Office Action dated Dec. 21, 2023, for CN Application No. 201980050714.4, with English translation, eighteen pages. [cited by applicant]
Japanese Notice of Allowance mailed Dec. 15, 2023, for JP Application No. 2020-571488, with English translation, eight pages. [cited by applicant]
Japanese Office Action mailed Jan. 30, 2024, for JP Application No. 2021-551538, with English translation, sixteen pages. [cited by applicant]
Notice of Allowance mailed Dec. 15, 2023, for U.S. Appl. No. 17/214,446, filed Mar. 26, 2021, seven 4 pages. [cited by applicant]
European Office Action dated Dec. 12, 2023, for EP Application No. 20766540.7, four pages. [cited by applicant]
International Preliminary Report and Written Opinion mailed Apr. 25, 2024, for PCT Application No. PCT/US2022/078063, seven pages. [cited by applicant]
International Preliminary Report on Patentability and Written Opinion mailed Apr. 25, 2024, for PCT Application No. PCT/US2022/078073, seven pages. [cited by applicant]
International Preliminary Report on Patentability and Written Opinion mailed May 2, 2024, for PCT Application No. PCT/US2022/078298, twelve pages. [cited by applicant]
International Search Report and Written Opinion mailed Jan. 11, 2023, for PCT Application No. PCT/US2022/078298, seventeen pages. [cited by applicant]
International Search Report and Written Opinion mailed Jan. 17, 2023, for PCT Application No. PCT/US2022/078073, thirteen pages. [cited by applicant]
International Search Report and Written Opinion mailed Jan. 25, 2023, for PCT Application No. PCT/US2022/078063, nineteen pages. [cited by applicant]
Japanese Office Action mailed May 2, 2024, for JP Application No. 2021-562002, with English translation, sixteen pages. [cited by applicant]
Notice of Allowance mailed Jun. 5, 2024, for U.S. Appl. No. 18/459,342, filed Aug. 31, 2023, eight pages. [cited by applicant]
Japanese Final Office Action mailed May 31, 2024, for JP Application No. 2021-551538, with English translation, eighteen pages. [cited by applicant]
Non-Final Office Action mailed Aug. 26, 2024, for U.S. Appl. No. 18/506,866, filed Nov. 10, 2023, twelve pages. [cited by applicant]
Non-Final Office Action mailed Aug. 26, 2024, for U.S. Appl. No. 18/418,131, filed Jan. 19, 2024, ten pages. [cited by applicant]
European Notice of Allowance dated Jul. 15, 2024, for EP Application No. 20791183.5 nine pages. [cited by applicant]
Japanese Notice of Allowance mailed Aug. 21, 2024, for JP Application No. 2021-562002, with English translation, 6 pages. [cited by applicant]
Chinese Notice of Allowance dated Feb. 18, 2025, for CN Application No. 202080044362.4, with English translation, 7 pages. [cited by applicant]
European Search Report dated Jan. 20, 2025, for EP Application No. 24221982.2 nine pages. [cited by applicant]
Non-Final Office Action mailed Mar. 4, 2025, for U.S. Appl. No. 18/029,355, filed Mar. 29, 2023, ten pages. [cited by applicant]
Chinese Notice of Allowance dated Nov. 7, 2024, for CN Application No. 201980050714.4, with English translation, 6 pages. [cited by applicant]
Hatashima, Takashi et al. “A study on a subband acoustic echo chaceller using a single microphone, Institute of Electronics, Information and Communication Engineers technical research report”, Feb. 1995, vol. 94, No. 52… [cited by applicant]
Japanese Notice of Allowance mailed Oct. 28, 2024, for JP Application No. 2021-551538 with English translation, 6 pages. [cited by applicant]
Japanese Office Action mailed Nov. 7, 2024, for JP Application No. 2023-142856, with English translation, 12 pages. [cited by applicant]
Notice of Allowance mailed Dec. 11, 2024, for U.S. Appl. No. 18/418,131, filed Jan. 19, 2024, seven pages. [cited by applicant]
Notice of Allowance mailed Dec. 19, 2024, for U.S. Appl. No. 18/506,866, filed Nov. 10, 2023, five pages. [cited by applicant]
Non-Final Office Action mailed Jan. 24, 2025, for U.S. Appl. No. 18/510,376, filed Nov. 15, 2023, seventeen pages. [cited by applicant]
Kaushik, Mayank, et al. “Three Dimensional Microphone and Source Position Estimation Using TDOA and TOF Measurements.” 2011 IEEE International Conference on Signal Processing, Communications and Computing (ICSPCC). IEEE… [cited by applicant]
Final Office Action mailed Sep. 7, 2023, for U.S. Appl. No. 17/214,446, filed Mar. 26, 2021, nineteen pages. [cited by applicant]
Notice of Allowance mailed Oct. 12, 2023, for U.S. Appl. No. 18/148,221, filed Dec. 29, 2022, five pages. [cited by applicant]
Notice of Allowance mailed Oct. 17, 2023, for U.S. Appl. No. 17/254,832, filed Dec. 21, 2020, sixteen pages. [cited by applicant]
Cited By (1)
US 12,417,766