IP Library Patent Application 19197282
Patent Application
App. No. 19/197,282

IDENTIFYING INPUT FOR SPEECH RECOGNITION ENGINE

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/197,282
Abstract

A method of presenting a signal to a speech recognition engine is disclosed. According to an example of the method, an audio signal is received from a user. A portion of the audio signal is identified, the portion having a first time and a second time. A pause in the portion of the audio signal, the pause comprising the second time, is identified. It is determined whether the pause indicates the completion of an utterance of the audio signal. In accordance with a determination that the pause indicates the completion of the utterance, the portion of the audio signal is presented as input to the speech recognition engine. In accordance with a determination that the pause does not indicate the completion of the utterance, the portion of the audio signal is not presented as input to the speech recognition engine.

Claims (74)

1 . A method comprising:

receiving, via a microphone, an audio signal, wherein the audio signal comprises voice activity;

receiving, via one or more sensors of a head-wearable device, non-verbal sensor data corresponding to a user of the head-wearable device, wherein:

the one or more sensors comprise a camera, and

the receiving the non-verbal sensor data comprises receiving information determined based on one or more of Simultaneous Localization and Mapping (SLAM) and visual odometry performed via visual data from the camera;

classifying audio data corresponding to the audio signal;

determining whether the audio signal comprises a pause in the voice activity;

responsive to determining that the audio signal comprises the pause in the voice activity, determining, based on the classifying of the audio data and based further on the non-verbal sensor data, whether the pause in the voice activity corresponds to an end point of the voice activity; and

responsive to determining that the pause in the voice activity corresponds to the end point of the voice activity, presenting a response to the user based on the voice activity,

wherein:

the determining whether the audio signal comprises the pause in the voice activity comprises:

determining, based on the information determined based on the one or more of SLAM and visual odometry, head poses of the user, and

determining whether the head poses comprise a head pose change, wherein the audio signal comprises the pause in accordance with a determination that the head poses comprise the head pose change, and

the determining whether the pause in the voice activity corresponds to the end point of the voice activity comprises:

determining a probability of interest based on the classifying of the audio data and based further on the non-verbal sensor data, the probability determined based on relative distances between the non-verbal sensor data and its neighbors in an N-dimensional space, and

in accordance with a determination that the probability of interest exceeds a threshold, determining that the pause in the voice activity corresponds to the end point of the voice activity.

2 . The method of claim 1 , further comprising:

in accordance with a determination that the probability of interest does not exceed the threshold:

determining that the pause in the voice activity does not correspond to the end point of the voice activity, and

forgoing presenting the response to the user based on the voice activity.

3 . The method of claim 1 , wherein the determining whether the audio signal comprises the pause in the voice activity further comprises determining whether an amplitude of the audio signal falls below a second threshold.

4 . The method of claim 1 further comprising:

in accordance with a determination that the probability of interest does not exceed the threshold:

determining that the pause in the voice activity does not correspond to the end point of the voice activity, and

determining whether the audio signal comprises a second pause corresponding to the end point of the voice activity.

5 . The method of claim 1 , wherein the determining whether the audio signal comprises the pause in the voice activity further comprises determining whether the audio signal comprises one or more verbal cues corresponding to the pause in the voice activity.

6 . The method of claim 5 , wherein the one or more verbal cues comprise a characteristic of the user's prosody.

7 . The method of claim 5 , wherein the one or more verbal cues comprise a terminating phrase.

8 . The method of claim 5 , wherein the one or more verbal cues further correspond to the end point of the voice activity.

9 . The method of claim 1 , wherein the non-verbal sensor data further comprises data indicative of the user's gaze.

10 . The method of claim 1 , wherein the non-verbal sensor data further comprises data indicative of the user's facial expression.

11 . The method of claim 1 , wherein the non-verbal sensor data further comprises data indicative of the user's heart rate.

12 . The method of claim 1 , wherein the determining whether the pause in the voice activity corresponds to the end point of the voice activity comprises identifying one or more interstitial sounds.

13 . The method of claim 1 , wherein the one or more sensors comprise the microphone.

14 . The method of claim 1 , wherein the determining whether the audio signal comprises the pause in the voice activity further comprises determining that a frequency component of the audio signal is indicative of the pause in the voice activity.

15 . The method of claim 1 , wherein the determining whether the audio signal comprises the pause in the voice activity further comprises determining whether an audio segment of the audio signal comprises the pause in the voice activity.

16 . The method of claim 1 , wherein the determining whether the pause in the voice activity corresponds to the end point of the voice activity further comprises determining whether an audio segment of the audio signal comprises the end point of the voice activity.

17 . A system, comprising:

a microphone of a head-wearable device;

one or more sensors of the head-wearable device; and

one or more processors configured to perform a method comprising:

receiving, via the microphone, an audio signal, wherein the audio signal comprises voice activity;

receiving, via the one or more sensors, non-verbal sensor data corresponding to a user of the head-wearable device, wherein:

the one or more sensors comprise a camera, and

the receiving the non-verbal sensor data comprises receiving information determined based on one or more of Simultaneous Localization and Mapping (SLAM) and visual odometry performed via visual data from the camera;

classifying audio data corresponding to the audio signal;

determining whether the audio signal comprises a pause in the voice activity;

responsive to determining that the audio signal comprises the pause in the voice activity, determining, based on the classifying of the audio data and based further on the non-verbal sensor data, whether the pause in the voice activity corresponds to an end point of the voice activity; and

responsive to determining that the pause in the voice activity corresponds to the end point of the voice activity, presenting a response to the user based on the voice activity,

wherein:

the determining whether the audio signal comprises the pause in the voice activity comprises:

 determining, based on the information determined based on the one or more of SLAM and visual odometry, head poses of the user, and

 determining whether the head poses comprise a head pose change, wherein the audio signal comprises the pause in accordance with a determination that the head poses comprise the head pose change, and

the determining whether the pause in the voice activity corresponds to the end point of the voice activity comprises:

 determining a probability of interest based on the classifying of the audio data and based further on the non-verbal sensor data, the probability determined based on relative distances between the non-verbal sensor data and its neighbors in an N-dimensional space, and

 in accordance with a determination that the probability of interest exceeds a threshold, determining that the pause in the voice activity corresponds to the end point of the voice activity.

18 . The system of claim 17 , further comprising a transmissive display of the head-wearable device, wherein response to the user is presented via the transmissive display.

19 . The system of claim 17 , wherein the non-verbal sensor data further comprises data indicative of the user's facial expression.

20 . A non-transitory computer-readable medium storing instructions, which, when executed by one or more processors, cause the one or more processors to perform a method comprising:

receiving, via a microphone, an audio signal, wherein the audio signal comprises voice activity;

receiving, via one or more sensors of a head-wearable device, non-verbal sensor data corresponding to a user of the head-wearable device, wherein:

the one or more sensors comprise a camera, and

the receiving the non-verbal sensor data comprises receiving information determined based on one or more of Simultaneous Localization and Mapping (SLAM) and visual odometry performed via visual data from the camera;

classifying audio data corresponding to the audio signal;

determining whether the audio signal comprises a pause in the voice activity;

responsive to determining that the audio signal comprises the pause in the voice activity, determining, based on the classifying of the audio data and based further on the non-verbal sensor data, whether the pause in the voice activity corresponds to an end point of the voice activity; and

responsive to determining that the pause in the voice activity corresponds to the end point of the voice activity, presenting a response to the user based on the voice activity,

wherein:

the determining whether the audio signal comprises the pause in the voice activity comprises:

determining, based on the information determined based on the one or more of SLAM and visual odometry, head poses of the user, and

determining whether the head poses comprise a head pose change, wherein the audio signal comprises the pause in accordance with a determination that the head poses comprise the head pose change, and

the determining whether the pause in the voice activity corresponds to the end point of the voice activity comprises:

determining a probability of interest based on the classifying of the audio data and based further on the non-verbal sensor data, the probability determined based on relative distances between the non-verbal sensor data and its neighbors in an N-dimensional space, and

in accordance with a determination that the probability of interest exceeds a threshold, determining that the pause in the voice activity corresponds to the end point of the voice activity.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2025
From: SHEEDER, ANTHONY ROBERT; ARORA, TUSHAR
To: MAGIC LEAP, INC.
Reel/Frame 072940/0445 →
SECURITY INTEREST Recorded Oct 31, 2025
From: MAGIC LEAP, INC.; MENTOR ACQUISITION ONE, LLC; MOLECULAR IMPRINTS, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 073439/0168 →