IP Library Granted Patent US 11,587,563
Granted Patent B2
US 11,587,563 · App. 16/805,337 · Granted Feb 21, 2023

Determining input for speech processing engine

Inventors: Anthony Robert Sheeder (Fort Lauderdale, FL); Colby Nelson Leider (Coral Gables, FL)
Assignee: Magic Leap, Inc.
G10L15/22G06F3/013G10L15/14G10L15/25G10L15/30G10L2015/223G10L2015/227
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,587,563
App. No.
16/805,337
Granted
Feb 21, 2023
Kind
B2
Abstract

A method of presenting a signal to a speech processing engine is disclosed. According to an example of the method, an audio signal is received via a microphone. A portion of the audio signal is identified, and a probability is determined that the portion comprises speech directed by a user of the speech processing engine as input to the speech processing engine. In accordance with a determination that the probability exceeds a threshold, the portion of the audio signal is presented as input to the speech processing engine. In accordance with a determination that the probability does not exceed the threshold, the portion of the audio signal is not presented as input to the speech processing engine.

Claims (41)

1. A method of presenting a signal to a speech processing engine, the method comprising:

receiving, via a first microphone, a first audio signal;

identifying a first portion of the first audio signal;

receiving first sensor data from one or more sensors associated with a wearable head unit configured to be worn by a user, wherein the one or more sensors do not include the first microphone;

determining, for the first portion of the first audio signal, a first probability that the first portion comprises speech directed by the user as input to the speech processing engine, wherein the first probability is determined based on the first sensor data and further based on a speech characteristic of the first portion of the first audio signal;

identifying a second portion of the first audio signal, the second portion subsequent to the first portion in the first audio signal;

receiving second sensor data from the one or more sensors;

determining, for the second portion of the first audio signal, based on the second sensor data and based further on a speech characteristic of the second portion of the first audio signal, a second probability that the second portion comprises speech directed by the user as input to the speech processing engine;

identifying a third portion of the first audio signal, the third portion subsequent to the second portion in the first audio signal;

receiving third sensor data from the one or more sensors;

determining, for the third portion of the first audio signal, based on the third sensor data and based further on a speech characteristic of the third portion of the first audio signal, a third probability that the third portion comprises speech directed by the user as input to the speech processing engine; and

in accordance with a determination that the first probability exceeds a first threshold, further in accordance with a determination that the second probability does not exceed the first threshold, and further in accordance with a determination that the third probability exceeds the first threshold:

presenting the first portion of the first audio signal and the third portion of the first audio signal as a first input to the speech processing engine, wherein the first input does not include the second portion of the first audio signal.

2. The method of claim 1 , wherein the first probability is determined based on a comparison of the first portion of the first audio signal to a plurality of audio signals in a database, each audio signal of the plurality of audio signals associated with a probability that its respective audio signal comprises speech directed as input to the speech processing engine.

3. The method of claim 1 , wherein the first probability is determined further based on a comparison of the first sensor data to a plurality of sensor data in a database, each sensor data of the plurality of sensor data in the database associated with an audio signal and further associated with a probability that its respective audio signal comprises speech directed as input to the speech processing engine.

4. The method of claim 1 , wherein each of the first sensor data, the second sensor data, and the third sensor data is indicative of one or more of a position, an orientation, an eye movement, an eye gaze target, or a vital sign of the user.

5. The method of claim 1 , further comprising determining based on the first sensor data whether the first portion of the first audio signal corresponds to the user and in accordance with a determination that the audio signal does not correspond to the user, discarding the first portion of the first audio signal.

6. The method of claim 1 , further comprising determining a query based on the first portion of the first audio signal, wherein the first portion of the first audio signal forms a portion of the query.

7. A system for providing input to a speech processing engine, the system including:

a microphone;

one or more sensors that do not include the microphone; and

circuitry configured to perform:

receiving, via the microphone, a first audio signal;

identifying a first portion of the first audio signal;

receiving, via the one or more sensors, a first sensor data;

determining, for the first portion of the first audio signal, a first probability that the first portion comprises speech directed by a user as input to the speech processing engine, wherein the first probability is determined based on the first sensor data and further based on a speech characteristic of the first portion of the first audio signal;

identifying a second portion of the first audio signal, the second portion subsequent to the first portion in the first audio signal;

receiving second sensor data from the one or more sensors;

determining, for the second portion of the first audio signal, based on the second sensor data and based further on a speech characteristic of the second portion of the first audio signal, a second probability that the second portion comprises speech directed by the user as input to the speech processing engine;

identifying a third portion of the first audio signal, the third portion subsequent to the second portion in the first audio signal;

receiving third sensor data from the one or more sensors;

determining, for the third portion of the first audio signal, based on the third sensor data and based further on a speech characteristic of the third portion of the first audio signal, a third probability that the third portion comprises speech directed by the user as input to the speech processing engine; and

in accordance with a determination that the first probability exceeds a first threshold, further in accordance with a determination that the second probability does not exceed the first threshold, and further in accordance with a determination that the third probability exceeds the first threshold:

presenting the first portion of the first audio signal and the third portion of the first audio signal as a first input to the speech processing engine, wherein the first input does not include the second portion of the first audio signal.

8. The system of claim 7 , wherein the first probability is determined further based on a comparison of the first portion of the first audio signal to a plurality of audio signals in a database, each audio signal of the plurality of audio signals associated with a probability that its respective audio signal comprises speech directed as input to the speech processing engine.

9. The system of claim 7 ,

wherein the first probability is determined further based on a comparison of the first sensor data to a plurality of sensor data in a database, each sensor data of the plurality of sensor data in the database associated with an audio signal and further associated with a probability that its respective audio signal comprises speech directed as input to the speech processing engine.

10. The system of claim 7 , wherein each of the first sensor data, the second sensor data, and the third sensor data is indicative of one or more of a position of the user, an orientation of the user, an eye movement of the user, an eye gaze target of the user, and a vital sign of the user.

11. The system of claim 7 , wherein the system includes a wearable head unit including the microphone, the one or more sensors, and the circuitry.

12. The system of claim 7 , wherein the system includes a vehicle including the microphone and the circuitry.

13. The system of claim 7 , wherein the system includes an electronic voice assistant including the microphone and the circuitry.

Assignments (4)
SECURITY INTEREST Recorded Oct 28, 2025
From: MAGIC LEAP, INC.; MENTOR ACQUISITION ONE, LLC; MOLECULAR IMPRINTS, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 073388/0027 →
SECURITY INTEREST Recorded Oct 15, 2025
From: MAGIC LEAP, INC.; MENTOR ACQUISITION ONE, LLC; MOLECULAR IMPRINTS, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 073109/0476 →
SECURITY INTEREST Recorded May 24, 2022
From: MOLECULAR IMPRINTS, INC.; MENTOR ACQUISITION ONE, LLC; MAGIC LEAP, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 060338/0665 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2020
From: SHEEDER, ANTHONY ROBERT; LEIDER, COLBY NELSON
To: MAGIC LEAP, INC.
Reel/Frame 052954/0165 →
Cited By (3)
US 12,417,766 US 12,688,845 US 12,696,045