IP Library › Granted Patent US 11,657,814
Granted Patent B2
US 11,657,814 · App. 17/066,433 · Granted May 23, 2023

Techniques for dynamic auditory phrase completion

Inventors: Stefan Marti (Oakland, CA); Joseph Verbeke (San Francisco, CA); Evgeny Burmistrov (Saratoga, CA); Priya Seshadri (San Francisco, CA)
Assignee: Harman International Industries, Incorporated
G10L15/22G10L15/02G10L15/183G10L15/1822
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,657,814
App. No.
17/066,433
Granted
May 23, 2023
Kind
B2
Abstract

Embodiments of the present disclosure set forth a computer-implemented method comprising detecting an initial phrase portion included in a first auditory signal generated by a user, identifying, based on the initial phrase portion, a supplemental phrase portion that complements the initial phrase portion to form a complete phrase, and providing a command signal that drives an output device to generate an audio output corresponding to the supplemental phrase portion.

Claims (57)

1. A computer-implemented method comprising:

utilizing an imaging device to capture image data of a user, or an input device to receive a physical input of the user, or a biometric sensor to obtain a biometric value of the user, wherein the image data includes a physical gesture of the user detected by the imaging device, and wherein the physical input includes a triggering input detected by the input device;

utilizing an audio capturing device to acquire a first auditory signal of the user and a processor to detect an initial phrase portion in the first auditory signal; and

using the processor to detect a request from the user to complete a phrase, wherein of the request includes (i) the physical gesture of the user, (ii) the triggering input of the user, or (iii) comparing the biometric value of the user to a threshold, wherein, in response to the request from the user to complete the phrase, the processor is configured to:

retrieve from a database a complete phrase that includes the initial phrase portion;

retrieve a supplemental phrase portion from the complete phrase that complements the initial phrase portion, the initial phrase portion and the supplemental phrase portion forming the complete phrase; and

utilize an audio output device to output an audio reproduction of the supplemental phrase portion.

2. The computer-implemented method of claim 1 , wherein the initial phrase portion is detected prior to detecting the request from the user to complete the phrase.

3. The computer-implemented method of claim 1 , further comprising:

prior to receiving the first auditory signal, recording a second auditory signal that includes the complete phrase; and

storing, in the database:

a first audio clip that corresponds to the initial phrase portion, and

a second audio clip that corresponds to the supplemental phrase portion.

4. The computer-implemented method of claim 3 , further comprising receiving a user input associated with the second auditory signal, wherein receiving the user input triggers storage of the first audio clip and the second audio clip.

5. The computer-implemented method of claim 3 , further comprising:

retrieving the second audio clip,

wherein the audio output device generates the audio reproduction of the second audio clip.

6. The computer-implemented method of claim 1 , further comprising:

prior to receiving the first auditory signal, receiving an input that includes the complete phrase;

parsing the complete phrase into the initial phrase portion and the supplemental phrase portion; and

storing, in a phrase completion table, an entry for the complete phrase including a mapping between the initial phrase portion and the supplemental phrase portion.

7. A system that completes a complete phrase that is partially spoken by a user, the system comprising:

at least one microphone that acquires an auditory signal generated by the user;

an imaging device configured to capture image data of the user, wherein the image data includes a physical gesture of the user;

an input device configured to receive a physical input of the user, wherein the physical input includes a triggering input;

a biometric sensor configured to receive a biometric value of the user;

an audio output device configured to output audio;

a memory storing instructions; and

one or more processors, that when executing the instructions, is configured to:

utilize the imaging device to capture the image data, or utilize the input device to receive the physical input, or utilize the biometric sensor to obtain the biometric value;

utilize the at least one microphone to acquire the auditory signal;

detect an initial phrase portion in the auditory signal;

detect a request from the user to complete a phrase, wherein the request includes (i) the physical gesture of the user, (ii) the triggering input of the user, or (iii) comparing the biometric value of the user to a threshold; and

in response to the request from the user to complete the phrase:

retrieve, from a database, a complete phrase that includes the initial phrase portion;

retrieve a supplemental phrase portion from the complete phrase that complements the initial phrase portion, the initial phrase portion and the supplemental phrase portion forming the complete phrase; and

utilize the audio output device to output an audio reproduction of the supplemental phrase portion.

8. The system of claim 7 , further comprising at least one of a visual sensor or a facial electromyography sensor that acquires biometric data associated with the user, wherein the biometric value is included in biometric data that includes at least one of eye gaze direction, muscle contraction, or facial movement.

9. The system of claim 7 , wherein:

the audio output device comprises a voice agent that synthesizes an audio output signal of a voice speaking the supplemental phrase portion, and

to utilize the audio output device to output the audio reproduction of the supplemental phrase portion, the one or more processors transmit a command signal to the voice agent to synthesize the audio output signal, the voice agent driving an output device to generate the audio reproduction.

10. The system of claim 7 , wherein the audio output device includes at least one speaker that generates a steerable beam that provides the audio reproduction to a target listener.

11. The system of claim 7 , wherein:

the memory further stores a phrase completion table; and

the phrase completion table includes at least an entry for the complete phrase including a mapping between the initial phrase portion and the supplemental phrase portion.

12. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of:

utilizing an imaging device to capture image data of a user, or an input device to receive a physical input of the user, or a biometric sensor to obtain a biometric value of the user, wherein the image data includes a physical gesture of the user detected by the imaging device, and wherein the physical input includes a triggering input detected by the input device;

utilizing an audio capturing device to acquire an auditory signal of the user and the processor to detect an initial phrase portion in the auditory signal;

detecting a request from the user to complete a phrase, wherein of the request includes (i) the physical gesture of the user, (ii) the triggering input of the user, or (iii) comparing the biometric value of the user to a threshold; and

in response to the request from the user to complete the phrase:

retrieving from a database, a complete phrase that includes the initial phrase portion;

retrieving a supplemental phrase portion from the complete phrase that complements the initial phrase portion, the initial phrase portion and the supplemental phrase portion forming the complete phrase; and

utilizing an audio output device to output an audio reproduction of the supplemental phrase portion.

13. The one or more non-transitory computer-readable media of claim 12 , further comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform the steps of:

prior to receiving the auditory signal, receiving an input that includes the complete phrase;

parsing the complete phrase into the initial phrase portion and the supplemental phrase portion; and

storing, in a phrase completion table, an entry for the complete phrase including a mapping between the initial phrase portion and the supplemental phrase portion.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 3, 2020
From: MARTI, STEFAN; VERBEKE, JOSEPH; BURMISTROV, EVGENY; SESHADRI, PRIYA
To: HARMAN INTERNATIONAL INDUSTRIES, INCORPORATED
Reel/Frame 054536/0016 →
Continuity (1)
Related Publication 20220115010A1 · Apr 14, 2022
Cited By (1)
US 12,347,135