IP Library Patent Application 17620005
Patent Application
App. No. 17/620,005

SYSTEMS AND METHODS FOR INTERPRETING A VOICE QUERY

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/620,005
Abstract

Systems and methods are described herein for enabling, on a local device, a voice control system that limits the amount of data needed to be transmitted to a remote server. A data structure to support a local speech-to-text model is built at the local device from stored transcriptions of previous queries and known commands, and associates actions with each transcription. The transcription of the particular query is generated using the local speech-to- text model and is used to identify an associated action to perform.

Claims (42)

1 . A method for interpreting a voice input received at a local device, the method comprising:

receiving the voice input via a voice-user interface at a local device;

generating a transcription of the voice input using a local speech processing model;

comparing the transcription to a data structure stored at the local device, wherein the data structure comprises a plurality of entries, and wherein each entry comprises an audio clip of a previously received voice input and a corresponding transcription;

determining whether the data structure comprises an entry that matches the voice input; and

in response to determining that the data structure comprises an entry that matches the voice input, identifying an action associated with the matching entry.

2 . The method of claim 1 , further comprising performing, at the local device, the identified action.

3 . The method of claim 1 , wherein each entry comprises an audio clip mapped to a phoneme, wherein the phoneme is mapped to a set of graphemes, wherein the set of graphemes is mapped to a sequence of graphemes, and wherein the sequence of graphemes is mapped to a transcription.

4 . The method of claim 1 , wherein comparing the voice input to the data structure stored at the local device comprises comparing the voice input to an audio clip associated with each entry in the data structure.

5 . The method of claim 1 , wherein comparing the voice input to the data structure stored at the local device comprises comparing the voice input to a plurality of graphemes associated with each entry in the data structure.

6 . The method of claim 1 , further comprising storing an audio clip of the voice input as a second clip associated with the matching voice input.

7 . The method of claim 1 , wherein the voice input corresponds to at least one of playing, pausing, skipping, exiting, tuning, fast-forwarding, rewinding, recording, increasing volume, decreasing volume, powering on, and powering off.

8 . The method of claim 1 , wherein the voice input corresponds to at least one of a title, a name, or an identifier.

9 . The method of claim 1 , further comprising:

determining that the local speech processing model cannot recognize the voice input;

transmitting, to a remote server, a request for transcription of the voice input into other data;

receiving the transcription of the voice input from the remote server; and

storing, in the data structure at the local device, an entry that associates an audio clip of the voice input with the corresponding transcription for use in recognition of a query subsequently received via the voice-user interface of the local device.

10 . The method of claim 1 , wherein the local device receives the transcription of the audio clip of the previously received voice input from a remote server prior to receiving the voice input via the voice-user interface at the local device.

11 . (canceled)

12 . A system for interpreting a voice input received at a local device, the system comprising:

the local device;

a control circuitry configured to:

receive the voice input via a voice-user interface at the local device;

generate a transcription of the voice input using a local speech processing model;

compare the transcription to a data structure stored at the local device, wherein the data structure comprises a plurality of entries, and wherein each entry comprises an audio clip of a previously received voice input and a corresponding transcription;

determine whether the data structure comprises an entry that matches the voice input; and

in response to determining that the data structure comprises an entry that matches the voice input, identifying an action associated with the matching entry.

13 . The system of claim 12 , wherein the control circuitry is configured to:

determine that the local speech processing model cannot recognize the voice input;

transmit, to a remote server, a request for transcription of the voice input into other data;

receive the transcription of the voice input from the remote server; and

store, in the data structure at the local device, an entry that associates an audio clip of the voice input with the corresponding transcription for use in recognition of a voice input subsequently received via the voice-user interface of the local device.

14 . The system of claim 12 , wherein the local device receives the transcription of the audio clip of the previously received voice input from a remote server over a communication network prior to receiving the voice input via the voice-user interface at the local device.

15 . (canceled)

16 . The system of claim 12 , wherein the control circuitry is further configured to perform, at the local device, the identified action.

17 . The system of claim 12 , wherein each entry comprises an audio clip mapped to a phoneme, wherein the phoneme is mapped to a set of graphemes, wherein the set of graphemes is mapped to a sequence of graphemes, and wherein the sequence of graphemes is mapped to a transcription.

18 . The system of claim 12 , wherein the control circuitry is configured to compare the voice input to the data structure stored at the local device by comparing the voice input to an audio clip associated with each entry in the data structure.

19 . The system of claim 12 , wherein the control circuitry is configured to compare the voice input to the data structure stored at the local device by comparing the voice input to a plurality of graphemes associated with each entry in the data structure.

20 . The system of claim 12 , wherein the control circuitry is further configured to store an audio clip of the voice input as a second clip associated with the matching voice input.

21 . The system of claim 12 , wherein the voice input corresponds to at least one of playing, pausing, skipping, exiting, tuning, fast-forwarding, rewinding, recording, increasing volume, decreasing volume, powering on, and powering off.

22 . The system of claim 12 , wherein the voice input corresponds to at least one of a title, a name, or an identifier.

Assignments (3)
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0207 →
SECURITY INTEREST Recorded May 3, 2023
From: ADEIA GUIDES INC.; ADEIA IMAGING LLC; ADEIA MEDIA HOLDINGS LLC; ADEIA MEDIA SOLUTIONS INC.; ADEIA SEMICONDUCTOR ADVANCED TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR BONDING TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR INC.; ADEIA SEMICONDUCTOR SOLUTIONS LLC; ADEIA SEMICONDUCTOR TECHNOLOGIES LLC; ADEIA SOLUTIONS LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 063529/0272 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2021
From: ROBERT JOSE, JEFFRY COPPS; GOYAL, AASHISH
To: ROVI GUIDES, INC.
Reel/Frame 058524/0690 →