IP Library Granted Patent US 11,657,801
Granted Patent B2
US 11,657,801 · App. 16/287,666 · Granted May 23, 2023

Voice command detection and prediction

Inventors: Rui Min (Vienna, VA); Hongcheng Wang (Falls Church, VA)
Assignee: Comcast Cable Communications, LLC
G10L15/16G06N20/00G10L15/22G10L15/30G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,657,801
App. No.
16/287,666
Granted
May 23, 2023
Kind
B2
Abstract

Methods, systems, and apparatuses for predicting an end of a command in a voice recognition input are described herein. The system may receive data comprising a voice input. The system may receive a signal comprising a voice input. The system may detect, in the voice input, data that is associated with a first portion of a command. The system may predict, based on the first portion and while the voice input is being received, a second portion of the command. The prediction may be generated by a machine learning algorithm that is trained based at least in part on historical data comprising user input data. The system may cause execution of the command, based on the first portion and the predicted second portion, prior to an end of the voice input.

Claims (44)

1. A method comprising:

receiving a signal comprising a voice input;

detecting, in the voice input, data that is associated with a first portion of a command;

predicting, based on the first portion and while the voice input is being received, and using a machine learning model trained based on historical user input data, a second portion of the command; and

causing execution of the command, based on the first portion and the predicted second portion, prior to an end of the voice input.

2. The method of claim 1 , further comprising:

storing the data indicative of the voice input; and

determining, based on the stored data, that the predicted second portion is incorrect, causing execution of a second command that is associated with the voice input.

3. The method of claim 1 , further comprising:

storing the predicted second portion; and

using the stored predicted second portion to train the machine learning model.

4. The method of claim 1 , wherein the predicting second portion is based in part on common input commands.

5. The method of claim 1 , wherein the predicting second portion is based on metadata, time information, location information, or demographic information.

6. The method of claim 1 , further comprising:

generating a connection between a plurality of predicted commands and a plurality of navigational commands to cause execution of a plurality of commands.

7. The method of claim 1 , wherein the predicted second portion is further based on differences between a format of the voice input and formats of previous inputs and based on changes in acoustic features.

8. The method of claim 1 , wherein the machine learning model uses a long short term memory (LSTM) network to predict the second portion.

9. The method of claim 8 , further comprising:

using available historical commands from a plurality of users to train the LSTM network.

10. The method of claim 8 , wherein the LSTM network is configured to learn relationships between actions and entities with temporal dependencies.

11. The method of claim 8 , wherein the LSTM network is configured to predict whether a voice command is complete and additional words that a user may speak.

12. The method of claim 11 , wherein the LSTM network is configured to determine whether an end of the command has been reached based on detecting a silence cue.

13. A system comprising:

a receiver configured to receive data indicative of a voice input;

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the device to:

receive a signal comprising a voice input;

detect, in the voice input, data that is associated with a first portion of a command;

predict, based on the first portion and while the voice input is being received, and using a machine learning model trained based on historical user input data, a second portion of the command; and

cause execution of the command, based on the first portion and the predicted second portion, prior to an end of the voice input.

14. A device, comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the device to:

receive a signal comprising a voice input;

detect, in the voice input, data that is associated with a first portion of a command;

predict, based on the first portion and while the voice input is being received, and using a machine learning model trained based on historical user input data, a second portion of the command; and

cause execution of the command, based on the first portion and the predicted second portion, prior to an end of the voice input.

15. The device of claim 14 , wherein the instructions, when executed by the one or more processors, further cause the device to:

store the data indicative of the voice input; and

determining, based on the stored data, that the predicted second portion is incorrect, cause execution of a second command that is associated with the voice input.

16. The device of claim 14 , wherein the machine learning model uses a long short term memory (LSTM) network to predict the second portion.

17. The device of claim 16 , further comprising:

the processor further configured to use available historical commands from a plurality of users to train the LSTM model.

18. The device of claim 16 , wherein the LSTM model is configured to learn relationships between actions and entities with temporal dependencies.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 28, 2019
From: MIN, RUI; WANG, HONGCHENG
To: COMCAST CABLE COMMUNICATIONS, LLC
Reel/Frame 048464/0122 →
Continuity (1)
Related Publication 20200273448A1 · Aug 27, 2020