IP Library Granted Patent US 12,706,087
Granted Patent B2
US 12,706,087 · App. 18/296,181 · Granted Aug 11, 2026

Methods and systems for voice control

Inventors: Scott Kurtz (Philadelphia, PA); Philip Stick (Philadelphia, PA); Gary Skrabutenas (Philadelphia, PA); Christian Buchter (Philadelphia, PA)
Assignee: Comcast Cable Communications, LLC
G10L15/08G10L15/05G10L15/22G10L15/30G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,706,087
App. No.
18/296,181
Granted
Aug 11, 2026
Kind
B2
Abstract

One or more portions of audio input may be detected. Timing data associated with the one or more portions of audio may be determined. Audio processing may be carried out based on the timing data.

Claims (35)

1 . A method comprising:

receiving, a first portion of audio data and a second portion of audio data;

determining, based on the first portion of audio data and the second portion of audio data, a first wake word confidence score;

determining, based on the second portion of audio data, a second wake word confidence score;

determining, based on the second wake word confidence score being less than the first wake word confidence score, that the first portion of audio data comprises a beginning of a wake word; and

based on the first portion of audio data comprising the beginning of the wake word, sending the first portion of audio data for processing and excluding the second portion of audio data.

2 . The method of claim 1 , wherein the first portion of audio data and the second portion of audio data comprise one or more user utterances received by a voice enabled device.

3 . The method of claim 1 , wherein the first wake word confidence score and the second wake word confidence score are configured to indicate the first portion of audio data or the second portion of audio data comprises a portion of the wake word.

4 . The method of claim 1 , further comprising opening a communication channel with a voice service based on the second wake word confidence score being less than the first wake word confidence score.

5 . The method of claim 1 , wherein sending the first portion of audio data and the second portion of audio data for processing comprises sending the first portion of audio data and the second portion of audio data to a cloud based voice service.

6 . The method of claim 1 , further comprising determining, based on the second wake word confidence score being less than the first wake word confidence score, that the second portion of audio data comprises an end of the wake word.

7 . The method of claim 1 , further comprising determining a wake word duration.

8 . A method comprising:

receiving a plurality of portions of audio data comprising a wake word;

determining, based on a first portion of audio data of the plurality of portions of audio data and a second portion of audio data of the plurality of portions of audio data, a first wake word confidence score;

determining, based on the second portion of audio data and a third portion of audio data of the plurality of portions of audio data, a second wake word confidence score;

determining, based on the second wake word confidence score being higher than the first wake word confidence score, that a time stamp associated with the second portion of audio data corresponds to a beginning of the wake word; and

based on determining the time stamp associated with the second portion of audio data corresponds to the beginning of the wake word, opening a communication channel with a service.

9 . The method of claim 8 , wherein the plurality of portions of audio data comprise one or more voice inputs received from a voice enabled device.

10 . The method of claim 8 , wherein the first wake word confidence score indicates a likelihood the first portion of audio data or the second portion of audio data comprise a wake word.

11 . The method of claim 8 , wherein determining that the time stamp associated with the second portion of audio data corresponds to the beginning of the wake word comprises determining an audio analysis window.

12 . The method of claim 8 , further comprising processing the plurality of portions of audio data, wherein processing the plurality of portions of audio data comprises one or more of: responding to a query, sending a message, or executing a command.

13 . The method of claim 8 , further comprising determining, based on the time stamp, a wake word duration.

14 . The method of claim 8 , further comprising receiving, from a voice service, one or more responses.

15 . A method comprising:

receiving an audio input comprising a plurality of audio frames;

determining, for each audio frame of the plurality of audio frames, a confidence score that an audio frame of the plurality of audio frames comprises a portion of a wake word;

determining an increase in the confidence score between a first two audio frames of the plurality of audio frames and a decrease in the confidence score between a second two audio frames of the plurality of audio frames;

determining, based on the increase in the confidence score, a first boundary of the wake word, and, based on the decrease in the confidence score, a second boundary of the wake word; and

sending only the audio input between the first boundary and the second boundary to a service.

16 . The method of claim 15 , wherein the audio input comprises one or more user utterances received by a voice enabled device.

17 . The method of claim 16 , further comprising receiving, based on the one or more user utterances, from a cloud based voice service, one or more responses.

18 . The method of claim 15 , further comprising determining, based on the first boundary and the second boundary, a wake word duration.

19 . The method of claim 15 , wherein the service comprises a cloud based voice service.

20 . The method of claim 19 , further comprising determining timing data associated with the first two audio frames and timing data associated with the second two audio frames.