IP Library Granted Patent US 12,334,065
Granted Patent B2
US 12,334,065 · App. 17/315,704 · Granted Jun 17, 2025

Voice command recognition system

Inventors: Ralph Birt (Santa Clara, CA); Robert Curtis (Los Gatos, CA); David H. Friedman (Austin, TX)
Assignee: ROKU, INC.
G10L15/22G10L15/08G10L21/034G10L21/0364G10L25/51G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,334,065
App. No.
17/315,704
Granted
Jun 17, 2025
Kind
B2
Abstract

Disclosed herein are system, apparatus, article of manufacture, method and/or computer program product embodiments, and/or combinations and sub-combinations thereof, for a voice command recognition system (VCR). An example embodiment operates by receiving a voice command directed to controlling a device, the voice command including a wake command and an action command. An amplitude of the wake command is determined. A gain adjustment for the voice command is calculated based on a comparison of the amplitude of the wake command to a target amplitude. An amplitude of the action command is adjusted based on the calculated gain adjustment for the voice command based on the comparison of the amplitude of the wake command to the target amplitude. A device command for controlling the device is identified based on the action command comprising the adjusted amplitude. The device command is provided to the device.

Claims (72)

1. A computer implemented method, comprising:

receiving a voice command directed to controlling a device, the voice command comprising a wake command and an action command;

determining an amplitude of the wake command;

determining a previous gain adjustment based on historical information identifying calculated gain adjustments for a plurality of previous voice commands;

applying the previous gain adjustment to only the wake command;

adjusting an amplitude of the action command to generate an amplitude-adjusted action command based on the previous gain adjustment applied to the wake command and a difference in the amplitude of the wake command and the amplitude of the action command;

adjusting an amplitude of the amplitude-adjusted action command to generate a second amplitude-adjusted action command based on a difference in the amplitude of the amplitude-adjusted action command and a target amplitude;

identifying a device command for controlling the device based on the second amplitude-adjusted action command; and

providing the device command to the device.

2. The method of claim 1 , wherein the previous gain adjustment is a median of the calculated gain adjustments for the plurality of previous voice commands.

3. The method of claim 1 , further comprising:

determining that the calculated gain adjustment for the wake command is an anomaly, wherein the calculated gain adjustment is excluded from the historical information when processing a subsequent voice command.

4. The method of claim 1 , further comprising:

calculating a correction based on the difference in the amplitude of the amplitude-adjusted action command and the target amplitude, wherein the correction is applied to an amplitude of a subsequent action command of a subsequent voice command.

5. The method of claim 4 , wherein calculating the correction comprises:

calculating a first correction for a first user; and

calculating a second correction for a second user different from the first user.

6. The method of claim 1 , further comprising:

detecting the wake command prior to receiving the action command, wherein the device is configured to output an audible beep upon detecting the wake command;

determining that the voice command comprises a continuous stream of speech after detecting the wake command;

suppressing the audible beep based on the continuous stream of speech determination; and

detecting the wake command from the voice command.

7. The computer implemented method of claim 1 , further comprising:

retrieving the historical information via a network.

8. A system, comprising:

a memory; and

at least one processor coupled to the memory and configured to perform operations comprising:

receiving a voice command directed to controlling a device, the voice command comprising a wake command and an action command;

determining an amplitude of the wake command;

determining a previous gain adjustment based on historical information identifying calculated gain adjustments for a plurality of previous voice commands;

applying the previous gain adjustment to only the wake command;

adjusting an amplitude of the action command to generate an amplitude-adjusted action command based on the previous gain adjustment applied to the wake command and a difference in the amplitude of the wake command and the amplitude of the action command;

adjusting an amplitude of the amplitude-adjusted action command to generate a second amplitude-adjusted action command based on a difference in the amplitude of the amplitude-adjusted action command and a target amplitude;

identifying a device command for controlling the device based on the second amplitude-adjusted action command; and

providing the device command to the device.

9. The system of claim 8 , wherein the previous gain adjustment is a median of the calculated gain adjustments for the plurality of previous voice commands.

10. The system of claim 8 , wherein the operations further comprise:

determining that the calculated gain adjustment for the voice command is an anomaly, wherein the calculated gain adjustment for the wake command is excluded from the historical information when processing a subsequent voice command.

11. The system of claim 8 , wherein the operations further comprise:

calculating a correction based on the difference in the amplitude of the amplitude-adjusted action command and the target amplitude, wherein the correction is applied to an amplitude of a subsequent action command of a subsequent voice command.

12. The system of claim 11 , wherein calculating the correction comprises:

calculating a first correction for a first user; and

calculating a second correction for a second user different from the first user.

13. The system of claim 8 , wherein the operations further comprise:

detecting the wake command prior to receiving the action command, wherein the device is configured to output an audible beep upon detecting the wake command;

determining that the voice command comprises a continuous stream of speech after detecting the wake command;

suppressing the audible beep based on the continuous stream of speech determination; and

detecting the wake command from the voice command.

14. The system of claim 9 , wherein the operations further comprise:

retrieving the historical information via a network.

15. A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising:

receiving a voice command directed to controlling a device, the voice command comprising a wake command and an action command;

determining an amplitude of the wake command;

determining a previous gain adjustment based on historical information identifying calculated gain adjustments for a plurality of previous voice commands;

applying the previous gain adjustment to only the wake command;

adjusting an amplitude of the action command to generate an amplitude-adjusted action command based on the previous gain adjustment applied to the wake command and a difference in the amplitude of the wake command and the amplitude of the action command;

adjusting an amplitude of the amplitude-adjusted action command to generate a second amplitude-adjusted action command based on a difference in the amplitude of the amplitude-adjusted action command and a target amplitude;

identifying a device command for controlling the device based on the second amplitude-adjusted action command; and

providing the device command to the device.

16. The non-transitory computer-readable medium of claim 15 , wherein the previous gain adjustment is a median of the calculated gain adjustments for the plurality of previous voice commands.

17. The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise:

determining that the calculated gain adjustment for the wake command is an anomaly, wherein the calculated gain adjustment is excluded from the historical information when processing a subsequent voice command.

18. The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise:

calculating a correction based on the difference in the amplitude of the amplitude-adjusted action command and the target amplitude, wherein the correction is applied to an amplitude of a subsequent action command of a subsequent voice command.

19. The non-transitory computer-readable medium of claim 18 , wherein calculating the correction comprises:

calculating a first correction for a first user; and

calculating a second correction for a second user different from the first user.

20. The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise:

detecting the wake command prior to receiving the action command, wherein the device is configured to output an audible beep upon detecting the wake command;

determining that the voice command comprises a continuous stream of speech after detecting the wake command;

suppressing the audible beep based on the continuous stream of speech determination; and

detecting the wake command from the voice command.

Assignments (2)
SECURITY INTEREST Recorded Sep 18, 2024
From: ROKU, INC.
To: CITIBANK, N.A.
Reel/Frame 068982/0377 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2021
From: BIRT, RALPH; CURTIS, ROBERT; FRIEDMAN, DAVID H.
To: ROKU, INC.
Reel/Frame 056197/0342 →
Continuity (1)
Related Publication 20220358915A1 · Nov 10, 2022
References Cited (17)
US 11164592B1 · Wu · 2021 [cited by examiner]
US 20010000534A1 · Matulich · 2001 [cited by examiner]
US 20020099300A1 · Kovtun · 2002 [cited by examiner]
US 20050221778A1 · Ishihara · 2005 [cited by examiner]
US 20080090617A1 · Sutardja · 2008 [cited by examiner]
US 20080200139A1 · Kobayashi · 2008 [cited by examiner]
US 20130325484A1 · Chakladar · 2013 [cited by examiner]
US 20140117852A1 · Zhai · 2014 [cited by examiner]
US 20200279575A1 · Unno · 2020 [cited by examiner]
US 20200380982A1 · Sereshki · 2020 [cited by examiner]
US 20220084519A1 · Kim · 2022 [cited by examiner]
JP 2002076996A · 2002 [cited by examiner]
WO WO2020141794A1 · 2020 [cited by applicant]
WO WO2020220345A1 · 2020 [cited by examiner]
Jakovljevic, N. et al., “Energy Normalization in Automatic Speech Recognition”; TSD 2008, LNAI 5246, pp. 341-347 (Sep. 8, 2008). [cited by applicant]
Régo, N., “How to Hear a Sound on Your Google Speaker After Saying OK Google”, Cool Blind Tech; Retrieved from the Internet at URL:https://coolblindtech.com/how-to-heara-sound-on-your-google-speaker-after-saying-ok-goog… [cited by applicant]
Extended European Search Report directed to related European Application No. 22172601.1, mailed Oct. 19, 2022; 9 pages. [cited by applicant]