IP Library Granted Patent US 11,804,227
Granted Patent B2
US 11,804,227 · App. 17/327,115 · Granted Oct 31, 2023

Local and cloud speech recognition

Inventors: Anthony John Wood (Los Gatos, CA); David Stern (Los Gatos, CA); Gregory Mack Garner (Springdale, AZ)
Assignee: Roku, Inc.
G10L15/30G06F3/167G10L15/22H04L67/10H04R1/326H04R27/00G10L15/20G10L21/0208G10L2015/223G10L2021/02082G10L2021/02166H04R3/005H04R2227/003H04R2227/005H04R2430/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,804,227
App. No.
17/327,115
Granted
Oct 31, 2023
Kind
B2
Abstract

Disclosed herein are system, apparatus, article of manufacture, method and/or computer program product embodiments, and/or combinations and sub-combinations thereof, for distributing the performance of speech recognition among a remote control device and a voice platform in the cloud. In some embodiments, the remote control device operates to receive a voice input from a user. The remote control device detects a trigger word in the voice input. The remote control device then processes the voice input. The remote control device then transmits the voice input to a voice platform based on the detecting in order to determine an intent associated with the voice input.

Claims (47)

1. A computer-implemented method for performing speech recognition, comprising:

receiving, by at least one processor at an electronic device, a voice input from a user;

detecting that a trigger word is in the voice input with a first confidence value;

transmitting the voice input to a voice platform in response to the detecting;

in response to the voice platform performing a secondary trigger word detection on the voice input, receiving a confirmation from the voice platform that the trigger word is in the voice input with a second confidence value; and

transmitting the voice input to the voice platform in response to the receiving the confirmation that the trigger word is in the voice input with the second confidence value.

2. The computer-implemented method of claim 1 , further comprising:

performing echo cancellation on the voice input.

3. The computer-implemented method of claim 1 , further comprising:

performing noise cancellation on the voice input using a position of the user, wherein the performing comprises adjusting a reception pattern for a microphone using the position of the user.

4. The computer-implemented method of claim 1 , wherein the second confidence value is higher than the first confidence value.

5. The computer-implemented method of claim 1 , wherein the transmitting the voice input to the voice platform in response to receiving the confirmation that the trigger word is in the voice input with the second confidence value further comprises:

transmitting a remainder of the voice input after the detected trigger word to the voice platform.

6. The computer-implemented method of claim 1 , wherein the voice platform converts the voice input into a text input using automated speech recognition.

7. The computer-implemented method of claim 1 , wherein the voice platform comprises a cloud computing platform.

8. An electronic device, comprising:

a microphone;

a memory; and

a processor coupled to the memory and configured to:

receive a voice input from a user via the microphone;

detect that a trigger word is in the voice input with a first confidence value;

transmit the voice input to a voice platform in response to the detecting;

in response to the voice platform performing a secondary trigger word detection on the voice input, receive a confirmation from the voice platform that the trigger word is in the voice input with a second confidence value; and

transmit the voice input to the voice platform in response to the receiving the confirmation that the trigger word is in the voice input with the second confidence value.

9. The electronic device of claim 8 , wherein the processor is further configured to:

perform echo cancellation on the voice input.

10. The electronic device of claim 8 , wherein the processor is further configured to:

perform noise cancellation on the voice input using a position of the user, wherein the performing comprises adjusting a reception pattern for the microphone using the position of the user.

11. The electronic device of claim 8 , wherein the second confidence value is higher than the first confidence value.

12. The electronic device of claim 8 , wherein to transmit the voice input to the voice platform in response to the receiving the confirmation that the trigger word is in the voice input with the second confidence value, the processor is further configured to:

transmit a remainder of the voice input after the detected trigger word to the voice platform.

13. The electronic device of claim 8 , wherein the voice platform converts the voice input into a text input using automated speech recognition.

14. The electronic device of claim 8 , wherein the voice platform comprises a cloud computing platform.

15. A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising:

receiving a voice input from a user;

detecting that a trigger word is in the voice input with a first confidence value;

transmitting the voice input to a voice platform in response to the detecting;

in response to the voice platform performing a secondary trigger word detection on the voice input, receiving a confirmation from the voice platform that the trigger word is in the voice input with a second confidence value; and

transmitting the voice input to the voice platform in response to the receiving the confirmation that the trigger word is in the voice input with the second confidence value.

16. The non-transitory computer-readable medium of claim 15 , the operations further comprising:

performing echo cancellation on the voice input.

17. The non-transitory computer-readable medium of claim 15 , the operations further comprising:

performing noise cancellation on the voice input using a position of the user, wherein the performing comprises adjusting a reception pattern for a microphone using the position of the user.

18. The non-transitory computer-readable medium of claim 15 , wherein the second confidence value is higher than the first confidence value.

19. The non-transitory computer-readable medium of claim 15 , wherein the transmitting the voice input to the voice platform in response to the receiving the confirmation that the trigger word is in the voice input with the second confidence value further comprises:

transmitting a remainder of the voice input after the detected trigger word to the voice platform.

20. The non-transitory computer-readable medium of claim 15 , wherein the voice platform converts the voice input into a text input using automated speech recognition.

Assignments (2)
SECURITY INTEREST Recorded Sep 18, 2024
From: ROKU, INC.
To: CITIBANK, N.A.
Reel/Frame 068982/0377 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2021
From: WOOD, ANTHONY JOHN; STERN, DAVID; GARNER, GREGORY MACK
To: ROKU, INC.
Reel/Frame 056316/0499 →
Continuity (3)
Continuation 16032868 · Jul 11, 2018
Provisional Application 62550935 · Aug 28, 2017
Related Publication 20210327433A1 · Oct 21, 2021
Cited By (4)
US 12,265,746 US 12,482,467 US 12,614,549 US 12,731,589