IP Library Granted Patent US 11,017,776
Granted Patent B2
US 11,017,776 · App. 16/032,868 · Granted May 25, 2021

Local and cloud speech recognition

Inventors: Anthony John Wood (Los Gatos, CA); David Stern (Los Gatos, CA); Gregory Mack Garner (Springdale, AZ)
Assignee: Roku, Inc.
G10L15/30G06F3/167G10L15/22H04L67/10H04R1/326H04R27/00G10L15/20G10L21/0208G10L2015/223G10L2021/02082G10L2021/02166H04R3/005H04R2227/003H04R2227/005H04R2430/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,017,776
App. No.
16/032,868
Granted
May 25, 2021
Kind
B2
Abstract

Disclosed herein are system, apparatus, article of manufacture, method and/or computer program product embodiments, and/or combinations and sub-combinations thereof, for distributing the performance of speech recognition among a remote control device and a voice platform in the cloud. In some embodiments, the remote control device operates to receive a voice input from a user. The remote control device detects a trigger word in the voice input. The remote control device then processes the voice input. The remote control device then transmits the voice input to a voice platform based on the detecting in order to determine an intent associated with the voice input.

Claims (48)

1. A computer implemented method for performing speech recognition for a digital assistant, comprising:

receiving, by at least one processor at an audio responsive electronic device, a voice input from a user;

detecting, by the at least one processor, that a trigger word is in the voice input with a first confidence value;

processing, by the at least one processor, the voice input;

transmitting, by the at least one processor, the voice input to a voice platform in response to the detecting the trigger word is in the voice input with the first confidence value;

in response to the voice platform performing a secondary trigger word detection on the voice input and determining an intent based on the voice input, receiving a confirmation from the voice platform that the trigger word is in the voice input with a second confidence value; and

transmitting, by the at least one processor, a remainder of the voice input to a digital assistant in the voice platform in response to the receiving the confirmation that the trigger word is in the voice input with the second confidence value.

2. The method of claim 1 , wherein the voice platform comprises a cloud computing platform.

3. The method of claim 1 , further comprising:

performing, by the at least one processor, echo cancellation on the voice input.

4. The method of claim 1 , the processing further comprising:

performing, by the at least one processor, noise cancellation on the voice input using a position of the user, wherein the performing comprises adjusting a reception pattern for a microphone using the position of the user.

5. The method of claim 1 , wherein the voice platform converts the voice input into a text input using automated speech recognition.

6. The method of claim 5 , wherein the voice platform converts the text input into the intent using natural language processing.

7. An audio responsive electronic device, comprising:

a microphone;

a memory; and

a processor coupled to the memory and configured to:

receive a voice input from a user via the microphone;

detect that a trigger word is in the voice input with a first confidence value;

process the voice input;

transmit the voice input to a voice platform in response to the detecting the trigger word is in the voice input with the first confidence value;

in response to the voice platform performing a secondary trigger word detection on the voice input and determining an intent based on the voice input, receive a confirmation from the voice platform that the trigger word is in the voice input with a second confidence value; and

transmit a remainder of the voice input to a digital assistant in the voice platform in response to the receiving the confirmation that the trigger word is in the voice input with the second confidence value.

8. The audio responsive electronic device of claim 7 , wherein the voice platform comprises a cloud computing platform.

9. The audio responsive electronic device of claim 7 , wherein the processor is further configured to:

perform echo cancellation on the voice input.

10. The audio responsive electronic device of claim 7 , wherein the processor is further configured to:

perform noise cancellation on the voice input using a position of the user, wherein the performing comprises adjusting a reception pattern for the microphone using the position of the user.

11. The audio responsive electronic device of claim 7 , wherein the voice platform converts the voice input into a text input using automated speech recognition.

12. The audio responsive electronic device of claim 11 , wherein the voice platform converts the text input into the intent using natural language processing.

13. A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, causes the at least one computing device to perform operations comprising:

receiving a voice input from a user;

detecting that a trigger word is in the voice input with a first confidence value;

processing the voice input;

transmitting the voice input to a voice platform in response to the detecting;

in response to the voice platform performing a secondary trigger word detection on the voice input and determining an intent based on the voice input, receiving a confirmation from the voice platform that the trigger word is in the voice input with a second confidence value; and

transmitting a remainder of the voice input to a digital assistant in the voice platform in response to the receiving the confirmation that the trigger word is in the voice input with the second confidence value.

14. The non-transitory computer-readable medium of claim 13 , the operations further comprising:

performing echo cancellation on the voice input.

15. The non-transitory computer-readable medium of claim 13 , the operations further comprising:

performing noise cancellation on the voice input using a position of the user, wherein the performing comprises adjusting a reception pattern for a microphone using the position of the user.

16. The non-transitory computer-readable medium of claim 13 , wherein the voice platform converts the voice input into a text input using automated speech recognition.

17. The non-transitory computer-readable medium of claim 16 , wherein the voice platform converts the text input into the intent using natural language processing.

18. The method of claim 1 , wherein the second confidence value is higher than the first confidence value.

19. The audio responsive electronic device of claim 7 , wherein the second confidence value is higher than the first confidence value.

20. The non-transitory computer-readable medium of claim 13 , wherein the transmitting the remainder of the voice input further comprises:

transmitting the remainder of the voice input after the detected trigger word to the digital assistant in the voice platform.

Assignments (4)
SECURITY INTEREST Recorded Sep 18, 2024
From: ROKU, INC.
To: CITIBANK, N.A.
Reel/Frame 068982/0377 →
TERMINATION AND RELEASE OF INTELLECTUAL PROPERTY SECURITY AGREEMENT (REEL/FRAME 048385/0375) Recorded Feb 22, 2023
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: ROKU, INC.
Reel/Frame 062826/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2021
From: WOOD, ANTHONY JOHN; STERN, DAVID; GARNER, GREGORY MACK
To: ROKU, INC.
Reel/Frame 056274/0740 →
PATENT SECURITY AGREEMENT Recorded Feb 20, 2019
From: ROKU, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 048385/0375 →
Continuity (2)
Provisional Application 62550935 · Aug 28, 2017
Related Publication 20190066687A1 · Feb 28, 2019