IP Library Granted Patent US 11,646,025
Granted Patent B2
US 11,646,025 · App. 17/347,021 · Granted May 9, 2023

Media system with multiple digital assistants

Inventors: Anthony John Wood (San Jose, CA); David Stern (Los Gatos, CA); Gregory Mack Garner (Springdale, AZ)
Assignee: Roku, Inc.
G10L15/22G06F3/167H04L67/1014G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,646,025
App. No.
17/347,021
Granted
May 9, 2023
Kind
B2
Abstract

Disclosed herein are system, apparatus, article of manufacture, method and/or computer program product embodiments, and/or combinations and sub-combinations thereof, for providing voice control using multiple digital assistants. In some embodiments, a voice platform operates to receive a voice input from a user. The voice platform selects a digital assistant from a plurality of digital assistants based on a trigger word. The voice platform then generates an intent from the voice input using the selected digital assistant. The voice platform then transmits the intent to a media device for processing.

Claims (47)

1. A computer implemented method for providing voice control using multiple digital assistants, comprising:

selecting, by at least one processor, a first digital assistant from a plurality of digital assistants to process a voice input using a trigger word in the voice input, wherein the selected first digital assistant is mapped to the trigger word;

generating an intent from the voice input;

determining that a second digital assistant from the plurality of digital assistants is unmapped to the trigger word, and that the second digital assistant processes a type of the intent from the voice input more often than the selected first digital assistant; and

selecting the second digital assistant from the plurality of digital assistants to process the voice input based on the determining.

2. The method of claim 1 , further comprising:

transmitting the intent to a voice adaptor at a media device, wherein the voice adaptor selects an application to process the intent based on a fixed rule, default application setting, a search result, or metadata in the intent.

3. The method of claim 1 , further comprising:

refining the intent based on information in a cloud computing platform.

4. The method of claim 1 , wherein the generating the intent further comprises:

generating the intent from the voice input using the selected first digital assistant.

5. The method of claim 1 , wherein the generating the intent further comprises:

converting the voice input into a text input using an automated speech recognizer associated with the selected first digital assistant; and

generating the intent from the text input using a natural language unit associated with the selected first digital assistant.

6. The method of claim 1 , wherein the determining further comprises:

determining that the second digital assistant processes the type of the intent from the voice input more often than the selected first digital assistant based on crowdsourced data, wherein the crowdsourced data indicates how often each digital assistant in the plurality of digital assistants is used to process the type of the intent.

7. The method of claim 6 , further comprising:

in response to selecting the second digital assistant, incrementing a count in the crowdsourced data that indicates a number of times the second digital assistant was selected.

8. A voice platform, comprising:

a memory; and

at least one processor coupled to the memory and configured to:

select a first digital assistant from a plurality of digital assistants to process a voice input using a trigger word in the voice input, wherein the selected first digital assistant is mapped to the trigger word;

generate an intent from the voice input;

determine that a second digital assistant from the plurality of digital assistants is unmapped to the trigger word, and that the second digital assistant processes a type of the intent from the voice input more often than the selected first digital assistant; and

select the second digital assistant from the plurality of digital assistants to process the voice input based on the determining.

9. The voice platform of claim 8 , wherein the at least one processor is further configured to:

transmit the intent to a voice adaptor at a media device, wherein the voice adaptor selects an application to process the intent based on a fixed rule, default application setting, a search result, or metadata in the intent.

10. The voice platform of claim 8 , wherein the at least one processor is further configured to:

refine the intent based on information in a cloud computing platform.

11. The voice platform of claim 8 , wherein to generate the intent, the at least one processor is further configured to:

generate the intent from the voice input using the selected first digital assistant.

12. The voice platform of claim 8 , wherein to generate the intent, the at least one processor is further configured to:

convert the voice input into a text input using an automated speech recognizer associated with the selected first digital assistant; and

generate the intent from the text input using a natural language unit associated with the selected first digital assistant.

13. The voice platform of claim 8 , wherein to determine that the second digital assistant processes the type of the intent from the voice input more often than the selected first digital assistant, the at least one processor is further configured to:

determine that the second digital assistant processes the type of the intent from the voice input more often than the selected first digital assistant based on crowdsourced data, wherein the crowdsourced data indicates how often each digital assistant in the plurality of digital assistants is used to process the type of the intent.

14. The voice platform of claim 13 , wherein the at least one processor is further configured to:

in response to selecting the second digital assistant, increment a count in the crowdsourced data that indicates a number of times the second digital assistant was selected.

15. A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising:

transmitting a voice input to a voice platform, wherein the voice platform selects a first digital assistant from a plurality of digital assistants to process the voice input using a trigger word in the voice input, generates an intent from the voice input, determines that a second digital assistant from the plurality of digital assistants is unmapped to the trigger word, and that the second digital assistant processes a type of the intent from the voice input more often than the selected first digital assistant, and selects the second digital assistant from the plurality of digital assistants to process the voice input based on the determining; and

receiving the intent from the voice platform.

16. The non-transitory computer-readable medium of claim 15 , wherein the receiving the intent from the voice platform further comprises:

receiving the intent at a voice adaptor, wherein the voice adaptor selects an application to process the intent based on a fixed rule, default application setting, a search result, or metadata in the intent.

17. The non-transitory computer-readable medium of claim 15 , wherein the voice platform refines the intent based on information in a cloud computing platform.

18. The non-transitory computer-readable medium of claim 15 , wherein the voice platform converts the voice input into a text input using an automated speech recognizer associated with the selected first digital assistant, and generates the intent from the text input using a natural language unit associated with the selected first digital assistant.

19. The non-transitory computer-readable medium of claim 15 , wherein the voice platform determines that the second digital assistant processes the type of the intent from the voice input more often than the selected first digital assistant based on crowdsourced data, wherein the crowdsourced data indicates how often each digital assistant in the plurality of digital assistants is used to process the type of the intent.

20. The non-transitory computer-readable medium of claim 19 , wherein the voice platform, in response to selecting the second digital assistant, increments a count in the crowdsourced data that indicates a number of times the second digital assistant was selected.

Assignments (2)
SECURITY INTEREST Recorded Sep 18, 2024
From: ROKU, INC.
To: CITIBANK, N.A.
Reel/Frame 068982/0377 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2021
From: WOOD, ANTHONY JOHN; STERN, DAVID; GARNER, GREGORY MACK
To: ROKU, INC.
Reel/Frame 056536/0022 →
Continuity (3)
Continuation 16032724 · Jul 11, 2018
Provisional Application 62550940 · Aug 28, 2017
Related Publication 20210304765A1 · Sep 30, 2021
Cited By (3)
US 12,265,746 US 12,482,467 US 12,614,549