IP Library Granted Patent US 11,232,797
Granted Patent B2
US 11,232,797 · App. 16/791,334 · Granted Jan 25, 2022

Voice to text conversion based on third-party agent content

Inventors: Barnaby James (Los Gatos, CA); Bo Wang (San Jose, CA); Sunil Vemuri (Pleasanton, CA); David Schairer (San Jose, CA); Ulas Kirazci (Mountain View, CA); Ertan Dogrultan (Belmont, CA); Petar Aleksic (Jersey City, NJ)
Assignee: Google LLC
G10L15/26G10L15/1815G10L15/22G10L15/30G10L2015/223G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,232,797
App. No.
16/791,334
Granted
Jan 25, 2022
Kind
B2
Abstract

Implementations relate to dynamically, and in a context-sensitive manner, biasing voice to text conversion. In some implementations, the biasing of voice to text conversions is performed by a voice to text engine of a local agent, and the biasing is based at least in part on content provided to the local agent by a third-party (3P) agent that is in network communication with the local agent. In some of those implementations, the content includes contextual parameters that are provided by the 3P agent in combination with responsive content generated by the 3P agent during a dialog that: is between the 3P agent, and a user of a voice-enabled electronic device; and is facilitated by the local agent. The contextual parameters indicate potential feature(s) of further voice input that is to be provided in response to the responsive content generated by the 3P agent.

Claims (45)

1. A method implemented by one or more processors of a local agent managed by a party, comprising:

storing an association of a plurality of contextual parameters to invocation of a third-party agent,

wherein the third-party agent is managed by an additional party that is distinct from the party that manages the local agent, and

wherein storing the association is responsive to the additional party providing the contextual parameters;

subsequent to storing the association:

receiving a voice input provided by a user via a voice-enabled electronic device;

converting a first segment of the voice input to text using a streaming voice to text model and without biasing of the streaming voice to text model based on the contextual parameters;

determining that the text, converted from the first segment of the voice input, conforms to an invocation keyword that is specific to invocation of the third-party agent;

in response to determining that the text conforms to the invocation keyword that is specific to invocation of the third-party agent, and in response to previously storing the association of the contextual parameters to invocation of the third-party agent:

converting a second segment of the voice input to additional text using the streaming voice to text model and using the contextual parameters to bias the streaming voice to text model, the second segment of the voice input being subsequent to the first segment of the voice input, and

transmitting at least a portion of the additional text to the third-party agent.

2. The method of claim 1 , wherein the contextual parameters comprise one or more particular tokens.

3. The method of claim 1 , wherein the contextual parameters comprise one or more semantic types of tokens.

4. The method of claim 3 , wherein the one or more semantic types of tokens identify one or more of a time semantic type and a date semantic type.

5. The method of claim 1 , wherein the contextual parameters comprise one or more semantic types of tokens and comprise one or more particular tokens.

6. The method of claim 1 , further comprising:

receiving, from the third-party agent and responsive to transmitting the portion of the additional text to the third-party agent, content that includes responsive content that is to be provided in response to the voice input; and

providing the responsive content as output for presentation to the user via the voice-enabled electronic device, the output being provided in response to the voice input.

7. The method of claim 6 , further comprising:

receiving an additional voice input provided by the user, the additional voice input provided via the voice-enabled device and being provided in response to the output; and

using the content received from the third-party agent to convert the additional voice input to additional text.

8. The method of claim 7 , wherein the content received from the third-party agent comprises an additional contextual parameter that is in addition to the contextual parameters with the stored association to invocation of the third-party agent, and wherein using the content received from the third-party agent to convert the additional voice input to additional text comprises using the additional contextual parameter to bias the streaming voice to text model.

9. A computing device, comprising:

a network interface;

memory storing instructions;

a storage subsystem that stores an association of a plurality of contextual parameters to invocation of a third-party agent,

wherein the third-party agent is managed by an additional party that is distinct from the party that manages the local agent, and

wherein the association is stored in the storage subsystem responsive to the additional party providing the contextual parameters;

one or more processors operable to execute instructions stored in the memory, comprising instructions to:

receive a voice input provided by a user;

convert a first segment of the voice input to text using a streaming voice to text model and without biasing of the streaming voice to text model based on the contextual parameters;

determine that the text, converted from the first segment of the voice input, conforms to an invocation keyword that is specific to invocation of the third-party agent;

in response to determining that the text conforms to the invocation keyword that is specific to invocation of the third-party agent, and in response to the storage subsystem storing the association of the contextual parameters to invocation of the third-party agent:

convert a second segment of the voice input to additional text using the streaming voice to text model and using the contextual parameters to bias the streaming voice to text model, the second segment of the voice input being subsequent to the first segment of the voice input, and

transmit, via the network interface, at least a portion of the additional text to the third-party agent.

10. The computing device of claim 9 , wherein the contextual parameters comprise one or more particular tokens.

11. The computing device of claim 10 , wherein the contextual parameters comprise one or more semantic types of tokens.

12. The computing device of claim 9 , wherein the contextual parameters comprise one or more semantic types of tokens.

13. The computing device of claim 9 , wherein the instructions further comprise instructions to:

receive, from the third-party agent and responsive to transmitting the portion of the additional text to the third-party agent, content that includes responsive content that is to be provided in response to the voice input; and

provide the responsive content as output for presentation in response to the voice input.

14. The computing device of claim 13 , wherein the instructions further comprise instructions to:

receiving an additional voice input provided by the user, the additional voice input provided in response to the output; and

use the content received from the third-party agent to convert the additional voice input to additional text.

15. The computing device of claim 14 , wherein the content received from the third-party agent comprises an additional contextual parameter that is in addition to the contextual parameters with the stored association to invocation of the third-party agent, and wherein the instructions to use the content received from the third-party agent to convert the additional voice input to additional text comprise instructions to use the additional contextual parameter to bias the streaming voice to text model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2020
From: JAMES, BARNABY; WANG, BO; VEMURI, SUNIL; SCHAIRER, DAVID; KIRAZCI, ULAS; DOGRULTAN, ERTAN; ALEKSIC, PETAR
To: GOOGLE INC.
Reel/Frame 052164/0427 →
CHANGE OF NAME Recorded Mar 19, 2020
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 052189/0804 →
Continuity (2)
Continuation 15372188 · Dec 7, 2016
Related Publication 20200184974A1 · Jun 11, 2020
Cited By (2)
US 12,217,759 US 12,548,568