IP Library Granted Patent US 12,217,759
Granted Patent B2
US 12,217,759 · App. 18/434,602 · Granted Feb 4, 2025

Voice to text conversion based on third-party agent content

Inventors: Barnaby James (Los Gatos, CA); Bo Wang (San Jose, CA); Sunil Vemuri (Pleasanton, CA); David Schairer (San Jose, CA); Ulas Kirazci (Mountain View, CA); Ertan Dogrultan (Belmont, CA); Petar Aleksic (Jersey City, NJ)
Assignee: GOOGLE LLC
G10L15/26G06F40/205G06F40/284G06F40/30G10L15/1815G10L15/183G10L15/22G10L15/30G10L2015/223G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,759
App. No.
18/434,602
Granted
Feb 4, 2025
Kind
B2
Abstract

Implementations relate to dynamically, and in a context-sensitive manner, biasing voice to text conversion. In some implementations, the biasing of voice to text conversions is performed by a voice to text engine of a local agent, and the biasing is based at least in part on content provided to the local agent by a third-party (3P) agent that is in network communication with the local agent. In some of those implementations, the content includes contextual parameters that are provided by the 3P agent in combination with responsive content generated by the 3P agent during a dialog that: is between the 3P agent, and a user of a voice-enabled electronic device; and is facilitated by the local agent. The contextual parameters indicate potential feature(s) of further voice input that is to be provided in response to the responsive content generated by the 3P agent.

Claims (47)

1. A voice-enabled electronic device comprising:

at least one processor; and

memory storing instructions that, when executed by the at least one processor, cause the at least one processor to be operable to:

receive, from a third-party agent, one or more contextual parameters associated with the third-party agent, wherein the third-party agent is managed by an additional party that is distinct from the party that manages a local agent of the voice-enabled electronic device;

receive, from a user of the voice-enabled electronic device, a voice input provided by the user;

in response to receiving the voice input:

convert, using a voice to text model, the voice input to text, wherein the instructions to convert the voice input to text comprise instructions to use one or more of the contextual parameters to bias the voice to text model in converting at least one segment of the voice input to the text; and

transmit at least a portion of the text to the third-party agent.

2. The voice-enabled electronic device of claim 1 , wherein one or more of the contextual parameters comprise one or more particular tokens.

3. The voice-enabled electronic device of claim 1 , wherein one or more of the contextual parameters comprise one or more semantic types of tokens.

4. The voice-enabled electronic device of claim 1 , wherein one or more of the semantic types of tokens identify one or more of a time semantic type and a date semantic type.

5. The voice-enabled electronic device of claim 1 , wherein one or more of the contextual parameters comprise one or more semantic types of tokens and comprise one or more particular tokens.

6. The voice-enabled electronic device of claim 1 , wherein the instructions further cause the voice-enabled electronic device to:

receive, from the third-party agent and responsive to transmitting at least the portion of the text to the third-party agent, content that includes responsive content that is to be provided in response to the voice input; and

provide the responsive content as output for presentation to the user via the voice-enabled electronic device, the output being provided in response to the voice input.

7. The voice-enabled electronic device of claim 6 , wherein the instructions further cause the voice-enabled electronic device to:

receive, from the user of the voice-enabled electronic device, an additional voice input provided by the user, the additional voice input being provided in response to the output; and

use the content received from the third-party agent to convert the additional voice input to additional text.

8. The voice-enabled electronic device of claim 7 , wherein the content received from the third-party agent comprises one or more additional contextual parameters associated with the third-party agent, and wherein using the content received from the third-party agent to convert the additional voice input to the additional text comprises using one or more of the additional contextual parameter to bias the voice to text model in converting at least one segment of the additional voice input to the additional text.

9. The voice-enabled electronic device of claim 1 , wherein receiving one or more of the contextual parameters associated with the third-party agent is in response to an invocation of the third-party agent.

10. The voice-enabled electronic device of claim 1 , wherein the voice to text model is a streaming voice to text model.

11. A system comprising:

at least one processor; and

memory storing instructions that, when executed by the at least one processor, cause the at least one processor to be operable to:

receive, from a third-party agent, one or more contextual parameters associated with the third-party agent, wherein the third-party agent is managed by an additional party that is distinct from the party that manages a local agent of a voice-enabled electronic device;

receive, from a user of the voice-enabled electronic device, a voice input provided by the user;

in response to receiving the voice input:

convert, using a voice to text model, the voice input to text, wherein the instructions to convert the voice input to text comprise instructions to use one or more of the contextual parameters to bias the voice to text model in converting at least one segment of the voice input to the text; and

transmit at least a portion of the text to the third-party agent.

12. The system of claim 11 , wherein one or more of the contextual parameters comprise one or more particular tokens.

13. The system of claim 11 , wherein one or more of the contextual parameters comprise one or more semantic types of tokens.

14. The system of claim 11 , wherein one or more of the semantic types of tokens identify one or more of a time semantic type and a date semantic type.

15. The system of claim 11 , wherein one or more of the contextual parameters comprise one or more semantic types of tokens and comprise one or more particular tokens.

16. The system of claim 11 , wherein the instructions further cause the voice-enabled electronic device to:

receive, from the third-party agent and responsive to transmitting at least the portion of the text to the third-party agent, content that includes responsive content that is to be provided in response to the voice input; and

provide the responsive content as output for presentation to the user via the voice-enabled electronic device, the output being provided in response to the voice input.

17. The system of claim 16 , wherein the instructions further cause the voice-enabled electronic device to:

receive, from the user of the voice-enabled electronic device, an additional voice input provided by the user, the additional voice input being provided in response to the output; and

use the content received from the third-party agent to convert the additional voice input to additional text.

18. The system of claim 17 , wherein the content received from the third-party agent comprises one or more additional contextual parameters associated with the third-party agent, and wherein using the content received from the third-party agent to convert the additional voice input to the additional text comprises using one or more of the additional contextual parameter to bias the voice to text model in converting at least one segment of the additional voice input to the additional text.

19. The system of claim 11 , wherein receiving one or more of the contextual parameters associated with the third-party agent is in response to an invocation of the third-party agent.

20. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to be operable to perform operations, the operations comprising:

receiving, from a third-party agent, one or more contextual parameters associated with the third-party agent, wherein the third-party agent is managed by an additional party that is distinct from the party that manages a local agent of a voice-enabled electronic device;

receiving, from a user of the voice-enabled electronic device, a voice input provided by the user;

in response to receiving the voice input:

converting, using a voice to text model, the voice input to text, wherein the instructions to convert the voice input to text comprise instructions to use one or more of the contextual parameters to bias the voice to text model in converting at least one segment of the voice input to the text; and

transmitting at least a portion of the text to the third-party agent.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 14, 2024
From: JAMES, BARNABY; WANG, BO; VEMURI, SUNIL; SCHAIRER, DAVID; KIRAZCI, ULAS; DOGRULTAN, ERTAN; ALEKSIC, PETER
To: GOOGLE INC.
Reel/Frame 066777/0145 →
CHANGE OF NAME Recorded Mar 14, 2024
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 066798/0417 →
Continuity (5)
Continuation 18125606 · Mar 23, 2023
Continuation 17582926 · Jan 24, 2022
Continuation 16791334 · Feb 14, 2020
Continuation 15372188 · Dec 7, 2016
Related Publication 20240274133A1 · Aug 15, 2024
References Cited (45)
US 7761499B2 · Hodjat et al. · 2010 [cited by applicant]
US 8204738B2 · Skuratovsky · 2012 [cited by applicant]
US 8862467B1 · Casado et al. · 2014 [cited by applicant]
US 10600418B2 · James et al. · 2020 [cited by applicant]
US 11232797B2 · James et al. · 2022 [cited by applicant]
US 11626115B2 · James et al. · 2023 [cited by applicant]
US 20050137868A1 · Epstein et al. · 2005 [cited by applicant]
US 20070016401A1 · Ehsani et al. · 2007 [cited by applicant]
US 20070150278A1 · Bates et al. · 2007 [cited by applicant]
US 20070150286A1 · Miller · 2007 [cited by applicant]
US 20090052636A1 · Webb et al. · 2009 [cited by applicant]
US 20090299745A1 · Kennewick et al. · 2009 [cited by applicant]
US 20100076843A1 · Ashton · 2010 [cited by applicant]
US 20120265528A1 · Gruber et al. · 2012 [cited by applicant]
US 20130197907A1 · Burke et al. · 2013 [cited by applicant]
US 20140040748A1 · Lemay et al. · 2014 [cited by applicant]
US 20140052445A1 · Beckford et al. · 2014 [cited by applicant]
US 20140278379A1 · Coccaro et al. · 2014 [cited by applicant]
US 20150066479A1 · Pasupalak et al. · 2015 [cited by applicant]
US 20150073790A1 · Steuble et al. · 2015 [cited by applicant]
US 20150302002A1 · Mathias et al. · 2015 [cited by applicant]
US 20150370787A1 · Akbacak et al. · 2015 [cited by applicant]
US 20160104482A1 · Aleksic et al. · 2016 [cited by applicant]
US 20160260433A1 · Sumner et al. · 2016 [cited by applicant]
US 20160351194A1 · Gao et al. · 2016 [cited by applicant]
US 20190122657A1 · James et al. · 2019 [cited by applicant]
US 20200184974A1 · James et al. · 2020 [cited by applicant]
US 20220148596A1 · James et al. · 2022 [cited by applicant]
US 20230260517A1 · James et al. · 2023 [cited by applicant]
CN 101297355 · 2008 [cited by applicant]
CN 104509080 · 2015 [cited by applicant]
CN 105027194 · 2015 [cited by applicant]
WO 2015179510 · 2015 [cited by applicant]
China National Intellectual Property Administration; Grant Notice issued in Application No. 201780076180.3; 4 pages; dated May 31, 2023. [cited by applicant]
China National Intellectual Property Administration; Notification of First Office Action issued in Application No. 201780076180.3; 21 pages; dated Nov. 28, 2022. [cited by applicant]
European Patent Office; Communication issued in Application No. 21190701.9; 7 pages; dated Nov. 18, 2021. [cited by applicant]
European Patent Office; Intention to Grant issued in Application No. 17778419.6; 45 pages; dated Feb. 11, 2021. [cited by applicant]
International Search Report and Written Opinion of PCT Ser. No. PCT/US2017/052730; 12 pages Dec. 14, 2017. [cited by applicant]
United Kingdom Intellectual Property Office; Examination Report issued in Application No. 1715619.1 dated Mar. 19, 2018. [cited by applicant]
European Patent Office; Written Opinion of the International Preliminary Examining Authority; 6 pages; dated Oct. 22, 2018. [cited by applicant]
European Patent Office; International Preliminary Report on Patentability of PCT/US2017/052730; 18 pages; dated Feb. 25, 2019. [cited by applicant]
United Kingdom Intellectual Property Office; Examination Report issued in Application No. GB1715619.1 dated Aug. 27, 2019. [cited by applicant]
European Patent Office; Examnation Report issued in Application No. 17778419.6 dated Oct. 10, 2019. [cited by applicant]
European Patent Office; Intention to Grant issued in Application No. 21190701.9; 50 pages; dated Jul. 7, 2023. [cited by applicant]
European Patent Office; Extended European Search Report issued for Application No. 24150123.8, 8 pages, dated Apr. 14, 2024. [cited by applicant]