IP Library › Granted Patent US 12,615,224
Granted Patent B2
US 12,615,224 · App. 18/370,239 · Granted Apr 28, 2026

Voice wrapper(s) for existing third-party text-based chatbot(s)

Inventors: Sasha Goldshtein (Tel Aviv, IL); Yoav Tzur (Tel Aviv, IL); Shlomo Fruchter (Ness Ziona, IL); Gal Moshitch (Jerusalem, IL); ChiTa Tsai (New York, NY); Sharon Sultan (Tel Aviv, IL)
Assignee: GOOGLE LLC
H04L51/02G10L13/08G10L15/22G10L2015/225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,615,224
App. No.
18/370,239
Granted
Apr 28, 2026
Kind
B2
Abstract

Implementations are directed to providing a voice wrapper to an existing third-party text-based chatbot to enable the existing third-party text-based chatbot to engage in corresponding voice-based conversations. The voice wrapper can include a plurality of components. For instance, the voice wrapper can include a plurality of input components for utilization in responding to a spoken utterance, and in lieu of the existing third-party text-based chatbot, and/or to modify input to be provided to the existing third-party text-based chatbot in responding to the spoken utterance. Also, for instance, the voice wrapper can include a plurality of output components for utilization in responding to the spoken utterance, to reduce perceived latency of the existing third-party text-based chatbot, and/or to modify output generated by the existing third-party text-based chatbot in responding to the spoken utterance.

Claims (72)

1 . A method implemented by one or more processors, the method comprising:

receiving, from a first-party entity, a voice wrapper for an existing third-party text-based chatbot that is managed by a third-party entity, the third-party entity being distinct from the first-party entity, and the voice wrapper including a plurality of components that enable the existing third-party text-based chatbot to engage in corresponding voice-based conversations with corresponding human users; and

causing the existing third-party text-based chatbot to engage in a given voice-based conversation with a given human user via a client device of the given human user, wherein causing the existing third-party text-based chatbot to engage in the given voice-based conversation with the given human user comprises:

receiving audio data that captures a spoken utterance provided by the given human user, wherein the spoken utterance is an interruption that interrupts a current response that is being audibly rendered for presentation to the given human user via one or more speakers of the client device of the given human user;

determining, based on processing the audio data that captures the spoken utterance and using one or more of the plurality of components of the voice wrapper, whether to utilize the voice wrapper in responding to the spoken utterance or the existing third-party text-based chatbot in responding to the spoken utterance, wherein determining whether to utilize the voice wrapper in responding to the spoken utterance or the existing third-party text-based chatbot in responding to the spoken utterance based on processing the audio data that captures the spoken utterance and using one or more of the plurality of components of the voice wrapper comprises:

processing, using an automatic speech recognition (ASR) component, of the plurality of components, of the voice wrapper, the audio data that captures the spoken utterance to generate ASR output that corresponds to the spoken utterance;

processing, using an interruption component, of the plurality of components, of the voice wrapper, the ASR output, or natural language understanding (NLU) output, that is generated by the voice wrapper based on processing the ASR output, to determine a type of the interruption, from among a plurality of disparate types of interruptions; and

determining, based on the type of interruption, whether to utilize the voice wrapper in responding to the spoken utterance or the existing third-party text-based chatbot in responding to the spoken utterance; and

in response to determining to utilize the voice wrapper in responding to the spoken utterance:

processing, using one or more of the plurality of components of the voice wrapper and without using the existing third-party text-based chatbot, the audio data that captures the spoken utterance to generate a voice wrapper response that is responsive to the spoken utterance; and

causing the voice wrapper response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user.

2 . The method of claim 1 , wherein the plurality of disparate types of interruptions include at least a non-critical interruption and a critical interruption.

3 . The method of claim 2 , wherein determining to utilize the voice wrapper in responding to the spoken utterance is in response to determining that the type of interruption is the non-critical interruption.

4 . The method of claim 3 , wherein processing the audio data that captures the spoken utterance to generate the voice wrapper response that is responsive to the spoken utterance and using one or more of the plurality of components of the voice wrapper and without using the existing third-party text-based chatbot comprises:

processing, using a fulfillment component, of the plurality of components, of the voice wrapper, the ASR output and/or the NLU output, to generate the voice wrapper response that is responsive to the spoken utterance.

5 . The method of claim 4 , wherein causing the voice wrapper response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user comprises:

processing, using a text-to-speech (TTS) component, of the plurality of components, of the voice wrapper, the voice wrapper response to generate synthesized speech audio data that captures synthesized speech corresponding to an audible response for the voice wrapper response;

ceasing the current response from being audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user; and

causing the audible response for the voice wrapper response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user.

6 . The method of claim 2 , wherein determining to utilize the existing third-party text-based chatbot in responding to the spoken utterance is in response to determining that the type of interruption is the critical interruption.

7 . The method of claim 6 , further comprising:

in response to determining to utilize the existing third-party text-based chatbot in responding to the spoken utterance:

processing, using the existing third-party text-based chatbot, the ASR output generated by the voice wrapper to generate an existing third-party text-based chatbot response that is responsive to the spoken utterance;

ceasing the current response from being audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user; and

causing the existing third-party text-based chatbot response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user.

8 . The method of claim 7 , wherein causing the existing third-party text-based chatbot response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user comprises:

processing, using a text-to-speech (TTS) component, of the plurality of components, of the voice wrapper, the existing third-party text-based chatbot response to generate synthesized speech audio data that captures synthesized speech corresponding to an audible response for the existing third-party text-based chatbot response; and

causing the audible response for the existing third-party text-based chatbot response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user.

9 . The method of claim 7 , further comprising:

while the existing third-party text-based chatbot is processing the ASR output generated by the voice wrapper to generate the existing third-party text-based chatbot response that is responsive to the spoken utterance:

determining, using a latency component, from among the plurality of components, of the voice wrapper, to cause pre-cached content to be audibly rendered for presentation to the given human user via one or more speakers of the client device of the given human user; and

causing the pre-cached content to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user.

10 . The method of claim 1 , wherein the plurality of components of the voice wrapper include a plurality of input components and a plurality of output components.

11 . A system comprising:

one or more hardware processors; and

memory storing instructions that, when executed by the one or more hardware processors, the one or more hardware processors are operable to:

receive, from a first-party entity, a voice wrapper for an existing third-party text-based chatbot that is managed by a third-party entity, the third-party entity being distinct from the first-party entity, and the voice wrapper including a plurality of components that enable the existing third-party text-based chatbot to engage in corresponding voice-based conversations with corresponding human users; and

cause the existing third-party text-based chatbot to engage in a given voice-based conversation with a given human user via a client device of the given human user, wherein, in causing the existing third-party text-based chatbot to engage in the given voice-based conversation with the given human user, the one or more hardware processors are operable to:

receive audio data that captures a spoken utterance provided by the given human user;

determine, based on processing the audio data that captures the spoken utterance and using one or more of the plurality of components of the voice wrapper, whether to utilize the voice wrapper in responding to the spoken utterance or the existing third-party text-based chatbot in responding to the spoken utterance, wherein, in determining whether to utilize the voice wrapper in responding to the spoken utterance or the existing third-party text-based chatbot in responding to the spoken utterance based on processing the audio data that captures the spoken utterance and using one or more of the plurality of components of the voice wrapper, the one or more hardware processors are further operable to:

process, using an automatic speech recognition (ASR) component, of the plurality of components, of the voice wrapper, the audio data that captures the spoken utterance to generate ASR output that corresponds to the spoken utterance;

process, using a disambiguation component, of the plurality of components, of the voice wrapper, the ASR output, or natural language understanding (NLU) output, that is generated by the voice wrapper based on processing the ASR output, to determine whether a need exists to disambiguate the spoken utterance; and

determine, based on whether the need exists to disambiguate the spoken utterance, whether to utilize the voice wrapper in responding to the spoken utterance or the existing third-party text-based chatbot in responding to the spoken utterance; and

in response to determining to utilize the voice wrapper in responding to the spoken utterance:

process, using one or more of the plurality of components of the voice wrapper and without using the existing third-party text-based chatbot, the audio data that captures the spoken utterance to generate a voice wrapper response that is responsive to the spoken utterance; and

cause the voice wrapper response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via one or more speakers of the client device of the given human user.

12 . The system of claim 11 , wherein determining to utilize the voice wrapper in responding to the spoken utterance is in response to determining that the need exists to disambiguate the spoken utterance.

13 . The system of claim 12 , wherein, in processing the audio data that captures the spoken utterance to generate the voice wrapper response that is responsive to the spoken utterance and using one or more of the plurality of components of the voice wrapper and without using the existing third-party text-based chatbot, the one or more hardware processors are further operable to:

process, using a fulfillment component, of the plurality of components, of the voice wrapper, the ASR output and/or the NLU output, to generate the voice wrapper response that is responsive to the spoken utterance.

14 . The system of claim 13 , wherein, in causing the voice wrapper response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user, the one or more hardware processors are further operable to:

process, using a text-to-speech (TTS) component, of the plurality of components, of the voice wrapper, the voice wrapper response to generate synthesized speech audio data that captures synthesized speech corresponding to an audible response for the voice wrapper response; and

cause the audible response for the voice wrapper response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user.

15 . The system of claim 11 , wherein determining to utilize the existing third-party text-based chatbot in responding to the spoken utterance is in response to determining that the need does not exist to disambiguate the spoken utterance.

16 . The system of claim 15 , wherein the one or more hardware processors are further operable to:

in response to determining to utilize the existing third-party text-based chatbot in responding to the spoken utterance:

process, using the existing third-party text-based chatbot, the ASR output generated by the voice wrapper to generate an existing third-party text-based chatbot response that is responsive to the spoken utterance; and

cause the existing third-party text-based chatbot response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user.

17 . The system of claim 16 , wherein, in causing the existing third-party text-based chatbot response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user, the one or more hardware processors are further operable to:

process, using a text-to-speech (TTS) component, of the plurality of components, of the voice wrapper, the existing third-party text-based chatbot response to generate synthesized speech audio data that captures synthesized speech corresponding to an audible response for the existing third-party text-based chatbot response; and

cause the audible response for the existing third-party text-based chatbot response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user.

18 . The system of claim 11 , wherein the plurality of components of the voice wrapper include a plurality of input components and a plurality of output components.

19 . A non-transitory computer-readable medium comprising computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform operations to:

receive, from a first-party entity, a voice wrapper for an existing third-party text-based chatbot that is managed by a third-party entity, the third-party entity being distinct from the first-party entity, and the voice wrapper including a plurality of components that enable the existing third-party text-based chatbot to engage in corresponding voice-based conversations with corresponding human users; and

cause the existing third-party text-based chatbot to engage in a given voice-based conversation with a given human user via a client device of the given human user, wherein the operations to cause the existing third-party text-based chatbot to engage in the given voice-based conversation with the given human user comprise operations to:

receive audio data that captures a spoken utterance provided by the given human user, wherein the spoken utterance is an interruption that interrupts a current response that is being audibly rendered for presentation to the given human user via one or more speakers of the client device of the given human user;

determine, based on processing the audio data that captures the spoken utterance and using one or more of the plurality of components of the voice wrapper, whether to utilize the voice wrapper in responding to the spoken utterance or the existing third-party text-based chatbot in responding to the spoken utterance, wherein, in determining whether to utilize the voice wrapper in responding to the spoken utterance or the existing third-party text-based chatbot in responding to the spoken utterance based on processing the audio data that captures the spoken utterance and using one or more of the plurality of components of the voice wrapper, the operations are further to:

process, using an automatic speech recognition (ASR) component, of the plurality of components, of the voice wrapper, the audio data that captures the spoken utterance to generate ASR output that corresponds to the spoken utterance;

process, using an interruption component, of the plurality of components, of the voice wrapper, the ASR output, or natural language understanding (NLU) output, that is generated by the voice wrapper based on processing the ASR output, to determine a type of the interruption, from among a plurality of disparate types of interruptions; and

determine, based on the type of interruption, whether to utilize the voice wrapper in responding to the spoken utterance or the existing third-party text-based chatbot in responding to the spoken utterance; and

in response to determining to utilize the voice wrapper in responding to the spoken utterance:

process, using one or more of the plurality of components of the voice wrapper and without using the existing third-party text-based chatbot, the audio data that captures the spoken utterance to generate a voice wrapper response that is responsive to the spoken utterance; and

cause the voice wrapper response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2023
From: GOLDSHTEIN, SASHA; TZUR, YOAV; FRUCHTER, SHLOMO; MOSHITCH, GAL; TSAI, CHITA; SULTAN, SHARON
To: GOOGLE LLC
Reel/Frame 065308/0135 →
Continuity (1)
Related Publication 20250097168A1 · Mar 20, 2025
References Cited (6)
US 20080147406A1 · Da Palma · 2008 [cited by examiner]
US 20220180874A1 · Pugliese · 2022 [cited by examiner]
WO 2023038654A1 · 2023 [cited by applicant]
European Patent Office; International Search Report and Written Opinion issued in Application No. PCT/US2024/045812; 17 pages; dated Jan. 7, 2025. [cited by applicant]
Wang, X. et al., “SwitchGPT: Adapting Large Language Models for Non-Text Outputs”; arXiv.org, Cornell University; arXiv:2309.07623; 15 pages; dated Sep. 14, 2023. [cited by applicant]
European Patent Office; Invitation to Pay Additional Fees issued in Application No. PCT/US2024/045812; 10 pages; dated Nov. 5, 2024. [cited by applicant]