Voice wrapper(s) for existing third-party text-based chatbot(s)
Implementations are directed to providing a voice wrapper to an existing third-party text-based chatbot to enable the existing third-party text-based chatbot to engage in corresponding voice-based conversations. The voice wrapper can include a plurality of components. For instance, the voice wrapper can include a plurality of input components for utilization in responding to a spoken utterance, and in lieu of the existing third-party text-based chatbot, and/or to modify input to be provided to the existing third-party text-based chatbot in responding to the spoken utterance. Also, for instance, the voice wrapper can include a plurality of output components for utilization in responding to the spoken utterance, to reduce perceived latency of the existing third-party text-based chatbot, and/or to modify output generated by the existing third-party text-based chatbot in responding to the spoken utterance.
1 . A method implemented by one or more processors, the method comprising:
receiving, from a first-party entity, a voice wrapper for an existing third-party text-based chatbot that is managed by a third-party entity, the third-party entity being distinct from the first-party entity, and the voice wrapper including a plurality of components that enable the existing third-party text-based chatbot to engage in corresponding voice-based conversations with corresponding human users; and
causing the existing third-party text-based chatbot to engage in a given voice-based conversation with a given human user via a client device of the given human user, wherein causing the existing third-party text-based chatbot to engage in the given voice-based conversation with the given human user comprises:
receiving audio data that captures a spoken utterance provided by the given human user, wherein the spoken utterance is an interruption that interrupts a current response that is being audibly rendered for presentation to the given human user via one or more speakers of the client device of the given human user;
determining, based on processing the audio data that captures the spoken utterance and using one or more of the plurality of components of the voice wrapper, whether to utilize the voice wrapper in responding to the spoken utterance or the existing third-party text-based chatbot in responding to the spoken utterance, wherein determining whether to utilize the voice wrapper in responding to the spoken utterance or the existing third-party text-based chatbot in responding to the spoken utterance based on processing the audio data that captures the spoken utterance and using one or more of the plurality of components of the voice wrapper comprises:
processing, using an automatic speech recognition (ASR) component, of the plurality of components, of the voice wrapper, the audio data that captures the spoken utterance to generate ASR output that corresponds to the spoken utterance;
processing, using an interruption component, of the plurality of components, of the voice wrapper, the ASR output, or natural language understanding (NLU) output, that is generated by the voice wrapper based on processing the ASR output, to determine a type of the interruption, from among a plurality of disparate types of interruptions; and
determining, based on the type of interruption, whether to utilize the voice wrapper in responding to the spoken utterance or the existing third-party text-based chatbot in responding to the spoken utterance; and
in response to determining to utilize the voice wrapper in responding to the spoken utterance:
processing, using one or more of the plurality of components of the voice wrapper and without using the existing third-party text-based chatbot, the audio data that captures the spoken utterance to generate a voice wrapper response that is responsive to the spoken utterance; and
causing the voice wrapper response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user.
2 . The method of claim 1 , wherein the plurality of disparate types of interruptions include at least a non-critical interruption and a critical interruption.
3 . The method of claim 2 , wherein determining to utilize the voice wrapper in responding to the spoken utterance is in response to determining that the type of interruption is the non-critical interruption.
4 . The method of claim 3 , wherein processing the audio data that captures the spoken utterance to generate the voice wrapper response that is responsive to the spoken utterance and using one or more of the plurality of components of the voice wrapper and without using the existing third-party text-based chatbot comprises:
processing, using a fulfillment component, of the plurality of components, of the voice wrapper, the ASR output and/or the NLU output, to generate the voice wrapper response that is responsive to the spoken utterance.
5 . The method of claim 4 , wherein causing the voice wrapper response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user comprises:
processing, using a text-to-speech (TTS) component, of the plurality of components, of the voice wrapper, the voice wrapper response to generate synthesized speech audio data that captures synthesized speech corresponding to an audible response for the voice wrapper response;
ceasing the current response from being audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user; and
causing the audible response for the voice wrapper response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user.
6 . The method of claim 2 , wherein determining to utilize the existing third-party text-based chatbot in responding to the spoken utterance is in response to determining that the type of interruption is the critical interruption.
7 . The method of claim 6 , further comprising:
in response to determining to utilize the existing third-party text-based chatbot in responding to the spoken utterance:
processing, using the existing third-party text-based chatbot, the ASR output generated by the voice wrapper to generate an existing third-party text-based chatbot response that is responsive to the spoken utterance;
ceasing the current response from being audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user; and
causing the existing third-party text-based chatbot response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user.
8 . The method of claim 7 , wherein causing the existing third-party text-based chatbot response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user comprises:
processing, using a text-to-speech (TTS) component, of the plurality of components, of the voice wrapper, the existing third-party text-based chatbot response to generate synthesized speech audio data that captures synthesized speech corresponding to an audible response for the existing third-party text-based chatbot response; and
causing the audible response for the existing third-party text-based chatbot response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user.
9 . The method of claim 7 , further comprising:
while the existing third-party text-based chatbot is processing the ASR output generated by the voice wrapper to generate the existing third-party text-based chatbot response that is responsive to the spoken utterance:
determining, using a latency component, from among the plurality of components, of the voice wrapper, to cause pre-cached content to be audibly rendered for presentation to the given human user via one or more speakers of the client device of the given human user; and
causing the pre-cached content to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user.
10 . The method of claim 1 , wherein the plurality of components of the voice wrapper include a plurality of input components and a plurality of output components.
11 . A system comprising:
one or more hardware processors; and
memory storing instructions that, when executed by the one or more hardware processors, the one or more hardware processors are operable to:
receive, from a first-party entity, a voice wrapper for an existing third-party text-based chatbot that is managed by a third-party entity, the third-party entity being distinct from the first-party entity, and the voice wrapper including a plurality of components that enable the existing third-party text-based chatbot to engage in corresponding voice-based conversations with corresponding human users; and
cause the existing third-party text-based chatbot to engage in a given voice-based conversation with a given human user via a client device of the given human user, wherein, in causing the existing third-party text-based chatbot to engage in the given voice-based conversation with the given human user, the one or more hardware processors are operable to:
receive audio data that captures a spoken utterance provided by the given human user;
determine, based on processing the audio data that captures the spoken utterance and using one or more of the plurality of components of the voice wrapper, whether to utilize the voice wrapper in responding to the spoken utterance or the existing third-party text-based chatbot in responding to the spoken utterance, wherein, in determining whether to utilize the voice wrapper in responding to the spoken utterance or the existing third-party text-based chatbot in responding to the spoken utterance based on processing the audio data that captures the spoken utterance and using one or more of the plurality of components of the voice wrapper, the one or more hardware processors are further operable to:
process, using an automatic speech recognition (ASR) component, of the plurality of components, of the voice wrapper, the audio data that captures the spoken utterance to generate ASR output that corresponds to the spoken utterance;
process, using a disambiguation component, of the plurality of components, of the voice wrapper, the ASR output, or natural language understanding (NLU) output, that is generated by the voice wrapper based on processing the ASR output, to determine whether a need exists to disambiguate the spoken utterance; and
determine, based on whether the need exists to disambiguate the spoken utterance, whether to utilize the voice wrapper in responding to the spoken utterance or the existing third-party text-based chatbot in responding to the spoken utterance; and
in response to determining to utilize the voice wrapper in responding to the spoken utterance:
process, using one or more of the plurality of components of the voice wrapper and without using the existing third-party text-based chatbot, the audio data that captures the spoken utterance to generate a voice wrapper response that is responsive to the spoken utterance; and
cause the voice wrapper response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via one or more speakers of the client device of the given human user.
12 . The system of claim 11 , wherein determining to utilize the voice wrapper in responding to the spoken utterance is in response to determining that the need exists to disambiguate the spoken utterance.
13 . The system of claim 12 , wherein, in processing the audio data that captures the spoken utterance to generate the voice wrapper response that is responsive to the spoken utterance and using one or more of the plurality of components of the voice wrapper and without using the existing third-party text-based chatbot, the one or more hardware processors are further operable to:
process, using a fulfillment component, of the plurality of components, of the voice wrapper, the ASR output and/or the NLU output, to generate the voice wrapper response that is responsive to the spoken utterance.
14 . The system of claim 13 , wherein, in causing the voice wrapper response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user, the one or more hardware processors are further operable to:
process, using a text-to-speech (TTS) component, of the plurality of components, of the voice wrapper, the voice wrapper response to generate synthesized speech audio data that captures synthesized speech corresponding to an audible response for the voice wrapper response; and
cause the audible response for the voice wrapper response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user.
15 . The system of claim 11 , wherein determining to utilize the existing third-party text-based chatbot in responding to the spoken utterance is in response to determining that the need does not exist to disambiguate the spoken utterance.
16 . The system of claim 15 , wherein the one or more hardware processors are further operable to:
in response to determining to utilize the existing third-party text-based chatbot in responding to the spoken utterance:
process, using the existing third-party text-based chatbot, the ASR output generated by the voice wrapper to generate an existing third-party text-based chatbot response that is responsive to the spoken utterance; and
cause the existing third-party text-based chatbot response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user.
17 . The system of claim 16 , wherein, in causing the existing third-party text-based chatbot response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user, the one or more hardware processors are further operable to:
process, using a text-to-speech (TTS) component, of the plurality of components, of the voice wrapper, the existing third-party text-based chatbot response to generate synthesized speech audio data that captures synthesized speech corresponding to an audible response for the existing third-party text-based chatbot response; and
cause the audible response for the existing third-party text-based chatbot response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user.
18 . The system of claim 11 , wherein the plurality of components of the voice wrapper include a plurality of input components and a plurality of output components.
19 . A non-transitory computer-readable medium comprising computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform operations to:
receive, from a first-party entity, a voice wrapper for an existing third-party text-based chatbot that is managed by a third-party entity, the third-party entity being distinct from the first-party entity, and the voice wrapper including a plurality of components that enable the existing third-party text-based chatbot to engage in corresponding voice-based conversations with corresponding human users; and
cause the existing third-party text-based chatbot to engage in a given voice-based conversation with a given human user via a client device of the given human user, wherein the operations to cause the existing third-party text-based chatbot to engage in the given voice-based conversation with the given human user comprise operations to:
receive audio data that captures a spoken utterance provided by the given human user, wherein the spoken utterance is an interruption that interrupts a current response that is being audibly rendered for presentation to the given human user via one or more speakers of the client device of the given human user;
determine, based on processing the audio data that captures the spoken utterance and using one or more of the plurality of components of the voice wrapper, whether to utilize the voice wrapper in responding to the spoken utterance or the existing third-party text-based chatbot in responding to the spoken utterance, wherein, in determining whether to utilize the voice wrapper in responding to the spoken utterance or the existing third-party text-based chatbot in responding to the spoken utterance based on processing the audio data that captures the spoken utterance and using one or more of the plurality of components of the voice wrapper, the operations are further to:
process, using an automatic speech recognition (ASR) component, of the plurality of components, of the voice wrapper, the audio data that captures the spoken utterance to generate ASR output that corresponds to the spoken utterance;
process, using an interruption component, of the plurality of components, of the voice wrapper, the ASR output, or natural language understanding (NLU) output, that is generated by the voice wrapper based on processing the ASR output, to determine a type of the interruption, from among a plurality of disparate types of interruptions; and
determine, based on the type of interruption, whether to utilize the voice wrapper in responding to the spoken utterance or the existing third-party text-based chatbot in responding to the spoken utterance; and
in response to determining to utilize the voice wrapper in responding to the spoken utterance:
process, using one or more of the plurality of components of the voice wrapper and without using the existing third-party text-based chatbot, the audio data that captures the spoken utterance to generate a voice wrapper response that is responsive to the spoken utterance; and
cause the voice wrapper response that is responsive to the spoken utterance to be audibly rendered for presentation to the given human user via the one or more speakers of the client device of the given human user.