Real time voice mode generative response engine with turn detection
Disclosed are systems, apparatuses, processes, and computer-readable media for real-time modes of a generative response engine. The present technology includes different technologies to improve real-time operation to create natural and dynamic interactions to user experiences. A method of the present technology includes receiving audio frames from a client device; determining that a first portion of the audio frames includes speech of a user during a speech instance; transmitting the first portion of the audio frames that includes the speech to a generative response engine; determining that the speech instance has concluded; and sending a signal to the generative response engine to perform an inference operation based on the first portion of the audio frames.
1 . A computing device associated with a cloud computing service, comprising:
at least one network interface; and
at least one processor coupled to the at least one network interface and configured to:
receive digital audio frames from a client device, wherein a respective digital audio frame has a fixed duration;
determine that a first portion of the digital audio frames includes speech of a user during a speech instance;
transmit the first portion of the digital audio frames that includes the speech to a generative response engine while the digital audio frames are being received;
when a single digital audio frame is identified as being silent, determine, using a conversation classifier, that the speech instance has concluded based on the single digital audio frame being silent and the first portion of the digital audio frames corresponding to an independent clause, wherein the conversation classifier is configured to identify the independent clause based on accumulated audio frames during the speech instance; and
send a start signal to the generative response engine indicating to begin an inference operation associated with the speech based on the first portion of the digital audio frames and the single digital audio frame being silent.
2 . The computing device of claim 1 , wherein the at least one processor is configured to:
prior to the determination that the speech instance has concluded, determine that the first portion of the digital audio frames does not correspond to an end to the speech instance; and
transmit a second portion of the digital audio frames to the generative response engine, wherein the determination that the speech instance has concluded is based on the first portion of the digital audio frames and the second portion of the digital audio frames, and the inference operation is performed based on the first portion of the digital audio frames and the second portion of the digital audio frames.
3 . The computing device of claim 1 , wherein transmitting of the first portion of the digital audio frames begins prior to the determination that the speech instance has concluded.
4 . The computing device of claim 1 , wherein the determination that the first portion of the digital audio frames includes speech is performed by a speech detector that is configured to detect a presence of speech in the digital audio frames, and the determination that the speech instance has concluded is performed by the conversation classifier that is configured to determine whether the digital audio frames include speech that indicates an end of the speech instance.
5 . The computing device of claim 1 , wherein the at least one processor is configured to:
receive a second portion of the digital audio frames that include speech after the start signal to perform the inference operation has been sent; and
transmit a cancellation signal to the generative response engine to cancel the inference operation based on the first portion of the digital audio frames.
6 . The computing device of claim 5 , wherein the at least one processor is configured to:
receive a response from performance of the inference operation based on the first portion of the digital audio frames from the generative response engine while receiving the second portion of the digital audio frames; and
discarding the response.
7 . The computing device of claim 1 , wherein the at least one processor is configured to:
receive a response from performance of the inference operation based on the first portion of the digital audio frames from the generative response engine; and
send the response to the client device.
8 . The computing device of claim 7 , wherein the at least one processor is configured to:
determine a first digital audio frame of the digital audio frames after the first portion of the digital audio frames corresponds to silence, wherein the first digital audio frame includes a nonce word.
9 . The computing device of claim 1 , wherein the conversation classifier is configured to indicate a second portion of the digital audio frames before the first portion of the digital audio frames correspond to an incomplete clause and the first portion of the digital audio frames and the second portion of the digital audio frames correspond to the independent clause.
10 . A method for detecting speech instances, comprising:
receiving digital audio frames from a client device, wherein a respective digital audio frame has a fixed duration;
determining that a first portion of the digital audio frames includes speech of a user during a speech instance;
transmitting the first portion of the digital audio frames that includes the speech to a generative response engine while the digital audio frames are being received;
when a single digital audio frame is identified as being silent, determining, using a conversation classifier, that the speech instance has concluded based on the single digital audio frame being silent and the first portion of the digital audio frames corresponding to an independent clause, wherein the conversation classifier is configured to identify the independent clause based on accumulated digital audio frames during the speech instance; and
sending a start signal to the generative response engine indicating to begin an inference operation associated with the speech based on the first portion of the digital audio frames and the single digital audio frame being silent.
11 . The method of claim 10 , further comprising:
prior to the determination that the speech instance has concluded, determining that the first portion of the digital audio frames does not correspond to an end to the speech instance; and
transmitting a second portion of the digital audio frames to the generative response engine, wherein the determination that the speech instance has concluded is based on the first portion of the digital audio frames and the second portion of the digital audio frames, and the inference operation is performed based on the first portion of the digital audio frames and the second portion of the digital audio frames.
12 . The method of claim 10 , wherein transmitting of the first portion of the digital audio frames begins prior to the determination that the speech instance has concluded.
13 . The method of claim 10 , wherein the determination that the first portion of the digital audio frames includes speech is performed by a speech detector that is configured to detect a presence of speech in the digital audio frames, and the determination that the speech instance has concluded is performed by the conversation classifier that is configured to determine whether the digital audio frames include speech that indicates an end of the speech instance.
14 . The method of claim 10 , further comprising:
receiving a second portion of the digital audio frames that include speech after the start signal to perform the inference operation has been sent; and
transmitting a cancellation signal to the generative response engine to cancel the inference operation based on the first portion of the digital audio frames.
15 . The method of claim 14 , further comprising:
receiving a response from performance of the inference operation based on the first portion of the digital audio frames from the generative response engine while receiving the second portion of the digital audio frames; and
discarding the response.
16 . The method of claim 10 , further comprising:
receiving a response from performance of the inference operation based on the first portion of the digital audio frames from the generative response engine; and
sending the response to the client device.
17 . The method of claim 16 , further comprising:
determining a first digital audio frame of the digital audio frames after the first portion of the digital audio frames corresponds to silence, wherein the first digital audio frame includes a nonce word.
18 . The method of claim 10 , wherein the conversation classifier is configured to indicate a second portion of the digital audio frames e-before the first portion of the digital audio frames correspond to an incomplete clause and the first portion of the digital audio frames and the second portion of the digital audio frames correspond to the independent clause.