IP Library › Granted Patent US 12,744,053
Granted Patent B1
US 12,744,053 · App. 19/256,715 · Granted Sep 22, 2026

Real time voice mode generative response engine with turn detection

Inventors: Wayne Chang (Menlo Park, CA); Mianna Chen (San Francisco, CA); Bogumil Kazimierz Giertler (San Francisco, CA); Tomer Kaftan (San Francisco, CA); Alex Kirillov (San Francisco, CA); Edede Oiwoh (San Francisco, CA); Olaoluwa Okelola (San Francisco, CA); Raul Puri (San Francisco, CA); Brendan Quinn (San Francisco, CA); Ievgen Varavva (San Mateo, CA); Tao Xu (Palo Alto, CA); Rowan Zellers (San Francisco, CA); Yu Zhang (Mountain View, CA)
Assignee: OpenAI OpCo, LLC
G10L25/78G10L15/22G10L15/04G10L15/16G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,744,053
App. No.
19/256,715
Granted
Sep 22, 2026
Kind
B1
Abstract

Disclosed are systems, apparatuses, processes, and computer-readable media for real-time modes of a generative response engine. The present technology includes different technologies to improve real-time operation to create natural and dynamic interactions to user experiences. A method of the present technology includes receiving audio frames from a client device; determining that a first portion of the audio frames includes speech of a user during a speech instance; transmitting the first portion of the audio frames that includes the speech to a generative response engine; determining that the speech instance has concluded; and sending a signal to the generative response engine to perform an inference operation based on the first portion of the audio frames.

Claims (48)

1 . A computing device associated with a cloud computing service, comprising:

at least one network interface; and

at least one processor coupled to the at least one network interface and configured to:

receive digital audio frames from a client device, wherein a respective digital audio frame has a fixed duration;

determine that a first portion of the digital audio frames includes speech of a user during a speech instance;

transmit the first portion of the digital audio frames that includes the speech to a generative response engine while the digital audio frames are being received;

when a single digital audio frame is identified as being silent, determine, using a conversation classifier, that the speech instance has concluded based on the single digital audio frame being silent and the first portion of the digital audio frames corresponding to an independent clause, wherein the conversation classifier is configured to identify the independent clause based on accumulated audio frames during the speech instance; and

send a start signal to the generative response engine indicating to begin an inference operation associated with the speech based on the first portion of the digital audio frames and the single digital audio frame being silent.

2 . The computing device of claim 1 , wherein the at least one processor is configured to:

prior to the determination that the speech instance has concluded, determine that the first portion of the digital audio frames does not correspond to an end to the speech instance; and

transmit a second portion of the digital audio frames to the generative response engine, wherein the determination that the speech instance has concluded is based on the first portion of the digital audio frames and the second portion of the digital audio frames, and the inference operation is performed based on the first portion of the digital audio frames and the second portion of the digital audio frames.

3 . The computing device of claim 1 , wherein transmitting of the first portion of the digital audio frames begins prior to the determination that the speech instance has concluded.

4 . The computing device of claim 1 , wherein the determination that the first portion of the digital audio frames includes speech is performed by a speech detector that is configured to detect a presence of speech in the digital audio frames, and the determination that the speech instance has concluded is performed by the conversation classifier that is configured to determine whether the digital audio frames include speech that indicates an end of the speech instance.

5 . The computing device of claim 1 , wherein the at least one processor is configured to:

receive a second portion of the digital audio frames that include speech after the start signal to perform the inference operation has been sent; and

transmit a cancellation signal to the generative response engine to cancel the inference operation based on the first portion of the digital audio frames.

6 . The computing device of claim 5 , wherein the at least one processor is configured to:

receive a response from performance of the inference operation based on the first portion of the digital audio frames from the generative response engine while receiving the second portion of the digital audio frames; and

discarding the response.

7 . The computing device of claim 1 , wherein the at least one processor is configured to:

receive a response from performance of the inference operation based on the first portion of the digital audio frames from the generative response engine; and

send the response to the client device.

8 . The computing device of claim 7 , wherein the at least one processor is configured to:

determine a first digital audio frame of the digital audio frames after the first portion of the digital audio frames corresponds to silence, wherein the first digital audio frame includes a nonce word.

9 . The computing device of claim 1 , wherein the conversation classifier is configured to indicate a second portion of the digital audio frames before the first portion of the digital audio frames correspond to an incomplete clause and the first portion of the digital audio frames and the second portion of the digital audio frames correspond to the independent clause.

10 . A method for detecting speech instances, comprising:

receiving digital audio frames from a client device, wherein a respective digital audio frame has a fixed duration;

determining that a first portion of the digital audio frames includes speech of a user during a speech instance;

transmitting the first portion of the digital audio frames that includes the speech to a generative response engine while the digital audio frames are being received;

when a single digital audio frame is identified as being silent, determining, using a conversation classifier, that the speech instance has concluded based on the single digital audio frame being silent and the first portion of the digital audio frames corresponding to an independent clause, wherein the conversation classifier is configured to identify the independent clause based on accumulated digital audio frames during the speech instance; and

sending a start signal to the generative response engine indicating to begin an inference operation associated with the speech based on the first portion of the digital audio frames and the single digital audio frame being silent.

11 . The method of claim 10 , further comprising:

prior to the determination that the speech instance has concluded, determining that the first portion of the digital audio frames does not correspond to an end to the speech instance; and

transmitting a second portion of the digital audio frames to the generative response engine, wherein the determination that the speech instance has concluded is based on the first portion of the digital audio frames and the second portion of the digital audio frames, and the inference operation is performed based on the first portion of the digital audio frames and the second portion of the digital audio frames.

12 . The method of claim 10 , wherein transmitting of the first portion of the digital audio frames begins prior to the determination that the speech instance has concluded.

13 . The method of claim 10 , wherein the determination that the first portion of the digital audio frames includes speech is performed by a speech detector that is configured to detect a presence of speech in the digital audio frames, and the determination that the speech instance has concluded is performed by the conversation classifier that is configured to determine whether the digital audio frames include speech that indicates an end of the speech instance.

14 . The method of claim 10 , further comprising:

receiving a second portion of the digital audio frames that include speech after the start signal to perform the inference operation has been sent; and

transmitting a cancellation signal to the generative response engine to cancel the inference operation based on the first portion of the digital audio frames.

15 . The method of claim 14 , further comprising:

receiving a response from performance of the inference operation based on the first portion of the digital audio frames from the generative response engine while receiving the second portion of the digital audio frames; and

discarding the response.

16 . The method of claim 10 , further comprising:

receiving a response from performance of the inference operation based on the first portion of the digital audio frames from the generative response engine; and

sending the response to the client device.

17 . The method of claim 16 , further comprising:

determining a first digital audio frame of the digital audio frames after the first portion of the digital audio frames corresponds to silence, wherein the first digital audio frame includes a nonce word.

18 . The method of claim 10 , wherein the conversation classifier is configured to indicate a second portion of the digital audio frames e-before the first portion of the digital audio frames correspond to an incomplete clause and the first portion of the digital audio frames and the second portion of the digital audio frames correspond to the independent clause.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2026
From: CHANG, WAYNE; CHEN, MIANNA; GIERTLER, BOGUMIL KAZIMIERZ; KAFTAN, TOMER; KIRILLOV, ALEX; OIWOH, EDEDE; OKELOLA, OLAOLUWA; PURI, RAUL; QUINN, BRENDAN; VARAVVA, IEVGEN; XU, TAO; ZELLERS, ROWAN; ZHANG, YU
To: OPENAI OPCO, LLC
Reel/Frame 074806/0516 →
Continuity (1)
Continuation 19098690 · Apr 2, 2025
References Cited (35)
US 6496799B1 · Pickering · 2002 [cited by applicant]
US 6882973B1 · Pickering · 2005 [cited by applicant]
US 7437286B2 · Pi et al. · 2008 [cited by applicant]
US 10121471B2 · Hoffmeister et al. · 2018 [cited by applicant]
US 11615239B2 · Brdiczka · 2023 [cited by examiner]
US 12243517B1 · Mehrabani · 2025 [cited by examiner]
US 12322384B1 · Ramachandran · 2025 [cited by examiner]
US 20160316059A1 · Nuta · 2016 [cited by examiner]
US 20160358598A1 · Williams · 2016 [cited by examiner]
US 20180330723A1 · Acero · 2018 [cited by examiner]
US 20190318759A1 · Doshi · 2019 [cited by examiner]
US 20200342874A1 · Teserra · 2020 [cited by examiner]
US 20230053341A1 · Konzelmann · 2023 [cited by examiner]
US 20230214597A1 · Lin · 2023 [cited by examiner]
US 20230230578A1 · Sharifi · 2023 [cited by examiner]
US 20230230585A1 · Anderson · 2023 [cited by examiner]
US 20230368781A1 · Choi · 2023 [cited by examiner]
US 20240312460A1 · Konzelmann · 2024 [cited by examiner]
US 20250094456A1 · Gu · 2025 [cited by examiner]
CN 117253485A · 2023 [cited by examiner]
Pinto, et al. “Predictive turn-taking: Leveraging language models to anticipate turn transitions in human-robot dialogue.” 2024 33rd IEEE International Conference on Robot and Human Interactive Communication (ROMAN). IE… [cited by examiner]
Alexander Veysov and Dimitrii Voronin, “One Voice Detector to Rule Them All”, The Gradient, url: <https://thegradient.pub/one-voice-detector-to-rule-them-all/%3E>, Feb. 19, 2022 (22 pages). [cited by applicant]
Domingo, Enric, “Claude 3.5 Sonnet API: This is howyou integrate the best LLM into yourAPP”, Medium, url: <https://medium.com/@enricdomingo/claude-sonnet-3-5-api-integrating-the-best-llm-into-our-app-7ec4623e2dac%3E>, J… [cited by applicant]
“Server-Sent Events: A WebSockets alternative ready for another look”, ably, url: <https://ably.com/topic/server-sent-events%3E>, Jun. 28, 2023 {7 pages). [cited by applicant]
“Server-Sent Events W3C Working Draft Apr. 23, 2009”, Ed. Ian Hickson, W3C, url: <https://www.w3.org/TR/2009/WD-eventsource-20090423/%3E>, Apr. 23, 2009 (15 pages). [cited by applicant]
“ETSI TS 126 071”, Digital cellular telecommunications system (Phase 2+) (GSM); Universal Mobile Telecommunications System (UMTS); LTE; Mandatory speech CODEC speech processing functions; AMR speech Codec; General descr… [cited by applicant]
“Genesys Audio Connector”, Genesys Cloud CX—Beta HQ, url: <https://community.genesys.com/discussion/genesys-audio-connector%3E>, Dec. 1, 2023 (18 pages). [cited by applicant]
“Genesys Cloud—Feb. 10, 2025”, Genesys Cloud Release Notes, Genesys Cloud Resource Center, url: <https://help.mypurecloud.com/releasenote/february-10-2025/%3E>, Feb. 10, 2025 (4 pages). [cited by applicant]
“Genesys Cloud—Jan. 13, 2025”, Genesys Cloud Release Notes, Genesys Cloud Resource Center, url: <https://help.mypurecloud.com/releasenote/january-13-2025/%3E>, Jan. 13, 2025 (3 pages). [cited by applicant]
“Genesys Cloud—Jan. 27, 2025”, Genesys Cloud Release Notes, Genesys Cloud Resource Center, url: <https://help.mypurecloud.com/releasenote/january-27-2025/%3E>, Jan. 27, 2025 (4 pages). [cited by applicant]
Mishra, Anubhav and Jessica Ho, “Enhancing customer service experiences using Conversational AI: Power your contact center with Amazon Lex and Genesys Cloud”, AWS Blogs, url: <https://aws.amazon.com/blogs/machine-learni… [cited by applicant]
Sharma, Shilpa, et al., “Voice Activity Detection”, International Journal of Creative Research Thoughts: An International Open Access, Peer Reviewed, Refereed Journal, vol. 9, Issue 5, May 2021 (4 pages). [cited by applicant]
Sjoberg, J. et al., “Real-Time Transport Protocol (RTP) Payload Format and File Storage Format for the Adaptive Multi-Rate (AMR) and Adaptive Multi-Rate Wideband (AMR-WB) Audio Codecs”, Network Working Group Request for… [cited by applicant]
Vardhini, Lakshmi, “Dialogflow CX for Effortless Chatbot Creation and Testing with an Intuitive UI”, cloudthat, url: <https://www.cloudthat.com/resources/blog/dialogflow-cx-for-effortless-chatbot-creation-and-testing-wi… [cited by applicant]
“Web Speech API Specification”, Eds. Glen Shires and Hans Wennborg, Contributors to the Web Speech API Specification, Speech API Community Group, url: <https://dvcs.w3.org/hg/speech-api/raw-file/tip/webspeechapi%3E>, Ju… [cited by applicant]