IP Library Granted Patent US 12,562,153
Granted Patent B2
US 12,562,153 · App. 17/579,059 · Granted Feb 24, 2026

Contemporaneous machine-learning analysis of audio streams

Inventors: Tianlin Shi (Menlo Park, CA); Kenneth George Oetzel (Belmont, CA)
Assignee: CRESTA INTELLIGENCE INC.
G10L15/16G06F21/6227G06N20/00G10L15/063G10L15/22G10L25/51G10L2015/088H04M3/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,562,153
App. No.
17/579,059
Granted
Feb 24, 2026
Kind
B2
Abstract

Described techniques select portions of an audio stream for transmission to a trained machine learning application, which generates response recommendations in real-time. This real-time response is facilitated by the system identifying, selecting and transmitting those portions of the audio stream likely to be most relevant to the conversation. Portions of an audio stream less likely to be relevant to the conversation are identified accordingly and not transmitted. The system may identify the relevant portions of an audio stream by detecting events in a contemporaneous event stream, use a trained machine learning model to identify events in an audio stream, or both.

Claims (80)

1 . One or more non-transitory computer-readable media storing instructions, which when executed by one or more hardware processors, cause performance of operations comprising:

enabling a real-time communication between a first communication device and a second communication device;

based on enabling the real-time communication, generating (a) an operating system audio stream corresponding to an audio record of the real-time communication and (b) an event stream recording a plurality of event descriptors corresponding to events detected by one or more applications executing on the first communication device concurrently with the operating system audio stream,

wherein the operating system audio stream comprises one or more of a first set of audio signals detected by a microphone associated with the first communication device or a second set of audio signals played by a speaker associated with the first communication device;

detecting a first event in the event stream at least by detecting a first set of one or more descriptors in the event stream;

responsive at least to determining the first event meets recommendation criteria:

generating a first marker in the operating system audio stream contemporaneous with the first event in the event stream;

using the first marker in the operating system audio stream, extracting a first portion of the operating system audio stream associated with the first event; and

transmitting the first portion of the operating system audio stream to a recommendation model; and

responsive to transmitting the first portion of the operating system audio stream to the recommendation model:

receiving, from the recommendation model, a recommendation for a particular operation to be performed on at least one device responsive to the first event; and

presenting, in a graphical user interface (GUI) displayed on the first communication device, the recommendation for the particular operation responsive to the first event.

2 . The one or more or more non-transitory computer-readable media of claim 1 , wherein the operations further comprise:

synchronizing the operating system audio stream with the event stream by simultaneously initiating the operating system audio stream and the event stream.

3 . The one or more or more non-transitory computer-readable media of claim 1 , wherein generating the first marker in the operating system audio stream comprises generating the first marker corresponding to a start time of the first event,

wherein the operations further comprise generating a second marker corresponding to an end time of the first event, and

wherein extracting the first portion of the operating system audio stream associated with the first event comprises extracting a portion of the operating system audio stream between the first marker and the second marker.

4 . The one or more or more non-transitory computer-readable media of claim 1 , wherein extracting the first portion of the operating system audio stream associated with the first event comprises extracting a portion of the operating system audio stream beginning at the first marker and including a predetermined duration of time subsequent to the first marker.

5 . The one or more or more non-transitory computer-readable media of claim 1 , wherein extracting the first portion of the operating system audio stream associated with the first event comprises extracting a portion of the operating system audio stream beginning at the first marker and including a predetermined duration of time prior to the first marker.

6 . The one or more or more non-transitory computer-readable media of claim 1 , wherein the event stream comprises an event stream associated with an Internet telephony application.

7 . The one or more or more non-transitory computer-readable media of claim 1 , wherein the event stream comprises a plurality of event streams from a corresponding plurality of applications, including at least:

a first event stream associated with an Internet telephony application; and

a second event stream associated with an information resource application.

8 . The one or more or more non-transitory computer-readable media of claim 1 , wherein transmitting the first portion of the operating system audio stream comprises transmitting the first portion of the operating system audio stream to a machine learning model, and

wherein the operations further comprise:

generating, by the machine learning model, the recommendation for the particular operation based on content contained in the first portion of the operating system audio stream.

9 . The one or more or more non-transitory computer-readable media of claim 1 , wherein the operations further comprise:

applying a set of rules to identify the first portion of the operating system audio stream associated with the first event to be transmitted.

10 . The one or more or more non-transitory computer-readable media of claim 1 , wherein the operations further comprise:

applying a rule to associate a call type and transmission profile with the first portion of the operating system audio stream; and

transmitting the first portion of the operation system audio stream with the call type and the transmission profile.

11 . The one or more or more non-transitory computer-readable media of claim 1 , wherein the operations further comprise:

generating the first marker in the operating system audio stream is further comprises:

identifying the first event as a particular event type;

applying a particular rule specifying the particular event type and a particular marker associated with the particular event type; and

generating the first marker based on (a) detecting the first event in the event stream, and (b) detecting the first event as the particular event type.

12 . A method comprising:

enabling a real-time communication between a first communication device and a second communication device;

based on enabling the real-time communication, generating (a) an operating system audio stream corresponding to an audio record of the real-time communication and (b) an event stream recording a plurality of event descriptors corresponding to events detected by one or more applications executing on the first communication device concurrently with the operating system audio stream,

wherein the operating system audio stream comprises one or more of a first set of audio signals detected by a microphone associated with the first communication device or a second set of audio signals played by a speaker associated with the first communication device;

detecting a first event in the event stream at least by detecting a first set of one or more descriptors in the event stream;

responsive at least to determining the first event meets recommendation criteria:

generating a first marker in the operating system audio stream contemporaneous with the first event in the event stream;

using the first marker in the operating system audio stream, extracting a first portion of the operating system audio stream associated with the first event; and

transmitting the first portion of the operating system audio stream to a recommendation model; and

responsive to transmitting the first portion of the operating system audio stream to the recommendation model:

receiving, from the recommendation model, a recommendation for a particular operation to be performed on at least one device responsive to the first event; and

presenting, in a graphical user interface (GUI) displayed on the first communication device, the recommendation for the particular operation responsive to the first event.

13 . The method of claim 12 , further comprising:

synchronizing the operating system audio stream with the event stream by simultaneously initiating the operating system audio stream and the event stream.

14 . The method of claim 12 , wherein generating the first marker in the operating system audio stream comprises generating the first marker corresponding to a start time of the first event,

wherein the method further comprises generating a second marker corresponding to an end time of the first event, and

wherein extracting the first portion of the operating system audio stream associated with the first event comprises extracting a portion of the operating system audio stream between the first marker and the second marker.

15 . The method of claim 12 , wherein extracting the first portion of the operating system audio stream associated with the first event comprises extracting a portion of the operating system audio stream beginning at the first marker and including a predetermined duration of time subsequent to the first marker.

16 . The method of claim 12 , wherein extracting the first portion of the operating system audio stream associated with the first event comprises extracting a portion of the operating system audio stream beginning at the first marker and including a predetermined duration of time prior to the first marker.

17 . The method of claim 12 , wherein the event stream comprises an event stream associated with an Internet telephony application.

18 . A system comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the system to perform operations comprising:

enabling a real-time communication between a first communication device and a second communication device;

based on enabling the real-time communication, generating (a) an operating system audio stream corresponding to an audio record of the real-time communication and (b) an event stream recording a plurality of event descriptors corresponding to events detected by one or more applications executing on the first communication device concurrently with the operating system audio stream,

wherein the operating system audio stream comprises one or more of a first set of audio signals detected by a microphone associated with the first communication device or a second set of audio signals played by a speaker associated with the first communication device;

detecting a first event in the event stream at least by detecting a first set of one or more descriptors in the event stream;

responsive at least to determining the first event meets recommendation criteria:

generating a first marker in the operating system audio stream contemporaneous with the first event in the event stream;

using the first marker in the operating system audio stream, extracting a first portion of the operating system audio stream associated with the first event; and

transmitting the first portion of the operating system audio stream to a recommendation model; and

responsive to transmitting the first portion of the operating system audio stream to the recommendation model:

receiving, from the recommendation model, a recommendation for a particular operation to be performed on at least one device responsive to the first event; and

presenting, in a graphical user interface (GUI) displayed on the first communication device, the recommendation for the particular operation responsive to the first event.

19 . The one or more or more non-transitory computer-readable media of claim 1 , wherein the one or more applications comprise at least one of (a) a conversation content extraction application that generates a stream of event descriptors based on events detected in the one or more of the first set of audio signals and the second set of audio signals and (b) a particular application running on a device accessed by a user, wherein the particular application generates the stream of event descriptors responsive to detecting interactions of the user with the particular application.

20 . The one or more or more non-transitory computer-readable media of claim 1 , wherein the operations further comprise:

detecting a second event in the event stream at least by detecting a second set of one or more descriptors in the event stream; and

responsive at least to determining the second event does not meet the recommendation criteria:

refraining from extracting a second portion of the operating system audio stream associated with the second event; and

refraining from transmitting the second portion of the operating system audio stream to the recommendation model.

21 . The one or more or more non-transitory computer-readable media of claim 1 , wherein determining the first event meets the recommendation criteria includes determining the first event is of a first event type.

22 . The one or more or more non-transitory computer-readable media of claim 1 , wherein detecting the first event in the event stream comprises:

monitoring requested and executed transactions in an operating system queue; and

comparing the requested and executed transactions to a set of event rules to detect a set of events, including the first event.

Assignments (4)
SECURITY INTEREST Recorded Jun 26, 2026
From: CRESTA INTELLIGENCE INC.
To: FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 075095/0172 →
RELEASE OF SECURITY INTEREST Recorded Aug 12, 2025
From: TRIPLEPOINT CAPITAL LLC; TRIPLEPOINT VENTURE GROWTH BDC CORP.; TRIPLEPOINT PRIVATE VENTURE CREDIT INC.
To: CRESTA INTELLIGENCE INC.
Reel/Frame 071996/0713 →
SECURITY INTEREST Recorded Jun 6, 2024
From: CRESTA INTELLIGENCE INC.
To: TRIPLEPOINT CAPITAL LLC, AS COLLATERAL AGENT
Reel/Frame 067650/0563 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2022
From: SHI, TIANLIN; OETZEL, KENNETH GEORGE
To: CRESTA INTELLIGENCE INC.
Reel/Frame 058697/0043 →
Continuity (4)
Continuation 17352543 · Jun 21, 2021
Continuation 17083486 · Oct 29, 2020
Continuation 17080100 · Oct 26, 2020
Related Publication 20220139382A1 · May 5, 2022
References Cited (17)
US 6570588B1 · Ando · 2003 [cited by examiner]
US 8145492B2 · Fujita · 2012 [cited by examiner]
US 9667794B2 · Xie · 2017 [cited by examiner]
US 10586237B2 · Coughlin et al. · 2020 [cited by applicant]
US 10916253B2 · Topaloglu · 2021 [cited by examiner]
US 11743378B1 · Johnston · 2023 [cited by examiner]
US 20040172252A1 · Aoki · 2004 [cited by examiner]
US 20110060587A1 · Phillips · 2011 [cited by examiner]
US 20120221330A1 · Thambiratnam · 2012 [cited by examiner]
US 20130138637A1 · Bachtiger · 2013 [cited by examiner]
US 20130195258A1 · Atef · 2013 [cited by examiner]
US 20160239259A1 · Lenchner · 2016 [cited by examiner]
US 20180124243A1 · Zimmerman · 2018 [cited by examiner]
US 20190341050A1 · Diamant · 2019 [cited by examiner]
US 20200094416A1 · Park · 2020 [cited by examiner]
US 20200174740A1 · Martay · 2020 [cited by examiner]
US 20210056966A1 · Bilac · 2021 [cited by examiner]