IP Library Granted Patent US 12,374,324
Granted Patent B2
US 12,374,324 · App. 17/964,341 · Granted Jul 29, 2025

Transcript tagging and real-time whisper in interactive communications

Inventors: Anup Shirodkar (Glen Allen, VA); Thomas Zahorik (Charlottesville, VA); Matthew C. Ford (North Chesterfield, VA); Miguel De La Rocha (Richmond, VA); Sara R. McDole (Midlothian, VA); Robert L. Scholtz, III (Midlothian, VA); Raghav Sahai (Plain City, OH); Kevin T. Shaffer (North Chesterfield, VA); Trent F. Hodges (Richmond, VA); Avinash Kawale (Glen Allen, VA); David Aber (Winter Garden, FL)
Assignee: Capital One Services, LLC
G10L15/1807G06F40/117G06F40/284G06F40/30G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,374,324
App. No.
17/964,341
Granted
Jul 29, 2025
Kind
B2
Abstract

Disclosed herein are system, method, and computer readable medium embodiments for machine learning systems to process interactive communications between at least two participants. Speech and text within the interactive communications are analyzed using machine learning models to infer insights located within the interactive communications. The inferred insights are converted to descriptive text or audio and tagged to the interactive communication as graphics or audio whispers reflecting the insights added to the interactive communication.

Claims (63)

1. A system for augmenting an interactive communication in a natural language processing environment, the system configured to:

convert, by a first machine learning model trained by a machine learning system, the interactive communication to a textual transcript;

extract and evaluate, by a second machine learning model trained by the machine learning system, textual cues in the interactive communication to infer a first insight located within the textual transcript;

generate first text corresponding to a description of the first insight;

tag a first section within the interactive communication with the first text corresponding to the description of the first insight;

display the first text proximate to the tagged first section of a graphic waveform of the interactive communication;

extract and evaluate, by a third machine learning model, audio cues in the interactive communication to infer a second insight located within the interactive communication;

generate second text corresponding to a description of the second insight;

tag a second section within the interactive communication with the second text corresponding to the description of the second insight;

display the second text proximate to the tagged second section of the graphic waveform of the interactive communication;

generate a first audio instance of the first text representing the first insight;

generate a second audio instance of the second text representing the second insight; and

overlay the first audio instance on the tagged first section as a first whisper voice and the second audio instance on the tagged second section as a second whisper voice of the interactive communication.

2. The system of claim 1 further configured to:

tag the first section within the textual transcript with the second text corresponding to the description of the second insight; and

display a graphic with the second text proximate to the first section of the textual transcript.

3. The system of claim 1 , wherein the third machine learning model comprises:

a prosodic cue model to extract and evaluate prosodic cues within the interactive communication.

4. The system of claim 3 , wherein the prosodic cues comprise any of:

frequency changes, pitch, pauses, length of sounds, volume, loudness, speech rate, voice quality, or stress placed on a specific utterance of speech.

5. The system of claim 1 further configured to:

display a graphic of the first text or the second text proximate to a selected section of the textual transcript.

6. The system of claim 1 further configured to:

superimpose the first audio instance or the second audio instance proximate to a selected section of the interactive communication.

7. The system of claim 1 , wherein the second machine learning model comprises any of:

a sentiment predictive model to extract and evaluate semantic cues within the textual transcript;

a key word model to extract and evaluate key words within the textual transcript;

a complaint predictive model to extract and evaluate semantic and key words within the textual transcript; or

a disclosure compliance predictive model to extract and evaluate disclosure key words within the textual transcript.

8. A computer-implemented method for processing a call in a natural language environment, comprising:

converting, by a first machine learning model trained by a machine learning system, an interactive communication to a textual transcript;

extracting and evaluating, by a second machine learning model trained by the machine learning system, textual cues in the interactive communication to infer a first insight located within the textual transcript;

generating first text corresponding to a description of the first insight;

tagging a first section within the interactive communication with the first text;

displaying the first text proximate to the tagged first section of a graphic waveform of the interactive communication;

extracting and evaluating, by a third machine learning model trained by the machine learning system, audio cues in the interactive communication to infer a second insight located within the interactive communication;

generating second text corresponding to a description of the second insight;

tagging a second section within the interactive communication with the second text corresponding to the description of the second insight;

displaying the second text proximate to the tagged second section of the graphic waveform of the interactive communication;

generating a first audio instance of the first text representing the first insight;

generating a second audio instance of the second text representing the second insight; and

overlaying the first audio instance proximate to the tagged first section as a first whisper voice and the second audio instance proximate to a tagged second section as a second whisper voice of the interactive communication.

9. The computer-implemented method of claim 8 , wherein the third machine learning model comprises a prosodic cue model to extract and evaluate prosodic cues within the interactive communication.

10. The computer-implemented method of claim 8 further comprising

displaying a graphic with the second text proximate to the second section of the textual transcript.

11. The computer-implemented method of claim 8 wherein the second machine learning model comprises any of:

a sentiment predictive model to extract and evaluate semantic cues within the textual transcript;

a key word model to extract and evaluate key words within the textual transcript;

a complaint predictive model to extract and evaluate semantic and key words within the textual transcript; or

a disclosure compliance predictive model to extract and evaluate disclosure key words within the textual transcript.

12. A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform natural language operations comprising:

converting, by a first machine learning model trained by a machine learning system, an interactive communication to a textual transcript;

extracting and evaluating, by a second machine learning model trained by the machine learning system, textual cues in the interactive communication to infer a first insight located within the textual transcript;

generating first text corresponding to a description of the first insight;

tagging a first section within the interactive communication with the first text;

displaying the first text proximate to the tagged first section of a graphic waveform of the interactive communication;

extracting and evaluating, by a third machine learning model trained by the machine learning system, audio cues in the interactive communication to infer a second insight located within the interactive communication;

generating second text corresponding to a description of the second insight;

tagging a second section within the interactive communication with the second text; and

displaying the second text proximate to the tagged second section of the graphic of the waveform of the interactive communication;

generating a first audio instance of the first text representing the first insight;

generating a second audio instance of the second text representing the second insight; and

overlaying the first audio instance proximate to the tagged first section as a first whisper voice and the second audio instance proximate to the tagged second section as a second whisper voice of the interactive communication.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 12, 2022
From: SHIRODKAR, ANUP; ZAHORIK, THOMAS; FORD, MATTHEW C.; DE LA ROCHA, MIGUEL; MCDOLE, SARA R.; SCHOLTZ, ROBERT L., III; SAHAI, RAGHAV; SHAFFER, KEVIN T.; HODGES, TRENT F.; KAWALE, AVINASH; ABER, DAVID
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 061394/0448 →
Continuity (1)
Related Publication 20240127804A1 · Apr 18, 2024
References Cited (36)
US 10311454B2 · McCord · 2019 [cited by applicant]
US 10623451B2 · Rathod · 2020 [cited by applicant]
US 10645224B2 · Dwyer et al. · 2020 [cited by applicant]
US 10938867B2 · Deole et al. · 2021 [cited by applicant]
US 11134153B2 · Megann et al. · 2021 [cited by applicant]
US 11170336B2 · Kaimal et al. · 2021 [cited by applicant]
US 11699113B1 · Pearson · 2023 [cited by examiner]
US 20060059120A1 · Xiong · 2006 [cited by examiner]
US 20090129565A1 · Hyndman · 2009 [cited by examiner]
US 20150024800A1 · Rodriguez · 2015 [cited by examiner]
US 20170169840A1 · Rubin · 2017 [cited by examiner]
US 20190325068A1 · Lai · 2019 [cited by examiner]
US 20190341050A1 · Diamant · 2019 [cited by examiner]
US 20200184307A1 · Lipka · 2020 [cited by examiner]
US 20200202268A1 · Retna · 2020 [cited by examiner]
US 20200380978A1 · Ahn · 2020 [cited by examiner]
US 20210004836A1 · Adibi et al. · 2021 [cited by applicant]
US 20210152880A1 · Marten · 2021 [cited by examiner]
US 20210157834A1 · Sivasubramanian · 2021 [cited by examiner]
US 20210158813A1 · Sivasubramanian · 2021 [cited by examiner]
US 20210326472A1 · Shelepov · 2021 [cited by examiner]
US 20220076424A1 · Shin · 2022 [cited by examiner]
US 20220108698A1 · Moritz · 2022 [cited by examiner]
US 20220229994A1 · Sharma · 2022 [cited by examiner]
US 20220238103A1 · Madhusudhan · 2022 [cited by examiner]
US 20230035155A1 · Huang · 2023 [cited by examiner]
US 20230052123A1 · Anaokar · 2023 [cited by examiner]
US 20230066100A1 · Cherukara · 2023 [cited by examiner]
US 20230367968A1 · Eisenstadt · 2023 [cited by examiner]
US 20230376970A1 · Henryson · 2023 [cited by examiner]
US 20240087547A1 · Ivers · 2024 [cited by examiner]
US 20240193069A1 · Kikuchi · 2024 [cited by examiner]
CA 3058928A1 · 2018 [cited by examiner]
CA 3032693A1 · 2019 [cited by examiner]
DE 112020002288T5 · 2022 [cited by examiner]
WO WO2022081930A1 · 2022 [cited by examiner]