IP Library Granted Patent US 11,238,884
Granted Patent B2
US 11,238,884 · App. 16/593,461 · Granted Feb 1, 2022

Systems and methods for recording quality driven communication management

Inventors: Simon Jolly (Nottingham, GB); Tony Commander (Manchester, GB); Kyrylo Zotkin (Kyiv, UA)
Assignee: RED BOX RECORDERS LIMITED
G10L21/0364G10L15/18G10L15/22G10L25/60H04L65/604G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,238,884
App. No.
16/593,461
Granted
Feb 1, 2022
Kind
B2
Abstract

A system, method and non-transitory computer readable medium for providing call quality driven communication management wherein an audio data stream of a communication session having one or more utterances is processed to generate a transcript of the communication session. The generated transcript is analyzed to determine whether a quality of the audio data stream, and one or more quality improvement measures when one or more audio artifacts are determined to be present in the audio data stream.

Claims (143)

1. A method comprising:

receiving an audio data stream of at least a portion of a communication session, the audio data stream comprising one or more utterances;

processing the audio data stream to generate a transcript of the communication session, the transcript comprising candidate text for each utterance in the audio data stream along with confidence measures;

analyzing the candidate text and confidence measures to determine a quality of the audio data stream, wherein analyzing the transcript comprises:

tagging, as recognized text, each candidate text having a confidence measure meeting certain criteria, and tagging, as unrecognized text, each candidate text having a confidence measure that fails to meet the certain criteria; and

calculating an audio quality score based on the ratio of recognized text to unrecognized text; and

initiating one or more quality improvement measures when the audio quality score indicates that one or more audio artifacts are present in the audio data stream; and

establishing an audio quality baseline by:

retrieving a previously captured audio data stream of a previous communications session, the previously captured audio data stream comprising one or more previous utterances;

processing the previously captured audio data stream to generate another transcript of the previous communication session, the another transcript comprising additional candidate text for each previous utterance in the previously captured audio data stream along with additional confidence measures;

calculating an average confidence measure for the additional candidate text as the audio quality baseline;

storing the audio quality baseline in association with the previously captured audio data stream; and

comparing the audio quality score to the audio quality baseline to determine whether the audio quality score indicates that one or more audio artifacts are present in the audio data stream.

2. The method of claim 1 , wherein analyzing the transcript further comprises:

determining whether one or more natural language rules are satisfied; and

based on a determination that a particular language rule is satisfied, updating the confidence measure for each utterance associated with the particular rule.

3. The method of claim 2 , wherein analyzing the transcript further comprises:

detecting whether one or more keywords are present in the transcript; and

updating the confidence measure for each utterance of the one or more keywords.

4. The method of claim 1 wherein the one or more quality improvement measures include modifying the transmission codec, altering network routing, and/or modifying transmission media type.

5. A call monitoring system comprising:

a non-transitory storage medium having a plurality of instructions stored thereon; and

at least one processor configure to execute the instructions to:

receive an audio data stream of at least a portion of a communication session, the audio data stream comprising one or more utterances;

process the audio data stream to generate a transcript of the communication session, the transcript comprising candidate text for each utterance in the audio data stream along with confidence measures;

analyze the candidate text and confidence measures to determine a quality of the audio data stream, wherein the processor is configured to execute the instructions to:

tag, as recognized text, each candidate text having a confidence measure meeting certain criteria, and tag, as unrecognized text, each candidate text having a confidence measure that fails to meet the certain criteria; and

calculate an audio quality score based on the ratio of recognized text to unrecognized text; and

initiate one or more quality improvement measures when the audio quality score indicates that one or more audio artifacts are present in the audio data stream; and

establish an audio quality baseline by:

retrieving a previously captured audio data stream of a previous communications session, the previously captured audio data stream comprising one or more previous utterances;

processing the previously captured audio data stream to generate another transcript of the previous communication session, the another transcript comprising additional candidate text for each previous utterance in the previously captured audio data stream along with additional confidence measures;

calculating an average confidence measure for the additional candidate text as the audio quality baseline;

storing the audio quality baseline in association with the previously captured audio data stream; and

compare the audio quality score to the audio quality baseline to determine whether the audio quality score indicates that one or more audio artifacts are present in the audio data stream.

6. The call monitoring system of claim 5 , wherein the processor is configured to execute the instructions to:

determine whether one or more natural language rules are satisfied; and

based on a determination that a particular language rule is satisfied, update the confidence measure for each utterance associated with the particular rule.

7. The call monitoring system of claim 6 , wherein the processor is configured to execute the instructions to:

detect whether one or more keywords are present in the transcript; and

update the confidence measure for each utterance of the one or more keywords.

8. The call monitoring system of claim 5 , wherein the one or more quality improvement measures include modifying the transmission codec, altering network routing, and/or modifying transmission media type.

9. A non-transitory computer-readable medium comprising a plurality of instructions, the instructions being executable by a processor to:

receive an audio data stream of at least a portion of a communication session, the audio data stream comprising one or more utterances;

process the audio data stream to generate a transcript of the communication session, the transcript comprising candidate text for each utterance in the audio data stream along with confidence measures;

analyze the candidate text and confidence measures to determine a quality of the audio data stream, wherein the processor is configured to execute the instructions to:

tag, as recognized text, each candidate text having a confidence measure meeting certain criteria, and tag, as unrecognized text, each candidate text having a confidence measure that fails to meet the certain criteria; and

calculate an audio quality score based on the ratio of recognized text to unrecognized text; and

initiate one or more quality improvement measures when the audio quality score indicates that one or more audio artifacts are present in the audio data stream; and

establish an audio quality baseline by:

retrieving a previously captured audio data stream of a previous communications session, the previously captured audio data stream comprising one or more previous utterances;

processing the previously captured audio data stream to generate another transcript of the previous communication session, the another transcript comprising additional candidate text for each previous utterance in the previously captured audio data stream along with additional confidence measures;

calculating an average confidence measure for the additional candidate text as the audio quality baseline;

storing the audio quality baseline in association with the previously captured audio data stream; and

comparing the audio quality score to the audio quality baseline to determine whether the audio quality score indicates that one or more audio artifacts are present in the audio data stream.

10. The non-transitory computer-readable medium of claim 9 , wherein the instructions are further executable by the processor to:

determine whether one or more natural language rules are satisfied; and

based on a determination that a particular language rule is satisfied, update the confidence measure for each utterance associated with the particular rule.

11. The non-transitory computer-readable medium of claim 10 , wherein the instructions are further executable by the processor to:

detect whether one or more keywords are present in the transcript; and

update the confidence measure for each utterance of the one or more keywords.

12. A method comprising:

receiving an audio data stream of at least a portion of a communication session, the audio data stream comprising one or more utterances;

processing the audio data stream to generate a transcript of the communication session, the transcript comprising candidate text for each utterance in the audio data stream along with confidence measures;

analyzing the candidate text and confidence measures to determine a quality of the audio data stream, wherein analyzing the transcript comprises:

tagging, as recognized text, each candidate text having a confidence measure meeting certain criteria, and tagging, as unrecognized text, each candidate text having a confidence measure that fails to meet the certain criteria; and

calculating an audio quality score based on the ratio of recognized text to unrecognized text; and

initiating one or more quality improvement measures when the audio quality score indicates that one or more audio artifacts are present in the audio data stream; and

establishing a baseline audio quality index by:

retrieving a plurality of previously captured audio data streams and corresponding metadata, the previously captured audio data streams each comprising one or more previous utterances;

processing each previously captured audio data stream to generate additional transcripts, the additional transcripts comprising additional candidate text along with additional confidence measures;

calculating a plurality of baseline audio quality scores as the average confidence measure the additional candidate text in each additional transcript; and

storing the plurality of baseline audio quality scores in association with the corresponding metadata;

selecting a particular baseline audio quality score from the baseline audio quality index based on a metadata of the audio stream; and

comparing the audio quality score to the particular baseline audio quality score to determine whether the audio quality score indicates that one or more audio artifacts are present in the audio data stream.

13. A method comprising:

receiving an audio data stream of at least a portion of a communication session, the audio data stream comprising one or more utterances;

processing the audio data stream to generate a transcript of the communication session, the transcript comprising candidate text for each utterance in the audio data stream along with confidence measures;

analyzing the candidate text and confidence measures to determine a quality of the audio data stream, wherein analyzing the transcript comprises:

tagging, as recognized text, each candidate text having a confidence measure meeting certain criteria, and tagging, as unrecognized text, each candidate text having a confidence measure that fails to meet the certain criteria; and

calculating an audio quality score based on the ratio of recognized text to unrecognized text; and

initiating one or more quality improvement measures when the audio quality score indicates that one or more audio artifacts are present in the audio data stream; and

establishing an audio quality baseline by:

synthesizing an ideal audio data stream free of any audio quality artifacts, the ideal audio data stream comprising one or more clear utterances;

processing the ideal audio data stream to generate another transcript, the another transcript comprising additional candidate text for each clear utterance in the ideal audio data stream along with additional confidence measures;

calculating an average confidence measure for the additional candidate text as the audio quality baseline; and

comparing the audio quality score to the audio quality baseline to determine whether the audio quality score indicates that one or more audio artifacts are present in the audio data stream.

14. A call monitoring system comprising:

a non-transitory storage medium having a plurality of instructions stored thereon; and

at least one processor configure to execute the instructions to:

receive an audio data stream of at least a portion of a communication session, the audio data stream comprising one or more utterances;

process the audio data stream to generate a transcript of the communication session, the transcript comprising candidate text for each utterance in the audio data stream along with confidence measures;

analyze the candidate text and confidence measures to determine a quality of the audio data stream, wherein the processor is configured to execute the instructions to:

tag, as recognized text, each candidate text having a confidence measure meeting certain criteria, and tag, as unrecognized text, each candidate text having a confidence measure that fails to meet the certain criteria; and

calculate an audio quality score based on the ratio of recognized text to unrecognized text; and

initiate one or more quality improvement measures when the audio quality score indicates that one or more audio artifacts are present in the audio data stream; and

establish a baseline audio quality index by:

retrieving a plurality of previously captured audio data streams and corresponding metadata, the previously captured audio data streams each comprising one or more previous utterances;

processing each previously captured audio data stream to generate additional transcripts, the additional transcripts comprising additional candidate text along with additional confidence measures;

calculating a plurality of baseline audio quality scores as the average confidence measure the additional candidate text in each additional transcript; and

storing the plurality of baseline audio quality scores in association with the corresponding metadata;

select a particular baseline audio quality score from the baseline audio quality index based on a metadata of the audio stream; and

compare the audio quality score to the particular baseline audio quality score to determine whether the audio quality score indicates that one or more audio artifacts are present in the audio data stream.

15. A call monitoring system comprising:

a non-transitory storage medium having a plurality of instructions stored thereon; and

at least one processor configure to execute the instructions to:

receive an audio data stream of at least a portion of a communication session, the audio data stream comprising one or more utterances;

process the audio data stream to generate a transcript of the communication session, the transcript comprising candidate text for each utterance in the audio data stream along with confidence measures;

analyze the candidate text and confidence measures to determine a quality of the audio data stream, wherein the processor is configured to execute the instructions to:

tag, as recognized text, each candidate text having a confidence measure meeting certain criteria, and tag, as unrecognized text, each candidate text having a confidence measure that fails to meet the certain criteria; and

calculate an audio quality score based on the ratio of recognized text to unrecognized text; and

initiate one or more quality improvement measures when the audio quality score indicates that one or more audio artifacts are present in the audio data stream; and

establish an audio quality baseline by:

synthesizing an ideal audio data stream free of any audio quality artifacts, the ideal audio data stream comprising one or more clear utterances;

processing the ideal audio data stream to generate another transcript, the another transcript comprising additional candidate text for each clear utterance in the ideal audio data stream along with additional confidence measures; and

calculating an average confidence measure for the additional candidate text as the audio quality baseline; and

compare the audio quality score to the audio quality baseline to determine whether the audio quality score indicates that one or more audio artifacts are present in the audio data stream.

16. A non-transitory computer-readable medium comprising a plurality of instructions, the instructions being executable by a processor to:

receive an audio data stream of at least a portion of a communication session, the audio data stream comprising one or more utterances;

process the audio data stream to generate a transcript of the communication session, the transcript comprising candidate text for each utterance in the audio data stream along with confidence measures;

analyze the candidate text and confidence measures to determine a quality of the audio data stream, wherein the processor is configured to execute the instructions to:

tag, as recognized text, each candidate text having a confidence measure meeting certain criteria, and tag, as unrecognized text, each candidate text having a confidence measure that fails to meet the certain criteria; and

calculate an audio quality score based on the ratio of recognized text to unrecognized text; and

initiate one or more quality improvement measures when the audio quality score indicates that one or more audio artifacts are present in the audio data stream; and

establish a baseline audio quality index by:

retrieving a plurality of previously captured audio data streams and corresponding metadata, the previously captured audio data streams each comprising one or more previous utterances;

processing each previously captured audio data stream to generate additional transcripts, the additional transcripts comprising additional candidate text along with additional confidence measures;

calculating a plurality of baseline audio quality scores as the average confidence measure the additional candidate text in each additional transcript; and

storing the plurality of baseline audio quality scores in association with the corresponding metadata;

select a particular baseline audio quality score from the baseline audio quality index based on a metadata of the audio stream; and

compare the audio quality score to the particular baseline audio quality score to determine whether the audio quality score indicates that one or more audio artifacts are present in the audio data stream.

17. A non-transitory computer-readable medium comprising a plurality of instructions, the instructions being executable by a processor to:

receive an audio data stream of at least a portion of a communication session, the audio data stream comprising one or more utterances;

process the audio data stream to generate a transcript of the communication session, the transcript comprising candidate text for each utterance in the audio data stream along with confidence measures;

analyze the candidate text and confidence measures to determine a quality of the audio data stream, wherein the processor is configured to execute the instructions to:

tag, as recognized text, each candidate text having a confidence measure meeting certain criteria, and tag, as unrecognized text, each candidate text having a confidence measure that fails to meet the certain criteria; and

calculate an audio quality score based on the ratio of recognized text to unrecognized text; and

initiate one or more quality improvement measures when the audio quality score indicates that one or more audio artifacts are present in the audio data stream; and

establish an audio quality baseline by:

synthesizing an ideal audio data stream free of any audio quality artifacts, the ideal audio data stream comprising one or more clear utterances;

processing the ideal audio data stream to generate another transcript, the another transcript comprising additional candidate text for each clear utterance in the ideal audio data stream along with additional confidence measures; and

calculating an average confidence measure for the additional candidate text as the audio quality baseline; and

compare the audio quality score to the audio quality baseline to determine whether the audio quality score indicates that one or more audio artifacts are present in the audio data stream.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2026
From: RED BOX RECORDERS LIMITED
To: UNIPHORE TECHNOLOGIES, INC.
Reel/Frame 073956/0941 →
RELEASE OF SECURITY INTEREST Recorded Oct 2, 2025
From: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY
To: UNIPHORE TECHNOLOGIES INC.
Reel/Frame 072454/0763 →
SECURITY INTEREST Recorded Dec 24, 2024
From: UNIPHORE TECHNOLOGIES INC.
To: FIRST-CITIZENS BANK & TRUST COMPANY
Reel/Frame 069674/0415 →
SECURITY INTEREST Recorded Aug 20, 2024
From: UNIPHORE TECHNOLOGIES INC.; UNIPHORE TECHNOLOGIES NORTH AMERICA INC.; UNIPHORE SOFTWARE SYSTEMS INC.; COLABO, INC.
To: HSBC VENTURES USA INC.
Reel/Frame 068335/0563 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2019
From: JOLLY, SIMON; COMMANDER, TONY; ZOTKIN, KYRYLO
To: RED BOX RECORDERS LIMITED
Reel/Frame 050637/0839 →
Continuity (1)
Related Publication 20210104253A1 · Apr 8, 2021