IP Library Granted Patent US 12,211,494
Granted Patent B2
US 12,211,494 · App. 17/735,410 · Granted Jan 28, 2025

System and method for automated processing of speech using machine learning

Inventor: Ashwin Subramanyam (Bangalore, IN)
Assignee: Conduent Business Services, LLC
G10L15/1815G10L15/063G10L15/32
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,211,494
App. No.
17/735,410
Granted
Jan 28, 2025
Kind
B2
Abstract

Automated systems and methods are provided for processing speech, comprising obtaining a digitally-encoded speech representation corresponding to a telecommunication interaction, wherein the digitally-encoded speech representation includes at least one of a voice recording or a transcript derived from audio of the telecommunication interaction; obtaining a digitally-encoded data set corresponding to at least one structured feature of the telecommunication interaction; obtaining a reference set, wherein the reference set includes a set of binary-classified existing satisfaction classifications; obtaining a trained machine learning algorithm, wherein the machine learning algorithm has been trained using a first plurality of reference telecommunication interactions which include user-provided satisfaction scores; extracting a feature set from the digitally-encoded speech representation; and by the machine learning algorithm and based on at least one structured feature and the feature set, generating a predicted satisfaction classification for the telecommunication interaction.

Claims (46)

1. A computer-implemented method for processing speech, the method comprising:

obtaining a digitally-encoded speech representation corresponding to a telecommunication interaction, wherein the digitally-encoded speech representation includes at least one of a voice recording or a transcript derived from audio of the telecommunication interaction;

obtaining a digitally-encoded data set corresponding to at least one structured feature of the telecommunication interaction;

obtaining a reference set, wherein the reference set includes a set of binary-classified existing satisfaction classifications;

obtaining a trained machine learning algorithm, wherein the machine learning algorithm has been trained using a first plurality of reference telecommunication interactions which include user-provided satisfaction scores;

extracting a feature set from the digitally-encoded speech representation, wherein the feature set comprises a first group of features corresponding to the at least one voice recording and a second group of features corresponding to the transcript;

generating a first subset of features as a function of a first univariate feature selection algorithm applied to the first group of features;

generating a second subset of features as a function of a second univariate feature selection algorithm applied to the second group of features; and

by the machine learning algorithm and based on at least one structured feature, the first subset of features and the second subset of features, generating a predicted satisfaction classification for the telecommunication interaction.

2. The method according to claim 1 , wherein the operation of generating the predicted satisfaction classification includes

applying the machine learning algorithm separately for each of the first and second subset of features.

3. The method according to claim 1 , further comprising applying a score normalization algorithm to the predicted satisfaction classification to generate a predicted satisfaction score.

4. The method according to claim 3 , wherein the score normalization algorithm is configured to bin the predicted satisfaction classification based on a user-provided threshold.

5. The method according to claim 1 , further comprising applying a response bias correction algorithm to the plurality of reference telecommunication interactions.

6. The method according to claim 1 , wherein the feature set includes at least one of a call silence ratio, an overtalk ratio, or a talk time ratio.

7. The method according to claim 1 , wherein the at least one structured feature includes at least one of a duration of the telecommunication interaction, a hold count, or a conference count.

8. The method according to claim 1 , wherein the machine learning algorithm has further been trained on a second plurality of reference telecommunications interactions which include operator-provided satisfaction scores.

9. A computing system for processing speech, the system comprising:

at least one electronic processor; and

a non-transitory computer readable medium storing instructions that, when executed by the at least one electronic processor, cause the electronic processor to perform operations comprising:

obtaining a digitally-encoded speech representation corresponding to a telecommunication interaction, wherein the digitally-encoded speech representation includes at least one of a voice recording or a transcript derived from audio of the telecommunication interaction,

obtaining a digitally-encoded data set corresponding to at least one structured feature of the telecommunication interaction,

obtaining a reference set, wherein the reference set includes a set of binary-classified existing satisfaction classifications,

obtaining a trained machine learning algorithm, wherein the machine learning algorithm has been trained using a first plurality of reference telecommunication interactions which respectively include user-provided satisfaction scores,

extracting a feature set from the digitally-encoded speech representation, wherein the feature set comprises a first group of features corresponding to the at least one voice recording and a second group of features corresponding to the transcript,

generating a first subset of features as a function of a first univariate feature selection algorithm applied to the first group of features,

generating a second subset of features as a function of a second univariate feature selection algorithm applied to the second group of features, and

by the machine learning algorithm and based on at least one structured feature, the first subset of features and the second subset of features, generating a predicted satisfaction classification for the telecommunication interaction.

10. The system according to claim 9 , wherein the operation of generating the predicted satisfaction classification includes

applying the machine learning algorithm separately for each of the first and second subset of features.

11. The system according to claim 9 , further comprising applying a score normalization algorithm to the predicted satisfaction classification to generate a predicted satisfaction score.

12. The system according to claim 11 , wherein the score normalization algorithm is configured to bin the predicted satisfaction classification based on a user-provided threshold.

13. The system according to claim 9 , further comprising applying a response bias correction algorithm to the plurality of reference telecommunication interactions.

14. The system according to claim 9 , wherein the feature set includes at least one of a call silence ratio, an overtalk ratio, or a talk time ratio.

15. The system according to claim 9 , wherein the at least one structured feature includes at least one of a duration of the telecommunication interaction, a hold count, or a conference count.

16. The system according to claim 9 , wherein the machine learning algorithm has further been trained on a second plurality of reference telecommunications interactions which respectively include operator-provided satisfaction scores.

17. A computer-implemented method for processing speech, the method comprising:

obtaining a digitally-encoded speech representation corresponding to a telecommunication interaction, wherein the digitally-encoded speech representation includes at least one of a voice recording or a transcript derived from audio of the telecommunication interaction;

obtaining a digitally-encoded data set corresponding to at least one structured feature of the telecommunication interaction; obtaining a reference set, wherein the reference set includes a set of binary-classified existing satisfaction classifications;

extracting a feature set from the digitally-encoded speech representation, generating a subset of features from the feature set using a univariate feature selection algorithm;

obtaining training data comprising a first plurality of reference telecommunication interactions which include user-provided satisfaction scores;

generating a weighted training data based on the training data, training a machine learning model using the weighted training data;

generating a predicted satisfaction classification for the telecommunication interaction, using the trained machine learning model, based on at least one structured feature and the subset of features.

18. The method of claim 17 , wherein generating the weighted training data comprises classifying the training data, using a classification model, based on whether the training data contains user-provided satisfaction scores.

19. The method of claim 18 , wherein the user-provided satisfaction scores comprises CSAT scores.

20. The method of claim 18 , wherein generating the weighted training data further comprises scoring the data in the training data classified as containing the user-provided satisfaction scores.

Assignments (3)
SECURITY AGREEMENT Recorded Oct 20, 2025
From: CONDUENT BUSINESS SERVICES, LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 073114/0679 →
SECURITY AGREEMENT Recorded Aug 26, 2025
From: CONDUENT BUSINESS SERVICES, LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 072556/0233 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 3, 2022
From: SUBRAMANYAM, ASHWIN
To: CONDUENT BUSINESS SERVICES, LLC
Reel/Frame 059798/0068 →
Continuity (1)
Related Publication 20230360644A1 · Nov 9, 2023
References Cited (10)
US 11854538B1 · Rozgic · 2023 [cited by examiner]
US 20200294528A1 · Perri · 2020 [cited by examiner]
US 20210044697A1 · Khafizov · 2021 [cited by examiner]
US 20210104245A1 · Aguilar Alas · 2021 [cited by examiner]
US 20210272040A1 · Johnson · 2021 [cited by examiner]
US 20220328064A1 · Shriberg · 2022 [cited by examiner]
US 20230034085A1 · Shalev · 2023 [cited by examiner]
US 20230297778A1 · Can · 2023 [cited by examiner]
Bockhorst et al. “Predicting Self-Reported Customer Satisfaction of Interactions with a Corporate Call Center.” Joint European Conf. on Machine Learning and Knowledge Discovery in Databases. Springer, Cham, pp. 179-190 … [cited by applicant]
Park & Gates, “Towards Real-Time Measurement of Customer Satisfaction Using Automatically Generated Call Transcripts”, CIKM '09: Proceedings of the 18th ACM Conf. on Information & Knowledge Mgmt. pp. 1387-1396 https://d… [cited by applicant]