IP Library Granted Patent US 11,120,812
Granted Patent B1
US 11,120,812 · App. 17/130,157 · Granted Sep 14, 2021

Application of machine learning techniques to select voice transformations

Inventor: Tianlin Shi (Menlo Park, CA)
Assignee: CRESTA INTELLIGENCE INC.
G10L21/013G06N20/00G10L15/063G10L15/22G10L25/51G10L25/90G10L2015/228G10L2021/0135
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,120,812
App. No.
17/130,157
Granted
Sep 14, 2021
Kind
B1
Abstract

Techniques for monitoring a conversation in real-time to detect attributes of a conversation, identifying a desired outcome of the conversation, and identifying voice modulations that may be applied to the agent's voice to help accomplish the desired outcome are disclosed. The system may identify voice modulations by comparing a current conversation to one or more prior conversations having desired outcomes similar to that of the current conversation. A trained machine learning model may select and apply voice modulations associated with accomplishing a desired outcome.

Claims (75)

1. One or more non-transitory computer-readable media storing instructions, which when executed by one or more hardware processors, cause performance of operations comprising:

monitoring a conversation in real-time to detect one or more attributes of the conversation;

identifying a desired outcome of a conversation based on the one or more attributes;

identifying a plurality of voice modulations for accomplishing the desired outcome of the conversation;

determining a confidence score for each voice modulation of the plurality of voice modulations, the confidence score associated with a likelihood that the desired outcome will be accomplished subsequent to using the voice modulation to a participant's voice in the conversation, wherein determining the confidence score further comprises;

identifying at least one prior conversation having a prior desired outcome similar to the desired outcome of the conversation;

determining that one or more voice modulations of the plurality of voice modulations were used in the at least one prior conversation;

determining one or more rates at which the prior desired outcome of the at least one prior conversation was accomplished;

determining the confidence score of the one or more voice modulations in proportion to the rate at which the desired outcome were accomplished in the at least one prior conversation;

based at least on a first confidence score of accomplishing the desired outcome subsequent to using a first voice modulation of the plurality of voice modulations: selecting the first voice modulation from the plurality of voice modulations; and

applying the first voice modulation to the participant's voice.

2. The one or more media of claim 1 , wherein:

a conversation type is identified based on the one or more attributes of the conversation; and

the desired outcome of the conversation is based, at least in part, on the conversation type.

3. The one or more media of claim 2 , wherein:

the conversation type comprises one of a sales conversation, a help desk conversation, a customer retention conversation, and a customer care conversation; and

the desired outcome of the conversation comprises completing a sale, answering a question, and resolving a customer complaint, respectively.

4. The one or more media of claim 1 , wherein the one or more voice modulations applied to the voice of the participant comprises one or more of altering a tone, a speed, and a regional accent.

5. The one or more media of claim 1 , wherein the one or more voice modulations applied to the voice of the participant comprises one or more of altering a content density.

6. The one or more media of claim 1 , wherein determining the confidence score further comprises:

identifying a set of voice modulations used in a prior conversation having a desired outcome similar to the desired outcome for the conversation;

determining similarity scores between the set of voice modulations used in the prior conversation and corresponding voice modulations of the one or more voice modulations; and

using the similarity scores as a factor in determining the confidence scores of the corresponding one or more voice modulations.

7. The one or more media of claim 1 , wherein determining the confidence score further comprises basing the confidence score on a rate at which desired outcomes are accomplished.

8. A method comprising:

monitoring a conversation in real-time to detect one or more attributes of the conversation;

identifying a desired outcome of a conversation based on the one or more attributes;

identifying a plurality of voice modulations for accomplishing the desired outcome of the conversation;

determining a confidence score for each voice modulation of the plurality of voice modulations, the confidence score associated with a likelihood that the desired outcome will be accomplished subsequent to using the voice modulation to a participant's voice in the conversation, wherein determining the confidence score further comprises;

identifying at least one prior conversation having a prior desired outcome similar to the desired outcome of the conversation;

determining that one or more voice modulations of the plurality of voice modulations were used in the at least one prior conversation;

determining one or more rates at which the prior desired outcome of the at least one prior conversation was accomplished;

determining the confidence score of the one or more voice modulations in proportion to the rate at which the desired outcome were accomplished in the at least one prior conversation;

based at least on a first confidence score of accomplishing the desired outcome subsequent to using a first voice modulation of the plurality of voice modulations: selecting the first voice modulation from the plurality of voice modulations; and

applying the first voice modulation to the participant's voice.

9. The method of claim 8 , wherein:

a conversation type is identified based on the one or more attributes of the conversation; and

the desired outcome of the conversation is based, at least in part, on the conversation type.

10. The method of claim 9 , wherein:

the conversation type comprises one of a sales conversation, a help desk conversation, a customer retention conversation, and a customer care conversation; and

the desired outcome of the conversation comprises completing a sale, answering a question, and resolving a customer complaint, respectively.

11. The method of claim 8 , wherein the one or more voice modulations applied to the voice of the participant comprises one or more of altering a tone, a speed, and a regional accent.

12. The method of claim 8 , wherein the one or more voice modulations applied to the voice of the participant comprises one or more of altering a content density.

13. The method of claim 8 , wherein determining the confidence score further comprises:

identifying a set of voice modulations used in a prior conversation having a desired outcome similar to the desired outcome for the conversation;

determining similarity scores between the set of voice modulations used in the prior conversation and corresponding voice modulations of the one or more voice modulations; and

using the similarity scores as a factor in determining the confidence scores of the corresponding one or more voice modulations.

14. The method of claim 8 , wherein determining the confidence score further comprises basing the confidence score on a rate at which desired outcomes are accomplished.

15. A system comprising:

at least one device including a hardware processor;

the system being configured to perform operations comprising:

monitoring a conversation in real-time to detect one or more attributes of the conversation;

identifying a desired outcome of a conversation based on the one or more attributes;

identifying a plurality of voice modulations for accomplishing the desired outcome of the conversation;

determining a confidence score for each voice modulation of the plurality of voice modulations, the confidence score associated with a likelihood that the desired outcome will be accomplished subsequent to using the voice modulation to a participant's voice in the conversation, wherein determining the confidence score further comprises;

identifying at least one prior conversation having a prior desired outcome similar to the desired outcome of the conversation;

determining that one or more voice modulations of the plurality of voice modulations were used in the at least one prior conversation;

determining one or more rates at which the prior desired outcome of the at least one prior conversation was accomplished;

determining the confidence score of the one or more voice modulations in proportion to the rate at which the desired outcome were accomplished in the at least one prior conversation;

based at least on a first confidence score of accomplishing the desired outcome subsequent to using a first voice modulation of the plurality of voice modulations:

selecting the first voice modulation from the plurality of voice modulations; and

applying the first voice modulation to the participant's voice.

16. The system of claim 15 , wherein:

a conversation type is identified based on the one or more attributes of the conversation; and

the desired outcome of the conversation is based, at least in part, on the conversation type.

17. The system of claim 16 , wherein:

the conversation type comprises one of a sales conversation, a help desk conversation, a customer retention conversation, and a customer care conversation; and

the desired outcome of the conversation comprises completing a sale, answering a question, and resolving a customer complaint, respectively.

18. The system of claim 15 , wherein the one or more voice modulations applied to the voice of the participant comprises one or more of altering a tone, a speed, and a regional accent.

19. The system of claim 15 , wherein the one or more voice modulations applied to the voice of the participant comprises one or more of altering a content density.

20. The system of claim 15 , wherein determining the confidence score further comprises:

identifying a set of voice modulations used in a prior conversation having a desired outcome similar to the desired outcome for the conversation;

determining similarity scores between the set of voice modulations used in the prior conversation and corresponding voice modulations of the one or more voice modulations; and

using the similarity scores as a factor in determining the confidence scores of the corresponding one or more voice modulations.

21. The system of claim 15 , wherein determining the confidence score further comprises basing the confidence score on a rate at which desired outcomes are accomplished.

Assignments (4)
SECURITY INTEREST Recorded Jun 26, 2026
From: CRESTA INTELLIGENCE INC.
To: FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 075095/0172 →
RELEASE OF SECURITY INTEREST Recorded Aug 12, 2025
From: TRIPLEPOINT CAPITAL LLC; TRIPLEPOINT VENTURE GROWTH BDC CORP.; TRIPLEPOINT PRIVATE VENTURE CREDIT INC.
To: CRESTA INTELLIGENCE INC.
Reel/Frame 071996/0713 →
SECURITY INTEREST Recorded Jun 6, 2024
From: CRESTA INTELLIGENCE INC.
To: TRIPLEPOINT CAPITAL LLC, AS COLLATERAL AGENT
Reel/Frame 067650/0563 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 5, 2021
From: SHI, TIANLIN
To: CRESTA INTELLIGENCE INC.
Reel/Frame 054808/0796 →
Continuity (1)
Continuation In Part 16944651 · Jul 31, 2020