IP Library Granted Patent US 9,711,167
Granted Patent B2
US 9,711,167 · App. 13/419,340 · Granted Jul 18, 2017

System and method for real-time speaker segmentation of audio interactions

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,711,167
App. No.
13/419,340
Granted
Jul 18, 2017
Kind
B2
Abstract

A system and method for real-time processing a signal of a voice interaction. In an embodiment, a digital representation of a portion of an interaction may be analyzed in real-time and a segment may be selected. The segment may be associated with a source based on a model of the source. The model may updated based on the segment. The updated model is used to associate subsequent segments with the source. Other embodiments are described and claimed.

Claims (54)

1. A method of associating a segment of a voice interaction with a speaker, the method comprising:

generating a plurality of first models by, for each of a plurality of speakers:

generating a first model used to identify the speaker based on analyzing one or more recordings of voice interactions that include the speaker, the first model comprising voice characteristics of the speaker;

while a voice interaction different than a voice interaction used to generate a model is in progress:

obtaining, by a controller, a digital representation of a portion of the voice interaction and selecting a segment therefrom;

associating, by the controller, the segment with a specific speaker of the plurality of speakers based on a first model of the plurality of first models, the first model unique to the specific speaker;

updating the first model based on the segment to produce an updated model; and

using the updated model to associate a subsequent segment of the voice interaction with the specific speaker of the plurality of speakers; and

wherein associating the segment with a speaker is performed when the identity of the speaker is unknown and comprises searching a set of models for a model matching the speaker, the matching based on an acoustic feature.

2. The method of claim 1 , comprising:

obtaining a second model related to a second speaker;

calculating for the segment a first score and a second score based on the first and second models; and

selecting to associate the segment with one of the first and second speakers based on the calculated scores.

3. The method of claim 2 , wherein the first and second scores are calculated based on calculating one or more probabilities of a respective one or more acoustic features given the first and second models.

4. The method of claim 1 , comprising iteratively and in real-time, updating the first model based on classified segments of the interaction.

5. The method of claim 1 , wherein the segment is selected based on a voice activity detection performed on at least part of the portion.

6. The method of claim 1 , comprising:

determining at least one acoustic feature in the segment;

calculating at least one probability value for the at least one acoustic feature based on the first model; and

associating the segment with the speaker if the at least one probability value meets a predefined criteria.

7. The method of claim 6 , comprising:

associating a plurality of segments with the speaker according to a respective plurality of probability values; and

updating the first model based on a subset of the segments, the subset selected based on the probability values.

8. The method of claim 1 , comprising determining at least one parameter related to the segment and selecting the first model, in real-time, based on the at least one parameter.

9. The method of claim 1 , comprising analyzing the segment to produce an analysis result and selecting the first model, in real-time, from a plurality of models based on the analysis result.

10. The method of claim 1 , comprising analyzing the segment to produce an analysis result and generating the first model, in real-time, based on the analysis result.

11. The method of claim 1 , comprising identifying the speaker by relating the updated model to a plurality of models associated with a respective plurality of known speakers.

12. The method of claim 1 , comprising continuously improving the model while the second voice interaction is in progress such that the model is better adapted to associate segments of the interaction with the speaker.

13. An article comprising a non-transitory computer-readable storage medium, having stored thereon instructions, that when executed on a computer, cause the computer to:

generate a plurality of first models by, for each of a plurality of speakers:

generate a first model used to identify the speaker based on analyzing one or more recordings of voice interactions that include the speaker, the first model comprising voice characteristics of the speaker;

while a voice interaction different than a voice interaction used to generate a model is in progress:

obtain a digital representation of a portion of the voice interaction;

analyze the portion and select a segment therefrom;

associate the segment with a specific speaker of the plurality of speakers based on a first model of the plurality of first models, the first model unique to the specific speaker;

update the first model based on the segment to produce an updated model; and

use the updated model to associate a subsequent segment of the voice interaction with the specific speaker of the plurality of speakers; and

wherein associating the segment with a speaker is performed when the identity of the speaker is unknown and comprises searching a set of models for a model matching the speaker, the matching based on an acoustic feature.

14. The article of claim 13 , wherein the instructions when executed further result in:

obtaining a second model related to a second speaker;

calculating for the segment a first score and a second score based on the first and second models; and

selecting to associate the segment with one of the first and second speakers based on the calculated scores.

15. The article of claim 14 , wherein the first and second scores are calculated based on calculating one or more probabilities for a respective one or more acoustic features given the first and second models.

16. The article of claim 13 , wherein the instructions when executed further result in iteratively and in real-time, updating the first model based on segments of the interaction.

17. The article of claim 13 , wherein the segment is selected based on a voice activity detection performed on at least part of the portion.

18. The article of claim 13 , wherein the instructions when executed further result in:

determining at least one acoustic feature in the segment;

calculating at least one probability value for the at least one acoustic feature based on the first model; and

associating the segment with the speaker if the at least one probability value meets a predefined criteria.

19. The article of claim 18 , wherein the instructions when executed further result in

associating a plurality of segments with the speaker according to a respective plurality of probability values; and

updating the first model based on a subset of the segments, the subset selected based on the probability scores.

20. The article of claim 13 , wherein the instructions when executed further result in analyzing the segment to produce an analysis result and selecting the first model, in real-time, from a plurality of models based on the analysis result.

21. The article of claim 13 , wherein the instructions when executed further result in identifying the speaker by relating an initial model to a plurality of models associated with a respective plurality of known speakers.

Assignments (4)
SECURITY INTEREST Recorded Feb 26, 2026
From: NICE LTD; NICE SYSTEMS INC.; NICE SYSTEMS TECHNOLOGIES INC.; INCONTACT, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 074986/0208 →
PATENT SECURITY AGREEMENT Recorded Dec 6, 2016
From: NICE LTD.; NICE SYSTEMS INC.; AC2 SOLUTIONS, INC.; ACTIMIZE LIMITED; INCONTACT, INC.; NEXIDIA, INC.; NICE SYSTEMS TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 040821/0818 →
CHANGE OF NAME Recorded Oct 18, 2016
From: NICE-SYSTEMS LTD.
To: NICE LTD.
Reel/Frame 040387/0527 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 29, 2013
From: WASSERBLAT, MOSHE; ASHKENAZI, TZACHI; BEN-ASHER, MERAV; PEREG, OREN
To: NICE-SYSTEMS LTD.
Reel/Frame 031872/0060 →