IP Library Granted Patent US 12,518,761
Granted Patent B2
US 12,518,761 · App. 18/475,599 · Granted Jan 6, 2026

Diarization using acoustic labeling

Inventors: Omer Ziv (Ramat Gan, IL); Ran Achituv (Hod Hasharon, IL); Ido Shapira (Tel Aviv, IL); Jeremie Dreyfuss (Tel Aviv, IL)
Assignee: VERINT SYSTEMS INC.
G10L17/00G10L17/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,518,761
App. No.
18/475,599
Granted
Jan 6, 2026
Kind
B2
Abstract

Systems and method of diarization of audio files use an acoustic voiceprint model. A plurality of audio files are analyzed to arrive at an acoustic voiceprint model associated to an identified speaker. Metadata associate with an audio file is used to select an acoustic voiceprint model. The selected acoustic voiceprint model is applied in a diarization to identify audio data of the identified speaker.

Claims (43)

1 . A method for diarization of audio files using acoustic voiceprints, the method comprising:

receiving an audio file for diarization by a processor, wherein the audio file includes at least two speakers speaking in the same audio file;

separating the audio file into a plurality of speaker segments, wherein each speaker segment is a segment of speech from at least one of the at least two speakers in the audio file, further wherein each segment is separated by a non-speech segment;

receiving a plurality of acoustic voiceprints from a voiceprint database server, wherein each acoustic voiceprint is trained from a plurality of speaker segments;

comparing the plurality of acoustic voiceprints to each speaker segment to determine which of the plurality of acoustic voiceprints match each speaker segment; and

applying a speaker label to each speaker segment based on the comparing, wherein the speaker label identifies a speaker associated with the speaker segment.

2 . The method of claim 1 , the method further comprising:

receiving speaker metadata for the audio file; and

selecting acoustic voiceprints from the plurality of acoustic voiceprints based upon the received speaker metadata for the comparing.

3 . The method of claim 1 , wherein each acoustic voice print is associated with a different known speaker.

4 . The method of claim 1 , wherein the audio file is real-time audio data of a customer service interaction including a customer service agent and at least one other speaker.

5 . The method of claim 1 , the method further comprising clustering similar speaker segments of the plurality of speaker segments, wherein the similar speaker segments have a high likelihood of containing speech from a single speaker.

6 . The method of claim 5 , wherein clustering the similar speaker segments of the plurality of speaker segments includes applying at least one metric to the speaker segments to label the segments of speech as belonging to a customer service agent or as belonging to an other speaker.

7 . The method of claim 6 , wherein the at least one metric is that of cluster size wherein the larger the cluster the more likely the segment belongs to the customer service agent.

8 . A system for diarization of audio files using acoustic voiceprints, the system comprising:

a memory comprising computer readable instructions;

a processor configured to read the computer readable instructions that when executed causes the system to:

receive an audio file for diarization by a processor, wherein the audio file includes at least two speakers speaking in the same audio file;

separate the audio file into a plurality of speaker segments, wherein each speaker segment is a segment of speech from at least one of the at least two speakers in the audio file, further wherein each segment is separated by a non-speech segment;

receive a plurality of acoustic voiceprints from a voiceprint database server, wherein each acoustic voiceprint is trained from a plurality of speaker segments;

compare the plurality of acoustic voiceprints to each speaker segment to determine which of the plurality of acoustic voiceprints match each speaker segment; and

apply a speaker label to each speaker segment based on the comparing, wherein the speaker label identifies a speaker associated with the speaker segment.

9 . The system of claim 8 , wherein the processor is further configured to cause the system to:

receive speaker metadata for the audio file; and

select acoustic voiceprints from the plurality of acoustic voiceprints based upon the received speaker metadata for the comparing.

10 . The system of claim 8 , wherein each acoustic voiceprint is associated with a different known speaker.

11 . The system of claim 8 , wherein the audio file is real-time audio data of a customer service interaction including a customer service agent and at least one other speaker.

12 . The system of claim 8 , wherein the processor is further configured to cause the system to: cluster similar speaker segments of the plurality of speaker segments, wherein the similar speaker segments have a high likelihood of containing speech from a single speaker.

13 . The system of claim 12 , wherein clustering the similar speaker segments of the plurality of speaker segments includes applying at least one metric to the speaker segments to label the segments of speech as belonging to a customer service agent or as belonging to an other speaker.

14 . The system of claim 13 , wherein the at least one metric is that of cluster size wherein the larger the cluster the more likely the segment belongs to the customer service agent.

15 . A non-transitory computer readable medium comprising computer readable code for diarization of audio files using acoustic voiceprints on a system that when executed by a processor, causes the system to:

receive an audio file for diarization by a processor, wherein the audio file includes at least two speakers speaking in the same audio file;

separate the audio file into a plurality of speaker segments, wherein each speaker segment is a segment of speech from at least one of the at least two speakers in the audio file, further wherein each segment is separated by a non-speech segment;

receive a plurality of acoustic voiceprints from a voiceprint database server, wherein each acoustic voiceprint is trained from a plurality of speaker segments;

compare the plurality of acoustic voiceprints to each speaker segment to determine which of the plurality of acoustic voiceprints match each speaker segment; and

apply a speaker label to each speaker segment based on the comparing, wherein the speaker label identifies a speaker associated with the speaker segment.

16 . The non-transitory computer readable medium of claim 15 , wherein the system is further caused to:

receive speaker metadata for the audio file; and

select acoustic voiceprints from the plurality of acoustic voiceprints based upon the received speaker metadata for the comparing.

17 . The non-transitory computer readable medium of claim 15 , wherein each acoustic voiceprint is associated with a different known speaker.

18 . The non-transitory computer readable medium of claim 15 , wherein the processor is further configured to cause the system to: cluster similar speaker segments of the plurality of speaker segments, wherein the similar speaker segments have a high likelihood of containing speech from a single speaker.

19 . The non-transitory computer readable medium of claim 18 , wherein clustering the similar speaker segments of the plurality of speaker segments includes applying at least one metric to the speaker segments to label the segments of speech as belonging to a customer service agent or as belonging to an other speaker.

20 . The non-transitory computer readable medium of claim 19 , wherein the at least one metric is that of cluster size wherein the larger the cluster the more likely the segment belongs to the customer service agent.

Assignments (3)
SECURITY INTEREST Recorded Dec 23, 2025
From: VERINT SYSTEMS INC.
To: ALTER DOMUS (US) LLC, AS COLLATERAL AGENT
Reel/Frame 074034/0919 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2024
From: VERINT SYSTEMS LTD.
To: VERINT SYSTEMS INC.,
Reel/Frame 066169/0209 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 27, 2023
From: ZIV, OMER; ACHITUV, RAN; SHAPIRA, IDO; DREYFUSS, JEREMIE
To: VERINT SYSTEMS LTD.
Reel/Frame 066130/0481 →
Continuity (8)
Continuation 17577238 · Jan 17, 2022
Continuation 16848385 · Apr 14, 2020
Continuation 16594812 · Oct 7, 2019
Continuation 16170306 · Oct 25, 2018
Continuation 14084974 · Nov 20, 2013
Provisional Application 61729064 · Nov 21, 2012
Provisional Application 61729067 · Nov 21, 2012
Related Publication 20240021206A1 · Jan 18, 2024
References Cited (156)
US 4653097A · Watanabe et al. · 1987 [cited by applicant]
US 4864566A · Chauveau · 1989 [cited by applicant]
US 5027407A · Tsunoda · 1991 [cited by applicant]
US 5222147A · Koyama · 1993 [cited by applicant]
US 5638430A · Hogan et al. · 1997 [cited by applicant]
US 5805674A · Anderson · 1998 [cited by applicant]
US 5907602A · Peel et al. · 1999 [cited by applicant]
US 5946654A · Newman et al. · 1999 [cited by applicant]
US 5963908A · Chadha · 1999 [cited by applicant]
US 5999525A · Krishnaswamy et al. · 1999 [cited by applicant]
US 6044382A · Martino · 2000 [cited by applicant]
US 6145083A · Shaffer et al. · 2000 [cited by applicant]
US 6266640B1 · Fromm · 2001 [cited by applicant]
US 6275806B1 · Pertrushin · 2001 [cited by applicant]
US 6427137B2 · Petrushin · 2002 [cited by applicant]
US 6480825B1 · Sharma et al. · 2002 [cited by applicant]
US 6510415B1 · Talmor et al. · 2003 [cited by applicant]
US 6587552B1 · Zimmerman · 2003 [cited by applicant]
US 6597775B2 · Lawyer et al. · 2003 [cited by applicant]
US 6915259B2 · Rigazio · 2005 [cited by examiner]
US 7006605B1 · Morganstein et al. · 2006 [cited by applicant]
US 7039951B1 · Chaudhari et al. · 2006 [cited by applicant]
US 7054811B2 · Barzilay · 2006 [cited by applicant]
US 7106843B1 · Gainsboro et al. · 2006 [cited by applicant]
US 7158622B2 · Lawyer et al. · 2007 [cited by applicant]
US 7212613B2 · Kim et al. · 2007 [cited by applicant]
US 7299177B2 · Broman et al. · 2007 [cited by applicant]
US 7386105B2 · Wasserblat et al. · 2008 [cited by applicant]
US 7403922B1 · Lewis et al. · 2008 [cited by applicant]
US 7539290B2 · Ortel · 2009 [cited by applicant]
US 7657431B2 · Hayakawa · 2010 [cited by applicant]
US 7660715B1 · Thambiratnam · 2010 [cited by examiner]
US 7668769B2 · Baker et al. · 2010 [cited by applicant]
US 7693965B2 · Rhoads · 2010 [cited by applicant]
US 7778832B2 · Broman et al. · 2010 [cited by applicant]
US 7822605B2 · Zigel et al. · 2010 [cited by applicant]
US 7908645B2 · Varghese et al. · 2011 [cited by applicant]
US 7940897B2 · Khor et al. · 2011 [cited by applicant]
US 8036892B2 · Broman et al. · 2011 [cited by applicant]
US 8073691B2 · Rajakumar · 2011 [cited by applicant]
US 8112278B2 · Burke · 2012 [cited by applicant]
US 8311826B2 · Rajakumar · 2012 [cited by applicant]
US 8510215B2 · Gutierrez · 2013 [cited by applicant]
US 8537978B2 · Jaiswal et al. · 2013 [cited by applicant]
US 9001976B2 · Arrowood · 2015 [cited by examiner]
US 10113440B2 · Fukaya · 2018 [cited by applicant]
US 10438592B2 · Ziv · 2019 [cited by examiner]
US 10650826B2 · Ziv · 2020 [cited by examiner]
US 11227603B2 · Ziv · 2022 [cited by examiner]
US 20010026632A1 · Tamai · 2001 [cited by applicant]
US 20020022474A1 · Blom et al. · 2002 [cited by applicant]
US 20020099649A1 · Lee et al. · 2002 [cited by applicant]
US 20030009333A1 · Sharma et al. · 2003 [cited by applicant]
US 20030050780A1 · Rigazio · 2003 [cited by examiner]
US 20030050816A1 · Givens et al. · 2003 [cited by applicant]
US 20030097593A1 · Sawa et al. · 2003 [cited by applicant]
US 20030147516A1 · Lawyer et al. · 2003 [cited by applicant]
US 20030208684A1 · Camacho et al. · 2003 [cited by applicant]
US 20040029087A1 · White · 2004 [cited by applicant]
US 20040111305A1 · Gavan et al. · 2004 [cited by applicant]
US 20040131160A1 · Mardirossian · 2004 [cited by applicant]
US 20040143635A1 · Galea · 2004 [cited by applicant]
US 20040167964A1 · Rounthwaite et al. · 2004 [cited by applicant]
US 20040203575A1 · Chin et al. · 2004 [cited by applicant]
US 20040240631A1 · Broman et al. · 2004 [cited by applicant]
US 20050010411A1 · Rigazio · 2005 [cited by examiner]
US 20050043014A1 · Hodge · 2005 [cited by applicant]
US 20050076084A1 · Loughmiller et al. · 2005 [cited by applicant]
US 20050125226A1 · Magee · 2005 [cited by applicant]
US 20050125339A1 · Tidwell et al. · 2005 [cited by applicant]
US 20050135595A1 · Bushey et al. · 2005 [cited by applicant]
US 20050185779A1 · Toms · 2005 [cited by applicant]
US 20060013372A1 · Russell · 2006 [cited by applicant]
US 20060098803A1 · Bushey et al. · 2006 [cited by applicant]
US 20060106605A1 · Saunders et al. · 2006 [cited by applicant]
US 20060149558A1 · Kahn · 2006 [cited by examiner]
US 20060161435A1 · Atef et al. · 2006 [cited by applicant]
US 20060212407A1 · Lyon · 2006 [cited by applicant]
US 20060212925A1 · Shull et al. · 2006 [cited by applicant]
US 20060248019A1 · Rajakumar · 2006 [cited by applicant]
US 20060251226A1 · Hogan et al. · 2006 [cited by applicant]
US 20060282660A1 · Varghese et al. · 2006 [cited by applicant]
US 20060285665A1 · Wasserblat et al. · 2006 [cited by applicant]
US 20060289622A1 · Khor et al. · 2006 [cited by applicant]
US 20060293891A1 · Pathuel · 2006 [cited by applicant]
US 20070041517A1 · Clarke et al. · 2007 [cited by applicant]
US 20070071206A1 · Gainsboro et al. · 2007 [cited by applicant]
US 20070074021A1 · Smithies et al. · 2007 [cited by applicant]
US 20070100608A1 · Gable et al. · 2007 [cited by applicant]
US 20070124246A1 · Lawyer et al. · 2007 [cited by applicant]
US 20070244702A1 · Kahn et al. · 2007 [cited by applicant]
US 20070250318A1 · Waserblat et al. · 2007 [cited by applicant]
US 20070280436A1 · Rajakumar · 2007 [cited by applicant]
US 20070282605A1 · Rajakumar · 2007 [cited by applicant]
US 20070288242A1 · Spengler · 2007 [cited by examiner]
US 20080010066A1 · Broman et al. · 2008 [cited by applicant]
US 20080181417A1 · Pereg · 2008 [cited by examiner]
US 20080195387A1 · Zigel et al. · 2008 [cited by applicant]
US 20080222734A1 · Redlich et al. · 2008 [cited by applicant]
US 20090046841A1 · Hodge · 2009 [cited by applicant]
US 20090103708A1 · Conway · 2009 [cited by examiner]
US 20090119106A1 · Rajakumar · 2009 [cited by applicant]
US 20090147939A1 · Morganstein et al. · 2009 [cited by applicant]
US 20090247131A1 · Champion et al. · 2009 [cited by applicant]
US 20090254971A1 · Herz et al. · 2009 [cited by applicant]
US 20090319269A1 · Aronowitz · 2009 [cited by applicant]
US 20100138282A1 · Kannan et al. · 2010 [cited by applicant]
US 20100228656A1 · Wasserblat et al. · 2010 [cited by applicant]
US 20100303211A1 · Hartig · 2010 [cited by applicant]
US 20100305946A1 · Gutierrez · 2010 [cited by applicant]
US 20100305960A1 · Gutierrez · 2010 [cited by applicant]
US 20100332287A1 · Gates · 2010 [cited by examiner]
US 20110004472A1 · Zlokarnik · 2011 [cited by applicant]
US 20110026689A1 · Metz et al. · 2011 [cited by applicant]
US 20110119060A1 · Aronowitz · 2011 [cited by applicant]
US 20110191106A1 · Khor et al. · 2011 [cited by applicant]
US 20110255676A1 · Marchand et al. · 2011 [cited by applicant]
US 20110282661A1 · Dobry et al. · 2011 [cited by applicant]
US 20110282778A1 · Wright et al. · 2011 [cited by applicant]
US 20110320484A1 · Smithies et al. · 2011 [cited by applicant]
US 20120053939A9 · Gutierrez et al. · 2012 [cited by applicant]
US 20120054202A1 · Rajakumar · 2012 [cited by applicant]
US 20120072453A1 · Guerra et al. · 2012 [cited by applicant]
US 20120130771A1 · Kannan · 2012 [cited by examiner]
US 20120253805A1 · Rajakumar et al. · 2012 [cited by applicant]
US 20120254243A1 · Zeppenfeld et al. · 2012 [cited by applicant]
US 20120263285A1 · Rajakumar et al. · 2012 [cited by applicant]
US 20120284026A1 · Cardillo et al. · 2012 [cited by applicant]
US 20130163737A1 · Dement et al. · 2013 [cited by applicant]
US 20130197912A1 · Hayakawa et al. · 2013 [cited by applicant]
US 20130253919A1 · Gutierrez et al. · 2013 [cited by applicant]
US 20130300939A1 · Chou et al. · 2013 [cited by applicant]
US 20140067394A1 · Abuzeina · 2014 [cited by applicant]
US 20140142940A1 · Ziv et al. · 2014 [cited by applicant]
US 20150055763A1 · Guerra et al. · 2015 [cited by applicant]
US 20160364606A1 · Conway et al. · 2016 [cited by applicant]
US 20160379032A1 · Mo et al. · 2016 [cited by applicant]
US 20160379082A1 · Rodriguez et al. · 2016 [cited by applicant]
EP 0598469 · 1994 [cited by applicant]
JP 2004193942 · 2004 [cited by applicant]
JP 2006038955 · 2006 [cited by applicant]
WO 2000077772 · 2000 [cited by applicant]
WO 2004079501 · 2004 [cited by applicant]
WO 2006013555 · 2006 [cited by applicant]
WO 2007001452 · 2007 [cited by applicant]
Baum, L.E., et al., “A Maximization Technique Occurring in the Statistical Analysis of Probabilistic Functions of Markov Chains,” The Annals of Mathematical Statistics, vol. 41, No. 1, 1970, pp. 164-171. [cited by applicant]
Cheng, Y., “Mean Shift, Mode Seeking, and Clustering,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 17, No. 8, 1995, pp. 790-799. [cited by applicant]
Cohen, I., “Noise Spectrum Estimation in Adverse Environment: Improved Minima Controlled Recursive Averaging,” IEEE Transactions On Speech and Audio Processing, vol. 11, No. 5, 2003, pp. 466-475. [cited by applicant]
Cohen, I., et al., “Spectral Enhancement by Tracking Speech Presence Probability in Subbands,” Proc. International Workshop in Hand-Free Speech Communication (HSC'01), 2001, pp. 95-98. [cited by applicant]
Coifman, R.R., et al., “Diffusion maps,” Applied and Computational Harmonic Analysis, vol. 21, 2006, pp. 5-30. [cited by applicant]
Hayes, M.H., “Statistical Digital Signal Processing and Modeling,” J. Wiley & Sons, Inc., New York, 1996, 200 pages. [cited by applicant]
Hermansky, H., “Perceptual linear predictive (PLP) analysis of speech,” Journal of the Acoustical Society of America, vol. 87, No. 4, 1990, pp. 1738-1752. [cited by applicant]
Lailler, C., et al., “Semi-Supervised and Unsupervised Data Extraction Targeting Speakers: From Speaker Roles to Fame?,” Proceedings of the First Workshop on Speech, Language and Audio in Multimedia (SLAM), Marseille, F… [cited by applicant]
Mermelstein, P., “Distance Measures for Speech Recognition—Psychological and Instrumental,” Pattern Recognition and Artificial Intelligence, 1976, pp. 374-388. [cited by applicant]
Schmalenstroeer, J., et al., “Online Diarization of Streaming Audio-Visual Data for Smart Environments,” IEEE Journal of Selected Topics in Signal Processing, vol. 4, No. 5, 2010, 12 pages. [cited by applicant]
Viterbi, A.J., “Error Bounds for Convolutional Codes and an Asymptotically Optimum Decoding Algorithm,” IEEE Transactions on Information Theory, vol. 13, No. 2, 1967, pp. 260-269. [cited by applicant]