IP Library Granted Patent US 12,417,776
Granted Patent B2
US 12,417,776 · App. 18/046,041 · Granted Sep 16, 2025

Online speaker diarization using local and global clustering

Inventors: Myungjong Kim (Milpitas, CA); Taeyeon Ki (Milpitas, CA); Vijendra Raj Apsingekar (San Jose, CA); Sungjae Park (Seoul, KR); SeungBeom Ryu (Suwon, KR); Hyuk Oh (Seoul, KR)
Assignee: Samsung Electronics Co., Ltd.
G10L21/028G10L17/02G10L17/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,776
App. No.
18/046,041
Granted
Sep 16, 2025
Kind
B2
Abstract

A method includes obtaining at least a portion of an audio stream containing speech activity. At least the portion of the audio stream includes multiple segments. The method also includes, for each of the segments, generating an embedding vector that represents the segment. The method further includes, within each of multiple local windows, clustering the embedding vectors into one or more clusters to perform speaker identification. Different clusters correspond to different speakers. The method also includes presenting at least one first sequence of speaker identities based on the speaker identification for the local windows. The method further includes, within each of multiple global windows, clustering the embedding vectors into one or more clusters to perform speaker identification. Each global window includes two or more of the local windows. The method further includes presenting at least one second sequence of speaker identities based on the speaker identification for the global windows.

Claims (60)

1. A method for performing online speaker diarization, the method comprising:

obtaining at least a portion of an audio stream containing speech activity, at least the portion of the audio stream comprising multiple segments;

for each of the multiple segments, generating an embedding vector that represents the segment;

within each of multiple local windows, performing monotonically-increasing chunk-based spectral clustering of the embedding vectors into one or more clusters to perform speaker identification, wherein different clusters correspond to different speakers, each local window having more than one of the segments;

presenting at least one first sequence of speaker identities based on the speaker identification performed for the local windows;

within each of multiple global windows, clustering the embedding vectors into one or more clusters to perform speaker identification, wherein each global window includes two or more of the local windows; and

presenting at least one second sequence of speaker identities based on the speaker identification performed for the global windows;

wherein performing the monotonically-increasing chunk-based spectral clustering comprises clustering different subsets of the embedding vectors in the local window such that each subsequent subset includes (i) the embedding vectors of at least one prior subset and (ii) at least one additional embedding vector not included in the at least one prior subset.

2. The method of claim 1 , further comprising:

matching entries in different sequences of speaker identities; and

replacing at least some of the entries in one sequence of speaker identities with at least some of the entries in another sequence of speaker identities so that clusters generated in different windows correspond to common speakers.

3. The method of claim 1 , wherein the local windows partially overlap one another.

4. The method of claim 1 , wherein:

the global windows comprise first and second global windows; and

the second global window is larger than and includes the first global window.

5. The method of claim 1 , wherein a final one of the global windows encompasses an entirety of the audio stream.

6. The method of claim 1 , further comprising:

determining a similarity between clusters of the embedding vectors in different ones of the local windows; and

in response to determining that the similarity exceeds a threshold, using a common cluster identifier for the clusters of the embedding vectors in the different local windows.

7. An apparatus comprising:

at least one processing device configured to:

obtain at least a portion of an audio stream containing speech activity, at least the portion of the audio stream comprising multiple segments;

for each of the multiple segments, generate an embedding vector that represents the segment;

within each of multiple local windows, perform monotonically-increasing chunk-based spectral clustering of the embedding vectors into one or more clusters to perform speaker identification, wherein different clusters correspond to different speakers, each local window having more than one of the segments;

initiate presentation of at least one first sequence of speaker identities based on the speaker identification performed for the local windows;

within each of multiple global windows, cluster the embedding vectors into one or more clusters to perform speaker identification, wherein each global window includes two or more of the local windows; and

initiate presentation of at least one second sequence of speaker identities based on the speaker identification performed for the global windows;

wherein, to perform the monotonically-increasing chunk-based spectral clustering, the at least one processing device is configured to cluster different subsets of the embedding vectors in the local window such that each subsequent subset includes (i) the embedding vectors of at least one prior subset and (ii) at least one additional embedding vector not included in the at least one prior subset.

8. The apparatus of claim 7 , wherein the at least one processing device is further configured to:

match entries in different sequences of speaker identities; and

replace at least some of the entries in one sequence of speaker identities with at least some of the entries in another sequence of speaker identities so that clusters generated in different windows correspond to common speakers.

9. The apparatus of claim 7 , wherein the local windows partially overlap one another.

10. The apparatus of claim 7 , wherein:

the global windows comprise first and second global windows; and

the second global window is larger than and includes the first global window.

11. The apparatus of claim 7 , wherein a final one of the global windows encompasses an entirety of the audio stream.

12. The apparatus of claim 7 , wherein the at least one processing device is further configured to:

determine a similarity between clusters of the embedding vectors in different ones of the local windows; and

in response to determining that the similarity exceeds a threshold, use a common cluster identifier for the clusters of the embedding vectors in the different local windows.

13. A non-transitory computer readable medium containing instructions that when executed cause at least one processor to:

obtain at least a portion of an audio stream containing speech activity, at least the portion of the audio stream comprising multiple segments;

for each of the multiple segments, generate an embedding vector that represents the segment;

within each of multiple local windows, perform monotonically-increasing chunk-based spectral clustering of the embedding vectors into one or more clusters to perform speaker identification, wherein different clusters correspond to different speakers, each local window having more than one of the segments;

initiate presentation of at least one first sequence of speaker identities based on the speaker identification performed for the local windows;

within each of multiple global windows, cluster the embedding vectors into one or more clusters to perform speaker identification, wherein each global window includes two or more of the local windows; and

initiate presentation of at least one second sequence of speaker identities based on the speaker identification performed for the global windows;

wherein the instructions that when executed cause the at least one processor to perform the monotonically-increasing chunk-based spectral clustering comprise instructions that when executed cause the at least one processor to cluster different subsets of the embedding vectors in the local window such that each subsequent subset includes (i) the embedding vectors of at least one prior subset and (ii) at least one additional embedding vector not included in the at least one prior subset.

14. The non-transitory computer readable medium of claim 13 , further containing instructions that when executed cause the at least one processor to:

match entries in different sequences of speaker identities; and

replace at least some of the entries in one sequence of speaker identities with at least some of the entries in another sequence of speaker identities so that clusters generated in different windows correspond to common speakers.

15. The non-transitory computer readable medium of claim 13 , wherein the local windows partially overlap one another.

16. The non-transitory computer readable medium of claim 13 , wherein:

the global windows comprise first and second global windows; and

the second global window is larger than and includes the first global window.

17. The non-transitory computer readable medium of claim 13 , further containing instructions that when executed cause the at least one processor to:

determine a similarity between clusters of the embedding vectors in different ones of the local windows; and

in response to determining that the similarity exceeds a threshold, use a common cluster identifier for the clusters of the embedding vectors in the different local windows.

18. The method of claim 1 , wherein portions of processing during the clustering of the embedding vectors in the local windows and portions of processing during the clustering of the embedding vectors in the global windows occur in parallel.

19. The apparatus of claim 7 , wherein the at least one processing device is configured to perform portions of processing during the clustering of the embedding vectors in the local windows and portions of processing during the clustering of the embedding vectors in the global windows in parallel.

20. The non-transitory computer readable medium of claim 13 , wherein the instructions when executed cause the at least one processor to perform portions of processing during the clustering of the embedding vectors in the local windows and portions of processing during the clustering of the embedding vectors in the global windows in parallel.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 12, 2022
From: KIM, MYUNGJONG; KI, TAEYEON; APSINGEKAR, VIJENDRA RAJ; PARK, SUNGJAE; RYU, SEUNGBEOM; OH, HYUK
To: SAMSUNG ELECTRONICS CO., LTD
Reel/Frame 061398/0549 →
Continuity (2)
Provisional Application 63356259 · Jun 28, 2022
Related Publication 20230419979A1 · Dec 28, 2023
References Cited (25)
US 10468031B2 · Church et al. · 2019 [cited by applicant]
US 11024291B2 · Castan Lavilla et al. · 2021 [cited by applicant]
US 11152013B2 · Song et al. · 2021 [cited by applicant]
US 11227605B2 · Grancharov et al. · 2022 [cited by applicant]
US 20190340944A1 · Liu et al. · 2019 [cited by applicant]
US 20200135204A1 · Robichaud et al. · 2020 [cited by applicant]
US 20200302939A1 · Khoury et al. · 2020 [cited by applicant]
US 20220115020A1 · Bradley et al. · 2022 [cited by applicant]
US 20220122615A1 · Chen et al. · 2022 [cited by applicant]
US 20220199091A1 · Kanda · 2022 [cited by examiner]
US 20230089308A1 · Wang · 2023 [cited by examiner]
CN 113571090A · 2021 [cited by applicant]
EP 1669980B1 · 2009 [cited by applicant]
KR 1020210149336A · 2021 [cited by applicant]
WO 2021026617A1 · 2021 [cited by applicant]
WO 2021225403A1 · 2021 [cited by applicant]
J. M. Coria, H. Bredin, S. Ghannay and S. Rosset, “Overlap-Aware Low-Latency Online Speaker Diarization Based on End-to-End Local Segmentation,” 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), … [cited by examiner]
Xia et al., “Turn-to-Diarize: Online Speaker Diarization Constrained by Transformer Transducer Speaker Turn Detection,” 2022 IEEE International Conference on Acoustics, Speech and Signal Processing, Jan. 2022, 8 pages. [cited by applicant]
Han et al., “BW-EDA-EEND: Streaming End-to-End Neural Speaker Diarization for a Variable No. of Speakers,” 2021 IEEE International Conference on Acoustics, Speech and Signal Processing, Feb. 2021, 5 pages. [cited by applicant]
Zhang et al., “Fully Supervised Speaker Diarization,” 2019 IEEE International Conference on Acoustics, Speech and Signal Processing, Feb. 2019, 5 pages. [cited by applicant]
Zhang et al., “Low-Latency Online Speaker Diarization with Graph-Based Label Generation,” Speaker and Language Recognition Workshop (Odyssey 2022), Jun. 2022, 8 pages. [cited by applicant]
International Search Report and Written Opinion of the International Searching Authority dated Aug. 30, 2023 in connection with International Patent Application No. PCT/KR2023/007432, 10 pages. [cited by applicant]
Lu et al., “SCAN: Learning Speaker Identity from Noisy Sensor Data,” Proceedings of the 16th ACM/IEEE International Conference on Information Processing in Sensor Networks, Apr. 2017, 12 pages. [cited by applicant]
Supplementary European Search Report dated Feb. 3, 2025 in connection with European Patent Application No. 23831747.3, 9 pages. [cited by applicant]
Wang et al., “Speaker Diarization with LSTM,” arXiv:1710.10468v7 [eess.AS], Jan. 2022, 5 pages. [cited by applicant]