IP Library Granted Patent US 12,315,517
Granted Patent B2
US 12,315,517 · App. 17/665,672 · Granted May 27, 2025

Method and system for correcting speaker diarization using speaker change detection based on text

Inventors: Namkyu Jung (Seongnam-si, KR); Geonmin Kim (Seongnam-si, KR); Youngki Kwon (Seongnam-si, KR); Hee Soo Heo (Seongnam-si, KR); Bong-Jin Lee (Seongnam-si, KR); Chan Kyu Lee (Seongnam-si, KR)
Assignees: NAVER CORPORATION; LINE WORKS CORP.
G10L17/14G10L17/22G10L21/028
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,315,517
App. No.
17/665,672
Granted
May 27, 2025
Kind
B2
Abstract

A method and system for correcting speaker diarization using a text-based speaker change detection. A speaker diarization correction method may include performing speaker diarization on an input audio stream; recognizing speech included in the input audio stream and converting the speech to text; detecting a speaker change based on the converted text; and correcting the speaker diarization based on the detected speaker change.

Claims (37)

1. A speaker diarization correction method of a computer apparatus comprising at least one processor, the method, which uses the at least one processor, comprising:

performing speaker diarization on an input audio stream;

recognizing a speech included in the input audio stream and converting the speech to text;

detecting a speaker change based on the converted text; and

correcting the speaker diarization based on the detected speaker change,

wherein the detecting of the speaker change comprises:

receiving a speech recognition result for each utterance section, wherein each utterance section consists of at least one word unit, and further wherein each word unit comprises a single word of text;

encoding text included in the speech recognition result for each utterance section to one or more word units of text, wherein the encoding of the text to the one or more word units of text comprises encoding an EndPoint Detection (EPD) unit text included in the speech recognition result for each utterance section to the one or more word units of text using sentence Bidirectional Encoder Representations from Transformers (sBERT);

encoding each of the word units of text to consider a conversation context; and

determining whether a speaker change compared to a previous word unit of text is present for each word unit of text, individually, in which the conversation context is considered.

2. The method of claim 1 , wherein the detecting of the speaker change comprises recognizing a speaker change status for each word unit of text using a module that is trained to receive a speech recognition result for each utterance section and to output a speaker change probability of a word unit.

3. The method of claim 1 , wherein the encoding of the word unit of text to consider the conversation context comprises encoding the word unit of text to consider the conversation context using dialog Bidirectional Encoder Representations from Transformers (dBERT).

4. The method of claim 1 , wherein the correcting comprises correcting the speaker diarization based on the word unit depending on whether the speaker change is present for each word unit of text.

5. A non-transitory computer-readable record medium storing instructions that, when executed by a processor, cause the processor to perform the following method:

performing speaker diarization on an input audio stream;

recognizing a speech included in the input audio stream and converting the speech to text;

detecting a speaker change based on the converted text; and

correcting the speaker diarization based on the detected speaker change,

wherein the detecting of the speaker change comprises:

receiving a speech recognition result for each utterance section, wherein each utterance section consists of at least one word unit, and further wherein each word unit comprises a single word of text;

encoding text included in the speech recognition result for each utterance section to one or more word units of text, wherein the encoding of the text to the one or more word units of text comprises encoding an EndPoint Detection (EPD) unit text included in the speech recognition result for each utterance section to the one or more word units of text using sentence Bidirectional Encoder Representations from Transformers (sBERT);

encoding each of the word units of text to consider a conversation context; and

determining whether a speaker change compared to a previous word unit of text is present for each word unit of text, individually, in which the conversation context is considered.

6. A computer apparatus comprising:

at least one processor configured to execute computer-readable instructions,

wherein the at least one processor causes the computer apparatus to:

perform speaker diarization on an input audio stream,

recognize speech included in the input audio stream and convert the speech to text,

detect a speaker change based on the converted text, and

correct the speaker diarization based on the detected speaker change,

wherein the detecting of the speaker change comprises:

receiving a speech recognition result for each utterance section wherein each utterance section consists of at least one word unit, and further wherein each word unit comprises a single word of text;

encoding text included in the speech recognition result for each utterance section to one or more word units of text, wherein the encoding of the text to the one or more word units of text comprises encoding an EndPoint Detection (EPD) unit text included in the speech recognition result for each utterance section to the one or more word units of text using sentence Bidirectional Encoder Representations from Transformers (sBERT);

encoding each of the word units of text to consider a conversation context; and

determining whether a speaker change compared to a previous word unit of text is present for each word unit of text, individually, in which the conversation context is considered.

7. The computer apparatus of claim 6 , wherein, to detect the speaker change, the at least one processor causes the computer apparatus to recognize a speaker change status for each word unit of text using a module that is trained to receive a speech recognition result for each utterance section and to output a speaker change probability of a word unit.

8. The computer apparatus of claim 6 , wherein the encoding of the word unit of text to consider the conversation context comprises encoding the word unit of text to consider the conversation context using dialog Bidirectional Encoder Representations from Transformers (dBERT).

Assignments (3)
CHANGE OF NAME Recorded Mar 7, 2024
From: WORKS MOBILE JAPAN CORPORATION
To: LINE WORKS CORP.
Reel/Frame 066684/0098 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2023
From: LINE CORPORATION
To: WORKS MOBILE JAPAN CORPORATION
Reel/Frame 064807/0656 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 7, 2022
From: JUNG, NAMKYU; KIM, GEONMIN; KWON, YOUNGKI; HEO, HEE SOO; LEE, BONG-JIN; LEE, CHAN KYU
To: NAVER CORPORATION; LINE CORPORATION
Reel/Frame 058906/0026 →
Priority Claims (1)
KR 10-2021-0017814 · Feb 8, 2021 · national
Continuity (1)
Related Publication 20220254351A1 · Aug 11, 2022
References Cited (8)
US 20160225374A1 · Rodriguez · 2016 [cited by applicant]
JP 5296455B2 · 2013 [cited by applicant]
JP 2020140169A · 2020 [cited by applicant]
KR 1020140014318 · 2015 [cited by applicant]
KR 102208387B1 · 2021 [cited by applicant]
KR 1020210009617A · 2021 [cited by applicant]
Meng et al. “Hierarchical RNN with Static Sentence-Level Attention for Text-based Speaker Change Detection” ArXiv:1703.07713v2 [cs.CL]Sep. 28, 2018 (Year: 2018). [cited by examiner]
Reimers, Nils, and Iryna Gurevych. âSentence-BERT: Sentence Embeddings Using Siamese BERT-Networks.â ArXiv (Cornell University). Ithaca: Cornell University Library, arXiv.org, 2019. Web. (Year: 2019). [cited by examiner]