IP Library Patent Application 17576492
Patent Application
App. No. 17/576,492

METHOD, SYSTEM, AND NON-TRANSITORY COMPUTER READABLE RECORD MEDIUM FOR SPEAKER DIARIZATION COMBINED WITH SPEAKER IDENTIFICATION

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/576,492
Abstract

Provided is a method, system, and non-transitory computer-readable record medium for speaker diarization combined with speaker identification. Provided is a speaker diarization method including setting a reference speech in relation to an audio file received as a speaker diarization target speech from a client; performing a speaker identification of identifying a speaker of the reference speech in the audio file using the reference speech; and performing a speaker diarization using clustering on a remaining utterance section unidentified in the audio file.

Claims (44)

1 . A speaker diarization method executed by a computer system comprising at least one processor configured to execute computer-readable instructions included in a memory, the speaker diarization method, which uses the at least one processor, comprising:

receiving an audio file including a diarization target speech from a client;

setting a reference speech in relation to the audio file received from the client;

performing a speaker identification of identifying a speaker of the reference speech in the audio file using the reference speech; and

performing a speaker diarization using clustering on any remaining unidentified utterance sections in the audio file.

2 . The speaker diarization method of claim 1 , wherein the setting of the reference speech comprises setting speech data including a label of a portion of speakers included in the audio file as the reference speech.

3 . The speaker diarization method of claim 1 , wherein the setting of the reference speech comprises receiving a selection on a speech of a portion of speakers included in the audio file from among speaker speeches pre-stored in a database related to the computer system and setting the selected speech as the reference speech.

4 . The speaker diarization method of claim 1 , wherein the setting of the reference speech comprises receiving an input of a speech of a portion of speakers included in the audio file through recording and setting the input speech as the reference speech.

5 . The speaker diarization method of claim 1 , wherein the performing of the speaker identification comprises:

verifying an utterance section corresponding to the reference speech from among utterance sections included in the audio file; and

mapping a speaker label of the reference speech to the utterance section corresponding to the reference speech.

6 . The speaker diarization method of claim 5 , wherein the verifying comprises verifying the utterance section corresponding to the reference speech based on a distance between an embedding extracted from the utterance section and an embedding extracted from the reference speech.

7 . The speaker diarization method of claim 5 , wherein the verifying comprises verifying the utterance section corresponding to the reference speech based on a distance between an embedding cluster that is a result of clustering an embedding extracted from the utterance section and an embedding extracted from the reference speech.

8 . The speaker diarization method of claim 5 , wherein the verifying comprises verifying the utterance section corresponding to the reference speech based on a result of clustering an embedding extracted from the reference speech with an embedding extracted from the utterance section.

9 . The speaker diarization method of claim 1 , wherein the performing of the speaker diarization comprises:

clustering an embedding extracted from the remaining utterance section; and

mapping an index of a cluster to the remaining utterance section.

10 . The speaker diarization method of claim 9 , wherein the clustering comprises:

calculating an affinity matrix based on the embedding extracted from the remaining utterance section;

extracting eigenvalues by performing an eigen decomposition on the affinity matrix;

sorting the extracted eigenvalues and determining a number of eigenvalues selected based on a difference between adjacent eigenvalues as a number of clusters; and

performing a speaker diarization clustering using the affinity matrix and the number of clusters.

11 . A non-transitory computer-readable record medium storing instructions that, when executed by a processor, cause the processor to computer-implement the speaker diarization method of claim. 1 .

12 . A computer system comprising:

at least one processor configured to execute computer-readable instructions included in a memory,

wherein the at least one processor comprises:

a reference setter configured to set a reference speech in relation to an audio file received as a speaker diarization target speech from a client;

a speaker identifier configured to perform speaker identification of identifying a speaker of the reference speech in the audio file using the reference speech; and

a speaker diarizer configured to perform speaker diarization using clustering on a remaining unidentified utterance section of the audio file.

13 . The computer system of claim 12 , wherein the reference setter is configured to set speech data including a label of a portion of speakers included in the audio file as the reference speech.

14 . The computer system of claim 12 , wherein the reference setter is configured to receive a selection on a speech of a portion of speakers included in the audio file from among speaker speeches pre-stored in a database related to the computer system and to set the selected speech as the reference speech.

15 . The computer system of claim 12 , wherein the reference setter is configured to receive an input of a speech of a portion of speakers included in the audio file through recording and to set the input speech as the reference speech.

16 . The computer system of claim 12 , wherein the speaker identifier is configured to:

verify an utterance section corresponding to the reference speech from among utterance sections included in the audio file, and

map a speaker label of the reference speech to the utterance section corresponding to the reference speech.

17 . The computer system of claim 16 , wherein the speaker identifier is configured to verify the utterance section corresponding to the reference speech based on a distance between an embedding extracted from the utterance section and an embedding extracted from the reference speech.

18 . The computer system of claim 16 , wherein the speaker identifier is configured to verify the utterance section corresponding to the reference speech based on a distance between an embedding cluster that is a result of clustering an embedding extracted from the utterance section and an embedding extracted from the reference speech.

19 . The computer system of claim 16 , wherein the speaker identifier is configured to verify the utterance section corresponding to the reference speech based on a result of clustering an embedding extracted from the reference speech with an embedding extracted from the utterance section.

20 . The computer system of claim 12 , wherein the speaker diarizer is configured to:

calculate an affinity matrix based on the embedding extracted from the remaining utterance section,

extract eigenvalues by performing an eigen decomposition on the affinity matrix,

sort the extracted eigenvalues and determine a number of eigenvalues selected based on a difference between adjacent eigenvalues as a number of clusters,

perform a speaker diarization clustering using the affinity matrix and the number of clusters, and

map an index of a cluster according to the speaker diarization clustering to the remaining utterance section.

Assignments (3)
CHANGE OF NAME Recorded Mar 7, 2024
From: WORKS MOBILE JAPAN CORPORATION
To: LINE WORKS CORP.
Reel/Frame 066684/0098 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2023
From: LINE CORPORATION
To: WORKS MOBILE JAPAN CORPORATION
Reel/Frame 064807/0656 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 14, 2022
From: KWON, YOUNGKI; KANG, HAN YONG; KIM, YOU JIN; KIM, HAN-GYU; LEE, BONG-JIN; JANG, JUNGHOON; HAN, ICKSANG; HEO, HEE SOO; CHUNG, JOON SON
To: NAVER CORPORATION; LINE CORPORATION
Reel/Frame 058663/0175 →