IP Library › Granted Patent US 12,230,277
Granted Patent B2
US 12,230,277 · App. 17/962,011 · Granted Feb 18, 2025

Method of generating summary based on main speaker

Inventors: Seongmin Park (Seoul, KR); Seungho Kwak (Seoul, KR)
Assignee: ActionPower Corp.
G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,230,277
App. No.
17/962,011
Granted
Feb 18, 2025
Kind
B2
Abstract

Disclosed are a method, a device, and a program for selecting a main speaker among speakers included in a sound source or a conversation record based on the sound source or the conversation record including conversation contents of at least one speaker and generating a summary based on the main speaker. A method of generating a summary for a sound source, the method being performed by at least one computing device, includes: generating a speak score for at least one speaker based on the sound source; determining a main speaker of the sound source based on a speak score for said at least one speaker; and generating a summary for the sound source in consideration of the determined main speaker.

Claims (63)

1. A method of generating a summary for a sound source data, the method being performed by at least one computing device, the method comprising:

generating a speak score for at least one speaker based on the sound source data, based on a weighted value sum of at least one of a degree of dispersion of the speak or a frequency of the speak of the speaker, wherein the frequency of the speak is determined by:

removing noise from the sound source data using a VAD module, wherein the VAD module receives the sound source and repeats a binary classification algorithm in which, when the speak of the speaker is recognized at a predetermined interval of the sound source, the VAD module outputs a numerical progression and wherein the VAD module further includes a probability distribution-based classification algorithm for determining whether the distribution is similar to the speak or the noise based on the distribution of the speak and the distribution of the noise;

extracting a plurality of feature vectors for the speaks included in the sound source data, wherein the feature vectors represent data points within the latent space that encapsulates the characteristics of the speak;

analyzing each of the plurality of feature vectors for similarities;

clustering each of the plurality of feature vectors based on the analyzed similarities;

distinguishing the plurality of speakers from each other based on the clustered feature vectors corresponding to each of the speaks; and

wherein the degree of dispersion of the speak is determined by:

identifying one or more speak indices for each of the at least one speaker based on the sound source data; and

calculating the degree of dispersion for the speaks of each speaker based on the identified one or more speak indices;

determining a main speaker of the sound source data based on the speak score for said at least one speaker; and

generating the summary for the sound source data in consideration of the determined main speaker.

2. The method of claim 1 , wherein the calculating of the degree of dispersion for one or more speaks of each speaker further includes:

identifying one or more speak indices for each speaker in a script generated based on the sound source data; and

calculating the degree of dispersion for one or more speaks of each speaker based on the identified one or more speak indices.

3. The method of claim 1 , wherein the generating of the speak score for said at least one speaker further includes distinguishing a plurality of speakers associated with the sound source data from each other.

4. The method of claim 1 , wherein the generating of the summary for the sound source data in consideration of the determined main speaker includes generating the summary of the sound source data by assigning a weighted value to the speaks associated with the determined main speaker among the speaks included in the sound source data.

5. The method of claim 4 , wherein the sound source data is converted into a plurality of tokens, and

the generating of the summary of the sound source data by assigning the weighted value to the speak associated with the determined main speaker further includes:

assigning a weighted value to the token associated with the determined main speaker;

selecting some tokens from among the plurality of tokens in consideration of the weighted value;

generating a sentence to be included in the summary based on the selected some tokens; and

generating the summary based on the generated sentence.

6. The method of claim 4 , wherein the sound source data is converted into a language graph including a plurality of nodes, and

the generating of the summary of the sound source data by assigning the weighted value to the speak associated with the determined main speaker further includes:

assigning a weighted value to a node associated with the determined main speaker;

selecting some nodes from among the plurality of nodes in consideration of the weighted value;

generating a sentence to be included in the summary based on the selected some nodes; and

generating the summary based on the generated sentence.

7. The method of claim 4 , wherein the sound source data is converted into a plurality of sentences, and

the generating of the summary of the sound source data by assigning the weighted value to the speak associated with the determined main speaker further includes:

assigning a weighted value to a sentence associated with the determined main speaker;

extracting some sentences from among the plurality of sentences in consideration of the weighted value; and

generating the summary based on the extracted some sentences.

8. A device, comprising:

at least one processor; and

a memory,

wherein the processor is configured to:

generate a speak score for at least one speaker based on a sound source data, based on a weighted value sum of at least one of a degree of dispersion of the speak or a frequency of the speak of the speaker;

wherein the frequency of the speak is determined by:

removing noise from the sound source data using a VAD module, wherein the VAD module receives the sound source and repeats a binary classification algorithm in which, when the speak of the speaker is recognized at a predetermined interval of the sound source, the VAD module outputs a numerical progression and wherein the VAD module further includes a probability distribution-based classification algorithm for determining whether the distribution is similar to the speak or the noise based on the distribution of the speak and the distribution of the noise;

extracting a plurality of feature vectors for the speaks included in the sound source data, wherein the feature vectors represent data points within the latent space that encapsulates the characteristics of the speak;

analyzing each of the plurality of feature vectors for similarities;

clustering each of the plurality of feature vectors based on the analyzed similarities;

distinguishing the plurality of speakers from each other based on the clustered feature vectors corresponding to each of the speaks; and

wherein the degree of dispersion of the speak is determined by:

identifying one or more speak indices for each of the at least one speaker based on the sound source data; and

calculating the degree of dispersion for the speaks of each speaker based on the identified one or more speak indices;

determine a main speaker of the sound source data based on the speak score for said at least one speaker; and

generate a summary for the sound source data in consideration of the determined main speaker.

9. A computer program stored in a non-transitory computer-readable storage medium, the computer program causing at least one processor to perform operations for generating a summary for a sound source, the operations comprising:

an operation of generating a speak score for at least one speaker based on a sound source data, based on a weighted value sum of at least one of a degree of dispersion of the speak or a frequency of the speak of the speaker;

wherein the frequency of the speak is determined by:

removing noise from the sound source data using a VAD module, wherein the VAD module receives the sound source and repeats a binary classification algorithm in which, when the speak of the speaker is recognized at a predetermined interval of the sound source, the VAD module outputs a numerical progression and wherein the VAD module further includes a probability distribution-based classification algorithm for determining whether the distribution is similar to the speak or the noise based on the distribution of the speak and the distribution of the noise;

extracting a plurality of feature vectors for the speaks included in the sound source data, wherein the feature vectors represent data points within the latent space that encapsulates the characteristics of the speak;

analyzing each of the plurality of feature vectors for similarities;

clustering each of the plurality of feature vectors based on the analyzed similarities;

distinguishing the plurality of speakers from each other based on the clustered feature vectors corresponding to each of the speaks; and

wherein the degree of dispersion of the speak is determined by:

identifying one or more speak indices for each of the at least one speaker based on the sound source data; and

calculating the degree of dispersion for the speaks of each speaker based on the identified one or more speak indices;

an operation of determining a main speaker of the sound source data based on the speak score for said at least one speaker; and

an operation of generating the summary for the sound source data in consideration of the determined main speaker.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 12, 2022
From: PARK, SEONGMIN; KWAK, SEUNGHO
To: ACTIONPOWER CORP.
Reel/Frame 061745/0146 →
Priority Claims (1)
KR 10-2022-0078041 · Jun 27, 2022 · national
Continuity (1)
Related Publication 20230419968A1 · Dec 28, 2023
References Cited (13)
US 10986046B2 · Kim · 2021 [cited by applicant]
US 11232266B1 · Biswas · 2022 [cited by examiner]
US 20150348538A1 · Donaldson · 2015 [cited by examiner]
US 20180342240A1 · Shellef et al. · 2018 [cited by applicant]
US 20220171936A1 · Wang · 2022 [cited by examiner]
JP 202263939A · 2022 [cited by applicant]
KR 101889809B1 · 2018 [cited by applicant]
KR 1020190096304 · 2019 [cited by applicant]
KR 102069695B1 · 2020 [cited by applicant]
KR 20200063346 · 2020 [cited by applicant]
KR 102298330B1 · 2021 [cited by applicant]
KR 102365611 · 2022 [cited by applicant]
Bak et al., Bert, A Leader's Final Decision Classification Model Tested on Meeting Records with Bert, ISSN 2383-630X (print) ISSN 2383-6296 (online), Journal of KIISE, vol. 48, No. 5, pp. 568-574, May 2021, 7 pgs. [cited by applicant]