IP Library › Granted Patent US 12,431,150
Granted Patent B2
US 12,431,150 · App. 18/118,682 · Granted Sep 30, 2025

Method and apparatus for reconstructing voice conversation

Inventors: Myeongjin Hwang (Seoul, KR); Suntae Kim (Seoul, KR); Changjin Ji (Seoul, KR)
Assignee: LLSOLLU CO., LTD.
G10L19/022G10L15/22G10L25/87
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,431,150
App. No.
18/118,682
Granted
Sep 30, 2025
Kind
B2
Abstract

A voice conversation reconstruction method performed by a voice conversation reconstruction apparatus is disclosed. The method includes acquiring speaker-specific voice recognition data about voice conversation, dividing the speaker-specific voice recognition data into a plurality of blocks using a boundary between tokens according to a predefined division criterion, arranging the plurality of blocks in chronological order irrespective of a speaker, merging blocks from continuous utterance of the same speaker among the arranged plurality of blocks, and reconstructing the plurality of blocks subjected to the merging in a conversation format in chronological order and based on a speaker.

Claims (44)

1. A voice conversation reconstruction method performed by a voice conversation reconstruction apparatus, the method comprising:

acquiring a plurality of speaker-specific voice recognition data corresponding to a plurality of speakers about voice conversation;

dividing each of the plurality of the speaker-specific voice recognition data into a plurality of blocks using a boundary between tokens such that each of the divided plurality of the speaker-specific voice recognition data includes voice data only by a single speaker, wherein the divided plurality of the speaker-specific voice recognition data are not in chronological order;

arranging the plurality of blocks of all the speaker-specific voice recognition data in chronological order without distinction of speaker;

among the arranged plurality of blocks, merging blocks when the blocks are neighbor and the speaker of the blocks are the same such that the speaker-specific voice recognition data in each of the merged blocks are in chronological order and include voice data only by the same speaker; and

reconstructing the plurality of blocks subjected to the merging in a conversation format in chronological order and based on a speaker step by step such that the speaker-specific voice recognition data in each of the reconstructed blocks are in chronological order and include voice data only by the same speaker,

wherein the steps are performed in order,

wherein acquiring the plurality of speaker-specific voice recognition data includes:

acquiring a first speaker-specific recognition result generated on an End Point Detection (EPD), and a second speaker-specific recognition result generated every preset time, and

collecting the first speaker-specific recognition result and the second speaker-specific recognition result without overlap and redundance therebetween to generate the speaker-specific voice recognition data, and

wherein the second speaker-specific recognition result is generated after a last EPD at which the first speaker-specific recognition result is generated occurs.

2. The method of claim 1 , wherein acquiring the speaker-specific voice recognition data includes:

acquiring a speaker-specific recognition result A generated on an EPD (End Point Detection) basis from the voice conversation and a speaker-specific recognition result B which is a partial result generated from the voice conversation; and

when the speaker of the A and the speaker of the B are the same, collecting the recognition result A and the recognition result B without overlap therebetween to generate the speaker-specific voice recognition data.

3. The method of claim 2 , wherein the second speaker-specific recognition result B is generated after the same speaker's last EPD occurs.

4. The method of claim 1 , wherein the merging is not performed when a silence period between neighboring tokens is longer than a predetermined time, or is not grammatically connected.

5. The method of claim 2 , wherein the method further comprises outputting the voice recognition data reconstructed in the conversation format on a screen, and wherein when the screen is updated, the speaker-specific voice recognition data is collectively updated or is updated based on the speaker-specific recognition result B.

6. A voice conversation reconstruction apparatus comprising:

an input unit configured to receive voice conversation input; and

a processor configured to process voice recognition of the voice conversation received through the input unit,

wherein the processor is configured to:

acquire a plurality of speaker-specific voice recognition data corresponding to a plurality of speakers about voice conversation;

divide each of the plurality of the speaker-specific voice recognition data into a plurality of blocks using a boundary between tokens such that each of the divided plurality of the speaker-specific voice recognition data includes voice data only by a single speaker, wherein the divided plurality of the speaker-specific voice recognition data are not in chronological order;

arrange the plurality of blocks of all the speaker-specific voice recognition data in chronological order without distinction of speaker;

merge blocks when the blocks are neighbor and the speaker of the blocks are the same such that the speaker-specific voice recognition data in each of the merged blocks are in chronological order and include voice data only by the same speaker; and

reconstruct the plurality of blocks subjected to the merging in a conversation format in chronological order and based on a speaker such that the speaker-specific voice recognition data in each of the reconstructed blocks are in chronological order and include voice data only by the same speaker,

wherein the processor is further configured to:

acquire a first speaker-specific recognition result generated on an End Point Detection (EPD), and a second speaker-specific recognition result generated every preset time, and

collect the first speaker-specific recognition result and the second speaker-specific recognition result without overlap and redundance therebetween to generate the speaker-specific voice recognition data, and

wherein the second speaker-specific recognition result is configured to be generated after a last EPD at which the first speaker-specific recognition result is generated occurs.

7. The apparatus of claim 6 , wherein the processor is further configured to:

acquire a speaker-specific recognition result A generated on an EPD (End Point Detection) basis from the voice conversation and a speaker-specific recognition result B which is a partial result generated from the voice conversation; and

collect the speaker-specific recognition result A and the speaker-specific recognition result B without overlap and redundance therebetween to generate the speaker-specific voice recognition data.

8. A non-transitory computer-readable recording medium storing instructions, when executed by one or more processors, that cause the one or more processors to perform a method comprising:

acquiring a plurality of speaker-specific voice recognition data corresponding to a plurality of speakers about voice conversation;

dividing each of the plurality of the speaker-specific voice recognition data into a plurality of blocks using a boundary between tokens such that each of the divided plurality of the speaker-specific voice recognition data includes voice data only by a single speaker, wherein the divided plurality of the speaker-specific voice recognition data are not in chronological order;

arranging the plurality of blocks of all the speaker-specific voice recognition data in chronological order without distinction of speaker;

among the arranged plurality of blocks, merging blocks when the blocks are neighbor and the speaker of the blocks are the same such that the speaker-specific voice recognition data in each of the merged blocks are in chronological order and include voice data only by the same speaker; and

reconstructing the plurality of blocks subjected to the merging in a conversation format in chronological order and based on a speaker step by step such that the speaker-specific voice recognition data in each of the reconstructed blocks are in chronological order and include voice data only by the same speaker,

wherein the steps are performed in order,

wherein acquiring the plurality of speaker-specific voice recognition data includes:

acquiring a first speaker-specific recognition result generated on an End Point Detection (EPD), and a second speaker-specific recognition result generated every preset time, and

collecting the first speaker-specific recognition result and the second speaker-specific recognition result without overlap and redundance therebetween to generate the speaker-specific voice recognition data, and

wherein the second speaker-specific recognition result is generated after a last EPD at which the first speaker-specific recognition result is generated occurs.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2023
From: HWANG, MYEONGJIN; KIM, SUNTAE; JI, CHANGJIN
To: LLSOLLU CO., LTD.
Reel/Frame 062925/0011 →
Priority Claims (1)
KR 10-2020-0029826 · Mar 10, 2020 · national
Continuity (2)
Continuation In Part 17198046 · Mar 10, 2021
Related Publication 20230223032A1 · Jul 13, 2023
References Cited (41)
US 6055495A · Tucker · 2000 [cited by examiner]
US 7295970B1 · Gorin · 2007 [cited by examiner]
US 7668718B2 · Kahn et al. · 2010 [cited by applicant]
US 10089067B1 · Abuelsaad et al. · 2018 [cited by applicant]
US 10636427B2 · Jung et al. · 2020 [cited by applicant]
US 12380910B2 · Kukde · 2025 [cited by examiner]
US 20040162724A1 · Hill et al. · 2004 [cited by applicant]
US 20060149558A1 · Kahn · 2006 [cited by examiner]
US 20080154594A1 · Itoh et al. · 2008 [cited by applicant]
US 20090265166A1 · Abe · 2009 [cited by applicant]
US 20100305945A1 · Krishnaswamy et al. · 2010 [cited by applicant]
US 20140074467A1 · Ziv · 2014 [cited by examiner]
US 20160283185A1 · McLaren · 2016 [cited by examiner]
US 20180020285A1 · Zass · 2018 [cited by applicant]
US 20190080688A1 · Itsui · 2019 [cited by applicant]
US 20190318743A1 · Reshef · 2019 [cited by examiner]
US 20190392837A1 · Jung et al. · 2019 [cited by applicant]
US 20200135204A1 · Robichaud · 2020 [cited by examiner]
US 20200175961A1 · Thomson · 2020 [cited by examiner]
US 20240126994A1 · Deilamsalehy · 2024 [cited by examiner]
CN 1629935A · 2005 [cited by applicant]
CN 1708997A · 2005 [cited by applicant]
CN 107430626A · 2017 [cited by applicant]
CN 107430853A · 2017 [cited by applicant]
CN 110851470A · 2020 [cited by applicant]
JP 2000112931A · 2000 [cited by applicant]
JP 2012003701A · 2012 [cited by applicant]
JP 2016085697A · 2016 [cited by applicant]
JP 2017161850A · 2017 [cited by applicant]
JP 2017182822A · 2017 [cited by applicant]
JP 6517419B1 · 2019 [cited by applicant]
KR 1020140078258 · 2014 [cited by applicant]
KR 1020190125154A · 2019 [cited by applicant]
KR 1020200011198A · 2020 [cited by applicant]
WO WO2009104332A1 · 2009 [cited by applicant]
WO WO2010113438A1 · 2010 [cited by applicant]
Hotta, et al., “Detecting Whether Incorrectly-Segmented Utterance Needs to be Restored or not”, SIG-SLUD-B303, pp. 45-52 (Feb. 26, 2014). [cited by applicant]
Office Action in Japanese Application No. 2021-038052 dated May 29, 2024 and English translation. [cited by applicant]
Office Action in Korean Application No. 10-2020-0029826 dated Jun. 25, 2020 and English translation. [cited by applicant]
First Office Action in CN Application No. 202110255584.7 dated Sep. 22, 2023. [cited by applicant]
Notice of Allowance in JP Application No. 2021-038052 dated Apr. 15, 2025. [cited by applicant]