IP Library › Patent Application 17198046
Patent Application
App. No. 17/198,046

METHOD AND APPARATUS FOR RECONSTRUCTING VOICE CONVERSATION

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/198,046
Abstract

A voice conversation reconstruction method performed by a voice conversation reconstruction apparatus is disclosed. The method includes acquiring speaker-specific voice recognition data about voice conversation, dividing the speaker-specific voice recognition data into a plurality of blocks using a boundary between tokens according to a predefined division criterion, arranging the plurality of blocks in chronological order irrespective of a speaker, merging blocks from continuous utterance of the same speaker among the arranged plurality of blocks, and reconstructing the plurality of blocks subjected to the merging in a conversation format in chronological order and based on a speaker.

Claims (37)

1 . A voice conversation reconstruction method performed by a voice conversation reconstruction apparatus, the method comprising:

acquiring a plurality of speaker-specific voice recognition data corresponding to a plurality of speakers about voice conversation;

dividing each of the plurality of the speaker-specific voice recognition data into a plurality of blocks using a boundary between tokens depending upon a predefined division criterion;

arranging the plurality of blocks of each of the plurality of the speaker-specific voice recognition data in chronological order irrespective of a speaker;

merging blocks from continuous utterance of the same speaker among the arranged plurality of blocks; and

reconstructing the plurality of blocks subjected to the merging in a conversation format in chronological order and based on a speaker.

2 . The method of claim 1 , wherein acquiring the speaker-specific voice recognition data includes:

acquiring a first speaker-specific recognition result generated on an EPD (End Point Detection) basis from the voice conversation and a second speaker-specific recognition result generated every preset time from the voice conversation; and

collecting the first speaker-specific recognition result and the second speaker-specific recognition result without overlap and redundance therebetween to generate the speaker-specific voice recognition data.

3 . The method of claim 2 , wherein the second speaker-specific recognition result is generated after a last EPD occurs.

4 . The method of claim 1 , wherein the predefined division criterion includes a silence period longer than or equal to a predetermined time duration or a morpheme feature related to a previous token.

5 . The method of claim 1 , wherein the merging include determining the continuous utterance from the same speaker based on a silence period shorter than or equal to a predetermined time duration or a syntax feature related to a previous block.

6 . The method of claim 2 , wherein the method further comprises outputting the voice recognition data reconstructed in the conversation format on a screen, wherein when the screen is updated, the speaker-specific voice recognition data is collectively updated or is updated based on the first speaker-specific recognition result.

7 . A voice conversation reconstruction apparatus comprising:

an input unit configured to receive voice conversation input; and

a processor configured to process voice recognition of the voice conversation received through the input unit,

wherein the processor is configured to:

acquire a plurality of speaker-specific voice recognition data corresponding to a plurality of speakers about voice conversation;

divide each of the plurality of the speaker-specific voice recognition data into a plurality of blocks using a boundary between tokens according to a predefined division criterion;

arrange the plurality of blocks of each of the plurality of the speaker-specific voice recognition data in chronological order irrespective of a speaker;

merge blocks from continuous utterance of the same speaker among the arranged plurality of blocks; and

reconstruct the plurality of blocks subjected to the merging in a conversation format in chronological order and based on a speaker.

8 . The apparatus of claim 7 , wherein the processor is further configured to:

acquire a first speaker-specific recognition result generated on an EPD (End Point Detection) basis from the voice conversation and a second speaker-specific recognition result generated every preset time from the voice conversation; and

collect the first speaker-specific recognition result and the second speaker-specific recognition result without overlap and redundance therebetween to generate the speaker-specific voice recognition data.

9 . A computer-readable recording medium storing therein a computer program, wherein the computer program includes instructions for enabling, when the instructions are executed by a processor, the processor to:

acquire a plurality of speaker-specific voice recognition data corresponding to a plurality of speakers about voice conversation;

divide each of the plurality of the speaker-specific voice recognition data into a plurality of blocks using a boundary between tokens according to a predefined division criterion;

arrange the plurality of blocks of each of the plurality of the speaker-specific voice recognition data in chronological order irrespective of a speaker;

merge blocks from continuous utterance of the same speaker among the arranged plurality of blocks; and

reconstruct the plurality of blocks subjected to the merging in a conversation format in chronological order and based on a speaker.

10 . A computer program stored in a computer-readable recording medium, wherein the computer program includes instructions for enabling, when the instructions are executed by a processor, the processor to:

acquire a plurality of speaker-specific voice recognition data corresponding to a plurality of speakers about voice conversation;

divide each of the plurality of the speaker-specific voice recognition data into a plurality of blocks using a boundary between tokens depending upon a predefined division criterion;

arrange the plurality of blocks of each of the plurality of the speaker-specific voice recognition data in chronological order irrespective of a speaker;

merge blocks from continuous utterance of the same speaker among the arranged plurality of blocks; and

reconstruct the plurality of blocks subjected to the merging in a conversation format in chronological order and based on a speaker.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2021
From: HWANG, MYEONGJIN; KIM, SUNTAE; JI, CHANGJIN
To: LLSOLLU CO., LTD.
Reel/Frame 055664/0342 →