IP Library › Granted Patent US 10,276,164
Granted Patent B2
US 10,276,164 · App. 15/823,937 · Granted Apr 30, 2019

Multi-speaker speech recognition correction system

Inventor: Munhak An (Seoul, KR)
Assignee: SORIZAVA CO., LTD.
G10L15/26G06F17/279G06F17/28G06F17/30746G10L15/08G10L15/32G10L21/0272G10L21/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,276,164
App. No.
15/823,937
Granted
Apr 30, 2019
Kind
B2
Abstract

The present invention relates to a multi-speaker speech recognition correction system for determining a speaker of an utterance with a simple method and easily correcting speech-recognized text during speech recognition for a plurality of speakers. According to the present invention, when speech signals are input to a multi-speaker speech recognition system from a plurality of microphones which are each provided to a corresponding one of a plurality of speakers, the multi-speaker speech recognition correction system may detect a speech session from a time point at which input of each of the speech signals is started to a time point at which the input of the speech signal is stopped, and a speech recognizer may convert only the detected speech sessions into text so that a speaker of an utterance can be identified by a simple method and speech recognition can be carried out at a low cost.

Claims (12)

1. A multi-speaker speech recognition correction system comprising:

a speech signal detector configured to, when a plurality of speech signals are received from a plurality of microphones which are each provided to a corresponding one of a plurality of speakers, detect a speech session from a time point at which input of each of the speech signals is started to a time point at which the input of the speech signal is stopped;

a speech recognizer configured to receive each of the speech sessions including time information and microphone identification information and convert the each of speech sessions to text;

a speech combiner configured to receive speech sessions from the speech signal detector and combine the speech sessions in an order of time points at which the inputs of the speech signals are started; and

a text corrector configured to receive pieces of speech-recognized text from the speech recognizer, receive speaker information for changing the microphone identification information, arrange and display the speaker information and the pieces of speech-recognized text in the order of time points at which the inputs of the speech signals are started, output an image obtained by capturing each of the plurality of speakers, display speaker tags for identifying each of the speakers in the image, output a speech combined by the speech combiner together with the speech-recognized text, and receive information for correcting the speech-recognized text,

wherein the text corrector includes a real-time input mode in which the speech-recognized text is displayed and a speaker tag matching speaker information of the displayed text is highlighted for recognition, a correction mode in which a speaker tag matching speaker information of text which is to be corrected is highlighted for recognition when information for correcting the speech-recognized text is input in the real-time input mode, and a speaker-specified play mode in which speech-recognized text or utterances of speech sessions of a speaker matching a selected speaker tag is output according to time when a selection signal is input for each of the speaker tags,

the text corrector pauses the display of the text when the information for correcting the speech-recognized text is received, and, when the correction is completed, the text corrector resumes the display of the text by returning to a point in time a predetermined amount of time in the past,

the text corrector transmits characteristic information, including a dialect, foreign words, exclamations, or fillers, of a speaker corresponding to each piece of the microphone identification information to the speech recognizer in advance, and

the speech recognizer converts the dialect to standard language, converts a foreign word to a native word, or removes an exclamation or a filler, which is a habit of the speaker, by applying the characteristic information received from the text corrector, and transmits a result thereof to the text corrector.

2. The multi-speaker speech recognition correction system of claim 1 , wherein the text corrector displays a punctuation mark by determining whether text received from the speech recognizer has an ending.

3. The multi-speaker speech recognition correction system of claim 1 , further comprising: a reviser configured to display a speech recognition result obtained by the speech recognizer and a correction result obtained by the text corrector to each of the plurality of speakers.

4. The multi-speaker speech recognition correction system of claim 3 , wherein the reviser receives information for correction or a revision completion signal, and transmits the signal to the text corrector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 28, 2017
From: AN, MUNHAK
To: SORIZAVA CO., LTD.
Reel/Frame 044234/0624 →
Priority Claims (1)
KR 10-2016-0176567 · Dec 22, 2016 · national
Continuity (1)
Related Publication 20180182396A1 · Jun 28, 2018