IP Library Granted Patent US 11,507,759
Granted Patent B2
US 11,507,759 · App. 16/824,110 · Granted Nov 22, 2022

Speech translation device, speech translation method, and recording medium

Inventors: Hiroki Furukawa (Osaka, JP); Atsushi Sakaguchi (Kyoto, JP); Tsuyoki Nishikawa (Osaka, JP)
Assignee: PANASONIC HOLDINGS CORPORATION
G06F40/58G06F3/04842G06F3/167G06F9/542G10L13/00G10L15/04G10L15/26G06F2203/04803
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,507,759
App. No.
16/824,110
Granted
Nov 22, 2022
Kind
B2
Abstract

A speech translation device, for conversation between a first speaker making an utterance in a first language and a second speaker making an utterance in a second language different from the first language, includes: a speech detector that detects, from sounds that are input, a speech segment in which the first speaker or the second speaker made an utterance; a display that, after speech recognition is performed on the utterance, displays a translation result obtained by translating the utterance from the first language to the second language or from the second language to the first language; and an utterance instructor that outputs, in the second language via the display, a message prompting the second speaker to make an utterance after a first speaker's utterance or outputs, in the first language via the display, a message prompting the first speaker to make an utterance after a second speaker's utterance.

Claims (90)

1. A speech translation device for conversation between a first speaker and a second speaker, the first speaker making an utterance in a first language, the second speaker making an utterance in a second language different from the first language, the speech translation device comprising:

a speech detector that detects, from sounds that are input to an audio input unit, a speech segment in which the first speaker or the second speaker has made an utterance;

a display that, after speech recognition is performed on the utterance in the speech segment detected by the speech detector, displays a translation result obtained by translating the utterance from the first language to the second language or a translation result obtained by translating the utterance from the second language to the first language;

an utterance circuit that outputs, in the second language via the display, a message prompting the second speaker to make an utterance after the first speaker has made an utterance or outputs, in the first language via the display, a message prompting the first speaker to make an utterance after the second speaker has made an utterance;

the audio input unit to which a voice of the utterance made by the first speaker or the second speaker in the conversation is input;

a speech recognizer that performs speech recognition on the utterance in the speech segment detected by the speech detector, to convert the utterance into text;

a translator that translates the text into which the utterance has been converted by the speech recognizer, from the first language to the second language or from the second language to the first language; and

an audio output unit that outputs by voice a result of the translation made by the translator, wherein

the audio input unit comprises a plurality of audio input units,

the speech translation device further comprises:

a first beam former that performs signal processing on a voice that is input to at least one of the plurality of audio input units, to cause directivity of sound collection to coincide with a sound source direction of the utterance made by the first speaker;

a second beam former that performs signal processing on the voice that is input to at least one of the plurality of audio input units, to cause directivity of sound collection to coincide with a sound source direction of the utterance made by the second speaker; and

an input switch that switches between obtaining an output signal from the first beam former and obtaining an output signal from the second beam former.

2. The speech translation device according to claim 1 , further comprising:

a priority utterance input unit that, when speech recognition is performed on the utterance made by the first speaker or the second speaker, causes speech recognition to be performed again on the utterance on which the speech recognition has been performed.

3. The speech translation device according to claim 1 , wherein

the utterance circuit:

outputs, in the first language via the display, the message prompting the first speaker to make an utterance when the speech translation device is activated; and

outputs, in the second language via the display, the message prompting the second speaker to make an utterance after the utterance made by the first speaker is translated from the first language to the second language and a result of the translation is displayed on the display.

4. The speech translation device according to claim 1 , wherein

after a start of the translation, the utterance circuit causes the audio output unit to output, a specified number of times, a voice message for prompting utterance, and

after the audio output unit has output the voice message the specified number of times, the utterance circuit causes the display to display a message for prompting utterance.

5. The speech translation device according to claim 1 , wherein

the speech recognizer outputs a result of the speech recognition performed on the utterance and a reliability score of the result, and

when the reliability score obtained from the speech recognizer is lower than or equal to a threshold, the utterance circuit outputs a message prompting utterance via at least one of the display or the audio output unit, without translating the utterance whose reliability score is lower than or equal to the threshold.

6. The speech translation device according to claim 1 , wherein

the speech translation device further comprises a sound source direction estimator that estimates a sound source direction by performing signal processing on the voice that is input to the plurality of audio input units, and

the utterance circuit causes the input switch to switch between the obtaining of an output signal from the first beam former and the obtaining of an output signal from the second beam former.

7. A speech translation device for conversation between a first speaker and a second speaker, the first speaker making an utterance in a first language, the second speaker making an utterance in a second language different from the first language, the speech translation device comprising:

a speech detector that detects, from sounds that are input to an audio input unit, a speech segment in which the first speaker or the second speaker has made an utterance;

a display that, after speech recognition is performed on the utterance in the speech segment detected by the speech detector, displays a translation result obtained by translating the utterance from the first language to the second language or a translation result obtained by translating the utterance from the second language to the first language;

an utterance circuit that outputs, in the second language via the display, a message prompting the second speaker to make an utterance after the first speaker has made an utterance or outputs, in the first language via the display, a message prompting the first speaker to make an utterance after the second speaker has made an utterance;

the audio input unit to which a voice of the utterance made by the first speaker or the second speaker in the conversation is input;

a speech recognizer that performs speech recognition on the utterance in the speech segment detected by the speech detector, to convert the utterance into text;

a translator that translates the text into which the utterance has been converted by the speech recognizer, from the first language to the second language or from the second language to the first language; and

an audio output unit that outputs by voice a result of the translation made by the translator, wherein

the audio input unit comprises a plurality of audio input units,

the speech translation device further comprises:

a sound source direction estimator that estimates a sound source direction by performing signal processing on a voice that is input to the plurality of audio input units; and

a controller that causes the display to display the first language in a display area corresponding to a location of the first speaker with respect to the speech translation device, and display the second language in a display area corresponding to a location of the second speaker with respect to the speech translation device, and

the controller:

compares a display direction and the sound source direction estimated by the sound source direction estimator, the display direction being a direction from the display of the speech translation device to the first speaker or the second speaker and being a direction for either one of the display areas of the display;

causes the speech recognizer and the translator to operate when the display direction substantially coincides with the sound source direction estimated; and

causes the speech recognizer and the translator to stop when the display direction is different from the sound source direction estimated.

8. The speech translation device according to claim 7 , further comprising:

a priority utterance input unit that, when speech recognition is performed on the utterance made by the first speaker or the second speaker, causes speech recognition to be performed again on the utterance on which the speech recognition has been performed.

9. The speech translation device according to claim 7 , wherein

when the controller causes the speech recognizer and the translator to stop, the utterance circuit outputs again a message prompting utterance in a specified language.

10. The speech translation device according to claim 7 , wherein

when the display direction is different from the sound source direction estimated, the utterance circuit outputs again a message prompting utterance in a specified language after a specified period of time has elapsed since the comparison made by the controller.

11. The speech translation device according to claim 7 , wherein

the utterance circuit:

outputs, in the first language via the display, the message prompting the first speaker to make an utterance when the speech translation device is activated; and

outputs, in the second language via the display, the message prompting the second speaker to make an utterance after the utterance made by the first speaker is translated from the first language to the second language and a result of the translation is displayed on the display.

12. The speech translation device according to claim 7 , wherein

after a start of the translation, the utterance circuit causes the audio output unit to output, a specified number of times, a voice message for prompting utterance, and

after the audio output unit has output the voice message the specified number of times, the utterance circuit causes the display to display a message for prompting utterance.

13. The speech translation device according to claim 7 , wherein

the speech recognizer outputs a result of the speech recognition performed on the utterance and a reliability score of the result, and

when the reliability score obtained from the speech recognizer is lower than or equal to a threshold, the utterance circuit outputs a message prompting utterance via at least one of the display or the audio output unit, without translating the utterance whose reliability score is lower than or equal to the threshold.

14. A speech translation method for conversation between a first speaker and a second speaker, the first speaker making an utterance in a first language, the second speaker making an utterance in a second language different from the first language, the speech translation method comprising:

detecting, from sounds that are input to an audio input unit, a speech segment in which the first speaker or the second speaker has made an utterance;

after performing speech recognition on the utterance in the speech segment detected, displaying on a display a translation result obtained by translating the utterance from the first language to the second language or a translation result obtained by translating the utterance from the second language to the first language;

outputting, in the second language via the display, a message prompting the second speaker to make an utterance after the first speaker has made an utterance, or outputting, in the first language via the display, a message prompting the first speaker to make an utterance after the second speaker has made an utterance;

performing speech recognition on the utterance in the speech segment detected, to convert the utterance into text;

translating the text into which the utterance has been converted, from the first language to the second language or from the second language to the first language; and

outputting by voice a result of the translation,

wherein

the audio input unit comprises a plurality of audio input units, and

the speech translation method further comprises:

performing signal processing by a first beam former on a voice that is input to at least one of the plurality of audio input units, to cause directivity of sound collection to coincide with a sound source direction of the utterance made by the first speaker;

performing signal processing by a second beam former on the voice that is input to at least one of the plurality of audio input units, to cause directivity of sound collection to coincide with a sound source direction of the utterance made by the second speaker; and

switching between obtaining an output signal from the first beam former and obtaining an output signal from the second beam former.

15. The speech translation method according to claim 14 , further comprising estimating a sound source direction by performing signal processing on the voice that is input to the plurality of audio input units.

16. A speech translation method for conversation between a first speaker and a second speaker, the first speaker making an utterance in a first language, the second speaker making an utterance in a second language different from the first language, the speech translation method comprising:

detecting, from sounds that are input to an audio input unit, a speech segment in which the first speaker or the second speaker has made an utterance;

after performing speech recognition on the utterance in the speech segment detected, displaying on a display a translation result obtained by translating the utterance from the first language to the second language or a translation result obtained by translating the utterance from the second language to the first language;

outputting, in the second language via the display, a message prompting the second speaker to make an utterance after the first speaker has made an utterance, or outputting, in the first language via the display, a message prompting the first speaker to make an utterance after the second speaker has made an utterance;

performing speech recognition, using a speech recognizer, on the utterance in the speech segment detected, to convert the utterance into text;

translating, using a translator, the text into which the utterance has been converted, from the first language to the second language or from the second language to the first language; and

outputting by voice a result of the translation made by the translator,

wherein

the audio input unit comprises a plurality of audio input units,

the speech translation method further comprises:

estimating a sound source direction by performing signal processing on a voice that is input to the plurality of audio input units; and

causing the display to display the first language in a display area corresponding to a location of the first speaker with respect to the display, and display the second language in a display area corresponding to a location of the second speaker with respect to the display, and

said causing comprises:

comparing a display direction and the sound source direction estimated, the display direction being a direction from the display to the first speaker or the second speaker and being a direction for either one of the display areas of the display;

causing the speech recognizer and the translator to operate when the display direction substantially coincides with the sound source direction estimated; and

causing the speech recognizer and the translator to stop when the display direction is different from the sound source direction estimated.

Assignments (2)
CHANGE OF NAME Recorded May 9, 2022
From: PANASONIC CORPORATION
To: PANASONIC HOLDINGS CORPORATION
Reel/Frame 059909/0607 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2020
From: FURUKAWA, HIROKI; SAKAGUCHI, ATSUSHI; NISHIKAWA, TSUYOKI
To: PANASONIC CORPORATION
Reel/Frame 052719/0516 →
Priority Claims (1)
JP JP2019-196078 · Oct 29, 2019 · national
Continuity (2)
Provisional Application 62823197 · Mar 25, 2019
Related Publication 20200311354A1 · Oct 1, 2020
Cited By (1)
US 12,682,895