IP Library › Granted Patent US 8,862,478
Granted Patent B2
US 8,862,478 · App. 13/499,311 · Granted Oct 14, 2014

Speech translation system, first terminal apparatus, speech recognition server, translation server, and speech synthesis server

Inventors: Satoshi Nakamura (Koganei, JP); Eiichiro Sumita (Koganei, JP); Yutaka Ashikari (Koganei, JP); Noriyuki Kimura (Koganei, JP); Chiori Hori (Koganei, JP)
Assignee: National Institute of Information and Communications Technology
G10L15/265G10L13/043G06F17/289G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,862,478
App. No.
13/499,311
Granted
Oct 14, 2014
Kind
B2
Abstract

In conventional network-type speech translation systems, devices or models for recognizing or synthesizing speech cannot be changed in accordance with speakers' attributes, and therefore, accuracy is reduced or inappropriate output occurs in each process of speech recognition, translation, and speech synthesis. Accuracy of each processing of speech translation, translation, or speech synthesis is improved and appropriate output is performed in a network-type speech translation system by, based on speaker attributes, appropriately changing the server to perform speech recognition or the speech recognition model, appropriately changing the translation server to perform translation or the translation model, or appropriately changing the speech synthesis server or speech synthesis model.

Claims (169)

1. A speech translation system including a first terminal apparatus for inputting speech, two or more speech recognition servers, one or more translation servers, and one or more speech synthesis servers,

wherein the first terminal apparatus comprises:

a first speaker attribute storage unit capable of having stored therein one or more speaker attributes which are attribute values of a speaker;

a first speech accepting unit that accepts speech;

a first speech recognition server selection unit that selects a speech recognition server from among the two or more speech recognition servers in accordance with the one or more speaker attributes; and

a first speech sending unit that sends speech information constituted from the speech accepted by the first speech accepting unit to the speech recognition server selected by the first speech recognition server selection unit,

each speech recognition server comprises:

a speech recognition model storage unit capable of having stored therein a speech recognition model for all two or more languages or a part of the two or more languages;

a speech information receiving unit that receives the speech information;

a speech recognition unit that performs speech recognition on the speech information received by the speech information receiving unit by using the speech recognition model in the speech recognition model storage unit, and acquires a speech recognition result; and

a speech recognition result sending unit that sends the speech recognition result,

each translation server comprises:

a translation model storage unit capable of having stored therein a translation model for all of the two or more languages or a part of the two or more languages;

a speech recognition result receiving unit that receives the speech recognition result;

a translation unit that translates into a target language, by using the translation model in the translation model storage unit, the speech recognition result received by the speech recognition result receiving unit, and acquires a translation result; and

a translation result sending unit that sends the translation result,

each speech synthesis server comprises:

a speech synthesis model storage unit capable of having stored therein a speech synthesis model for all of the two or more languages or a part of the two or more languages;

a translation result receiving unit that receives the translation result;

a speech synthesis unit that performs, by using the speech synthesis model in the speech synthesis model storage unit, speech synthesis on the translation result received by the translation result receiving unit, and acquires a speech synthesis result; and

a speech synthesis result sending unit that sends the speech synthesis result to a second terminal apparatus; wherein

the one or more speaker attributes are selected from the group of speaker class and dynamic speaker attribute information,

wherein speaker class is determined based on one or more of degree of difficulty in used words,

information indicating a degree of politeness of used terms,

information indicating the degree of grammatical correctness, and

information indicating a multiple degree of these elements, and

information indicating whether or not the speaker is a native speaker.

2. A speech translation system including a first terminal apparatus for inputting speech, one or more speech recognition servers, one or more translation servers, and one or more speech synthesis servers,

wherein the first terminal apparatus comprises:

a first speech accepting unit that accepts speech; and

a first speech sending unit that sends speech information constituted from the speech accepted by the first speech accepting unit to the speech recognition server,

each speech recognition server comprises:

a third speaker attribute storage unit capable of having stored therein one or more speaker attributes which are attribute values of a speaker;

a speech recognition model storage unit capable of having stored therein two or more speech recognition models for all two or more languages or a part of the two or more languages;

a speech information receiving unit that receives the speech information;

a speech recognition model selection unit that selects a speech recognition model from among the two or more speech recognition models in accordance with the one or more speaker attributes;

a speech recognition unit that performs, by using a speech recognition model selected by the speech recognition model selection unit, speech recognition on the speech information received by the speech information receiving unit, and acquires a speech recognition result; and

a speech recognition result sending unit that sends the speech recognition result,

each translation server comprises:

a translation model storage unit capable of having stored therein a translation model for all of the two or more languages or a part of the two or more languages;

a speech recognition result receiving unit that receives the speech recognition result;

a translation unit that translates into a target language, by using the translation model in the translation model storage unit, the speech recognition result received by the speech recognition result receiving unit, and acquires a translation result; and

a translation result sending unit that sends the translation result,

each speech synthesis server comprises:

a speech synthesis model storage unit capable of having stored therein a speech synthesis model for all of the two or more languages or a part of the two or more languages;

a translation result receiving unit that receives the translation result;

a speech synthesis unit that performs, by using the speech synthesis model in the speech synthesis model storage unit, speech synthesis on the translation result received by the translation result receiving unit, and acquires a speech synthesis result; and

a speech synthesis result sending unit that sends the speech synthesis result to a second terminal apparatus; wherein

the one or more speaker attributes are selected from the group of speaker class and dynamic speaker attribute information,

wherein speaker class is determined based on one or more of degree of difficulty in used words,

information indicating a degree of politeness of used terms,

information indicating the degree of grammatical correctness, and

information indicating a multiple degree of these elements, and

information indicating whether or not the speaker is a native speaker.

3. A speech translation system including one or more speech recognition servers, two or more translation servers, and one or more speech synthesis servers,

wherein each speech recognition server comprises:

a third speaker attribute storage unit capable of having stored therein one or more speaker attributes which are attribute values of a speaker;

a speech recognition model storage unit capable of having stored therein a speech recognition model for all two or more languages or a part of the two or more languages;

a speech information receiving unit that receives speech information;

a speech recognition unit that performs, by using the speech recognition model in the speech recognition model storage unit, speech recognition on the speech information received by the speech information receiving unit, and acquires a speech recognition result;

a translation server selection unit that selects a translation server from among the two or more translation servers in accordance with the one or more speaker attributes; and

a speech recognition result sending unit that sends the speech recognition result to the translation server selected by the translation server selection unit,

each translation server comprises:

a translation model storage unit capable of having stored therein a translation model for all of the two or more languages or a part of the two or more languages;

a speech recognition result receiving unit that receives the speech recognition result;

a translation unit that translates into a target language, by using the translation model in the translation model storage unit, the speech recognition result received by the speech recognition result receiving unit, and acquires a translation result; and

a translation result sending unit that sends the translation result,

each speech synthesis server comprises:

a speech synthesis model storage unit capable of having stored therein a speech synthesis model for all of the two or more languages or a part of the two or more languages;

a translation result receiving unit that receives the translation result;

a speech synthesis unit that performs, by using the speech synthesis model in the speech synthesis model storage unit, speech synthesis on the translation result received by the translation result receiving unit, and acquires a speech synthesis result; and

a speech synthesis result sending unit that sends the speech synthesis result to a second terminal apparatus; wherein

the one or more speaker attributes are selected from the group of speaker class and dynamic speaker attribute information,

wherein speaker class is determined based on one or more of degree of difficulty in used words,

information indicating a degree of politeness of used terms,

information indicating the degree of grammatical correctness, and

information indicating a multiple degree of these elements, and

information indicating whether or not the speaker is a native speaker.

4. A speech translation system including one or more speech recognition servers, one or more translation servers, and one or more speech synthesis servers,

wherein each speech recognition server comprises:

a speech recognition model storage unit capable of having stored therein a speech recognition model for all two or more languages or a part of the two or more languages;

a speech information receiving unit that receives speech information;

a speech recognition unit that performs, by using the speech recognition model in the speech recognition model storage unit, speech recognition on the speech information received by the speech information receiving unit, and acquires a speech recognition result; and

a speech recognition result sending unit that sends the speech recognition result to the translation server,

each translation server comprises:

a translation model storage unit capable of having stored therein two or more translation models for all of the two or more languages or a part of the two or more languages;

a fourth speaker attribute storage unit capable of having stored therein one or more speaker attributes;

a speech recognition result receiving unit that receives the speech recognition result;

a translation model selection unit that selects a translation model from among the two or more translation models in accordance with the one or more speaker attributes;

a translation unit that translates into a target language, by using the translation model selected by the translation model selection unit, the speech recognition result received by the speech recognition result receiving unit, and acquires a translation result; and

a translation result sending unit that sends the translation result,

each speech synthesis server comprises:

a speech synthesis model storage unit capable of having stored therein a speech synthesis model for all of the two or more languages or a part of the two or more languages;

a translation result receiving unit that receives the translation result;

a speech synthesis unit that performs, by using the speech synthesis model in the speech synthesis model storage unit, speech synthesis on the translation result received by the translation result receiving unit, and acquires a speech synthesis result; and

a speech synthesis result sending unit that sends the speech synthesis result to a second terminal apparatus; wherein

the one or more speaker attributes are selected from the group of speaker class and dynamic speaker attribute information,

wherein speaker class is determined based on one or more of degree of difficulty in used words,

information indicating a degree of politeness of used terms,

information indicating the degree of grammatical correctness, and

information indicating a multiple degree of these elements, and

information indicating whether or not the speaker is a native speaker.

5. A speech translation system including one or more speech recognition servers, one or more translation servers, and two or more speech synthesis servers,

wherein each speech recognition server comprises:

a speech recognition model storage unit capable of having stored therein a speech recognition model for all two or more languages or a part of the two or more languages;

a speech information receiving unit that receives speech information;

a speech recognition unit that performs, by using the speech recognition model in the speech recognition model storage unit, speech recognition on the speech information received by the speech information receiving unit, and acquires a speech recognition result; and

a speech recognition result sending unit that sends the speech recognition result to the translation server,

each translation server comprises:

a translation model storage unit capable of having stored therein a translation model for all of the two or more languages or a part of the two or more languages;

a fourth speaker attribute storage unit capable of having stored therein one or more speaker attributes;

a speech recognition result receiving unit that receives the speech recognition result;

a translation unit that translates into a target language, by using the translation model in the translation model storage unit, the speech recognition result received by the speech recognition result receiving unit, and acquires a translation result;

a speech synthesis server selection unit that selects a speech synthesis server from among the two or more speech synthesis servers in accordance with the one or more speaker attributes; and

a translation result sending unit that sends the translation result to the speech synthesis server selected by the speech synthesis server selection unit,

each speech synthesis server comprises:

a speech synthesis model storage unit capable of having stored therein a speech synthesis model for all of the two or more languages or a part of the two or more languages;

a translation result receiving unit that receives the translation result;

a speech synthesis unit that performs, by using the speech synthesis model in the speech synthesis model storage unit, speech synthesis on the translation result received by the translation result receiving unit, and acquires a speech synthesis result; and

a speech synthesis result sending unit that sends the speech synthesis result to a second terminal apparatus; wherein

the one or more speaker attributes are selected from the group of speaker class and dynamic speaker attribute information,

wherein speaker class is determined based on one or more of degree of difficulty in used words,

information indicating a degree of politeness of used terms,

information indicating the degree of grammatical correctness, and

information indicating a multiple degree of these elements, and

information indicating whether or not the speaker is a native speaker.

6. A speech translation system including one or more speech recognition servers, one or more translation servers, and one or more speech synthesis servers,

wherein each speech recognition server comprises:

a speech recognition model storage unit capable of having stored therein a speech recognition model for all two or more languages or a part of the two or more languages;

a speech information receiving unit that receives speech information;

a speech recognition unit that performs, by using the speech recognition model in the speech recognition model storage unit, speech recognition on the speech information received by the speech information receiving unit, and acquires a speech recognition result; and

a speech recognition result sending unit that sends the speech recognition result to the translation server,

each translation server comprises:

a translation model storage unit capable of having stored therein a translation model for all of the two or more languages or a part of the two or more languages;

a speech recognition result receiving unit that receives the speech recognition result;

a translation unit that translates into a target language, by using the translation model in the translation model storage unit, the speech recognition result received by the speech recognition result receiving unit, and acquires a translation result; and

a translation result sending unit that sends the translation result to the speech synthesis server,

each speech synthesis server comprises:

a speech synthesis model storage unit capable of having stored therein two or more speech synthesis models for all of the two or more languages or a part of the two or more languages,

a fifth speaker attribute storage unit capable of having stored therein one or more speaker attributes,

a translation result receiving unit that receives the translation result;

a speech synthesis model selection unit that selects a speech synthesis model from among the two or more speech synthesis models in accordance with the one or more speaker attributes;

a speech synthesis unit that performs, by using the speech synthesis model selected by the speech synthesis model selection unit, speech synthesis on the translation result received by the translation result receiving unit, and acquires a speech synthesis result; and

a speech synthesis result sending unit that sends the speech synthesis result to a second terminal apparatus; wherein

the one or more speaker attributes are selected from the group of speaker class and dynamic speaker attribute information,

wherein speaker class is determined based on one or more of degree of difficulty in used words,

information indicating a degree of politeness of used terms,

information indicating the degree of grammatical correctness, and

information indicating a multiple degree of these elements, and

information indicating whether or not the speaker is a native speaker.

7. The speech translation system according to claim 1 ,

wherein the first terminal apparatus comprises:

a first speaker attribute accepting unit that accepts one or more speaker attributes; and

a first speaker attribute accumulation unit that accumulates the one or more speaker attributes in the first speaker attribute storage unit.

8. The speech translation system according to claim 2 ,

wherein each speech recognition server further comprises:

a speech speaker attribute acquiring unit that acquires one or more speaker attributes related to speech from the speech information received by the speech information receiving unit; and

a third speaker attribute accumulation unit that accumulates, in the third speaker attribute storage unit, one or more speaker attributes acquired by the speech speaker attribute acquiring unit.

9. The speech translation system according to claim 4 ,

wherein each translation server comprises:

a language speaker attribute acquiring unit that acquires one or more speaker attributes related to language from the speech recognition result received by the speech recognition result receiving unit; and

a fourth speaker attribute accumulation unit that accumulates, in the fourth speaker attribute storage unit, the one or more speaker attributes acquired by the language speaker attribute acquiring unit.

10. The speech translation system according to claim 1 ,

wherein a source language identifier for identifying a source language which is a language used by the speaker, a target language identifier for identifying a target language which is a language into which translation is performed, and speech translation control information containing one or more speaker attributes are sent from the speech recognition server via the one or more translation servers to the speech synthesis server, and

the speech recognition server selection unit, the speech recognition unit, the speech recognition model selection unit, the translation server selection unit, the translation unit, or the translation model selection unit, the speech synthesis server selection unit, the speech synthesis unit, or the speech synthesis model selection unit performs respective processing by using the speech translation control information.

11. The first terminal apparatus that constitutes the speech translation system according to claim 1 .

12. The speech recognition server that constitutes the speech translation system according to claim 2 .

13. The translation server that constitutes the speech translation system according to claim 4 .

14. The speech synthesis server that constitutes the speech translation system according to claim 6 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 16, 2012
From: NAKAMURA, SATOSHI; SUMITA, EIICHIRO; ASHIKARI, YUTAKA; KIMURA, NORIYUKI; HORI, CHIORI
To: NATIONAL INSTITUTE OF INFORMATION AND COMMUNICATIONS TECHNOLOGY
Reel/Frame 028052/0700 →
Priority Claims (1)
JP 2009-230442 · Oct 2, 2009 · national
Continuity (1)
Related Publication 20120197629A1 · Aug 2, 2012