IP Library › Granted Patent US 10,580,400
Granted Patent B2
US 10,580,400 · App. 15/891,625 · Granted Mar 3, 2020

Method for controlling artificial intelligence system that performs multilingual processing

Inventor: Gyuhyeok Jeong (Seoul, KR)
Assignee: LG ELECTRONICS INC.
G10L15/005G10L15/187G10L15/22G10L15/30G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,580,400
App. No.
15/891,625
Granted
Mar 3, 2020
Kind
B2
Abstract

This specification relates to a method for controlling an artificial intelligence system which performs a multilingual processing based on artificial intelligence technology. The method for controlling an artificial intelligence system which performs a multilingual processing includes: receiving voice information through a microphone; determining a language of the voice information, based on a preset reference; selecting a specific voice recognition server from a plurality of voice recognition servers which process different languages, based on a result of the determination; and transmitting the voice information to the selected specific voice recognition server.

Claims (73)

1. A method for controlling a multilingual audio processing system, the method comprising:

receiving voice information via a microphone;

separating the voice information into a plurality of phonemes to determine the language of the received voice information;

determining a language of each of the plurality of phonemes;

determining at least one language of the received voice information based on a preset reference language information;

selecting a specific voice recognition server from a plurality of voice recognition servers based on the determined at least one language, wherein the plurality of voice recognition servers correspond to different languages and the specific voice recognition server corresponds to the at least one determined language; and

generating a query comprising the received voice information and transmitting the query to the selected specific voice recognition server,

wherein:

the plurality of phonemes is determined to correspond to a plurality of languages;

the selected specific voice recognition server is configured to process mixed language voice information; and

the selected specific voice recognition server is configured to process voice information comprising specific mixed languages based on repeated machine learning training using training data comprising voice information with the specific mixed languages.

2. The method of claim 1 , wherein the preset reference language information is stored in a memory at a client of the system, and the client determines the at least one language of the received voice information to select the specific voice recognition server.

3. The method of claim 1 , wherein the determined at least one language corresponds to a single language, and the selected specific voice recognition server is configured to process only voice information in the single language.

4. The method of claim 1 , further comprising:

receiving a response to the generated query from the specific voice recognition server;

generating reply information to the received voice information based on the received response; and

outputting the generated reply in response to the received voice information.

5. The method of claim 4 , wherein the outputted generated reply is in the form of an audio output.

6. The method of claim 5 , wherein the audio output of the generated reply is performed in the determined language of the received voice information.

7. The method of claim 4 , wherein the outputted generated reply is displayed on a display of a client terminal of the system.

8. The method of claim 4 , further comprising:

storing the received voice information in a memory at a client terminal of the system;

receiving another voice information via the microphone;

retrieving the stored voice information from the memory;

generating a similarity value between the another voice information and the retrieved stored voice information; and

outputting a reply stored in the memory and associated with the stored voice information when the generated similarity value is equal to or greater than a threshold value.

9. The method of claim 4 , wherein when the voice information comprises a plurality of different languages, the generated reply to the voice information is in one of the plurality of languages determined to be a preferred language.

10. The method of claim 1 , further comprising:

requesting a language translation with respect to the voice information to the specific voice recognition server for translating the voice information into a second language from a first language;

receiving language translation data for the voice information from the specific voice recognition server;

generating reply information to the received voice information based on the received language translation data; and

outputting the generated reply in response to the received voice information.

11. The method of 10 , wherein the outputted generated reply is in the form of an audio output.

12. The method of 10 , wherein the outputted generated reply is displayed on a display of a client terminal of the system.

13. A multilingual audio processing terminal, the terminal comprising:

a microphone configured to receive audio information;

a transceiver configured to transmit and receive information; and

a controller configured to:

receive voice information via the microphone;

separate the voice information into a plurality of phonemes to determine the language of the received voice information;

determine a language of each of the plurality of phonemes;

determine at least one language of the received voice information based on a preset reference language information;

select a specific voice recognition server from a plurality of voice recognition servers based on the determined at least one language, wherein the plurality of voice recognition servers correspond to a different languages and the specific voice recognition server corresponds to the at least one determined language; and

transmit, via the transceiver, a query comprising the received voice information to the selected specific voice recognition server,

wherein:

the plurality of phonemes is determined to correspond to a plurality of languages;

the selected specific voice recognition server is configured to process mixed language voice information; and

the selected specific voice recognition server is configured to process voice information comprising specific mixed languages based on repeated machine learning training using training data comprising voice information with the specific mixed languages.

14. The terminal of claim 13 , further comprising a memory, wherein:

the preset reference language information is stored in the memory; and

the controller is further configured to determine the at least one language of the received voice information to select the specific voice recognition server.

15. The terminal of claim 13 , wherein the determined at least one language corresponds to a single language, and the selected specific voice recognition server is configured to process only voice information in the single language.

16. The terminal of claim 13 , further comprising an output configured to output information, wherein the controller is further configured to:

receive, via the transceiver, a response to the generated query from the specific voice recognition server;

generate reply information to the received voice information based on the received response; and

output, via the output, the generated reply in response to the received voice information.

17. The terminal of claim 16 , wherein the output comprises a speaker and the outputted generated reply is in the form of an audio output.

18. The terminal of claim 17 , wherein the audio output of the generated reply is performed in the determined language of the received voice information.

19. The terminal of claim 16 , wherein the output comprises a display configured to display information and the outputted generated reply is displayed on the display.

20. The terminal of claim 16 , further comprising a memory, wherein the controller is further configured to:

store the received voice information in the memory;

receive another voice information via the microphone;

retrieve the stored voice information from the memory;

generate a similarity value between the another voice information and the retrieved stored voice information; and

output, via the output, a reply stored in the memory and associated with the stored voice information when the generated similarity value is equal to or greater than a threshold value.

21. The terminal of claim 16 , wherein when the voice information comprises a plurality of different languages, the generated reply to the voice information is in one of the plurality of languages determined to be a preferred language.

22. The terminal of claim 13 , wherein the controller is further configured to:

transmit, via the transceiver, a request for language translation with respect to the voice information to the specific voice recognition server for translating the voice information into a second language from a first language;

receive, via the transceiver, language translation data for the voice information from the specific voice recognition server;

generate reply information to the received voice information based on the received language translation data; and

output, via the output, the generated reply in response to the received voice information.

23. The terminal of claim 22 , wherein the output comprises a speaker and the outputted generated reply is in the form of an audio output.

24. The terminal of claim 22 , wherein the output comprises a display configured to display information and the outputted generated reply is displayed on the display.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 8, 2018
From: JEONG, GYUHYEOK
To: LG ELECTRONICS INC.
Reel/Frame 044869/0530 →
Priority Claims (1)
KR 10-2017-0022530 · Feb 20, 2017 · national
Continuity (1)
Related Publication 20180240456A1 · Aug 23, 2018
Cited By (1)
US 12,494,195