IP Library › Granted Patent US 11,631,400
Granted Patent B2
US 11,631,400 · App. 16/786,654 · Granted Apr 18, 2023

Electronic apparatus and controlling method thereof

Inventors: Beomseok Lee (Suwon-si, KR); Sangha Kim (Suwon-si, KR); Yoonjin Yoon (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L15/16G10L15/063G10L15/18G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,631,400
App. No.
16/786,654
Filed
Feb 10, 2020
Granted
Apr 18, 2023
Kind
B2
Art Unit
2655
USPC
704/232
Abstract

An electronic apparatus configured to acquire information on a plurality of candidate texts corresponding to input speech of a user through a general speech recognition module, determine text corresponding to the input speech from among the plurality of candidate texts using a trained personal language model, and output the text as a result of speech recognition of the input speech.

Claims (61)

1. An electronic apparatus comprising:

an input circuitry comprising a touch panel and a microphone;

a memory storing at least one instruction and a personal language model trained to recognize speech of a user of the electronic apparatus; and

a processor configured to execute at least one instruction, wherein the processor when executing the at least one instruction is configured to:

based on a text input by the user of the electronic apparatus being received through the touch panel, train the personal language model corresponding to the user from among at least one personal language model using the received text,

based on a first input speech of the user of being received through the microphone, acquire information on a plurality of candidate texts corresponding to the first input speech and a plurality of scores corresponding to the plurality of candidate texts through a general speech recognition module including a general language model different from the personal language model,

identify text corresponding to the first input speech from among the plurality of candidate texts by correcting at least one score of the plurality of scores corresponding to the plurality of candidate texts using the trained personal language model, and

output the identified text as a result of speech recognition of the first input speech,

wherein the processor when executing the at least one instruction is further configured to:

acquire user information comprising user preference information of the user and user location information of the user,

train the personal language model based on the received text, the user preference information of the user and the user location information of the user,

identify the text corresponding to the first input speech using the personal language model based on the received text, the user preference information of the user and the user location information of the user,

based on a candidate text having a highest score among the plurality of candidate texts corresponding to the first input speech and the text identified by correcting the at least one score of the plurality of scores corresponding to the plurality of candidate texts using the trained personal language model being different from each other, output a message requesting the user to input a user feedback,

based on the highest score of the candidate text being less than a threshold score, output the message requesting the user to input the user feedback, and

in response to the user feedback being input, train the personal language model based on the user feedback.

2. The electronic apparatus as claimed in claim 1 , wherein the memory stores a personal acoustic model of the user of the electronic apparatus, and

wherein the processor when executing the at least one instruction is further configured to train the personal acoustic model based on the identified text and the first input speech.

3. The electronic apparatus as claimed in claim 1 , wherein the general speech recognition module comprises a general acoustic model and the general language model, and

wherein the processor when executing the at least one instruction is further configured to:

acquire the plurality of candidate texts corresponding to the first input speech and the plurality of scores corresponding to the plurality of candidate texts through the general acoustic model and the general language model,

adjust the plurality of scores using the personal language model, and

select the candidate text among the plurality of candidate texts having the highest score among the plurality of scores adjusted using the personal language model as the text corresponding to the first input speech.

4. The electronic apparatus as claimed in claim 3 , wherein the processor when executing the at least one instruction is further configured to identify whether the highest score of the candidate text is greater than or equal to the threshold score, and

based on the highest score being greater or equal to the threshold score, select the candidate text as the text corresponding to the first input speech.

5. The electronic apparatus as claimed in claim 4 , wherein the processor when executing the at least one instruction is further configured to control the electronic apparatus to output a message requesting the user to repeat the first input speech, based on the highest score being less than the threshold score.

6. The electronic apparatus as claimed in claim 1 , further comprising:

a communicator,

wherein the general speech recognition module is stored in an external server, and

wherein the processor when executing the at least one instruction is further configured to:

based on the first input speech, control the communicator to transmit the first input speech to the external server, and

obtain information on the plurality of candidate texts corresponding to the first input speech from the external server.

7. The electronic apparatus as claimed in claim 1 , wherein the processor when executing the at least one instruction is further configured to, based on a user confirmation of the identified text as the result of the speech recognition of the first input speech, retrain the personal language model based on the result of the speech recognition of the first input speech.

8. A method of controlling an electronic apparatus, the method comprising:

based on a text input by a user of the electronic apparatus being received through a touch panel, training a personal language model corresponding to the user from among at least one personal language model using the received text;

based on a first input speech of the user being received through a microphone, acquiring information on a plurality of candidate texts corresponding to the first input speech and a plurality of scores corresponding to the plurality of candidate texts through a general speech recognition module including a general language model different from the personal language model;

identifying text corresponding to the first input speech from among the plurality of candidate texts by correcting at least one score of the plurality of scores corresponding to the plurality of candidate texts using the personal language model trained to recognize speech of the user; and

outputting the identified text as a result of speech recognition of the first input speech,

wherein the method further comprises:

acquiring user information comprising user preference information of the user and user location information of the user,

wherein the training of the personal language model corresponding to the user comprises training the personal language model based on the received text, the user preference information of the user and the user location information of the user, and

wherein the identifying of the text corresponding to the first input speech comprises identifying the text corresponding to the first input speech using the personal language model based on the received text, the user preference information of the user and the user location information of the user,

wherein the method further comprises:

based on a candidate text having a highest score among the plurality of candidates texts corresponding to the first input speech and the text identified by correcting the at least one score of the plurality of scores corresponding to the plurality of candidate texts using the trained personal language model being different from each other, outputting a message requesting the user to input a user feedback,

based on the highest score of the candidate text being less than a threshold score, outputting the message requesting the user to input the user feedback, and

in response to the user feedback being input, training the personal language model based on the user feedback.

9. The method as claimed in claim 8 , further comprising:

training a personal acoustic model based on the identified text and the first input speech.

10. The method as claimed in claim 8 , wherein the general speech recognition module comprises a general acoustic model and the general language model,

wherein the acquiring of the information comprises acquiring the plurality of candidate texts corresponding to the first input speech and the plurality of scores corresponding to the plurality of candidate texts through the general acoustic model and the general language model,

wherein the identifying comprises:

adjusting the plurality of scores corresponding using the personal language model, and

selecting, among the plurality of candidate texts, the candidate text having the highest score among the plurality of scores adjusted using the personal language model as the text corresponding to the first input speech.

11. The method as claimed in claim 10 , wherein the selecting of the candidate text among the plurality of candidate texts comprises:

identifying whether the highest score of the candidate text is greater than or equal to the threshold score, and

based on the highest score being greater than or equal to the threshold score, selecting the candidate text as the text corresponding to the first input speech.

12. The method as claimed in claim 11 , wherein the selecting of the candidate text among the plurality of candidate texts comprises providing a message requesting the user to repeat the first input speech, based on the highest score being less than the threshold score.

13. The method as claimed in claim 8 , wherein the general speech recognition module is stored in an external server, and

wherein the acquiring of the information comprises transmitting the first input speech to the external server; and

obtaining information on the plurality of candidate texts corresponding to the input speech from the external server.

14. The method as claimed in claim 8 , further comprising:

based on user confirmation of the identified text as the result of the speech recognition of the first input speech, retraining the personal language model based on the result of the speech recognition of the first input speech.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2020
From: LEE, BEOMSEOK; KIM, SANGHA; YOON, YOONJIN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 051773/0545 →
Priority Claims (1)
KR 10-2019-0015516 · Feb 11, 2019 · national
Continuity (1)
Related Publication 20200258504A1 · Aug 13, 2020
Cited By (2)
US 12,248,757 US 12,562,158