IP Library Granted Patent US 10,403,267
Granted Patent B2
US 10,403,267 · App. 15/544,198 · Granted Sep 3, 2019

Method and device for performing voice recognition using grammar model

Inventors: Chi-youn Park (Gyeonggi-do, KR); Il-hwan Kim (Gyeongsangbuk-do, KR); Kyung-min Lee (Gyeonggi-do, KR); Nam-hoon Kim (Gyeonggi-do, KR); Jae-won Lee (Seoul, KR)
Assignee: Samsung Electronics Co., Ltd
G10L15/063G10L15/02G10L15/14G10L15/187G10L15/197G10L2015/025G10L2015/0633G10L2015/0635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,403,267
App. No.
15/544,198
Granted
Sep 3, 2019
Kind
B2
Abstract

A method of updating speech recognition data including a language model used for speech recognition, the method including obtaining language data including at least one word; detecting a word that does not exist in the language model from among the at least one word; obtaining at least one phoneme sequence regarding the detected word; obtaining components constituting the at least one phoneme sequence by dividing the at least one phoneme sequence into predetermined unit components; determining information regarding probabilities that the respective components constituting each of the at least one phoneme sequence appear during speech recognition; and updating the language model based on the determined probability information.

Claims (49)

1. A method of updating speech recognition data comprising a language model used for speech recognition, the method comprising:

obtaining language data comprising at least one word;

detecting a word that does not exist in the language model from among the at least one word;

obtaining at least one phoneme sequence regarding the detected word;

obtaining components constituting the at least one phoneme sequence by dividing the at least one phoneme sequence into predetermined unit components;

determining information about probabilities that the respective components constituting each of the at least one phoneme sequence appear during the speech recognition; and

updating the language model based on the determined probability information.

2. The method of claim 1 , wherein the language model comprises a first language model and a second language model, and

the updating of the language model comprises updating the second language model based on the determined probability information.

3. The method of claim 2 , further comprising:

updating the first language model based on at least one appearance probability information included in the second language model; and

updating a pronunciation dictionary comprising information about phoneme sequences of words based on the at least one phoneme sequence of the detected word.

4. The method of claim 1 , wherein the information about probabilities comprises information about an appearance probability of each of the components under a condition that a word or another component precedes the corresponding component.

5. The method of claim 2 , wherein the determining the information about probabilities comprises:

obtaining situation information about a surrounding situation corresponding to the detected word; and

selecting one of the first language model and the second language model to add appearance probability information regarding the detected word based on the situation information.

6. The method of claim 5 , wherein the updating of the language model comprises updating the second language model regarding a module corresponding to the situation information based on the determined information.

7. A method of performing speech recognition, the method comprising:

obtaining speech data for performing speech recognition;

obtaining at least one phoneme sequence from the speech data;

obtaining information about probabilities that predetermined unit components constituting the at least one phoneme sequence appear during the speech recognition;

determining one of the at least one phoneme sequence based on the information about the probabilities that the predetermined unit components appear during the speech recognition; and

obtaining a word corresponding to the determined phoneme sequence based on segment information for converting the predetermined unit components included in the determined phoneme sequence into a word.

8. The method of claim 7 , wherein the obtaining of the at least one phoneme sequence comprises obtaining a phoneme sequence regarding which information about the word corresponding to the determined phoneme sequence exists in a pronunciation dictionary including information about at least one of phoneme sequences of words and a phoneme sequence, regarding which information about a word corresponding to the phoneme sequence does not exist in the pronunciation dictionary.

9. The method of claim 7 , wherein the obtaining of the information about probabilities comprises:

identifying a first language model and a second language model including appearance probability information regarding the predetermined unit components;

determining weights with respect to the first language model and the second language model;

obtaining at least one appearance probability information regarding the predetermined unit components from the first language model and the second language model; and

obtaining the appearance probability information regarding the predetermined unit components by applying the determined weights to the obtained at least one appearance probability information according to each language model to which the respective at least one appearance probability information belongs.

10. The method of claim 9 , wherein the obtaining of the information about probabilities comprises:

obtaining situation information regarding the speech data;

determining the second language model based on the situation information; and

obtaining the appearance probability information regarding predetermined unit components from the determined second language model.

11. The method of claim 10 , wherein the second language model corresponds to a module or a group comprising at least one module, and

if the obtained situation information comprises an identifier of the module, the second language model corresponds to the identifier.

12. The method of claim 10 , wherein the situation information comprises personalized model information comprising at least one of acoustic information by classes and information about preferred languages by classes, and

the determining the second language model comprises:

determining a class regarding the speech data based on the at least one of the acoustic information and the information about the preferred languages by classes; and

determining the second language model based on the determined class.

13. The method of claim 7 , further comprising:

obtaining text that is a result of speech recognition of the speech data;

detecting information about content from the text or situation information;

detecting acoustic information from the speech data;

determining a class corresponding to the information about the content and the acoustic information; and

updating information about a language model corresponding to the determined class based on at least one of the information about the content and the situation information.

14. A device for performing speech recognition, the device comprising:

a user inputter, which obtains speech data for performing speech recognition; and

a controller, which obtains at least one phoneme sequence from the speech data, obtains information about probabilities that predetermined unit components constituting the at least one phoneme sequence appear during speech recognition, determines one of the at least one phoneme sequence based on the information about the probabilities that the predetermined unit components appear, and obtains a word corresponding to the determined phoneme sequence based on segment information for converting the predetermined unit components included in the determined phoneme sequence into a word.

15. A non-transitory computer-readable recording medium storing a program for implementing the method of claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 20, 2017
From: PARK, CHI-YOUN; KIM, IL-HWAN; LEE, KYUNG-MIN; KIM, NAM-HOON; LEE, JAE-WON
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 043052/0170 →
Continuity (1)
Related Publication 20170365251A1 · Dec 21, 2017