IP Library Granted Patent US 12,444,402
Granted Patent B2
US 12,444,402 · App. 17/847,469 · Granted Oct 14, 2025

Speech recognition device and operating method thereof

Inventors: Jakub Hoscilowicz (Warsaw, PL); Kornel Jankowski (Warsaw, PL)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L15/02G06F16/68G10L15/04G10L15/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,444,402
App. No.
17/847,469
Granted
Oct 14, 2025
Kind
B2
Abstract

Provided are a method and device for speech recognition. The speech recognition method includes: receiving a speech signal generated by an utterance of a user; identifying a named entity from the received speech signal; determining a speech signal portion, which corresponds to the identified named entity, from the received speech signal; generating a first acoustic embedding vector corresponding to the speech signal portion, based on an acoustic embedding model; determining a second acoustic embedding vector that is one of a plurality of acoustic embedding vectors corresponding to a plurality of named entities included in an acoustic embedding database (DB), based on distances between the plurality of acoustic embedding vectors and the first acoustic embedding vector; determining a corrected named entity corresponding to the second acoustic embedding vector; and providing a result of speech recognition with respect to the speech signal, based on the corrected named entity.

Claims (60)

1. A speech recognition method comprising:

receiving a speech signal generated by an utterance of a user;

identifying a named entity from the received speech signal;

determining a speech signal portion, which corresponds to the identified named entity, from the received speech signal;

generating a first acoustic embedding vector corresponding to the speech signal portion, based on an acoustic embedding model;

determining an acoustic embedding vector closet to the first acoustic embedding vector from among a plurality of acoustic embedding vectors corresponding to a plurality of named entities included in an acoustic embedding database (DB), as a second acoustic embedding vector, based on the acoustic embedding model;

determining a corrected named entity corresponding to the second acoustic embedding vector, from among the plurality of named entities included in the acoustic embedding DB;

displaying the corrected named entity and a result of speech recognition with respect to the speech signal, based on the corrected named entity;

determining at least one candidate embedding vector including an acoustic embedding vector second closest to the first acoustic embedding vector and an acoustic embedding vector third closest to the first acoustic embedding vector from among the plurality of acoustic embedding vectors, based on the acoustic embedding model;

displaying at least one candidate named entity corresponding to the at least one candidate embedding vector from among the plurality of named entities included in the acoustic embedding DB; and

based on receiving a user input for selecting one of the at least one candidate named entity, displaying a result of the speech recognition corresponding to the selected candidate named entity and storing the first acoustic embedding vector in the acoustic embedding DB to correspond to the selected candidate named entity,

wherein the acoustic embedding model is trained using training data, the training data comprising texts representing one or more named entities, respective phoneme labels corresponding to the texts, and one or more speech signals corresponding to utterances of the one or more named entities; and wherein the training comprises:

performing a first training of the acoustic embedding model using the training data including the texts, the respective phoneme labels, and the one or more speech signals; and

performing a second training of the acoustic embedding model using a subset of the training data, the subset including the texts and the one or more speech signals.

2. The speech recognition method of claim 1 , wherein the determining of the second acoustic embedding vector based on the distances between the plurality of acoustic embedding vectors and the first acoustic embedding vector comprises:

determining the second acoustic embedding vector which is closest in distance to the first acoustic embedding vector, from among the plurality of acoustic embedding vectors.

3. The speech recognition method of claim 1 , wherein the acoustic embedding DB includes a named entity corresponding to two or more embedding vectors that are converted from different speech signals.

4. The speech recognition method of claim 3 , wherein the different speech signals are speech signals according to pronunciations of the named entity in different linguistic spheres.

5. The speech recognition method of claim 1 , wherein the determining of the speech signal portion corresponding to the identified named entity, from the received speech signal, comprises:

identifying time periods of phonemes represented by the received speech signal;

determining a time period corresponding to the identified named entity, based on the identified time periods of the phonemes; and

determining a portion of the received speech signal, which corresponds to the time period, to be the speech signal portion corresponding to the identified named entity.

6. The speech recognition method of claim 1 , wherein, as the result of the speech recognition, an application providing a content regarding the corrected named entity is executed, and

the speech recognition method further comprises:

when a user input for the provided content is received, storing the first acoustic embedding vector in the acoustic embedding DB to correspond to the corrected named entity.

7. The speech recognition method of claim 1 , further comprising:

displaying a menu for selecting the named entity identified from the received speech signal, in addition to the result of the speech recognition based on the corrected named entity;

when a user input for selecting the identified named entity is received, storing the first acoustic embedding vector in the acoustic embedding DB to correspond to the identified named entity; and

providing the result of the speech recognition with respect to the speech signal, based on the identified named entity.

8. A speech recognition device comprising:

a microphone;

at least one memory storing one or more instructions; and

at least one processor configured to execute the one or more instructions to:

receive, via the microphone, a speech signal generated by an utterance of a user;

identify a named entity from the received speech signal;

determine a speech signal portion, which corresponds to the identified named entity, from the received speech signal;

generate a first acoustic embedding vector corresponding to the speech signal portion, based on an acoustic embedding model;

determine an acoustic embedding vector closest to the first acoustic embedding vector from among a plurality of acoustic embedding vectors, corresponding to a plurality of named entities included in an acoustic embedding database (DB), as a second acoustic embedding vector, based on the acoustic embedding model;

determine a corrected named entity, which corresponds to the second acoustic embedding vector, from among the plurality of named entities included in the acoustic embedding DB;

display the corrected named entity and a result of speech recognition with respect to the speech signal, based on the corrected named entity;

display at least one candidate named entity corresponding to the at least one candidate embedding vector from among the plurality of named entities included in the acoustic embedding DB; and

based on receiving a user input for selecting one of the at least one candidate named entity, display a result of the speech recognition corresponding to the selected candidate named entity and store the first acoustic embedding vector in the acoustic embedding DB to correspond to the selected candidate named entity,

wherein the acoustic embedding model is trained using training data, the training data comprising texts representing one or more named entities, respective phoneme labels corresponding to the texts, and one or more speech signals corresponding to utterances of the one or more named entities; and wherein the training comprises:

performing a first training of the acoustic embedding model using the training data including the texts, the respective phoneme labels, and the one or more speech signals; and

performing a second training of the acoustic embedding model using a subset of the training data, the subset including the texts and the one or more speech signals.

9. The speech recognition device of claim 8 , wherein the at least one processor is configured to execute the one or more instructions to:

determine, the second acoustic embedding vector which is closest in distance to the first acoustic embedding vector, from among the plurality of acoustic embedding vectors.

10. The speech recognition device of claim 8 , wherein the acoustic embedding DB includes a named entity corresponding to two or more embedding vectors that are converted from different speech signals.

11. The speech recognition device of claim 10 , wherein the different speech signals are speech signals according to pronunciations of the named entity in different linguistic spheres.

12. The speech recognition device of claim 8 , wherein the at least one processor is configured to execute the one or more instructions to determine the speech signal portion, which corresponds to the identified named entity, from the received speech signal by:

identifying time periods of phonemes represented by the received speech signal;

determining a time period corresponding to the identified named entity, based on the identified time periods of the phonemes; and

determining a portion of the received speech signal, which corresponds to the time period, to be the speech signal portion corresponding to the identified named entity.

13. The speech recognition device of claim 8 , wherein, as the result of the speech recognition, an application providing a content regarding the corrected named entity is executed, and

the at least one processor is further configured to execute the one or more instructions to:

when a user input for the provided content is received, store the first acoustic embedding vector in the acoustic embedding DB to correspond to the corrected named entity.

14. The speech recognition device of claim 8 , wherein the at least one processor is further configured to execute the one or more instructions to:

display a menu for selecting the named entity identified from the received speech signal, in addition to the result of the speech recognition based on the corrected named entity;

when a user input for selecting the identified named entity is received, store the first acoustic embedding vector in the acoustic embedding DB to correspond to the identified named entity; and

provide the result of the speech recognition with respect to the speech signal, based on the identified named entity.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2022
From: HOSCILOWICZ, JAKUB; JANKOWSKI, KORNEL
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 060294/0094 →
Priority Claims (1)
KR 10-2021-0126707 · Sep 24, 2021 · national
Continuity (2)
Continuation PCTKR2022008311 · Jun 13, 2022
Related Publication 20230115538A1 · Apr 13, 2023
References Cited (30)
US 9454957B1 · Mathias et al. · 2016 [cited by applicant]
US 9779730B2 · Hong et al. · 2017 [cited by applicant]
US 9922025B2 · Cross, III et al. · 2018 [cited by applicant]
US 10055489B2 · Haviv et al. · 2018 [cited by applicant]
US 10073887B2 · Chehreghani · 2018 [cited by applicant]
US 10255907B2 · Nallasamy · 2019 [cited by examiner]
US 10672391B2 · Maergner et al. · 2020 [cited by applicant]
US 10839159B2 · Yang et al. · 2020 [cited by applicant]
US 11232785B2 · Seo et al. · 2022 [cited by applicant]
US 11410642B2 · Scuderi et al. · 2022 [cited by applicant]
US 20090077122A1 · Fume · 2009 [cited by examiner]
US 20160027437A1 · Hong et al. · 2016 [cited by applicant]
US 20170084267A1 · Shin · 2017 [cited by examiner]
US 20190377747A1 · Fan et al. · 2019 [cited by applicant]
US 20200020327A1 · Chae et al. · 2020 [cited by applicant]
US 20210216722A1 · Dai et al. · 2021 [cited by applicant]
US 20220129632A1 · Wei et al. · 2022 [cited by applicant]
CN 111737979A · 2020 [cited by applicant]
CN 112257422A · 2021 [cited by applicant]
CN 112836513A · 2021 [cited by applicant]
CN 111798840B · 2023 [cited by examiner]
KR 1020180062003A · 2018 [cited by applicant]
KR 1020190083629A · 2019 [cited by applicant]
KR 1020190098928A · 2019 [cited by applicant]
KR 1020210001937A · 2021 [cited by applicant]
KR 102332729B1 · 2021 [cited by applicant]
WO 2016048350A1 · 2016 [cited by applicant]
WO WO2020069051A1 · 2020 [cited by examiner]
WO WO2020256749A1 · 2020 [cited by examiner]
International Search Report (PCT/ISA/220 and PCT/ISA/210) and Written Opinion (PCT/ISA/237) issued Sep. 22, 2022 by the International Searching Authority in International Application No. PCT/KR2022/008311. [cited by applicant]
Cited By (1)
US 12,718,029