IP Library › Granted Patent US 11,355,101
Granted Patent B2
US 11,355,101 · App. 16/814,869 · Granted Jun 7, 2022

Artificial intelligence apparatus for training acoustic model

Inventor: Jeehye Lee (Seoul, KR)
Assignee: LG ELECTRONICS INC.
G10L15/063G06N3/04G06N3/08G10L15/005G10L15/02G10L15/22G10L15/30G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,355,101
App. No.
16/814,869
Granted
Jun 7, 2022
Kind
B2
Abstract

Disclosed is an artificial intelligence (AI) apparatus for training an acoustic model, and more particularly, an AI apparatus for training an acoustic model including a shared network and a branch network connected to the shared network using speech data and phonemes corresponding to the speech data.

Claims (30)

1. An artificial intelligence (AI) apparatus for speech recognition, comprising:

an input interface configured to: obtain speech data corresponding to user speech, receive second speech data, and receive user input for selecting a native speaker or a foreign speaker; and

a processor configured to:

train an acoustic model including a shared network and a branch network connected to the shared network using the speech data and phonemes corresponding to the speech data, wherein the branch network includes a native branch network and a foreign branch network,

change an updated acoustic model to connect the shared network to the native branch network or the foreign branch network corresponding to the received user input for selecting the native speaker or the foreign speaker,

input the received second speech data to the changed updated acoustic model, and

perform speech recognition using phonemes output by the changed updated acoustic model.

2. The AI apparatus of claim 1 , wherein the shared network includes a first AI model configured to apply common acoustic information of a plurality of speech data used during training to output hidden representation and configured to transfer the hidden representation to a layer of the branch network.

3. The AI apparatus of claim 2 , wherein the plurality of speech data includes native speech data or foreign speech data; and

wherein, when obtaining the native speech data, the processor is further configured to train the shared network and a native branch network connected to the shared network, and when obtaining the foreign speech data, the processor is further configured to train the shared network and a foreign branch network connected to the shared network.

4. The AI apparatus of claim 2 , wherein the branch network includes a second AI model configured to output the phonemes corresponding to the speech data based on the hidden representation transferred by a layer of the shared network.

5. The AI apparatus of claim 1 , wherein, when the speech data is speech data uttered by a native speaker, the processor is further configured to train the shared network and a native branch network connected to the shared network using native speech data and phonemes of a native language corresponding to native speech data.

6. The AI apparatus of claim 1 , wherein, when the speech data is uttered by a foreign speaker, the processor is further configured to train the shared network and a foreign branch network connected to the shared network using foreign speech data and phonemes of a foreign language corresponding to the foreign speech data.

7. The AI apparatus of claim 1 , wherein the second speech data includes foreign speech data; and

wherein the processor is further configured to: input the second speech data to the updated acoustic model and perform speech recognition using the phonemes of native language output by the updated acoustic model.

8. An operation method of an artificial intelligence (AI) apparatus for speech recognition, the method comprising:

obtaining speech data corresponding to user speech, second speech data, and user input for selecting a native speaker or a foreign speaker;

training an acoustic model that includes a shared network and a branch network connected to the shared network using the speech data and phonemes corresponding to the speech data, wherein the branch network includes a native branch network and a foreign branch network;

changing an updated acoustic model to connect the shared network to the native branch network or the foreign branch network corresponding to the obtained user input for selecting the native speaker or the foreign speaker;

input the second speech data to the changed updated acoustic model; and

perform speech recognition using phonemes output by the changed updated acoustic model.

9. The method of claim 8 , wherein the training the acoustic model includes applying common acoustic information of a plurality of speech data used when the shared network is trained to output hidden representation and transferring the hidden representation to a layer of the branch network.

10. The method of claim 9 , wherein the plurality of speech data includes native speech data or foreign speech data; and

wherein the training the acoustic model includes:

when obtaining the native speech data, training the shared network and a native branch network connected to the shared network; and

when obtaining the foreign speech data, training the shared network and a native branch network connected to the shared network.

11. The method of claim 9 , wherein the training the acoustic model includes outputting the phonemes corresponding to the speech data based on the hidden representation transferred by the layer of the branch network.

12. The method of claim 8 , wherein the training the acoustic model includes, when the speech data is speech data uttered by a native speaker, training the shared network and a native branch network connected to the shared network using native speech data and phonemes of a native language corresponding to native speech data.

13. The method of claim 8 , wherein the training the acoustic model includes, when the speech data is uttered by a foreign speaker, training the shared network and a foreign branch network connected to the shared network using foreign speech data and phonemes of a foreign language corresponding to the foreign speech data.

14. The method of claim 8 , wherein, the performing the speech recognition includes, when the second speech data includes foreign speech data, performing speech recognition using the phonemes of a native language output by the updated acoustic model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2020
From: LEE, JEEHYE
To: LG ELECTRONICS INC.
Reel/Frame 052074/0105 →
Priority Claims (1)
KR 10-2019-0171684 · Dec 20, 2019 · national
Continuity (1)
Related Publication 20210193119A1 · Jun 24, 2021