IP Library Granted Patent US 11,935,523
Granted Patent B2
US 11,935,523 · App. 16/684,619 · Granted Mar 19, 2024

Detection of correctness of pronunciation

Inventor: Aleksandr Diment (Tampere, FI)
Assignee: Master English Oy
G10L15/187G06N3/044G06N3/08G10L15/02G10L15/16G10L15/22G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,935,523
App. No.
16/684,619
Granted
Mar 19, 2024
Kind
B2
Abstract

There is provided automatic detection of pronunciation errors in spoken words utilizing a neural network model that is trained for a target phoneme. The target phoneme may be a phoneme in English language. The pronunciation errors may be detected in English words.

Claims (57)

1. A method, comprising:

receiving, in an acoustic branch of a system for automatic detection of pronunciation errors, an audio waveform of a recording comprising a known pronounced word comprising a target phoneme;

transforming the audio waveform to a sequence of frames;

extracting one or more acoustic features from the sequence of frames;

inputting the one or more acoustic features to a recurrent block of the acoustic branch;

receiving, in a phonetic branch of the system for automatic detection of pronunciation errors, a sequence of phonetic symbols representing ideal pronunciation of the pronounced word;

mapping the sequence of phonetic symbols using a neural network trained for the target phoneme to obtain an embedded phonetic sequence;

inputting the embedded phonetic sequence to a recurrent block of the phonetic branch; and

classifying, by a neural network trained for the target phoneme, a pronunciation in the recording as correct or as comprising one or more pronunciation errors in the target phoneme based on a representation of the one or more acoustic features as obtained from the recurrent block of the acoustic branch and a representation of the sequence of phonetic symbols as obtained from the recurrent block of the phonetic branch, wherein the classifying comprises: concatenating the representation of the one or more acoustic features as obtained from the recurrent block of the acoustic branch and the representation of the sequence of phonetic symbols as obtained from the recurrent block of the phonetic branch to learn the neural network for detecting one or more pronunciation errors in the target phoneme from the one or more acoustic features.

2. The method according to claim 1 , further comprising:

classifying, in response to detecting that a likelihood of the one or more pronunciation errors is below a pre-determined threshold, a pronunciation in the recording as correct.

3. The method according to claim 1 , further comprising:

providing feedback to a user based on the classified pronunciation.

4. The method according to claim 1 , further comprising:

selecting, in response to classifying the pronunciation as comprising one or more pronunciation errors in the target phoneme, a video showing correct pronunciation of the word comprising the target phoneme; and

providing the selected video for display.

5. The method according to claim 1 , further comprising:

selecting, in response to classifying the pronunciation as comprising one or more pronunciation errors in the target phoneme, one or more further words comprising the target phoneme, wherein the one or more further words are different than the pronounced word in the recording; and

providing a user with the one or more further words to be pronounced.

6. The method according to claim 1 , wherein the pronounced word comprises a second target phoneme; and wherein the method further comprises:

mapping the sequence of phonetic symbols using a neural network trained for the second target phoneme to obtain a second embedded phonetic sequence; and

classifying, by a neural network trained for the second target phoneme, a pronunciation in the recording as correct or comprising one or more pronunciation errors in the second target phoneme based on the one or more acoustic features and the second embedded phonetic sequence.

7. An apparatus comprising at least one processor; and at least one memory including computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to perform:

receiving, in an acoustic branch of a system for automatic detection of pronunciation errors, an audio waveform of a recording comprising a known pronounced word comprising a target phoneme;

transforming the audio waveform to a sequence of frames;

extracting one or more acoustic features from the sequence of frames;

inputting the one or more acoustic features to a recurrent block of the acoustic branch;

receiving, in a phonetic branch of the system for automatic detection of pronunciation errors, a sequence of phonetic symbols representing ideal pronunciation of the pronounced word;

mapping the sequence of phonetic symbols using a neural network trained for the target phoneme to obtain an embedded phonetic sequence;

inputting the embedded phonetic sequence to a recurrent block of the phonetic branch; and

classifying, by a neural network trained for the target phoneme, a pronunciation in the recording as correct or as comprising one or more pronunciation errors in the target phoneme based on a representation of the one or more acoustic features as obtained from the recurrent block of the acoustic branch and a representation of the sequence of phonetic symbols as obtained from the recurrent block of the phonetic branch, wherein the classifying comprises: concatenating the representation of the one or more acoustic features as obtained from the recurrent block of the acoustic branch and the representation of the sequence of phonetic symbols as obtained from the recurrent block of the phonetic branch to learn the neural network for detecting one or more pronunciation errors in the target phoneme from the one or more acoustic features.

8. The apparatus according to claim 7 , further comprising:

classifying, in response to detecting that a likelihood of the one or more pronunciation errors is below a pre-determined threshold, a pronunciation in the recording as correct.

9. The apparatus according to claim 7 , further comprising:

providing feedback to a user based on the classified pronunciation.

10. The apparatus according to claim 7 , further comprising:

selecting, in response to classifying the pronunciation as comprising one or more pronunciation errors, a video showing correct pronunciation of the word comprising the target phoneme; and providing the selected video for display.

11. The apparatus according to claim 7 , further comprising:

selecting, in response to classifying the pronunciation as comprising one or more pronunciation errors in the target phoneme, one or more further words comprising the target phoneme, wherein the one or more further words are different than the pronounced word in the recording; and

providing a user with the one or more further words to be pronounced.

12. The apparatus according to claim 7 , wherein the pronounced word comprises a second target phoneme; and wherein the apparatus further comprises:

mapping the sequence of phonetic symbols using a neural network trained for the second target phoneme to obtain a second embedded phonetic sequence;

inputting the second embedded phonetic sequence to a recurrent block of the phonetic branch of the system for automatic detection of pronunciation errors; and

classifying, by a neural network trained for the second target phoneme, a pronunciation in the recording as correct or comprising one or more pronunciation errors in the second target phoneme based on the one or more acoustic features as obtained from the recurrent block of the acoustic branch and the second embedded phonetic sequence as obtained from the recurrent block of the phonetic branch.

13. A non-transitory computer readable medium comprising program instructions that, when executed by at least one processor, cause an apparatus to at least perform:

receiving, in an acoustic branch of a system for automatic detection of pronunciation errors, an audio waveform of a recording comprising a known pronounced word comprising a target phoneme;

transforming the audio waveform to a sequence of frames;

extracting one or more acoustic features from the sequence of frames;

inputting the one or more acoustic features to a recurrent block of the acoustic branch;

receiving, in a phonetic branch of the system for automatic detection of pronunciation errors, a sequence of phonetic symbols representing ideal pronunciation of the pronounced word;

mapping the sequence of phonetic symbols using a neural network trained for the target phoneme to obtain an embedded phonetic sequence;

inputting the embedded phonetic sequence to a recurrent block of the phonetic branch; and

classifying, by a neural network trained for the target phoneme, a pronunciation in the recording as correct or as comprising one or more pronunciation errors in the target phoneme based on a representation of the one or more acoustic features as obtained from the recurrent block of the acoustic branch and a representation of the sequence of phonetic symbols as obtained from the recurrent block of the phonetic branch, wherein the classifying comprises: concatenating the representation of the one or more acoustic features as obtained from the recurrent block of the acoustic branch and the representation of the sequence of phonetic symbols as obtained from the recurrent block of the phonetic branch to learn the neural network for detecting one or more pronunciation errors in the target phoneme from the one or more acoustic features.

14. The non-transitory computer readable medium according to claim 13 , wherein the pronounced word comprises a second target phoneme; and wherein the apparatus is further caused to perform:

mapping the sequence of phonetic symbols using a neural network trained for the second target phoneme to obtain a second embedded phonetic sequence;

inputting the second embedded phonetic sequence to a recurrent block of the phonetic branch of the system for automatic detection of pronunciation errors; and

classifying, by a neural network trained for the second target phoneme, a pronunciation in the recording as correct or comprising one or more pronunciation errors in the second target phoneme based on the one or more acoustic features as obtained from the recurrent block of the acoustic branch and the second embedded phonetic sequence as obtained from the recurrent block of the phonetic branch.

Assignments (3)
CHANGE OF NAME Recorded Sep 13, 2023
From: WORDDIVE OY
To: MASTER ENGLISH OY
Reel/Frame 064884/0380 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2020
From: TAMPEREEN YLIOPISTO
To: WORDDIVE OY
Reel/Frame 051595/0736 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2020
From: DIMENT, ALEKSANDR
To: TAMPEREEN YLIOPISTO
Reel/Frame 051490/0056 →
Continuity (1)
Related Publication 20210151036A1 · May 20, 2021
Cited By (1)
US 12,265,790