Method and system using phoneme embedding
A system and method for creating an embedded phoneme map from a corpus of speech in accordance with a multiplicity of acoustic features of the speech. The embedded phoneme map is used to determine how to pronounce borrowed words from a lending language in the borrowing language, using the phonemes of the borrowing language that are closest to the phonemes of the lending language. The embedded phoneme map is also used to help linguists visualize the phonemes being pronounced by a speaker in real-time and to help non-native speakers practice pronunciation by displaying the differences between proper pronunciation and actual pronunciation for open-ended speech by the speaker.
1. A method, comprising:
receiving speech audio segments;
extracting acoustic features from the speech audio segments;
determining, according to the acoustic features, phoneme vectors within a multidimensional embedding space in a memory;
projecting the phoneme vectors onto a two dimensional space in the memory representing vowels centrally and consonants peripherally;
providing a two-dimensional display of the locations of the projected phoneme vectors within the two dimensional space;
identifying, within the multidimensional embedding space, respective phoneme-specific Voronoi partitions in which the phoneme vectors are present; and
performing the following one or more times for respective phoneme vectors, to create an effect of animating locations of the respective phoneme vectors on the two dimensional display while a user is speaking:
highlighting the location of the respective phoneme vector in the two-dimensional display according to a respective distance between the respective phoneme vector and a respective center of the Voronoi partition in which the respective phoneme vector is present.
2. The method of claim 1 , further comprising:
unhighlighting the location of the respective phoneme vector after the location is highlighted and before another location is highlighted.
3. The method of claim 1 , where the locations are highlighted in real time as a human speaker is speaking the speech audio segments.
4. The method of claim 1 , where the locations are highlighted in semi-real time as a human speaker is speaking the speech audio segments.
5. The method of claim 1 wherein a user interface includes a start button and a stop button to indicate the start and end of when the speech audio segment is to be included in the received speech audio segments.
6. The method of claim 1 , wherein the plurality of locations are highlighted for different durations, in accordance with the duration of the phonemes in the speech audio.
7. The method of claim 1 , wherein providing a two dimensional display further comprises:
displaying a user interface comprising a plurality of phonemes representing the speech audio and further indicating phonemes for a native speaker's pronunciation of the speech audio.
8. The method of claim 1 , wherein the phoneme-specific Voronoi partitions are of differing sizes.
9. The method of claim 1 , further comprising:
training a model, using the extracted acoustic features, to define contiguous regions within the multidimensional embedding space, each region corresponding to a label,
whereby a given new segment of speech audio of a phoneme can be mapped to a single regions with the multidimensional embedding space.
10. The method of claim 1 , further comprising:
determining a midpoint value for each region of a the multidimensional embedding space, wherein the multidimensional embedding space is a phonetic embedding space.
11. The method of claim 1 , further comprising:
displaying printed phonemes representing the received speech audio segments in a sentence format; and
indicating a degree of correctness of the displayed printed phonemes based on the distance between the phoneme vector and the midpoint value of the center of the Voronoi partition, wherein the distance indicates a degree of incorrectness of a pronunciation of the phoneme in the received speech audio.
12. The method of claim 1 , further comprising:
adding phonemes corresponding to the phoneme vectors to a borrowing language dictionary as a new word from a borrowed language.
13. A device, comprising:
a memory storing executable instructions;
a processor, executing the instructions stored in the memory to perform the following:
receiving speech audio segments;
extracting acoustic features from the speech audio segments;
determining, according to the acoustic features, phoneme vectors within a multidimensional embedding space in a memory;
projecting the phoneme vectors onto a two dimensional space in the memory, representing vowels centrally and consonants peripherally;
providing a two-dimensional display of locations of the projected phoneme vectors within the two dimensional space;
identifying, within the multidimensional embedding space, respective phoneme-specific Voronoi partitions in which the phoneme vectors are present; and
highlighting the location of the phoneme vector in the two-dimensional display according to respective distances between the phoneme vectors and centers of the Voronoi partitions in which the phoneme vectors are present.