IP Library › Granted Patent US 12,266,350
Granted Patent B1
US 12,266,350 · App. 17/583,812 · Granted Apr 1, 2025

Pronunciation features for language models

Inventors: Siddha Ganju (Santa Clara, CA); Ruthie Lyle (Durham, NC); Steven Dalton (Cary, NC)
Assignee: Nvidia Corporation
G10L15/16G10L13/08G10L15/02G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,266,350
App. No.
17/583,812
Filed
Jan 25, 2022
Granted
Apr 1, 2025
Kind
B1
Art Unit
2654
USPC
704/232
Abstract

Systems and methods are directed toward evaluating auditory inputs against a range of tolerance to provide feedback regarding pronunciation. An auditory input may be evaluated using a trained machine learning system and evaluated for similarity against a target word. Similarity may be scored and then evaluated to determine whether the similarity falls within a range of tolerance, wherein the range of tolerance may be adjusted or modified for particular uses. A score within the range of tolerance is indicative of a word that has been pronounced such that it would be perceptible.

Claims (46)

1. A computer-implemented method, comprising:

receiving an auditory input including at least one phoneme forming at least a portion of a word;

determining, using a trained machine learning system, a similarity between the at least one phoneme and a target word;

obtaining one or more properties of a user providing the auditory input;

determining, based at least in part on the one or more properties, one or more tuning parameters;

determining the similarity is within a range of tolerance, the range of tolerance being tunable based at least on the one or more properties; and

providing confirmation of the at least one phoneme.

2. The computer-implemented method of claim 1 , wherein the trained machine learning system includes at least a Siamese neural network.

3. The computer-implemented method of claim 1 , further comprising:

adjusting the range of tolerance based, at least in part, on the one or more tuning parameters.

4. The computer-implemented method of claim 1 , wherein the one or more properties include at least one of a user native language, a user language competence, a user age, a user education level, or a user performance level.

5. The computer-implemented method of claim 1 , further comprising:

determining, based at least in part on a first user score, a first user performance level, the first user score associated with a first time period;

determining, based at least in part on a second user score, a second user performance level, the second user score associated with a second time period, later than the first time period; and

determining a change in user performance level based at least in part on the first user performance level and the second user performance level.

6. The computer-implemented method of claim 1 , further comprising:

storing the auditory input and the similarity; and

updating one or more parameters of the trained machine learning system using, at least in part, the auditory input and the similarity.

7. The computer-implemented method of claim 1 , wherein data used to adjust one or more tuning parameters, including the range of tolerance, is locally stored on the user device.

8. The computer-implemented method of claim 1 , wherein the range of tolerance corresponds to a learned distance metric between the at least one phoneme and at least one of the target word or the target phoneme.

9. A method, comprising:

generating one or more warped words based, at least in part, on a ground truth word, the one or more warped words including at least one of a misspelling or a mispronunciation of the ground truth word;

generating for each of the one or more warped words and the ground truth word, respective audio and text samples;

training, using at least the respective audio and text samples, a machine learning system;

generating one or more embeddings for the ground truth word and the one or more warped words; and

determining, based at least in part on the one or more embeddings, a range of tolerance.

10. The method of claim 9 , wherein the machine learning system includes at least a Siamese neural network.

11. The method of claim 9 , wherein the respective audio samples are produced using a text to speech system.

12. The method of claim 9 , further comprising:

modifying the range of tolerance based, at least in part, on one or more tuning parameters.

13. The method of claim 12 , wherein the one or more tuning parameters are based, at least in part, on one or more user properties.

14. A computer-implemented method, comprising:

providing an interface for a language learning system to a user, the interface to prompt the user to perform one or more actions;

receiving, responsive to a first prompt, an auditory input from the user, the auditory input corresponding to an utterance associated with a target word;

determining the utterance is within a range of tolerance for the target word, the range of tolerance corresponding to a learned distance metric associated with a perceptibility of the target word, wherein the range of tolerance is based, at least in part, on a user competency level; and

providing, via the interface, a notification of the utterance being within the range of tolerance.

15. The computer-implemented method of claim 14 , further comprising:

receiving, responsive to a second prompt, a user input, the second prompt presenting at least one of a word spelling or an auditory word pronunciation;

determining the user input corresponds to a correct answer associated with the second prompt; and

providing, via the interface, a second notification of the user input being the correct answer.

16. The computer-implemented method of claim 14 , wherein the learned distance metric is generated using at least one of triplet loss or contrastive loss.

17. The computer-implemented method of claim 14 , further comprising:

determining a first competency level for the user;

collecting performance metrics for the user;

determining a second competency level for the user; and

adjusting the range of tolerance based, at least in part, on the second competency level.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2022
From: GANJU, SIDDHA; LYLE, RUTHIE; DALTON, STEVEN
To: NVIDIA CORPORATION
Reel/Frame 058834/0088 →
Continuity (2)
Provisional Application 63272952 · Oct 28, 2021
Provisional Application 63181934 · Apr 29, 2021
References Cited (13)
US 10923111B1 · Fan · 2021 [cited by examiner]
US 11928764B2 · Wang · 2024 [cited by examiner]
US 20020138265A1 · Stevens · 2002 [cited by examiner]
US 20100246837A1 · Krause · 2010 [cited by examiner]
US 20100299148A1 · Krause · 2010 [cited by examiner]
US 20120069131A1 · Abelow · 2012 [cited by examiner]
US 20170270919A1 · Parthasarathi · 2017 [cited by examiner]
US 20190189026A1 · Daniels · 2019 [cited by examiner]
US 20200035231A1 · Parthasarathi · 2020 [cited by examiner]
US 20230142339A1 · Getselevich · 2023 [cited by examiner]
US 20230394823A1 · Weng · 2023 [cited by examiner]
US 20240037756A1 · Huang · 2024 [cited by examiner]
WO WO2022047311A1 · 2022 [cited by examiner]
Cited By (1)
US 12,586,350