IP Library Granted Patent US 12,488,784
Granted Patent B2
US 12,488,784 · App. 17/897,308 · Granted Dec 2, 2025

System and method for adapting natural language understanding (NLU) engines optimized on text to audio input

Inventor: Jean-Francois Lavallee (Quebec, CA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/063G06F40/242G10L15/02G10L15/187G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,784
App. No.
17/897,308
Granted
Dec 2, 2025
Kind
B2
Abstract

A method, computer program product, and computing system for generating a plurality of potential vocalizations of a plurality of text samples. A plurality of phonemes associated with the plurality of potential vocalizations are identified. A plurality of phonetically-related text portions are generated based upon, at least in part, the plurality of phonemes. A natural language understanding (NLU) engine is trained using the plurality of phonetically-related text portions.

Claims (25)

1 . A computer-implemented method, executed on a computing device, comprising:

generating a plurality of potential vocalizations of a plurality of text samples, wherein generating the plurality of potential vocalizations includes processing the plurality of text samples using a text-to-speech tokenizer that converts each text sample into a token representing the speech content of the respective text sample;

identifying a plurality of phonemes associated with the plurality of potential vocalizations, wherein identifying the plurality of phonemes associated with the plurality of potential vocalizations includes matching at least a portion of the plurality of potential vocalizations to the plurality of phonemes from a phonetic dictionary:

generating a plurality of phonetically-related text portions based upon, at least in part, the plurality of phonemes, wherein generating the plurality of phonetically-related text portions includes:

identifying a plurality of sequences of text portions with at least a threshold similarity to the plurality of phonemes; and

selecting a sequence of text portions from the plurality of sequences of text portions that covers each of the plurality of text samples and is in a same order as the plurality of text samples, thus defining the plurality of phonetically-related text portions; and

training a natural language understanding (NLU) engine using the plurality of phonetically-related text portions.

2 . The computer-implemented method of claim 1 , wherein the NLU engine is a machine learning model.

3 . The computer-implemented method of claim 1 , wherein training the NLU engine using the plurality of phonetically-related text portions includes generating a training batch including a predefined amount of the plurality of text samples and a predefined amount of the plurality of phonetically-related text portions.

4 . The computer-implemented method of claim 3 , wherein training the NLU engine using the plurality of phonetically-related text portions includes training the NLU engine with the training batch including the predefined amount of the plurality of text samples and the predefined amount of the plurality of phonetically-related text portions.

5 . A computing system comprising: a memory; and

a processor to generate a plurality of potential vocalizations of a plurality of text samples, wherein generating the plurality of potential vocalizations includes processing the plurality of text samples using a text-to-speech tokenizer that converts each text sample into a token representing the speech content of the respective text sample, to match at least a portion of the plurality of potential vocalizations to a plurality of phonemes from a phonetic dictionary, to generate a plurality of phonetically-related text portions based upon, at least in part, the plurality of phonemes, wherein generating the plurality of phonetically-related text portions includes: identifying a plurality of sequences of text portions with at least a threshold similarity to the plurality of phonemes, and selecting a sequence of text portions from the plurality of sequences of text portions that covers each of the plurality of text samples and is in a same order as the plurality of text samples, thus defining the plurality of phonetically-related text portions, and to train a natural language understanding (NLU) engine with a predefined amount of the plurality of text samples and a predefined amount of the plurality of phonetically-related text portions.

6 . The computing system of claim 5 , wherein the NLU engine is a machine learning model.

7 . The computing system of claim 6 , wherein generating the plurality of phonetically-related text portions includes generating the plurality of phonetically-related text portions at each epoch during fine-tuning of the NLU engine.

8 . The computing system of claim 5 , wherein training the NLU engine using the plurality of phonetically-related text portions includes generating a training batch including the predefined amount of the plurality of text samples and the predefined amount of the plurality of phonetically-related text portions.

9 . The computing system of claim 8 , wherein the predefined amount of the plurality of text samples is less than or equal to fifty percent of the training batch.

10 . A computer program product residing on a hardware computer readable medium having a plurality of instructions stored thereon which, when executed by a processor, cause the processor to perform operations comprising:

generating a plurality of potential vocalizations of a plurality of text samples, wherein generating the plurality of potential vocalizations includes processing the plurality of text samples using a text-to-speech tokenizer that converts each text sample into a token representing the speech content of the respective text sample;

identifying a plurality of phonemes associated with the plurality of potential vocalizations, wherein identifying the plurality of phonemes associated with the plurality of potential vocalizations includes matching at least a portion of the plurality of potential vocalizations to the plurality of phonemes from a phonetic dictionary:

generating a plurality of phonetically-related text portions based upon, at least in part, the plurality of phonemes, wherein generating the plurality of phonetically-related text portions includes:

identifying a plurality of sequences of text portions with at least a threshold similarity to the plurality of phonemes; and

selecting a sequence of text portions from the plurality of sequences of text portions that covers each of the plurality of text samples and is in a same order as the plurality of text samples, thus defining the plurality of phonetically-related text portions; and

training a natural language understanding (NLU) engine using the plurality of phonetically-related text portions, wherein training the NLU engine using the plurality of phonetically-related text portions includes generating a training batch including a predefined amount of the plurality of text samples and a predefined amount of the plurality of phonetically-related text portions.

11 . The computer program product of claim 10 , wherein the NLU engine is a machine learning model.

12 . The computer program product of claim 11 , wherein generating the plurality of phonetically-related text portions includes generating the plurality of phonetically-related text portions at each epoch during fine-tuning of the NLU engine.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 29, 2022
From: LAVALLEE, JEAN-FRANCOIS
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 060922/0500 →