IP Library Granted Patent US 12,488,784
Granted Patent B2
US 12,488,784 · App. 17/897,308 · Granted Dec 2, 2025

System and method for adapting natural language understanding (NLU) engines optimized on text to audio input

Inventor: Jean-Francois Lavallee (Quebec, CA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/063G06F40/242G10L15/02G10L15/187G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,784
App. No.
17/897,308
Granted
Dec 2, 2025
Kind
B2
Abstract

A method, computer program product, and computing system for generating a plurality of potential vocalizations of a plurality of text samples. A plurality of phonemes associated with the plurality of potential vocalizations are identified. A plurality of phonetically-related text portions are generated based upon, at least in part, the plurality of phonemes. A natural language understanding (NLU) engine is trained using the plurality of phonetically-related text portions.

Claims (25)

1 . A computer-implemented method, executed on a computing device, comprising:

generating a plurality of potential vocalizations of a plurality of text samples, wherein generating the plurality of potential vocalizations includes processing the plurality of text samples using a text-to-speech tokenizer that converts each text sample into a token representing the speech content of the respective text sample;

identifying a plurality of phonemes associated with the plurality of potential vocalizations, wherein identifying the plurality of phonemes associated with the plurality of potential vocalizations includes matching at least a portion of the plurality of potential vocalizations to the plurality of phonemes from a phonetic dictionary:

generating a plurality of phonetically-related text portions based upon, at least in part, the plurality of phonemes, wherein generating the plurality of phonetically-related text portions includes:

identifying a plurality of sequences of text portions with at least a threshold similarity to the plurality of phonemes; and

selecting a sequence of text portions from the plurality of sequences of text portions that covers each of the plurality of text samples and is in a same order as the plurality of text samples, thus defining the plurality of phonetically-related text portions; and

training a natural language understanding (NLU) engine using the plurality of phonetically-related text portions.

2 . The computer-implemented method of claim 1 , wherein the NLU engine is a machine learning model.

3 . The computer-implemented method of claim 1 , wherein training the NLU engine using the plurality of phonetically-related text portions includes generating a training batch including a predefined amount of the plurality of text samples and a predefined amount of the plurality of phonetically-related text portions.

4 . The computer-implemented method of claim 3 , wherein training the NLU engine using the plurality of phonetically-related text portions includes training the NLU engine with the training batch including the predefined amount of the plurality of text samples and the predefined amount of the plurality of phonetically-related text portions.

5 . A computing system comprising: a memory; and

a processor to generate a plurality of potential vocalizations of a plurality of text samples, wherein generating the plurality of potential vocalizations includes processing the plurality of text samples using a text-to-speech tokenizer that converts each text sample into a token representing the speech content of the respective text sample, to match at least a portion of the plurality of potential vocalizations to a plurality of phonemes from a phonetic dictionary, to generate a plurality of phonetically-related text portions based upon, at least in part, the plurality of phonemes, wherein generating the plurality of phonetically-related text portions includes: identifying a plurality of sequences of text portions with at least a threshold similarity to the plurality of phonemes, and selecting a sequence of text portions from the plurality of sequences of text portions that covers each of the plurality of text samples and is in a same order as the plurality of text samples, thus defining the plurality of phonetically-related text portions, and to train a natural language understanding (NLU) engine with a predefined amount of the plurality of text samples and a predefined amount of the plurality of phonetically-related text portions.

6 . The computing system of claim 5 , wherein the NLU engine is a machine learning model.

7 . The computing system of claim 6 , wherein generating the plurality of phonetically-related text portions includes generating the plurality of phonetically-related text portions at each epoch during fine-tuning of the NLU engine.

8 . The computing system of claim 5 , wherein training the NLU engine using the plurality of phonetically-related text portions includes generating a training batch including the predefined amount of the plurality of text samples and the predefined amount of the plurality of phonetically-related text portions.

9 . The computing system of claim 8 , wherein the predefined amount of the plurality of text samples is less than or equal to fifty percent of the training batch.

10 . A computer program product residing on a hardware computer readable medium having a plurality of instructions stored thereon which, when executed by a processor, cause the processor to perform operations comprising:

generating a plurality of potential vocalizations of a plurality of text samples, wherein generating the plurality of potential vocalizations includes processing the plurality of text samples using a text-to-speech tokenizer that converts each text sample into a token representing the speech content of the respective text sample;

identifying a plurality of phonemes associated with the plurality of potential vocalizations, wherein identifying the plurality of phonemes associated with the plurality of potential vocalizations includes matching at least a portion of the plurality of potential vocalizations to the plurality of phonemes from a phonetic dictionary:

generating a plurality of phonetically-related text portions based upon, at least in part, the plurality of phonemes, wherein generating the plurality of phonetically-related text portions includes:

identifying a plurality of sequences of text portions with at least a threshold similarity to the plurality of phonemes; and

selecting a sequence of text portions from the plurality of sequences of text portions that covers each of the plurality of text samples and is in a same order as the plurality of text samples, thus defining the plurality of phonetically-related text portions; and

training a natural language understanding (NLU) engine using the plurality of phonetically-related text portions, wherein training the NLU engine using the plurality of phonetically-related text portions includes generating a training batch including a predefined amount of the plurality of text samples and a predefined amount of the plurality of phonetically-related text portions.

11 . The computer program product of claim 10 , wherein the NLU engine is a machine learning model.

12 . The computer program product of claim 11 , wherein generating the plurality of phonetically-related text portions includes generating the plurality of phonetically-related text portions at each epoch during fine-tuning of the NLU engine.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 29, 2022
From: LAVALLEE, JEAN-FRANCOIS
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 060922/0500 →
Continuity (1)
Related Publication 20240071368A1 · Feb 29, 2024
References Cited (11)
US 11080336B2 · Van Dusen · 2021 [cited by examiner]
US 20090068625A1 · Petro · 2009 [cited by examiner]
US 20170221475A1 · Bruguier · 2017 [cited by examiner]
US 20190096390A1 · Kurata · 2019 [cited by examiner]
US 20200175961A1 · Thomson · 2020 [cited by examiner]
US 20210151042A1 · Park · 2021 [cited by examiner]
WO WO2021216299A1 · 2021 [cited by examiner]
WO WO2024178262A1 · 2024 [cited by examiner]
N. T. Rudrappa and M. V. Reddy, “Using Machine Learning for Speech Extraction and Translation: HiTEK Languages, ” 2022 9th International Conference on Computing for Sustainable Global Development (INDIACom), New Delhi, … [cited by examiner]
D. Govind, R. Vishnu and D. Pravena, “Improved method for epoch estimation in telephonic speech signals using zero frequency filtering,” 2015 IEEE International Conference on Signal and Image Processing Applications (IC… [cited by examiner]
K. W. Gamage, V. Sethu and E. Ambikairajah, “Modeling variable length phoneme sequences—A step towards linguistic information for speech emotion recognition in wider world,” 2017 Seventh International Conference on Affe… [cited by examiner]