IP Library › Granted Patent US 12,482,483
Granted Patent B2
US 12,482,483 · App. 18/057,967 · Granted Nov 25, 2025

Length perturbation techniques for improving generalization of deep neural network acoustic models

Inventors: Xiaodong Cui (Chappaqua, NY); Brian E. D. Kingsbury (Cortlandt Manor, NY); George Andrei Saon (Stamford, CT)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G10L21/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,482,483
App. No.
18/057,967
Filed
Nov 22, 2022
Granted
Nov 25, 2025
Kind
B2
Examiner
VO, HUYEN X
Art Unit
2656
USPC
704/503
Abstract

One or more systems, devices, computer program products and/or computer-implemented methods of use provided herein relate to length perturbation techniques for improving generalization of DNN acoustic models. A computer-implemented system can comprise a memory that can store computer executable components. The computer-implemented system can further comprise a processor that can execute the computer executable components stored in the memory, wherein the computer executable components can comprise a frame skipping component that can remove one or more frames from an acoustic utterance via frame skipping. The computer executable components can further comprise a frame insertion component that can insert one or more replacement frames into the acoustic utterance via frame insertion to replace the one or more frames with the one or more replacement frames to enable length perturbation of the acoustic utterance.

Claims (43)

1 . A computer-implemented system, comprising:

a memory that stores computer executable components; and

a processor that executes at least one of the computer executable components that:

performs length perturbation of an acoustic utterance comprising a group of frames in a sequence, wherein the performing the length perturbation comprises:

random sampling a first defined percentage of frames from the group of frames resulting in a first subset of drop frames;

for each drop frame, removing a first defined quantity of consecutive frames from the group of frames starting with the drop frame;

random sampling a second defined percentage of frames from the group of frames resulting in a second subset of insert frames; and

for each insert frame, inserting a second defined quantity of replacement frames into the acoustic utterance after the insert frame.

2 . The computer-implemented system of claim 1 , wherein the at least one of the computer executable components further:

receives the acoustic utterance from a speech recognition system.

3 . The computer-implemented system of claim 1 , wherein the length perturbation of the acoustic utterance assists with generalization of a deep neural network (DNN) acoustic model.

4 . The computer-implemented system of claim 1 , wherein the length perturbation of the acoustic utterance is a one-way perturbation enabled by at least one of the removing the first defined quantity of consecutive frames or the inserting of the second defined quantity of replacement frames.

5 . The computer-implemented system of claim 1 , wherein the length perturbation of the acoustic utterance is a two-way perturbation enabled by both the removing the first defined quantity of consecutive frames and the frame inserting of the second defined quantity of replacement frames.

6 . The computer-implemented system of claim 1 , wherein at least one of the computer executable components applies the length perturbation of the acoustic utterance as individual data augmentation technique.

7 . The computer-implemented system of claim 1 , wherein at least one of the computer executable components applies the length perturbation of the acoustic utterance as a first data augmentation technique in combination with at least a second data augmentation technique that is different from the length perturbation.

8 . The computer-implemented system of claim 1 , wherein the length perturbation of the acoustic utterance is applied to one or more domains of feature representations of the acoustic utterance.

9 . The computer-implemented system of claim 1 , wherein the the removing the first defined quantity of consecutive frames and the frame inserting of the second defined quantity of replacement frames are applied towards the length perturbation using one or more sequence-to-sequence models, to assist with generalization of the one or more sequence-to-sequence models.

10 . The computer-implemented system of claim 1 , wherein the the removing the first defined quantity of consecutive frames and the frame inserting of the second defined quantity of replacement frames are applied towards the length perturbation using one or more sequence-to-one models, to assist with generalization of the one or more sequence-to-one models.

11 . A computer-implemented method, comprising:

performing, by a system operatively coupled to a processor, length perturbation of an acoustic utterance comprising a group of frames in a sequence, wherein the performing the length perturbation comprises:

random sampling a first defined percentage of frames from the group of frames resulting in a first subset of drop frames;

for each drop frame, removing a first defined quantity of consecutive frames from the group of frames starting with the drop frame;

random sampling a second defined percentage of frames from the group of frames resulting in a second subset of insert frames; and

for each insert frame, inserting a second defined quantity of replacement frames into the acoustic utterance after the insert frame.

12 . The computer-implemented method of claim 11 , further comprising:

receiving, by the system, the acoustic utterance from a speech recognition system.

13 . The computer-implemented method of claim 11 , wherein the length perturbation of the acoustic utterance assists with generalization of a DNN acoustic model.

14 . The computer-implemented method of claim 11 , wherein the length perturbation of the acoustic utterance is a one-way perturbation enabled by at least one of the removing the first defined quantity of consecutive frames or the inserting of the second defined quantity of replacement frames.

15 . The computer-implemented method of claim 11 , wherein the length perturbation of the acoustic utterance is a two-way perturbation enabled by both the removing the first defined quantity of consecutive frames and the frame inserting of the second defined quantity of replacement frames.

16 . The computer-implemented method of claim 11 , further comprising:

applying, by the system, the length perturbation of the acoustic utterance as an individual data augmentation technique.

17 . The computer-implemented method of claim 11 , further comprising:

applying, by the system, the length perturbation of the acoustic utterance as a first data augmentation technique in combination with at least a second data augmentation technique that is different from the length perturbation.

18 . The computer-implemented method of claim 11 , further comprising:

applying, by the system, the length perturbation of the acoustic utterance to one or more domains of feature representations of the acoustic utterance.

19 . A computer program product for improving generalization of a DNN acoustic model, the computer program product comprising a non-transitory computer readable medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to:

performing, by the processor, length perturbation of an acoustic utterance comprising a group of frames in a sequence, wherein the performing the length perturbation comprises:

random sampling a first defined percentage of frames from the group of frames resulting in a first subset of drop frames;

for each drop frame, removing a first defined quantity of consecutive frames from the group of frames starting with the drop frame;

random sampling a second defined percentage of frames from the group of frames resulting in a second subset of insert frames; and

for each insert frame, inserting a second defined quantity of replacement frames into the acoustic utterance after the insert frame.

20 . The computer program product of claim 19 , wherein the program instructions are further executable by the processor to cause the processor to:

receive, by the processor, the acoustic utterance from a speech recognition system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 22, 2022
From: CUI, XIAODONG; KINGSBURY, BRIAN E. D.; SAON, GEORGE ANDREI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 061854/0632 →
Continuity (1)
Related Publication 20240170005A1 · May 23, 2024
References Cited (37)
US 9715870B2 · Hwang et al. · 2017 [cited by applicant]
US 10930268B2 · Yoo et al. · 2021 [cited by applicant]
US 11227579B2 · Nagano et al. · 2022 [cited by applicant]
US 20100023315A1 · Quirk · 2010 [cited by applicant]
US 20140146695A1 · Kim · 2014 [cited by examiner]
US 20140210652A1 · Bartnik et al. · 2014 [cited by applicant]
US 20170040016A1 · Cui et al. · 2017 [cited by applicant]
US 20170040022A1 · Sung · 2017 [cited by examiner]
US 20170103740A1 · Hwang et al. · 2017 [cited by applicant]
US 20180234721A1 · Yang · 2018 [cited by examiner]
US 20190371301A1 · Yoo et al. · 2019 [cited by applicant]
US 20200027444A1 · Prabhavalkar et al. · 2020 [cited by applicant]
US 20200175335A1 · Li et al. · 2020 [cited by applicant]
US 20200243102A1 · Schmidt · 2020 [cited by examiner]
US 20210043186A1 · Nagano et al. · 2021 [cited by applicant]
US 20210141995A1 · Lundgaard et al. · 2021 [cited by applicant]
US 20210350238A1 · Lei · 2021 [cited by applicant]
US 20220108220A1 · Qin et al. · 2022 [cited by applicant]
US 20230107741A1 · Saraf et al. · 2023 [cited by applicant]
US 20230230602A1 · Mariager · 2023 [cited by examiner]
CN 109767782A · 2019 [cited by applicant]
CN 109767782B · 2020 [cited by applicant]
CN 114154519A · 2022 [cited by applicant]
CN 114154519B · 2022 [cited by applicant]
WO 2021217619A1 · 2021 [cited by applicant]
Huang, Sh. et al. | “Context-Aware Selective Label Smoothing for Calibrating Sequence Recognition Model.” MM '21, Oct. 20-24, 2021, Virtual Event, China, 9 pages. [cited by applicant]
Cui, X. et al. | “Improving Generalization of Deep Neural Network Acoustic Models with Length Perturbation and N-best Based Label Smoothing.” arXiv:2203.15176v1 [cs.CL] Mar. 29, 2022, 5 pages. [cited by applicant]
Gong et al. | “MaxUp: A Simple Way to Improve Generalization of Neural Network Training.” arXiv:2002.09024v1 [cs.LG] Feb. 20, 2020, 10 pages. [cited by applicant]
Zhang, X. et al. | “Improving Deep Neural Network Acoustic Models Using Generalized Maxout Networks.” Center for Language and Speech Processing & Human Language Technology Center of Excellence, The Johns Hopkins Univers… [cited by applicant]
Zheng, Y. et al. | “Regularizing Neural Networks via Adversarial Model Perturbation.” CVPR 2021, public version of IEEE Xplore publication, published 2021, 10 pages. [cited by applicant]
ip.com | “Uncertainty Modeling for Neural-Network-Based Speaker Comparison.” Disclosed Anonymously, IP.com No. IPCOM000270265D, IP.com Electronic Publication Date: Jun. 22, 2022, 4 pages. [cited by applicant]
ip.com | “Deep Learning, Linguistic Processing and Dynamic Optimization based Revision Recommendation for optimum learning Path.” Disclosed Anonymously, IP.com No. IPCOM000267766D, IP.com Electronic Publication Date: No… [cited by applicant]
ip.com | “Class Identification in Deep-Learning Models.” Disclosed Anonymously, IP.com No. PCOM000265288D, IP.com Electronic Publication Date: Mar. 23, 2021, 10 pages. [cited by applicant]
Jain et al., “Spliceout: A Simple and Efficient Audio Augmentation Method”, Oct. 13, 2021, 25 pages. [cited by applicant]
Sunder et al., “Towards End-to-End Integration of Dialog History for Improved Spoken Language Understanding”, Apr. 11, 2022, 05 pages. [cited by applicant]
Serai et al., “Improving Speech Recognition Error Prediction for Modern and Off-The-Shelf Speech Recognizers”, IEEE, May 2019, 12 pages. [cited by applicant]
United States Non-Final Rejection dated Aug. 25, 2025, 13 pages in U.S. Appl. No. 18/057,983. [cited by applicant]