IP Library Granted Patent US 12,451,119
Granted Patent B2
US 12,451,119 · App. 17/656,757 · Granted Oct 21, 2025

Systems and methods for employing alternate spellings for improving the recognition of rare words

Inventors: Jennifer Drexler Fox (Austin, TX); Danny Chen (Austin, TX); Natalie Delworth (Austin, TX)
Assignee: Rev.com, Inc.
G10L15/063G06F40/284G06F40/47G10L15/16G10L15/193G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,451,119
App. No.
17/656,757
Granted
Oct 21, 2025
Kind
B2
Abstract

A method of adding a custom vocabulary to a transcription system includes receiving a custom vocabulary at an ASIRW module. The method further includes tokenizing the custom vocabulary with the ASIRW module. The method further includes creating a new WFST (weighted finite-state transducer) with the ASIRW module. The method further includes transcribing audio using the new WFST with the ASIRW module. The tokenizing includes performing a translation model on each word of the custom vocabulary.

Claims (51)

1. A method of adding a custom vocabulary to a transcription system, the method comprising:

receiving a custom vocabulary at an ASIRW module, the custom vocabulary is a list of words which are likely to appear in audio, the custom vocabulary list provided by a user to the ASIRW module, the user providing the audio to the transcription system;

tokenizing the custom vocabulary with the ASIRW module;

creating a new WFST (weighted finite-state transducer) with the ASIRW module;

transcribing the audio using the new WFST with the ASIRW module.

2. The method of claim 1 , wherein the tokenizing includes performing a translation model on each word of the custom vocabulary.

3. The method of claim 1 , wherein the custom vocabulary includes phrases.

4. The method of claim 2 , wherein the tokenizing includes creating predicted tokenizations for alternate spellings of the custom vocabulary.

5. The method of claim 4 , wherein the new WFST includes tokenizations for every word in a lexicon, plus tokenizations for the custom vocabulary, plus the predicted tokenizations for the alternate spellings of the custom vocabulary.

6. The method of claim 5 , wherein the creating the new WFST includes running CTC (connectionist temporal classification) decoding.

7. A system for adding a custom vocabulary to a transcription system, the system comprising:

a ASIRW module executing code and configured to:

receive a custom vocabulary at the ASIRW module, the custom vocabulary is a list of words which are likely to appear in audio, the custom vocabulary list provided by a user to the ASIRW module, the user providing the audio to the transcription system;

tokenize the custom vocabulary with the ASIRW module;

create a new WFST (weighted finite-state transducer) with the ASIRW module;

transcribe the audio using the new WFST with the ASIRW module.

8. The system of claim 7 , wherein the tokenizing includes performing a translation model on each word of the custom vocabulary.

9. The system of claim 7 , wherein the custom vocabulary includes phrases.

10. The system of claim 8 , wherein the tokenizing includes creating predicted tokenizations for alternate spellings of the custom vocabulary.

11. The system of claim 10 , wherein the new WFST includes tokenizations for every word in a lexicon, plus tokenizations for the custom vocabulary, plus the predicted tokenizations for the alternate spellings of the custom vocabulary.

12. The system of claim 11 , wherein the creating the new WFST includes running CTC (connectionist temporal classification) decoding.

13. A system for adding a custom vocabulary to a transcription system, the system comprising:

a ASIRW module executing code and configured to:

receive a custom vocabulary at the ASIRW module, the custom vocabulary is a list of words which are likely to appear in audio, the custom vocabulary list provided by a user to the ASIRW module, the user providing the audio to the transcription system;

tokenize the custom vocabulary with the ASIRW module including performing a translation model on each word of the custom vocabulary;

create a new WFST (weighted finite-state transducer) with the ASIRW module, wherein the new WFST includes tokenizations for every word in a lexicon, plus tokenizations for the custom vocabulary, plus the predicted tokenizations for the alternate spellings of the custom vocabulary;

transcribe the audio using the new WFST with the ASIRW module.

14. A non-transitory digital storage medium having a computer program stored thereon to perform a method of adding a custom vocabulary to a transcription system, the method comprising:

receiving a custom vocabulary at an ASIRW module, the custom vocabulary is a list of words which are likely to appear in audio, the custom vocabulary list provided by a user to the ASIRW module, the user providing the audio to the transcription system;

tokenizing the custom vocabulary with the ASIRW module;

creating a new WFST (weighted finite-state transducer) with the ASIRW module;

transcribing the audio using the new WFST with the ASIRW module.

15. A method of alternative spelling training, the method comprising:

providing a ground-truth text, the ground-truth text resulting from a live user verifying and/or creating a transcription for an audio signal including words;

providing ASR-generated text;

aligning the ground-truth text and the ASR-generated text with a text alignment tool;

generating error pairs from the aligning;

filtering the error pairs according to uncommon words to create filtered error pairs;

adding the filtered error pairs to the translation model training along with the previously existing training configuration files, to create a trained translation model;

using the trained translation model to generate a transcript from the audio signal.

16. The method of claim 15 , wherein the ground-truth text results resulting from a live user creating a transcription for the audio signal including words.

17. The method of claim 15 , wherein the ASR-generated text results from automatic speech recognition.

18. A system of alternative spelling training, the system comprising:

an alternative spelling training module executing code and configured to:

receive a ground-truth text, the ground-truth text resulting from a live user verifying and/or creating a transcription for an audio signal including words;

receive ASR-generated text;

align the ground-truth text and the ASR-generated text with a text alignment tool;

generate error pairs from the aligning;

filter the error pairs according to uncommon words to create filtered error pairs;

add the filtered error pairs to the translation model training along with the previously existing training configuration files, to create a trained translation model;

generate a transcript from the audio signal using the trained translation model.

Assignments (4)
SECURITY INTEREST Recorded Aug 8, 2024
From: REV.COM, INC.
To: FIRST-CITIZENS BANK & TRUST COMPANY
Reel/Frame 068224/0202 →
SECURITY INTEREST Recorded Jan 5, 2024
From: REV.COM, INC.
To: SILICON VALLEY BANK
Reel/Frame 066026/0887 →
SECURITY INTEREST Recorded Jan 5, 2024
From: REV.COM, INC.
To: SILICON VALLEY BANK, AS AGENT
Reel/Frame 066026/0949 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2022
From: FOX, JENNIFER DREXLER; CHEN, DANNY; DELWORTH, NATALIE
To: REV.COM, INC.
Reel/Frame 059414/0349 →
Continuity (1)
Related Publication 20230326450A1 · Oct 12, 2023
References Cited (8)
US 20160203819A1 · Ma · 2016 [cited by examiner]
US 20170148431A1 · Catanzaro · 2017 [cited by examiner]
US 20170323638A1 · Malinowski · 2017 [cited by examiner]
US 20180025723A1 · Nagao · 2018 [cited by examiner]
US 20180233134A1 · Yoon · 2018 [cited by examiner]
US 20190139540A1 · Kanda · 2019 [cited by examiner]
CN 108417202B · 2020 [cited by examiner]
WO WO2019219968A1 · 2019 [cited by examiner]