IP Library Granted Patent US 10,210,860
Granted Patent B1
US 10,210,860 · App. 16/108,107 · Granted Feb 19, 2019

Augmented generalized deep learning with special vocabulary

Inventors: Jeff Ward (San Francisco, CA); Adam Sypniewski (Ypsilanti, MI); Scott Stephenson (San Francisco, CA)
Assignee: Deepgram, Inc.
G10L15/063G06K9/6256G06N3/0481G06N3/084G10L15/16G10L15/197G10L2015/0635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,210,860
App. No.
16/108,107
Granted
Feb 19, 2019
Kind
B1
Abstract

Systems and methods are disclosed for customizing a neural network for a custom dataset, when the neural network has been trained on data from a general dataset. The neural network may comprise an output layer including one or more nodes corresponding to candidate outputs. The values of the nodes in the output layer may correspond to a probability that the candidate output is the correct output for an input. The values of the nodes in the output layer may be adjusted for higher performance when the neural network is used to process data from a custom dataset.

Claims (34)

1. A method for customizing a neural network trained on a general dataset to a custom dataset, the method comprising:

providing a trained speech recognition neural network, the trained speech recognition neural network including a plurality of layers each having a plurality of nodes, the trained speech recognition neural network including an output layer with nodes corresponding to words of a vocabulary, the nodes of the output layer outputting values, wherein the values output by the nodes in the output layer correspond to a probability of the corresponding word in the vocabulary being a correct transcription of an input;

for a plurality of words in the vocabulary, determining the frequency of occurrence of the word in a general training set and the frequency of occurrence of the word in a custom dataset;

during inference using the trained speech recognition neural network, for each word in the plurality of words, multiplying the value output by the output node for the word by the frequency of occurrence of the word in the custom dataset divided by the frequency of occurrence of the word in the general training set to obtain a custom model probability; and

generating a transcription of a spoken input based on the custom model probability.

2. The method of claim 1 , wherein the plurality of words comprises all of the words in the vocabulary.

3. The method of claim 1 , wherein the frequency of occurrence of the word in the general training set is set to a threshold minimum value if the word does not appear in the general training set.

4. The method of claim 1 , wherein one or more nodes of the trained speech recognition neural network include a sigmoid activation function.

5. The method of claim 1 , wherein the trained speech recognition neural network includes one or more locally connected neural network layers.

6. The method of claim 5 , wherein the trained speech recognition neural network includes one or more recurrent neural network layers.

7. The method of claim 6 , wherein the trained speech recognition neural network has been trained in an end-to-end training process including backpropagation through each of its layers.

8. A non-transitory computer-readable medium comprising instructions for:

providing a trained speech recognition neural network, the trained speech recognition neural network including a plurality of layers each having a plurality of nodes, the trained speech recognition neural network including an output layer with nodes corresponding to words of a vocabulary, the nodes of the output layer outputting values, wherein the values output by the nodes in the output layer correspond to a probability of the corresponding word in the vocabulary being a correct transcription of an input;

for a plurality of words in the vocabulary, determining the frequency of occurrence of the word in a general training set and the frequency of occurrence of the word in a custom dataset;

during inference using the trained speech recognition neural network, for each word in the plurality of words, multiplying the value output by the output node for the word by the frequency of occurrence of the word in the custom dataset divided by the frequency of occurrence of the word in the general training set to obtain a custom model probability; and

generating a transcription of a spoken input based on the custom model probability.

9. The non-transitory computer-readable medium of claim 8 , wherein the plurality of words comprises all of the words in the vocabulary.

10. The non-transitory computer-readable medium of claim 8 , wherein the frequency of occurrence of the word in the general training set is set to a threshold minimum value if the word does not appear in the general training set.

11. The non-transitory computer-readable medium of claim 8 , wherein one or more nodes of the trained speech recognition neural network include a sigmoid activation function.

12. The non-transitory computer-readable medium of claim 8 , wherein the trained speech recognition neural network includes one or more locally connected neural network layers.

13. The non-transitory computer-readable medium of claim 12 , wherein the trained speech recognition neural network includes one or more recurrent neural network layers.

14. The non-transitory computer-readable medium of claim 13 , wherein the trained speech recognition neural network has been trained in an end-to-end training process including backpropagation through each of its layers.

15. A system comprising:

a processor;

a non-transitory computer-readable medium comprising instructions for:

providing a trained speech recognition neural network, the trained speech recognition neural network including a plurality of layers each having a plurality of nodes, the trained speech recognition neural network including an output layer with nodes corresponding to words of a vocabulary, the nodes of the output layer outputting values, wherein the values output by the nodes in the output layer correspond to a probability of the corresponding word in the vocabulary being a correct transcription of an input;

for a plurality of words in the vocabulary, determining the frequency of occurrence of the word in a general training set and the frequency of occurrence of the word in a custom dataset;

during inference using the trained speech recognition neural network, for each word in the plurality of words, multiplying the value output by the output node for the word by the frequency of occurrence of the word in the custom dataset divided by the frequency of occurrence of the word in the general training set to obtain a custom model probability; and

generating a transcription of a spoken input based on the custom model probability.

16. The system of claim 15 , wherein the plurality of words comprises all of the words in the vocabulary.

17. The system of claim 15 , wherein the frequency of occurrence of the word in the general training set is set to a threshold minimum value if the word does not appear in the general training set.

18. The system of claim 15 , wherein one or more nodes of the trained speech recognition neural network include a sigmoid activation function.

19. The system of claim 15 , wherein the trained speech recognition neural network includes one or more locally connected neural network layers.

20. The system of claim 19 , wherein the trained speech recognition neural network has been trained in an end-to-end training process including backpropagation through each of its layers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2018
From: SYPNIEWSKI, ADAM; WARD, JEFF; STEPHENSON, SCOTT
To: DEEPGRAM, INC.
Reel/Frame 047560/0053 →
Continuity (1)
Provisional Application 62703892 · Jul 27, 2018
Cited By (44)
US 12,197,712 US 12,197,817 US 12,200,297 US 12,204,932 US 12,211,502 US 12,216,894 US 12,219,314 US 12,223,282 US 12,236,952 US 12,242,853 US 12,254,063 US 12,254,887 US 12,260,234 US 12,271,732 US 12,272,377 US 12,277,954 US 12,293,763 US 12,301,635 US 12,304,081 US 12,333,404 US 12,353,971 US 12,361,943 US 12,367,879 US 12,386,434 US 12,386,491 US 12,394,407 US 12,400,137 US 12,431,128 US 12,475,372 US 12,477,470 US 12,488,798 US 12,499,875 US 12,511,543 US 12,525,227 US 12,530,876 US 12,547,875 US 12,556,890 US 12,591,412 US 12,608,171 US 12,613,730 US 12,619,452 US 12,651,158 US 12,682,232 US 12,710,926