IP Library › Granted Patent US 10,515,301
Granted Patent B2
US 10,515,301 · App. 15/000,852 · Granted Dec 24, 2019

Small-footprint deep neural network

Inventors: Jinyu Li (Redmond, WA); Yifan Gong (Sammamish, WA); Yongqiang Wang (Kirkland, WA)
Assignee: Microsoft Technology Licensing, LLC
G06N3/04G06N3/02G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,515,301
App. No.
15/000,852
Granted
Dec 24, 2019
Kind
B2
Abstract

Conversion of a large-footprint DNN to a small-print DNN is performed using a variety of techniques, including split-vector quantization. The small-foot print DNN may be distributed to a variety of devices, including mobile devices. Further, the small-footprint DNN may aid a digital assistant on a device in interpreting speech input.

Claims (52)

1. A system comprising:

at least one processor; and

a memory storage device, the memory storage device comprising instructions, that, when executed by the at least one processor, perform a method of generating a small-footprint matrix based on a large-footprint matrix of a neural network, the method comprising:

receiving a first matrix of a neural network, wherein the first matrix comprises a vector, and the vector including a sub-vector;

generating a code book based on the first matrix, wherein the code book comprises a codeword and an index, wherein the first codeword is a representation of the sub-vector, and wherein the index refers to the codeword;

identifying a first codeword and a first index from the code book, wherein the first codeword approximates the sub-vector and the first index refers to the first codeword in the code book; and

generating a second matrix of the neural network based on the first matrix, wherein the second matrix comprises the first index in place of the sub-vector, the second matrix being smaller in footprint than the first matrix.

2. The system of claim 1 , further comprising:

identifying the first codeword and the first index for the first codeword from the code book, wherein the first matrix includes a second vector, the second vector including a second sub-vector, and the first codeword approximating a second sub-vector in the first matrix; and

generating the second matrix of the neural network based on the first matrix, wherein the second matrix comprises the first index in place of the second sub-vector.

3. The system of claim 2 , wherein the second matrix refers to the first codeword in the code book based on the first index.

4. The system of claim 1 , further comprising,

fine-tuning the codeword, wherein the codeword is a numerical representation of the sub-vector.

5. The system of claim 1 , wherein the first matrix of the neural network is shared remotely over a network by a mobile device.

6. The system of claim 1 , wherein the method further comprises:

sending the second matrix to a mobile device.

7. The system of claim 1 further comprising:

receiving an audio input; and

determining a phoneme that approximates the audio input based on the second matrix.

8. A computer implemented method of generating a small-footprint matrix based on a large-footprint matrix of a neural network, the method comprising:

receiving a first matrix of a neural network, wherein the first matrix comprises a vector and the vector including a sub-vector;

generating a code book, wherein the code book comprises a codeword and an index, wherein the codeword is a representation of the sub-vector, and wherein the index refers to the codeword;

identifying a first codeword and a first index from the code book, wherein the first codeword approximates the sub-vector and the first index refers to the first codeword; and

generating a second matrix of the neural network based on the first matrix, wherein the second matrix comprises the first index in place of the sub-vector, the second matrix being smaller in footprint than the first matrix.

9. The method of claim 8 , further comprising:

identifying the first codeword and the first index for the first codeword from the code book, wherein the first matrix includes a second vector, the second vector including a second sub-vector, and the first codeword approximating a second sub-vector in the first matrix; and

generating the second matrix of the neural network based on the first matrix, wherein the second matrix comprises the first index in place of the second sub-vector.

10. The method of claim 9 , further comprising

fine-tuning the first codeword, wherein the codeword is a numerical representation of the sub-vector.

11. The method of claim 9 , wherein the second matrix refers to the first codeword in the code book based on the first index.

12. The method of claim 8 , wherein the method further comprises:

sending the second matrix to a mobile device.

13. The method of claim 8 further comprising:

receiving an audio input; and

determining a phoneme that approximates the audio input based on the second matrix.

14. A computer readable storage device storing instructions that perform a method of generating a small-footprint matrix based on a large-footprint matrix of a neural network, when executed, the method comprising:

receiving a first matrix of a neural network, wherein the first matrix comprises a vector and the vector including a sub-vector;

generating a code book, wherein the code book comprises codeword and an index, wherein the first codeword is a representation of the sub-vector, and wherein the index refers to the codeword;

identifying a first codeword and a first index from the code book, wherein the first codeword approximates the sub-vector and the first index refers to the first codeword in the code book; and

generating a second matrix of the neural network based on the first matrix, wherein the second matrix comprises the first index in place of the sub-vector, the second is matrix being smaller in footprint than the first matrix.

15. The computer readable storage device of claim 14 , the method further comprising:

identifying the first codeword and the first index for the first codeword from the code book, wherein the first matrix includes a second vector, the second vector including a second sub-vector, and the first codeword approximating the second sub-vector in the first matrix; and

generating the second matrix of the neural network based on the first matrix, wherein the second matrix comprises the first index in place of the second sub-vector in the second matrix.

16. The computer readable storage device of claim 15 , the method further comprising,

fine-tuning the codeword, wherein the codeword is a numerical representation of the sub-vector.

17. The computer readable storage device of claim 16 , wherein fine tuning the codeword comprises using an adaptive learning rate.

18. The computer readable storage device of claim 15 , wherein the second matrix refers to the first codeword in the code book based on the first index.

19. The computer readable storage device of claim 14 , wherein the method further comprises:

sending the second matrix to a mobile device.

20. The computer readable storage device of claim 14 further comprising:

receiving an audio input; and

determining a phoneme that approximates the audio input based on the second matrix.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2016
From: GONG, YIFAN; LI, JINYU; WANG, YONGQIANG
To: MICROSOFT TECHNOLOGY LICENSING,
Reel/Frame 037525/0371 →
Continuity (2)
Provisional Application 62149395 · Apr 17, 2015
Related Publication 20160307095A1 · Oct 20, 2016
Cited By (1)
US 12,346,803