IP Library › Granted Patent US 11,210,565
Granted Patent B2
US 11,210,565 · App. 16/206,714 · Granted Dec 28, 2021

Machine learning model with depth processing units

Inventors: Jinyu Li (Redmond, WA); Liang Lu (Redmond, WA); Changliang Liu (Bothell, WA); Yifan Gong (Sammamish, WA)
Assignee: Microsoft Technology Licensing, LLC
G06K9/6262G06K9/46G06N3/08G06N20/00G10L15/063G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,210,565
App. No.
16/206,714
Filed
Nov 30, 2018
Granted
Dec 28, 2021
Kind
B2
Art Unit
2661
USPC
382/157
Abstract

Representative embodiments disclose machine learning classifiers used in scenarios such as speech recognition, image captioning, machine translation, or other sequence-to-sequence embodiments. The machine learning classifiers have a plurality of time layers, each layer having a time processing block and a depth processing block. The time processing block is a recurrent neural network such as a Long Short Term Memory (LSTM) network. The depth processing blocks can be an LSTM network, a gated Deep Neural Network (DNN) or a maxout DNN. The depth processing blocks account for the hidden states of each time layer and uses summarized layer information for final input signal feature classification. An attention layer can also be used between the top depth processing block and the output layer.

Claims (47)

1. A system comprising:

a trained machine learning classifier that is configured to receive an input signal at a first time step and output a classified posterior at the first time step, the trained machine learning classifier comprising a plurality of processing blocks arranged in a plurality of layers, wherein a processing block in the plurality of processing blocks at the first time step and a first layer comprises:

a first time processing block connected so that:

an output of the first time processing block is input into a second time processing block at a second time step that is next in time to the first time step;

an output of a third time processing block at a third time step is received as input to the first time processing block, the third time step being immediately prior in time to the first time step;

an output of a fourth time processing block at the first time step is received as input to the first time processing block, wherein the first time processing block comprising a recurrent neural network; and

a first layer processing block connected so that:

an output of a second layer processing block in a second layer that is adjacent the first layer is input to the first layer processing block; and

an output of the first time processing block is received as input to the first layer processing block;

the input signal comprising at least one of speech data, audio data, handwriting data, or textual data.

2. The system of claim 1 , wherein the recurrent neural network is an LSTM network.

3. The system of claim 1 , wherein the first layer processing block is an LSTM network.

4. The system of claim 1 , wherein the first layer processing block is a gated DNN.

5. The system of claim 1 , wherein the first layer processing block is a maxout DNN.

6. The system of claim 3 , the trained machine learning classifier further comprising an attention layer between a top layer processing block and an output layer of the trained machine learning classifier.

7. The system of claim 4 , the trained machine learning classifier further comprising an attention layer between a top layer processing block and an output layer of the trained machine learning classifier.

8. The system of claim 5 , the trained machine learning classifier further comprising an attention layer between a top layer processing block and an output layer of the trained machine learning classifier.

9. The system of claim 1 , the trained machine learning classifier further comprising an output layer that is a softmax layer.

10. A method comprising:

providing an input signal to a trained machine learning classifier at a first time step; and

outputting a classified posterior at the first time step, the trained machine learning classifier comprising a plurality of processing blocks arranged in a plurality of layers, wherein a processing block in the plurality of processing blocks at the first time step and a first layer comprises:

a first time processing block connected so that:

an output of the first time processing block is input into a second time processing block at a second time step that is next in time to the first time step;

an output of a third time processing block at a third time step is received as input to the first time processing block, the third time step being immediately prior in time to the first time step;

an output of a fourth time processing block at the first time step is received as input to the first time processing block, wherein the first time processing block comprising a recurrent neural network; and

a first layer processing block connected so that:

an output of a second layer processing block in a second layer that is adjacent the first layer is input to the first layer processing block; and

an output of the first time processing block is received as input to the first layer processing block, wherein the input signal comprises at least one of speech data, audio data, handwriting data, or textual data.

11. The method of claim 10 , wherein the recurrent neural network is an LSTM network.

12. The method of claim 10 , wherein the first layer processing block is an LSTM network.

13. The method of claim 12 , the trained machine learning classifier further comprising an attention layer between a top layer processing block and an output layer of the trained machine learning classifier.

14. The method of claim 10 , wherein the first layer processing block is a gated DNN.

15. The method of claim 14 , the trained machine learning classifier further comprising an attention layer between a top layer processing block and an output layer of the trained machine learning classifier.

16. The method of claim 10 , wherein the first layer processing block is a maxout DNN.

17. The method of claim 16 , the trained machine learning classifier further comprising an attention layer between a top layer processing block and an output layer of the trained machine learning classifier.

18. The method of claim 10 , the trained machine learning classifier further comprising an output layer that is a softmax layer.

19. A non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause the processor to perform acts comprising:

providing an input signal to a trained machine learning classifier at a first time step; and

outputting a classified posterior at the first time step, the trained machine learning classifier comprising a plurality of processing blocks arranged in a plurality of layers, wherein a processing block in the plurality of processing blocks at the first time step and a first layer comprises:

a first time processing block connected so that:

an output of the first time processing block is input into a second time processing block at a second time step that is next in time to the first time step;

an output of a third time processing block at a third time step is received as input to the first time processing block, the third time step being immediately prior in time to the first time step;

an output of a fourth time processing block at the first time step is received as input to the first time processing block, wherein the first time processing block comprising a recurrent neural network; and

a first layer processing block connected so that:

an output of a second layer processing block in a second layer that is adjacent the first layer is input to the first layer processing block; and

an output of the first time processing block is received as input to the first layer processing block, wherein the input signal comprises at least one of speech data, audio data, handwriting data, or textual data.

20. The non-transitory computer-readable medium of claim 19 , wherein the recurrent neural network is an LSTM network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 30, 2018
From: GONG, YIFAN; LI, JINYU; LIU, CHANGLIANG; LU, LIANG
To: MICROSOFT TECHNOLOGY LICENSING LLC
Reel/Frame 047645/0605 →
Continuity (1)
Related Publication 20200175335A1 · Jun 4, 2020
Cited By (1)
US 12,327,173