IP Library Granted Patent US 9,483,265
Granted Patent B2
US 9,483,265 · App. 13/956,696 · Granted Nov 1, 2016

Vectorized lookup of floating point values

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,483,265
App. No.
13/956,696
Granted
Nov 1, 2016
Kind
B2
Abstract

Systems and techniques disclosed herein include methods for de-quantization of feature vectors used in automatic speech recognition. A SIMD vector processor is used in one embodiment for efficient vectorized lookup of floating point values in conjunction with fMPE processing for increasing the discriminative power of input signals. These techniques exploit parallelism to effectively reduce the latency of speech recognition in a system operating in a high dimensional feature space. In one embodiment, a bytewise integer lookup operation effectively performs a floating point or a multiple byte lookup.

Claims (49)

1. A computer-implemented method for performing automatic speech recognition, the method comprising, using at least one processor programmed to perform:

loading a first set of data elements into a lookup table having a base address;

loading a vector register with a translation constant corresponding to an arrangement of the first set of data elements;

loading a second set of data elements into an index register;

invoking a vector addition instruction to combine the translation constant and the index register;

invoking a vector permute instruction on the first set of data elements at the base address and the combined translation constant and index register to decode the second set of data elements into a set of segments in a destination memory;

reconstructing a plurality of quantized values into a plurality of dequantized values in parallel;

producing, based on the plurality of dequantized values, a plurality of enhanced feature vectors; and

outputting the plurality of enhanced feature vectors to a speech recognizer that performs automatic speech recognition based on the plurality of enhanced feature vectors.

2. The computer-implemented method of claim 1 , wherein the lookup table comprises a plurality of vector registers.

3. The computer-implemented method of claim 1 , wherein the destination memory comprises a plurality of vector registers.

4. The computer-implemented method of claim 3 , wherein the set of segments comprise a representation of a floating point constant; and

further comprising invoking at least one vector instruction operating on the floating point constant.

5. The computer-implemented method of claim 4 , wherein the at least one vector instruction is a multiplication of a feature vector and row of an fMPE matrix.

6. The computer-implemented method of claim 1 , wherein the set of segments are decoded from non-consecutive data elements in the lookup table.

7. The computer-implemented method of claim 1 , wherein the permute instruction comprises a single SIMD instruction.

8. The computer-implemented method of claim 7 , wherein the single SIMD instruction vector permute instruction comprises a TBL instruction in a NEON instruction set.

9. The computer-implemented method of claim 1 , wherein the second set of data elements represent quantized values of a speech feature vector.

10. The computer-implemented method of claim 1 , wherein loading a second set of data elements into an index register further comprises replicating each element of the second set of data elements to populate the index register.

11. A system for performing automatic speech recognition, the system comprising:

one or more processors;

a memory including one or more instructions stored thereon that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

loading a first set of data elements into a lookup table having a base address;

loading a first vector register with a translation constant corresponding to an arrangement of the first set of data elements;

loading a second set of data elements into an index register;

invoking a vector addition instruction to combine the translation constant and the index register;

invoking a vector permute instruction on the first set of data elements at the base address and the combined translation constant and index register to decode the second set of data elements into a set of segments in a destination memory;

reconstructing a plurality of quantized values into a plurality of dequantized values in parallel;

producing, based on the plurality of dequantized values, a plurality of enhanced feature vectors; and

outputting the plurality of enhanced feature vectors to a speech recognizer that performs automatic speech recognition based on the plurality of enhanced feature vectors.

12. The system of claim 11 , wherein the lookup table comprises a plurality of vector registers; and

wherein the destination memory comprises a plurality of vector registers.

13. The system of claim 12 , wherein the set of segments comprise a representation of a floating point constant; and

further comprising invoking at least one vector instruction operating on the floating point constant.

14. The system of claim 13 , wherein the at least one vector instruction is a multiplication of a feature vector and a row of an fMPE matrix.

15. The system of claim 11 , wherein the set of segments are decoded from non-consecutive data elements in the lookup table.

16. The system of claim 15 , wherein the permute instruction comprises a single SIMD instruction.

17. The system of claim 11 , wherein the single SIMD instruction vector permute instruction comprises a TBL instruction in a NEON instruction set.

18. The system of claim 11 , wherein the second set of data elements represents quantized values of a speech feature vector.

19. The system of claim 11 , wherein loading a second set of data elements into an index register further comprises replicating each element of the second set of data elements to populate the index register.

20. A computer program product including a non-transitory computer-storage medium having instructions stored thereon for processing data information for performing automatic speech recognition, such that the instructions, when carried out by a processing device, cause the processing device to perform:

loading a first set of data elements into a lookup table having a base address;

loading a first vector register with a translation constant corresponding to an arrangement of the first set of data elements;

loading a second set of data elements into an index register;

invoking a vector addition instruction to combine the translation constant and the index register;

invoking a vector permute instruction on the first set of data elements at the base address and the combined translation constant and index register to decode the second set of data elements into a set of segments in a destination memory;

reconstructing a plurality of quantized values into a plurality of dequantized values in parallel;

producing, based on the plurality of dequantized values, a plurality of enhanced feature vectors; and

outputting the plurality of enhanced feature vectors to a speech recognizer that performs automatic speech recognition based on the plurality of enhanced feature vectors.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 1, 2013
From: WICK, JUSTIN VAUGHN
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 030924/0171 →