IP Library Granted Patent US 11,720,789
Granted Patent B2
US 11,720,789 · App. 16/672,352 · Granted Aug 8, 2023

Fast nearest neighbor search for output generation of convolutional neural networks

Inventors: Hessam Bagherinezhad (Seattle, WA); Dmitry Belenko (Redmond, WA)
Assignee: Apple Inc.
G06N3/08G06F16/90335G06F17/16G06F40/30G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,720,789
App. No.
16/672,352
Granted
Aug 8, 2023
Kind
B2
Abstract

In one embodiment, a method includes receiving an input vector corresponding to a query at a neural network model comprising a plurality of layers, wherein the plurality of layers comprise a last layer associated with a mapping matrix, generating a binary matrix based on the mapping matrix, an identity matrix, and one or more Gaussian vectors, generating an integer vector based on the binary matrix and a binary vector associated with the input vector, identifying a plurality of indices corresponding to a plurality of top values of the integer vector for the integer vector, generating an output vector based on the input vector and a plurality of rows of the mapping matrix, wherein the plurality of rows is associated with the plurality of identified indices, respectively, and determining the query is associated with one or more classes based on the output vector.

Claims (54)

1. A method comprising, by one or more computing systems:

receiving, at a neural network model comprising a plurality of layers, an input vector corresponding to a data item, wherein the plurality of layers comprise a last-layer associated with a mapping matrix;

generating a binary matrix based on the mapping matrix;

generating an integer vector based on the binary matrix and a binary vector associated with the input vector;

identifying, for the integer vector, a plurality of indices corresponding to a plurality of top values of the integer vector;

generating an output vector based on the input vector and a plurality of rows of the mapping matrix, wherein the plurality of rows is associated with the plurality of identified indices, respectively; and

determining, based on the output vector, that the data item is associated with one or more classes.

2. The method of claim 1 , wherein the binary matrix is further generated based on an identity matrix and one or more Gaussian vectors, and the method further comprising:

generating an expansion matrix by concatenating the identity matrix and the one or more Gaussian vectors.

3. The method of claim 2 , wherein the expansion matrix is associated with a hyper parameter, the hyper parameter being determined based on a trade-off between speed and accuracy of the neural network model.

4. The method of claim 2 , wherein generating the binary matrix comprises:

multiplying the mapping matrix with the expansion matrix; and

applying a sign operation on the multiplication result.

5. The method of claim 2 , wherein the binary vector is generated by:

multiplying the input vector with the expansion matrix; and

applying a sign operation on the multiplication result.

6. The method of claim 1 , wherein generating the integer vector comprises:

multiplying the binary matrix with the binary vector.

7. The method of claim 1 , wherein a number of the plurality of top values of the binary vector is determined based on a trade-off between speed and accuracy of the neural network model.

8. The method of claim 1 , wherein generating the output vector comprises:

generating a compression matrix based on the plurality of rows of the mapping matrix; and

multiplying the input vector with the compression matrix.

9. The method of claim 1 , wherein determining the data item is associated with the one or more classes comprises:

identifying one or more indices corresponding to one or more top values of the output vector; and

identifying the one or more classes corresponding to the identified indices, respectively.

10. The method of claim 1 , wherein the binary matrix is further generated based on one or more Gaussian vectors that are randomly generated.

11. The method of claim 1 , wherein the data item comprises one or more of a text, an audio clip, an image, or a video.

12. One or more computer-readable non-transitory storage media embodying software that is operable when executed to:

receive, at a neural network model comprising a plurality of layers, an input vector corresponding to a data item, wherein the plurality of layers comprise a layer associated with a mapping matrix;

generate a binary matrix based on the mapping matrix, an identity matrix, and one or more Gaussian vectors;

generate an output vector based on the input vector and a plurality of rows of the mapping matrix; and

determine, based on the output vector, that the data item is associated with one or more classes.

13. The media of claim 12 , wherein the software is further operable when executed to:

generate an expansion matrix by concatenating the identity matrix and the one or more Gaussian vectors.

14. The media of claim 13 , wherein the expansion matrix is associated with a hyper parameter, the hyper parameter being determined based on a trade-off between speed and accuracy of the neural network model.

15. The media of claim 13 , wherein generating the binary matrix comprises:

multiplying the mapping matrix with the expansion matrix; and

applying a sign operation on the multiplication result.

16. The media of claim 13 , wherein a binary vector associated with the input vector is generated by:

multiplying the input vector with the expansion matrix; and

applying a sign operation on the multiplication result.

17. A system comprising:

one or more processors; and

a non-transitory memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to:

receive, at a neural network model comprising a plurality of layers, an input vector corresponding to a data item, wherein the plurality of layers comprise a layer associated with a mapping matrix;

generate a binary matrix based on the mapping matrix;

generate an output vector based on the input vector and a plurality of rows of the mapping matrix; and

determine, based on the output vector, that the data item is associated with one or more classes.

18. The system of claim 17 , wherein the binary matrix is further generated based on an identity matrix and one or more Gaussian vectors and the processors are further operable when executing the instructions to:

generate an expansion matrix by concatenating the identity matrix and the one or more Gaussian vectors.

19. The system of claim 18 , wherein the expansion matrix is associated with a hyper parameter, the hyper parameter being determined based on a trade-off between speed and accuracy of the neural network model.

20. The system of claim 18 , wherein generating the binary matrix comprises:

multiplying the mapping matrix with the expansion matrix; and

applying a sign operation on the multiplication result.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2021
From: XNOR.AI, INC.
To: APPLE INC.
Reel/Frame 058390/0589 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2019
From: BAGHERINEZHAD, HESSAM; BELENKO, DMITRY
To: XNOR.AI, INC.
Reel/Frame 050895/0118 →
Continuity (2)
Provisional Application 62858911 · Jun 7, 2019
Related Publication 20200387783A1 · Dec 10, 2020