IP Library Granted Patent US 12,367,346
Granted Patent B2
US 12,367,346 · App. 16/033,259 · Granted Jul 22, 2025

Natural language processing with k-NN

Inventor: Avidan Akerib (Tel Aviv, IL)
Assignee: GSI Technology Inc.
G06F40/30G06N3/04G06N3/042G06N3/044G06N3/045G06N3/08G06N5/041G06N3/048G06N20/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,346
App. No.
16/033,259
Granted
Jul 22, 2025
Kind
B2
Abstract

A system for natural language processing includes a memory array and a processor. The memory array is divided into a similarity section storing a plurality of feature vectors, a SoftMax section in which to determine probabilities of occurrence of the feature vectors, a value section storing a plurality of modified feature vectors, and a marker section. The processor activates the array to perform parallel operations in each column indicated by the marker section: a similarity operation in the similarity section between a vector question and feature vectors stored in indicated columns; a SoftMax operation in the SoftMax section to determine an associated SoftMax probability value for indicated feature vectors; a multiplication operation in the value section to multiply the associated SoftMax value by modified feature vectors stored in indicated columns; and a vector sum in the value section to accumulate an attention vector of output of the multiplication operation.

Claims (29)

1. A system for natural language processing, the system comprising:

a memory array having rows and columns, said memory array being divided into a similarity section initially storing a plurality of feature or key vectors in columns thereof, wherein each of said vectors has a fixed size, a SoftMax section in which to determine probabilities of occurrence of said feature or key vectors, a value section initially storing a plurality of modified feature vectors in columns thereof, and a marker section storing a marker vector in a row thereof specifying columns to be operated upon, wherein said sections being contiguous to one another such that a column of one section of said sections is contiguous with a column of a neighboring section and operations in one or more columns of said memory array are associated with one feature vector to be processed, and wherein said memory array comprises a bit line processor per column of each said section, each said bit line processor operating on one bit of data of its associated section; and

an in-memory processor operating in a constant time as a function of said fixed size and irrespective of the number of said vectors, said processor to activate said memory array to perform the following operations in parallel in each column indicated by said marker vector:

a similarity operation in said similarity section between a vector question and each said feature vector stored in each said indicated column to generate a similarity output in each said indicated column;

a SoftMax operation in said SoftMax section on each said similarity output in said similarity section to determine an associated SoftMax value for each said indicated feature vector, wherein an intermediate output of exponential operations of said SoftMax operation is stored in said bit-line processor of said SoftMax section of each said indicated column; and

a multiplication operation in said value section to multiply each said associated SoftMax value in said SoftMax section by each said modified feature vector stored in each said indicated column to generate a multiplication output in each said indicated column;

said in-memory processor to also perform a horizontal vector sum in said value section of said multiplication output in each said indicated column to accumulate an attention vector sum, said vector sum to be used to generate a new vector question for a further iteration or to generate an output value in a final iteration.

2. The system according to claim 1 wherein said memory array comprises operational portions, one portion per iteration of a natural language processing operation, each portion being divided into said similarity, SoftMax, and value sections.

3. The system according to claim 1 wherein said memory array is one of: an SRAM, a non-volatile, a volatile, and a non-destructive array.

4. The system according to claim 1 and also comprising a neural network feature extractor to generate said feature and modified feature vectors.

5. The system according to claim 1 and wherein said feature vectors comprise features of a word, a sentence, or a document.

6. The system according to claim 1 wherein said feature vectors are the output of a pre-trained neural network.

7. The system according to claim 1 and also comprising a pre-trained neural network to generate an initial vector question.

8. The system according to claim 7 and also comprising a question generator to generate a further question from said initial vector question and said attention vector sum.

9. The system according to claim 8 wherein said question generator is a neural network.

10. The system according to claim 8 and wherein said question generator is implemented as a matrix multiplier on bit lines of said memory array.

11. A method for natural language processing, the method comprising:

having a memory array having rows and columns, said memory array being divided into a similarity section initially storing a plurality of feature or key vectors in columns thereof, wherein each of said vectors has a fixed size, a SoftMax section in which to determine probabilities of occurrence of said feature or key vectors, a value section initially storing a plurality of modified feature vectors in columns thereof, and a marker section storing a marker vector in a row thereof specifying columns to be operated upon, said sections being contiguous to one another such that a column of one section of said sections is contiguous with a column of a neighboring section and operations in one or more columns of said memory array are associated with one feature vector to be processed, and wherein said memory array comprises a bit line processor per column of each said section, each said bit line processor operating on one bit of data of its associated section; and

activating said memory array to operate in a constant time as a function of said fixed size and irrespective of the number of said vectors to perform the following operations in parallel in each column indicated by said marker vector:

performing a similarity operation in said similarity section between a vector question and each said feature vector stored in each said indicated column to generate a similarity output in each said indicated column;

performing a SoftMax operation in said SoftMax section on each said similarity output in said similarity section to determine an associated SoftMax value for each said indicated feature vector, wherein an intermediate output of exponential operations of said SoftMax operation is stored in said bit-line processor of said SoftMax section of each said indicated column; and

performing a multiplication operation in said value section to multiply each said associated SoftMax value in said SoftMax section by each said modified feature vector stored in each said indicated column to generate a multiplication output in each said indicated column;

performing a horizontal vector sum operation in said value section of said multiplication output in each said indicated column to accumulate an attention vector sum, said vector sum to be used to generate a new vector question for a further iteration or to generate an output value in a final iteration.

12. The method according to claim 11 and also comprising generating said feature and modified feature vectors with a neural network and storing them into said similarity and value sections, respectively.

13. The method according to claim 11 and wherein said feature vectors comprise features of a word, a sentence, or a document.

14. The method according to claim 11 and also comprising generating an initial vector question using a pre-trained neural network.

15. The method according to claim 14 and also comprising generating a further question from said initial vector question and said attention vector sum.

16. The method according to claim 15 wherein generating a further question utilizes a neural network.

17. The method according to claim 15 and wherein said generating a further question comprises performing matrix multiplication on bit lines of said memory array.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 3, 2019
From: AKERIB, AVIDAN
To: GSI TECHNOLOGY INC.
Reel/Frame 048224/0679 →
Continuity (3)
Provisional Application 62686114 · Jun 18, 2018
Provisional Application 62533076 · Jul 16, 2017
Related Publication 20180341642A1 · Nov 29, 2018
References Cited (45)
US 5014327A · Potter · 1991 [cited by applicant]
US 5799300A · Agrawal · 1998 [cited by applicant]
US 5983224A · Singh · 1999 [cited by applicant]
US 8099380B1 · Shahabi · 2012 [cited by applicant]
US 8238173B2 · Akerib · 2012 [cited by applicant]
US 9418719B2 · Akerib · 2016 [cited by applicant]
US 9558812B2 · Akerib · 2017 [cited by applicant]
US 9653166B2 · Akerib · 2017 [cited by applicant]
US 9859005B2 · Akerib · 2018 [cited by applicant]
US 10153042B2 · Ehrman · 2018 [cited by applicant]
US 10210935B2 · Akerib · 2019 [cited by applicant]
US 10249362B2 · Shu · 2019 [cited by applicant]
US 10402165B2 · Lazer · 2019 [cited by applicant]
US 10489480B2 · Akerib · 2019 [cited by applicant]
US 10514914B2 · Lazer · 2019 [cited by applicant]
US 10521229B2 · Shu · 2019 [cited by applicant]
US 10534836B2 · Shu · 2020 [cited by applicant]
US 10635397B2 · Lazer · 2020 [cited by applicant]
US 10725777B2 · Shu · 2020 [cited by applicant]
US 10777262B1 · Haig · 2020 [cited by applicant]
US 20130080490A1 · Plondke · 2013 [cited by applicant]
US 20150131383A1 · Akerib · 2015 [cited by examiner]
US 20150146491A1 · Akerib · 2015 [cited by examiner]
US 20150200009A1 · Akerib · 2015 [cited by examiner]
US 20150332126A1 · Hikida · 2015 [cited by applicant]
US 20160086222A1 · Kurapati · 2016 [cited by applicant]
US 20160275876A1 · Hagood · 2016 [cited by applicant]
US 20170103324A1 · Weston · 2017 [cited by examiner]
JP 2008276344 · 2008 [cited by applicant]
JP 201633806A · 2016 [cited by applicant]
KR 2014008270 · 2014 [cited by applicant]
KR 101612605B1 · 2016 [cited by applicant]
Miller et al. “Key-value memory networks for directly reading documents.” arXiv preprint arXiv: 1606.03126 (Year: 2016). [cited by examiner]
Sukhbaatar et al., “End-to-end memory networks.” Advances in neural information processing systems 28 (Year: 2015). [cited by examiner]
Li, Shuangchen, et al. “Pinatubo: A processing-in-memory architecture for bulk bitwise operations in emerging non-volatile memories.” Proceedings of the 53rd Annual Design Automation Conference. (Year: 2016). [cited by examiner]
Piombo et al., “Analog soft max circuit with dynamic gain control.” Research in Microelectronics and Electronics, 2005 PhD. vol. 1. IEEE (Year: 2005). [cited by examiner]
International Search Report for corresponding PCT application PCT/IB2017/54233 mailed on Jan. 12, 2018. [cited by applicant]
Machine Translation of Korean Publication 10-1612605 downloaded from the Korean Patent Office website on Oct. 15, 2020. [cited by applicant]
English Abstract of JP2008276344 downloaded from Google Patents on May 26, 2019. [cited by applicant]
Chatzimilioudis, Distributed In-Memory Processing of All k Nearest Neighbor Queries, IEEE Transaction on Knowledge and Data Engineering, Apr. 2016. [cited by applicant]
Rinkus, “A Cortex-inspired Associative Memory with O(1) Time Complexity Learning, Recall and Recognition of Sequences”, Brandeis University, 2006. [cited by applicant]
Canal, “Memory Structure”, Department d'Arquitectura de Computadors, Universitat Politecnica de Catalunya, Jul. 6, 2016. [cited by applicant]
ICS33 Lecture Note, “The Complexity of Python Operators/Functions”—ICS UCI, Donald Bren School of Information & Computer Sciences, 2014. [cited by applicant]
BIONB330 Reading Material, “Associative Memory”, Cornell University, 2013. [cited by applicant]
Zhang, “k-nearest neighbors associative memory model for face recognition”, AI 2005, LNAI 3809, pp. 540-549, 2005. [cited by applicant]