IP Library Granted Patent US 11,675,973
Granted Patent B2
US 11,675,973 · App. 17/102,679 · Granted Jun 13, 2023

Electronic device and operation method for embedding an input word using two memory operating speeds

Inventors: Sejung Kwon (Suwon-si, KR); Dongsoo Lee (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06F40/205G06F12/0813G06F40/237G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,675,973
App. No.
17/102,679
Granted
Jun 13, 2023
Kind
B2
Abstract

An electronic device is provided. The electronic device includes a first memory configured to operate at a first speed and store compressed vectors corresponding to words, and scaling factors corresponding to the compressed vectors; a second memory that is faster than the first memory and is configured to store a first group of the compressed vectors, and store a first group of the scaling factors; and a processor configured to obtain a first compressed vector and a first scaling factor corresponding to an input word from the first memory or the second memory and process the obtained first compressed vector and the obtained first scaling factor by using a neural network.

Claims (54)

1. An electronic device for processing a word by using a language model, the electronic device comprising:

a first memory configured to operate at a first speed and store a compressed embedding matrix, which includes a plurality of compressed vectors corresponding to a plurality of words, and scaling factors corresponding to the plurality of compressed vectors;

a second memory configured to operate at a second speed that is faster than the first speed, store a first group of the plurality of compressed vectors identified based on first frequency information of the plurality of compressed vectors, and store a first group of the scaling factors identified based on second frequency information of the scaling factors; and

a processor configured to obtain a first compressed vector and a first scaling factor corresponding to an input word from the first memory or the second memory and process the obtained first compressed vector and the obtained first scaling factor by using a neural network.

2. The electronic device of claim 1 , wherein the first memory is separate from the processor, and the processor comprises the second memory.

3. The electronic device of claim 1 , wherein each of the plurality of compressed vectors is provided in a corresponding row of the compressed embedding matrix, and

wherein a smaller value of an index representing the corresponding row in the compressed embedding matrix corresponds to an increased frequency of a corresponding word.

4. The electronic device of claim 1 , wherein a second scaling factor is assigned to k compressed vectors among the plurality of compressed vectors, and a third scaling factor is assigned to m compressed vectors with a lower frequency than the k compressed vectors, wherein k is less than m.

5. The electronic device of claim 1 , wherein the processor is further configured to store compressed vectors in the second memory with a frequency greater than or equal to a preset first value based on the first frequency information of the plurality of compressed vectors and store scaling factors in the second memory with a frequency greater than or equal to a preset second value based on the second frequency information of the scaling factors.

6. The electronic device of claim 1 , wherein the second memory comprises a first cache memory configured to store the first group of the plurality of compressed vectors and a second cache memory configured to store the first group of the scaling factors, and

wherein the processor is further configured to:

identify whether the first compressed vector is stored in the first cache memory;

based on the first compressed vector not being stored in the first cache memory, read the first compressed vector from the first cache memory;

based on the first compressed vector being stored in the first cache memory, read the first compressed vector from the first memory;

identify whether the first scaling factor exists in the second cache memory;

based on the first scaling factor being stored in the second cache memory, read the first scaling factor from the second cache memory; and

based on the first scaling factor not being stored in the second cache memory, read the first scaling factor from the first memory.

7. The electronic device of claim 1 , wherein the processor is further configured to identify address information of the first scaling factor based on address information of the first compressed vector and obtain the first scaling factor from the first memory or the second memory based on whether the address information of the first scaling factor indicates the first memory or the second memory.

8. The electronic device of claim 1 , wherein the processor is further configured to identify the first scaling factor corresponding to the first compressed vector based on mapping information indicating a mapping relationship between the plurality of compressed vectors and the scaling factors.

9. The electronic device of claim 1 , further comprising an input interface configured to receive the input word.

10. The electronic device of claim 1 , wherein the processor is further configured to obtain a result value based on data output from the neural network, and

wherein the electronic device further comprises an output interface configured to output the obtained result value.

11. An operation method of an electronic device for processing a word by using a language model, the operation method comprising:

storing, in a first memory configured to operate and a first speed, a compressed embedding matrix including a plurality of compressed vectors respectively corresponding to a plurality of words and scaling factors corresponding to the plurality of compressed vectors;

storing, in a second memory configured to operate at a second speed that is faster than the first speed, a first group of the plurality of compressed vectors identified based on first frequency information of the plurality of compressed vectors;

storing, in the second memory, a first group of the scaling factors identified based on second frequency information of the scaling factors;

obtaining, from the first memory or the second memory, a first compressed vector and a first scaling factor corresponding to an input word; and

processing the obtained first compressed vector and the obtained first scaling factor by using a neural network.

12. The operation method of claim 11 , wherein the first memory is separate from a processor of the electronic device that is configured to process the word, and the processor comprises the second memory.

13. The operation method of claim 11 , wherein each of the plurality of compressed vectors is provided in a corresponding row of the compressed embedding matrix, and

wherein a smaller value of an index representing the corresponding row in the compressed embedding matrix corresponds to an increased frequency of a corresponding word.

14. The operation method of claim 11 , wherein a second scaling factor is assigned to k compressed vectors among the plurality of compressed vectors, and a third scaling factor is assigned to m compressed vectors with a lower frequency than the k compressed vectors, wherein k is less than m.

15. The operation method of claim 11 , wherein the storing of the first group of the plurality of compressed vectors and the first group of the scaling factors in the second memory comprises storing, in the second memory, compressed vectors with a frequency greater than or equal to a preset first value based on the first frequency information of the plurality of compressed vectors and storing, in the second memory, scaling factors with a frequency greater than or equal to a preset second value based on the second frequency information of the scaling factors.

16. The operation method of claim 11 , wherein the second memory comprises a first cache memory configured to store the first group of the plurality of compressed vectors and a second cache memory configured to store the first group of the scaling factors, and

wherein the obtaining of the first compressed vector and the first scaling factor from the first memory or the second memory comprises:

identifying whether the first compressed vector is stored in the first cache memory;

based on the first compressed vector being stored in the first cache memory, reading the first compressed vector from the first cache memory;

based on the first compressed vector not being stored in the first cache memory, reading the first compressed vector from the first memory;

identifying whether the first scaling factor is stored in the second cache memory;

based on the first scaling factor being stored in the second cache memory, reading the first scaling factor from the second cache memory; and

based on the first scaling factor not being stored in the second cache memory, reading the first scaling factor from the first memory.

17. The operation method of claim 11 , wherein the obtaining of the first compressed vector and the first scaling factor from the first memory or the second memory comprises:

identifying address information of the first scaling factor based on address information of the first compressed vector; and

obtaining the first scaling factor from the first memory or the second memory based on whether the address information of the first scaling factor indicates the first memory or the second memory.

18. The operation method of claim 11 , wherein the obtaining of the first compressed vector and the first scaling factor from the first memory or the second memory comprises identifying the first scaling factor corresponding to the first compressed vector based on mapping information indicating a mapping relationship between the plurality of compressed vectors and the scaling factors.

19. The operation method of claim 11 , further comprising:

obtaining a result value based on data output from the neural network; and

outputting the obtained result value.

20. One or more non-transitory computer-readable recording media having stored therein a program for controlling an electronic device to perform a method, the method comprising:

storing, in a first memory configured to operate and a first speed, a compressed embedding matrix including a plurality of compressed vectors respectively corresponding to a plurality of words and scaling factors corresponding to the plurality of compressed vectors;

storing, in a second memory configured to operate at a second speed that is faster than the first speed, a first group of the plurality of compressed vectors identified based on first frequency information of the plurality of compressed vectors;

storing, in the second memory, a first group of the scaling factors identified based on second frequency information of the scaling factors;

obtaining, from the first memory or the second memory, a first compressed vector and a first scaling factor corresponding to an input word; and

processing the obtained first compressed vector and the obtained first scaling factor by using a neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 24, 2020
From: KWON, SEJUNG; LEE, DONGSOO
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 054511/0283 →
Priority Claims (1)
KR 10-2020-0012189 · Jan 31, 2020 · national
Continuity (1)
Related Publication 20210240925A1 · Aug 5, 2021