IP Library › Granted Patent US 12,645,389
Granted Patent B2
US 12,645,389 · App. 17/664,905 · Granted Jun 2, 2026

Computational storage device for deep-learning recommendation system and method of operating the same

Inventors: Minho Kim (Seongnam-si, KR); Wijik Lee (Suwon-si, KR); Sooyoung Ji (Seoul, KR); Sanghwa Jin (Seongnam-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06F3/0655G06F3/0604G06F3/0679G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,389
App. No.
17/664,905
Filed
May 25, 2022
Granted
Jun 2, 2026
Kind
B2
Art Unit
2142
USPC
706/12
Abstract

A computational storage device includes a nonvolatile memory configured to store a plurality of embedding tables for a deep-learning recommendation system (DLRS), and a storage controller configured to control an operation of the nonvolatile memory, store a plurality of applications that are off-loaded from a host device executing the DLRS, and support an execution of the DLRS by executing the plurality of applications and performing a plurality of calculations based on the plurality of embedding tables. The storage controller includes a machine learning engine configured to determine a management scheme of at least one embedding table of the plurality of embedding tables and the plurality of applications by analyzing the at least one embedding table and the plurality of applications.

Claims (84)

1 . A computational storage device comprising:

a nonvolatile memory configured to store a plurality of embedding tables for a deep-learning recommendation system (DLRS); and

a storage controller configured to:

control an operation of the nonvolatile memory;

store a plurality of applications that are off-loaded from a host device executing the DLRS; and

support an execution of the DLRS by executing the plurality of applications and performing a plurality of calculations based on the plurality of embedding tables,

wherein the storage controller includes a machine learning engine configured to determine a management scheme of at least one embedding table of the plurality of embedding tables and the plurality of applications by analyzing the at least one embedding table and the plurality of applications, and

wherein the machine learning engine is further configured to perform an analysis of usage of the at least one embedding table by determining a number of times the at least one embedding table is used.

2 . The computational storage device of claim 1 , wherein the machine learning engine is further configured to:

classify the at least one embedding table as a first embedding table or a second embedding table based on the analysis and a reference usage number.

3 . The computational storage device of claim 2 , wherein the nonvolatile memory includes:

a first block including single-level memory cells (SLCs), each SLC storing one data bit; and

a second block including multi-level memory cells (MLCs), each MLC storing two or more data bits, and

wherein first embedding tables are stored in the first block and second embedding tables are stored in the second block.

4 . The computational storage device of claim 1 , further comprising a buffer memory configured to temporarily store data stored in or to be stored in the nonvolatile memory,

wherein the machine learning engine is configured to:

analyze usage of a plurality of vector data included in the plurality of embedding tables;

detect first vector data from among the plurality of vector data based on a result of analyzing the usage of the plurality of vector data; and

temporarily store the first vector data in the buffer memory, and

wherein the first vector data is used by loading from the buffer memory.

5 . The computational storage device of claim 1 , further comprising a reconfigurable hardware configured to store and execute at least one of the plurality of applications,

wherein the machine learning engine is configured to:

analyze usage of a plurality of operators associated with the plurality of applications;

detect a first application from among the plurality of applications; and

store the first application in the reconfigurable hardware, and

wherein the first application is used by the reconfigurable hardware.

6 . The computational storage device of claim 5 , wherein the reconfigurable hardware is disposed inside the storage controller.

7 . The computational storage device of claim 5 , wherein the reconfigurable hardware is disposed outside the storage controller.

8 . The computational storage device of claim 1 , wherein the machine learning engine includes:

a mode configuration module configured to receive operation mode information;

a retrieving module configured to receive a plurality of internal information from the nonvolatile memory;

a feature analyzing module configured to perform an analyzing operation on the plurality of internal information based on the plurality of embedding tables;

a clustering module configured to classify at least some of the plurality of internal information into a first cluster and a second cluster by performing a clustering operation based on a result of the analyzing an operation on the plurality of internal information;

a categorizing module configured to:

receive a plurality of pattern information,

select a first reference value and a second reference value from the first cluster and the second cluster, respectively, based on a result of the clustering operation and the plurality of pattern information, and

categorize the at least some of the plurality of internal information into a first reference group and a second reference group based on the first reference value and the second reference value; and

an estimating module configured to perform an estimating operation for analyzing the plurality of pattern information using the first reference group and the second reference group,

wherein the clustering module is configured to select an optimal clustering value by calculating a lowest similarity value between clusters.

9 . The computational storage device of claim 8 , wherein:

the clustering module is further configured to classify remaining information of the plurality of internal information that is not classified into the first and second clusters into a third cluster;

the categorizing module is further configured to select a third reference value from the third cluster and categorize the remaining information into a third reference group based on the third reference value; and

the estimating module is further configured to perform the estimating operation using the first, second, and third reference groups.

10 . The computational storage device of claim 8 , wherein the clustering module is further configured to:

classify the at least some of the plurality of internal information into the first cluster and the second cluster by performing a k-means clustering operation; and

determine the optimal clustering value representing an optimal quantity of clusters using an elbow technique.

11 . The computational storage device of claim 10 , wherein the categorizing module is configured to:

select internal information having a largest k-mean among internal information included in the first cluster as the first reference value;

select internal information having a largest k-mean among internal information included in the second cluster as the second reference value; and

categorize the at least some of the plurality of internal information into the first reference group and the second reference group by performing a support vector machine (SVM) operation on the internal information included in the first cluster and the internal information included in the second cluster based on the first reference value and the second reference value.

12 . The computational storage device of claim 8 , wherein the plurality of internal information include workload information associated with features of the plurality of applications, input/output (I/O) pattern information associated with features of accessing the computational storage device by the host device, and nonvolatile memory information associated with features of the nonvolatile memory.

13 . The computational storage device of claim 8 , wherein the feature analyzing module includes:

a first feature analyzer configured to perform a first feature analyzing operation without the plurality of embedding tables;

a second feature analyzer configured to perform a second feature analyzing operation using the plurality of embedding tables; and

a third feature analyzer configured to perform a third feature analyzing operation based on a result of the first feature analyzing operation and a result of the second feature analyzing operation.

14 . The computational storage device of claim 8 , wherein the machine learning engine further includes a reporting module configured to update the plurality of internal information based on a result of the estimating operation.

15 . The computational storage device of claim 1 , wherein the storage controller further includes:

a host interface configured to communicate with the host device; and

a program slot configured to store the plurality of applications.

16 . The computational storage device of claim 15 , wherein the host interface is configured to operate based on a nonvolatile memory express (NVMe) technical proposal (TP) 4091 protocol.

17 . The computational storage device of claim 15 , further comprising a buffer memory configured to temporarily store data stored in or to be stored in the nonvolatile memory, and

wherein the storage controller further includes a buffer manager configured to manage the buffer memory.

18 . The computational storage device of claim 17 , wherein the host interface is configured to operate based on compute express link (CXL) protocol.

19 . A method of operating a computational storage device, the method comprising:

storing a plurality of applications that are off-loaded from a host device executing a deep-learning recommendation system (DLRS);

storing a plurality of embedding tables that are used in the DLRS;

supporting an execution of the DLRS by executing the plurality of applications and performing a plurality of calculations based on the plurality of embedding tables; and

determining, by use of a machine learning engine, a management scheme of at least one of the plurality of embedding tables and the plurality of applications by analyzing the at least one of the plurality of embedding tables and the plurality of applications,

wherein the machine learning engine is configured to perform an analysis of usage of the at least one embedding table by determining a number of times the at least one embedding table is used.

20 . A computational storage device comprising:

a nonvolatile memory configured to store a plurality of embedding tables for a deep-learning recommendation system (DLRS);

a buffer memory configured to temporarily store data stored in or to be stored in the nonvolatile memory;

a storage controller configured to:

control an operation of the nonvolatile memory,

store a plurality of applications that are off-loaded from a host device executing the DLRS, and

support an execution of the DLRS by executing the plurality of applications and performing a plurality of calculations based on the plurality of embedding tables; and

a reconfigurable hardware configured to store and execute at least one of the plurality of applications, the reconfigurable hardware being disposed inside or outside the storage controller,

wherein the storage controller includes a machine learning engine configured to:

analyze at least one of the plurality of embedding tables and the plurality of applications by performing a k-means clustering operation, a support vector machine (SVM) operation, and an estimating operation based on a plurality of internal information; and

determine a management scheme of the at least one of the plurality of embedding tables and the plurality of applications based on a result of analyzing the at least one of the plurality of embedding tables and the plurality of applications, and

wherein the machine learning engine is further configured to:

analyze usage of the plurality of embedding tables and classify the plurality of embedding tables into first embedding tables stored in single-level memory cells (SLCs) and second embedding tables stored in multi-level memory cells (MLCs), or

analyze usage of a plurality of vector data included in the plurality of embedding tables, temporarily store at least one of the plurality of vector data in the buffer memory, and use the at least one of the plurality of vector data stored in the buffer memory, or

analyze usage of a plurality of operators associated with the plurality of applications, store at least one of the plurality of applications in the reconfigurable hardware, and use the at least one of the plurality of applications stored in the reconfigurable hardware.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 25, 2022
From: KIM, MINHO; LEE, WIJIK; JI, SOOYOUNG; JIN, SANGHWA
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 060010/0324 →
Priority Claims (1)
KR 10-2021-0128011 · Sep 28, 2021 · national
Continuity (1)
Related Publication 20230102226A1 · Mar 30, 2023
References Cited (23)
US 7930322B2 · MaClennan · 2011 [cited by applicant]
US 10963503B2 · Skiles et al. · 2021 [cited by applicant]
US 10984337B2 · Bai et al. · 2021 [cited by applicant]
US 10984344B2 · Yamaguchi et al. · 2021 [cited by applicant]
US 20180204113A1 · Galron et al. · 2018 [cited by applicant]
US 20190065486A1 · Lin · 2019 [cited by examiner]
US 20190391976A1 · Tsai et al. · 2019 [cited by applicant]
US 20200089769A1 · Crossley et al. · 2020 [cited by applicant]
US 20200401344A1 · Bazarsky · 2020 [cited by examiner]
US 20210150338A1 · Semenov · 2021 [cited by applicant]
US 20210264220A1 · Wei et al. · 2021 [cited by applicant]
US 20210383067A1 · Reisswig · 2021 [cited by examiner]
US 20220172825A1 · Bhagalia · 2022 [cited by examiner]
CN 112114968A · 2020 [cited by applicant]
JP 2013228933 · 2013 [cited by applicant]
JP 2014112283 · 2014 [cited by applicant]
KR 101629178 · 2016 [cited by applicant]
What is NVMe and Why is it Important, by Gupta, https://blog.westerndigital.com/nvme-important-data-driven-businesses/ (Year: 2020). [cited by examiner]
Compute Express Link: A Coherent Interface for Ultra-High-Speed Transfers, Kurt Lender, https://www.dmtf.org/sites/default/files/CXL_Overview_Virtual_DMTF_APTS_July_2020.pdf (Year: 2020). [cited by examiner]
Attention in Recommendation Systems, Lee (Year: 2020). [cited by examiner]
Monolith: Real Time Recommendation System With Collisionless Embedding Table, Liu et al., Sep. 27, 2022 (Year: 2022). [cited by examiner]
Yao, et al., “K-SVM: An Effective SVM Algorithm Based on K-means Clustering”, Journal of Computers, vol. 8, No. 10, Oct. 2013, pp. 2632-2639. [cited by applicant]
Office Action dated May 29, 2025 issued in corresponding Korean Patent Application No. 10-2021-0128011. [cited by applicant]