IP Library Granted Patent US 12,443,534
Granted Patent B2
US 12,443,534 · App. 18/440,263 · Granted Oct 14, 2025

Reference file management for artificial intelligence models

Inventor: Chao Sun (San Jose, CA)
Assignee: Sandisk Technologies, Inc.
G06F12/0862
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,443,534
App. No.
18/440,263
Granted
Oct 14, 2025
Kind
B2
Abstract

A Data Storage Device (DSD) includes a first memory storing reference files used to derive vector embeddings in a vector database. A query vector embedding is received from a host and one or more vector embeddings similar to the query vector embedding are identified in the vector database. One or more reference files from which the one or more vector embeddings were derived are identified and stored in a second memory for faster access. In another aspect, a query vector embedding is received by a host that identifies one or more vector embeddings that are similar to the query vector embedding and retrieves the one or more vector embeddings from a DSD to provide to an Artificial Intelligence (AI) model. One or more reference files are identified from which the one or more vector embeddings were derived and are prefetched from the DSD for storage in a host memory.

Claims (53)

1. A Data Storage Device (DSD), comprising:

a first memory configured to store a plurality of reference files that have been used to derive a plurality of vector embeddings in a vector database stored in the DSD; and

circuitry configured to:

receive a query vector embedding from a host;

perform an Approximate Nearest Neighbor (ANN) search of the vector database to identify one or more vector embeddings that are close to the query vector embedding;

send the one or more vector embeddings to the host;

identify one or more reference files from which the one or more vector embeddings were derived using vector metadata for the one or more vector embeddings; and

store the one or more reference files in a second memory of the host or of the DSD, the second memory configured to provide access to data faster than the first memory.

2. The DSD of claim 1 , wherein the circuitry is further configured to create a zone in the second memory for caching reference files from the plurality of reference files stored in the first memory.

3. The DSD of claim 1 , wherein the circuitry is further configured to prioritize storage of the one or more reference files in the second memory over at least one other operation performed by the DSD.

4. The DSD of claim 1 , wherein the circuitry is further configured to:

associate the one or more reference files stored in the second memory with the respective one or more vector embeddings derived from the one or more reference files;

receive a request from the host for a reference file corresponding to the one or more vector embeddings;

identify the requested reference file stored in the second memory using the association; and

send the requested reference file stored in the second memory to the host.

5. The DSD of claim 1 , wherein the circuitry is further configured to perform the ANN search by comparing the query vector embedding with a plurality of candidate vector embeddings in the vector database to identify the one or more vector embeddings that are closest to the query vector embedding.

6. The DSD of claim 5 , wherein the circuitry is further configured to perform the ANN search of the vector database by at least one of determining a cosine of an angle between candidate vector embeddings and the query vector embedding, determining a Euclidian distance between candidate vector embeddings and the query vector embedding and determining a dot product between candidate vector embeddings and the query vector embedding to identify the one or more vector embeddings.

7. The DSD of claim 1 , wherein the circuitry is further configured to use the vector metadata as part of a filtering operation in a search for the one or more vector embeddings.

8. The DSD of claim 1 , wherein the second memory is further configured to store an index that includes vector metadata for the plurality of vector embeddings in the vector database, and wherein the circuitry is further configured to identify the vector metadata for the one or more vector embeddings using the index.

9. A method performed by a host for an Artificial Intelligence (AI) model, the method comprising:

receiving a query vector embedding;

performing an Approximate Nearest Neighbor (ANN) search of a vector database to identify one or more vector embeddings in the vector database that are similar to the query vector embedding;

retrieving the one or more vector embeddings from the vector database;

providing the one or more vector embeddings to the AI model;

identifying one or more reference files from which the one or more vector embeddings were derived using vector metadata for the one or more vector embeddings;

prefetching the one or more reference files from a Data Storage Device (DSD); and

storing the prefetched one or more reference files in at least one memory of the host.

10. The method of claim 9 , further comprising creating a zone in the at least one memory of the host for caching reference files from a plurality of reference files stored in the DSD.

11. The method of claim 9 , further comprising prioritizing the prefetching of the one or more reference files from the DSD over at least one other operation performed by at least one of the host and the DSD.

12. The method of claim 9 , further comprising:

associating the one or more reference files stored in the at least one memory with the respective one or more vector embeddings derived from the one or more reference files;

receiving a request from an application for a reference file corresponding to the one or more vector embeddings;

identifying the requested reference file stored in the at least one memory using the association; and

sending the requested reference file stored in the at least one memory to the application.

13. The method of claim 9 , further comprising performing the ANN search by comparing the query vector embedding with a plurality of candidate vector embeddings in the vector database to identify the one or more vector embeddings that are closest to the query vector embedding.

14. The method of claim 13 , further comprising performing the ANN search of the vector database by at least one of determining a cosine of an angle between candidate vector embeddings and the query vector embedding, determining a Euclidian distance between candidate vector embeddings and the query vector embedding and determining a dot product between candidate vector embeddings and the query vector embedding.

15. The method of claim 9 , further comprising using the vector metadata as part of a filtering operation in a search for the one or more vector embeddings.

16. The method of claim 9 , wherein the at least one memory of the host further stores an index that includes vector metadata for a plurality of vector embeddings in the vector database, and wherein the method further comprises identifying the vector metadata for the one or more vector embeddings using the index.

17. A Data Storage Device (DSD), comprising:

a first memory configured to store a plurality of reference files that have been used to derive a plurality of vector embeddings in a vector database stored in the DSD; and

means for:

receiving a query vector embedding from a host;

performing an Approximate Nearest Neighbor (ANN) search of the vector database to identify one or more vector embeddings that are close to the query vector embedding;

sending the one or more vector embeddings to the host;

identifying one or more reference files from which the one or more vector embeddings were derived using vector metadata for the one or more vector embeddings; and

storing the one or more reference files in a second memory of the host or of the DSD, the second memory configured to provide access to data faster than the first memory.

18. The DSD of claim 17 , further comprising means for prioritizing the storage of the one or more reference files in the second memory over at least one other operation performed by the DSD.

19. The DSD of claim 17 , further comprising means for:

associating the one or more reference files stored in the second memory with the respective one or more vector embeddings derived from the one or more reference files;

receiving a request from the host for a reference file corresponding to the one or more vector embeddings;

identifying the requested reference file stored in the second memory using the association; and

sending the requested reference file stored in the second memory to the host.

20. The DSD of claim 17 , wherein the second memory is further configured to store an index that includes vector metadata for the plurality of vector embeddings in the vector database, and wherein the DSD further comprises means for identifying the vector metadata for the one or more vector embeddings using the index.

Assignments (7)
PARTIAL RELEASE OF SECURITY INTERESTS Recorded Apr 25, 2025
From: JPMORGAN CHASE BANK, N.A., AS AGENT
To: SANDISK TECHNOLOGIES, INC.
Reel/Frame 071382/0001 →
SECURITY AGREEMENT Recorded Apr 25, 2025
From: SANDISK TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 071050/0001 →
PATENT COLLATERAL AGREEMENT Recorded Aug 23, 2024
From: SANDISK TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A., AS THE AGENT
Reel/Frame 068762/0494 →
CHANGE OF NAME Recorded Jun 27, 2024
From: SANDISK TECHNOLOGIES, INC.
To: SANDISK TECHNOLOGIES, INC.
Reel/Frame 067982/0032 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 29, 2024
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: SANDISK TECHNOLOGIES, INC.
Reel/Frame 067567/0682 →
PATENT COLLATERAL AGREEMENT (AR) Recorded May 15, 2024
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A., AS THE AGENT
Reel/Frame 067417/0329 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 14, 2024
From: SUN, CHAO
To: WESTERN DIGITAL TECHNOLOGIES, INC.
Reel/Frame 066596/0377 →
Continuity (1)
Related Publication 20250258774A1 · Aug 14, 2025
References Cited (18)
US 11514028B2 · Pathak et al. · 2022 [cited by applicant]
US 20040098650A1 · Rajsuman · 2004 [cited by examiner]
US 20200272566A1 · Saeki · 2020 [cited by applicant]
CN 102129472B · 2012 [cited by applicant]
CN 115310605A · 2022 [cited by applicant]
KR 101975101B1 · 2019 [cited by applicant]
Dave Bergstein, Solving Complex Problems with Vector Databases; Mar. 8, 2022; New Tech Forum. [cited by applicant]
Yang et al., “SGDP: A Stream-Graph Neural Network Based Data Prefetcher”; Apr. 7, 2023; International Joint Conference on Neural Networks 2023. [cited by applicant]
Selvaganapathy C, “Transforming Enterprise Knowledge Management Using Large Language Models (LLM)”; Nov. 6, 2023; The AI Discovery. [cited by applicant]
Kim et al., “Q-Selector-Based Prefetching Method for DRAM/NVM Hybrid Main Memory System”; Dec. 16, 2020; MDPI. [cited by applicant]
Chakraborttii et al., “Learning I/O Access Patterns to Improve Prefetching in SSDs”; Sep. 2020; The European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases. [cited by applicant]
Wang et al. “An Efficient Memory System Design with Specialized Caching Mechanism for Recommendation Inference”; Sep. 9, 2023; ACM Transactions on Embedded Computing Systems, vol. 22, Issue 56, Article No. 100, pp. 1-22. [cited by applicant]
James Luan, “ChatGPT+Vector Database+Prompt-As-Code—The CVP Stack”; Apr. 4, 2023. [cited by applicant]
Levine et al., “Investing in Pinecone”; Apr. 27, 2023. [cited by applicant]
Mark Hinkle, “Vector Databases: Long-Term Memory for Artificial Intelligence”; May 12, 2023; The New Stack. [cited by applicant]
Microsoft, “What Is a Vector Database?”; Dec. 13, 2023. [cited by applicant]
Pending U.S. Appl. No. 18/449,116 by Sun et al., “Computational SSD Supporting Rapid File Semantic Search”, filed Aug. 14, 2023. [cited by applicant]
Pending U.S. Appl. No. 18/449,165 by Sun et al., “Error Correction Methods for Computational SSD Supporting Rapid File Semantic Search”, filed Aug. 14, 2023. [cited by applicant]