IP Library › Granted Patent US 12,640,753
Granted Patent B2
US 12,640,753 · App. 18/301,816 · Granted May 26, 2026

Selective and flexible compression of data associated with a machine learning model

Inventors: Kaushal Gandhi (South San Francisco, CA); Olivia Wu (Los Altos, CA); Soheil Gharahi (Austin, TX); Thomas Mark Ulrich (Sunnyvale, CA); Abdulkadir Utku Diril (Menlo Park, CA); Khasim S. Dudekula (Scarsdale, NY); Eda Sahin (Sunnyvale, CA)
Assignee: Meta Platforms, Inc.
H03M7/42
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,640,753
App. No.
18/301,816
Granted
May 26, 2026
Kind
B2
Abstract

Systems, apparatuses and methods provide technology that compresses first data based on a first compression scheme to generate second data, where the first data is associated with a first machine learning model. The technology stores the second data into a memory, adjusts a first entry of a lookup table to correspond to the first compression scheme based on the first data being compressed based on the first compression scheme, provide the second data from the memory to processing elements of a processing array during execution of the first machine learning model, and decompresses, at the processing array, the second data based on the lookup table to obtain the first data.

Claims (86)

1 . At least one non-transitory computer readable storage medium comprising a set of instructions, which when executed by a computing device, cause the computing device to:

compress first data based on a first compression scheme to generate second data, wherein the first data is associated with a first machine learning model;

store the second data into a memory;

adjust a first entry of a lookup table to correspond to the first compression scheme based on the first data being compressed based on the first compression scheme;

provide the second data from the memory to processing elements of a processing array during execution of the first machine learning model; and

decompress, at the processing array, the second data based on the lookup table to obtain the first data;

wherein: the first data includes one or more of weights or biases; the first compression scheme is an Asymmetric Numeral Systems compression scheme; the first data is to be compressed based on the first compression scheme prior to the first machine model being executed; and to decompress the second data, the processing array is to decompress the second data with the processing elements,

wherein the instructions, when executed, cause the computing device to:

determine that compression of the first data with the first compression scheme reduces a size of the first data to be less than or equal to an available memory space of the memory;

predict that the first compression scheme is to be applied at least a first number of times;

determine that the at least the first number of times meets a reuse threshold;

select the first compression scheme based on the at least the first number of times meeting the reuse threshold, and compression of the first data with the first compression scheme reducing the size of the first data to be less than or equal to the available memory space;

generate metadata to associate the first compression scheme with the second data; and

store the second data in association with the metadata into the memory.

2 . The at least one non-transitory computer readable storage medium of claim 1 , wherein the instructions, when executed, cause the computing device to:

compress third data based on a second compression scheme to generate fourth data, wherein the third data is associated with a second machine learning model;

store the fourth data into the memory;

adjust a second entry of the lookup table to correspond to the second compression scheme based on the third data being compressed based on the second compression scheme;

provide the third data from the memory to the processing array during execution of the second machine learning model; and

decompress, at the processing array, the fourth data based on the lookup table to obtain the third data.

3 . The at least one non-transitory computer readable storage medium of claim 2 , wherein the instructions, when executed, cause the computing device to:

replace the first entry in the lookup table with the second entry.

4 . The at least one non-transitory computer readable storage medium of claim 2 , wherein the instructions, when executed, cause the computing device to:

insert the second entry into the lookup table so that the lookup table includes the first entry and the second entry.

5 . The at least one non-transitory computer readable storage medium of claim 1 , wherein the instructions, when executed, cause the computing device to:

identify that a first portion of the first data is compressible by an amount based on the first compression scheme; and

determine that the first portion is to be compressed with the first compression scheme based on the amount meeting a compression threshold.

6 . A system comprising:

one or more processors; and

first memory coupled to the one or more processors, the first memory comprising instructions executable by the one or more processors, the one or more processors being operable when executing the instructions to:

compress first data based on a first compression scheme to generate second data, wherein the first data is associated with a first machine learning model;

store the second data into a second memory;

adjust a first entry of a lookup table to correspond to the first compression scheme based on the first data being compressed based on the first compression scheme;

provide the second data from the second memory to processing elements of a processing array during execution of the first machine learning model;

decompress, at the processing array, the second data based on the lookup table to obtain the first data;

determine that compression of the first data with the first compression scheme reduces a size of the first data to be less than or equal to an available memory space of the second memory;

predict that the first compression scheme is to be applied at least a first number of times;

determine that the at least the first number of times meets a reuse threshold;

select the first compression scheme based on the at least the first number of times meeting the reuse threshold, and compression of the first data with the first compression scheme reducing the size of the first data to be less than or equal to the available memory space;

generate metadata to associate the first compression scheme with the second data; and

store the second data in association with the metadata into the second memory.

7 . The system of claim 6 , wherein the one or more processors are further operable when executing the instructions to:

compress third data based on a second compression scheme to generate fourth data, wherein the third data is associated with a second machine learning model;

store the fourth data into the second memory;

adjust a second entry of the lookup table to correspond to the second compression scheme based on the third data being compressed based on the second compression scheme;

provide the third data from the second memory to the processing array during execution of the second machine learning model; and

decompress, at the processing array, the fourth data based on the lookup table to obtain the third data.

8 . The system of claim 7 , wherein the one or more processors are further operable when executing the instructions to:

replace the first entry in the lookup table with the second entry.

9 . The system of claim 7 , wherein the one or more processors are further operable when executing the instructions to:

insert the second entry into the lookup table so that the lookup table includes the first entry and the second entry.

10 . The system of claim 6 , wherein the one or more processors are further operable when executing the instructions to:

identify that a first portion of the first data is compressible by an amount based on the first compression scheme; and

determine that the first portion is to be compressed with the first compression scheme based on the amount meeting a compression threshold.

11 . The system of claim 6 , wherein:

the first data includes one or more of weights or biases;

the first compression scheme is an Asymmetric Numeral Systems compression scheme;

the first data is to be compressed based on the first compression scheme prior to the first machine model being executed; and

to decompress the second data, the processing array is to decompress the second data with the processing elements.

12 . A method comprising:

compressing first data based on a first compression scheme to generate second data, wherein the first data is associated with a first machine learning model;

storing the second data into a memory;

adjusting a first entry of a lookup table to correspond to the first compression scheme based on the first data being compressed based on the first compression scheme;

providing the second data from the memory to processing elements of a processing array during execution of the first machine learning model; and

decompressing, at the processing array, the second data based on the lookup table to obtain the first data;

wherein: the first data includes one or more of weights or biases; the first compression scheme is an Asymmetric Numeral Systems compression scheme; the first data is to be compressed based on the first compression scheme prior to the first machine model being executed; and to decompress the second data, the processing array is to decompress the second data with the processing elements,

wherein the method, further comprising:

determining that compression of the first data with the first compression scheme reduces a size of the first data to be less than or equal to an available memory space of the memory;

predicting that the first compression scheme is to be applied at least a first number of times;

determining that the at least the first number of times meets a reuse threshold;

selecting the first compression scheme based on the at least the first number of times meeting the reuse threshold, and compression of the first data with the first compression scheme reducing the size of the first data to be less than or equal to the available memory space;

generating metadata to associate the first compression scheme with the second data; and

storing the second data in association with the metadata into the memory.

13 . The method of claim 12 , further comprising:

compressing third data based on a second compression scheme to generate fourth data, wherein the third data is associated with a second machine learning model;

storing the fourth data into the memory;

adjusting a second entry of the lookup table to correspond to the second compression scheme based on the third data being compressed based on the second compression scheme;

providing the third data from the memory to the processing array during execution of the second machine learning model; and

decompressing, at the processing array, the fourth data based on the lookup table to obtain the third data.

14 . The method of claim 13 , further comprising:

replacing the first entry in the lookup table with the second entry.

15 . The method of claim 13 , further comprising:

inserting the second entry into the lookup table so that the lookup table includes the first entry and the second entry.

16 . The method of claim 12 , further comprising:

identifying that a first portion of the first data is compressible by an amount based on the first compression scheme; and

determining that the first portion is to be compressed with the first compression scheme based on the amount meeting a compression threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2025
From: GHARAHI, SOHEIL; GANDHI, KAUSHAL; WU, OLIVIA; SAHIN, EDA; ULRICH, THOMAS MARK; DIRIL, ABDULKADIR UTKU; DUDEKULA, KHASIM S
To: META PLATFORMS, INC.
Reel/Frame 073173/0784 →
Continuity (1)
Related Publication 20240348263A1 · Oct 17, 2024
References Cited (10)
US 20130275396A1 · Condict · 2013 [cited by examiner]
US 20210027148A1 · Meng et al. · 2021 [cited by applicant]
US 20220013153A1 · Zafar · 2022 [cited by examiner]
US 20220404887A1 · Cruise · 2022 [cited by examiner]
US 20240236295A1 · Martinelli · 2024 [cited by examiner]
US 20240340125A1 · Marzban · 2024 [cited by examiner]
US 20250175192A1 · Galvin · 2025 [cited by examiner]
EP 4096100A1 · 2022 [cited by applicant]
European Search Report for European Patent Application No. 24162238.0, dated Aug. 15, 2024, 7 pages. [cited by applicant]
Racape F., et al., “Adaptive and Conditional Arithmetic Coding for Lossless Compression of Deep Neural Networks,” Interdigital, [NNR] CE3: Report on Arithmetic Coding Results, 130. MPEG Meeting, Apr. 10, 2020, 10 pages,… [cited by applicant]