IP Library › Granted Patent US 12,387,053
Granted Patent B2
US 12,387,053 · App. 17/585,619 · Granted Aug 12, 2025

Large-scale text data encoding and compression

Inventors: Zhong Fang Yuan (Xi'an, CN); Tong Liu (Xi'an, CN); Wen Wang (Beijing, CN); Chen Gao (Xi'an, CN); Xiang Yu Yang (Xi'an, CN)
Assignee: International Business Machines Corporation
G06F40/40G06F40/205
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,053
App. No.
17/585,619
Granted
Aug 12, 2025
Kind
B2
Abstract

Embodiments of the present invention provide an approach for compressing data, and more particularly, to large-scale text data encoding and compression using absolute overfitting on pre-trained language models. Large-scale data is parsed into sentences. A unique token is generated for each sentence to form a token list. A generative (or compression) model is trained from the tokens in the token list to produce the corresponding sentence of each token through absolute overfitting of a pre-trained language model. The compressed text data is stored as the token list and generative model, resulting in a storage space savings.

Claims (35)

1. A computer-implemented method for encoding and compressing text data, comprising computer-implemented steps of:

parsing, using a pre-trained language model, received text data to be compressed into a set of sentences;

generating, using the pre-trained language model, a unique token for each sentence among the set of sentences to form a token list;

training a generative model using absolute overfitting on the pre-trained language model, wherein the training further comprises training a decoder part of the generative model such that an output generated for a given input that includes the unique token matches an original sentence in the received text data corresponding to the unique token; and

compressing the received text data using the trained generative model and each unique token within the token list, wherein the compressing further comprises using the trained generative model to produce a corresponding sentence of each unique token through the absolute overfitting of the pre-trained language model, wherein a combination of the token list and the trained generative model that is based on the absolute overfitting of the pre-trained language model represent the compressed text data.

2. The computer-implemented method of claim 1 , wherein the unique token includes an embedding, a length of the corresponding sentence, and a hash identifier.

3. The computer-implemented method of claim 1 , wherein the pre-trained language model is a Generative Pre-trained Transformer 3 (GPT-3) or Bidirectional Encoder Representations from Transformers (BERT) model.

4. The computer-implemented method of claim 1 , wherein training the generative model further comprises using the unique token within the token list as input to produce the output and ensuring, using a length of the corresponding sentence stored in the unique token as a constraint, the output exactly matches a text of the corresponding sentence.

5. The computer-implemented method of claim 1 , further comprising decompressing each sentence in the text data using the trained generative model and the token list.

6. The computer-implemented method of claim 5 , further comprising organizing, using a hash identifier of each unique token in the token list, the decompressed sentences in an exact order of the received text data, wherein the hash identifier of each unique token references a text unit of the corresponding sentence.

7. The computer-implemented method of claim 2 , wherein the embedding of the unique token includes a sentence vector of the sentence corresponding to the token.

8. A system for encoding and compressing text data, comprising:

a memory medium comprising program instructions;

a bus coupled to the memory medium; and

a processor, for executing the program instructions, coupled to the memory medium that when executing the program instructions causes the system to:

parse, using a pre-trained language model, received text data to be compressed into a set of sentences;

generate, using the pre-trained language model, a unique token for each sentence among the set of sentences to form a token list;

train a generative model using absolute overfitting on the pre-trained language model, wherein the training further comprises training a decoder part of the generative model such that an output generated for a given input that includes the unique token matches an original sentence in the received text data corresponding to the unique token; and

compress the received text data using the trained generative model and each unique token within the token list, wherein the compressing further comprises using the trained generative model to produce a corresponding sentence of each unique token through the absolute overfitting of the pre-trained language model, wherein a combination of the token list and the trained generative model that is based on the absolute overfitting of the pre-trained language model represent the compressed text data.

9. The system of claim 8 , wherein the unique token includes an embedding, a length of the corresponding sentence, and a hash identifier.

10. The system of claim 8 , wherein the pre-trained language model is a Generative Pre-trained Transformer 3 (GPT-3) or Bidirectional Encoder Representations from Transformers (BERT) model.

11. The system of claim 8 , the memory medium further comprising instructions to train the generative model using the unique token within the token list as input to produce the output and ensuring, using a length of the corresponding sentence stored in the unique token as a constraint, the output exactly matches a text of the corresponding sentence.

12. The system of claim 8 , the memory medium further comprising instructions to decompress each sentence in the text data using the trained generative model and the token list.

13. The system of claim 12 , the memory medium further comprising instructions to organize, using a hash identifier of each unique token in the token list, the decompressed sentences in an exact order of the received text data, wherein the hash identifier of each unique token references a text unit of the corresponding sentence.

14. The system of claim 9 , wherein the embedding of the unique token includes a sentence vector of the sentence corresponding to the unique token.

15. A computer program product for encoding and compressing text data, the computer program product comprising a computer readable storage device, and program instructions stored on the computer readable storage device, to:

parse, using a pre-trained language model, received text data to be compressed into a set of sentences;

generate, using the pre-trained language model, a unique token for each sentence among the set of sentences to form a token list;

train a generative model using absolute overfitting on the pre-trained language model, wherein the training further comprises training a decoder part of the generative model such that an output generated for a given input that includes the unique token matches an original sentence in the received text data corresponding to the unique token; and

compress the received text data using the trained generative model and each unique token within the token list, wherein the compressing further comprises using the trained generative model to produce a corresponding sentence of each unique token through the absolute overfitting of the pre-trained language model, wherein a combination of the token list and the trained generative model that is based on the absolute overfitting of the pre-trained language model represent the compressed text data.

16. The computer program product of claim 15 , wherein the unique token includes an embedding, a length of the corresponding sentence, and a hash identifier.

17. The computer program product of claim 15 , wherein the pre-trained language model is a Generative Pre-trained Transformer 3 (GPT-3) or Bidirectional Encoder Representations from Transformers (BERT) model.

18. The computer program product of claim 15 , further comprising program instructions stored on the computer readable storage device to train the generative model using the unique token within the token list as input to produce the output and ensuring, using a length of the corresponding sentence stored in the unique token as a constraint, the output exactly matches a text of the corresponding sentence.

19. The computer program product of claim 15 , further comprising program instructions stored on the computer readable storage device to decompress each sentence in the text data using the trained generative model and the token list.

20. The computer program product of claim 19 , further comprising program instructions stored on the computer readable storage device to organize, using a hash identifier of each unique token in the token list, the decompressed sentences in an exact order of the received text data, wherein the hash identifier of each unique token references a text unit of the corresponding sentence.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 27, 2022
From: YUAN, ZHONG FANG; LIU, TONG; WANG, WEN; GAO, CHEN; YANG, XIANG YU
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 058785/0350 →
Continuity (1)
Related Publication 20230237278A1 · Jul 27, 2023
References Cited (21)
US 5793869A · Claflin, Jr. · 1998 [cited by applicant]
US 8775805B2 · von Mueller et al. · 2014 [cited by applicant]
US 10229111B1 · Filippova et al. · 2019 [cited by applicant]
US 12045272B2 · Mahapatra · 2024 [cited by examiner]
US 20190018933A1 · Oono · 2019 [cited by examiner]
US 20200042547A1 · Prakash · 2020 [cited by examiner]
US 20210117617A1 · Blaya · 2021 [cited by examiner]
US 20210374338A1 · Shrivastava · 2021 [cited by examiner]
US 20220067529A1 · Wagner · 2022 [cited by examiner]
US 20220318255A1 · Fei · 2022 [cited by examiner]
US 20220343076A1 · Saito · 2022 [cited by examiner]
US 20220405461A1 · Lempel · 2022 [cited by examiner]
US 20230020886A1 · Mahapatra · 2023 [cited by examiner]
CN 112395891A · 2021 [cited by applicant]
CN 117827111A · 2024 [cited by examiner]
WO 2021074272A1 · 2021 [cited by applicant]
Fevry et al., “Unsupervised sentence compression using denoising auto-encoders”, 2018, arXiv preprint arXiv:1809.02669. Sep. 7, 2018. (Year: 2018). [cited by examiner]
Hermansson, “Using pre-trained language models for extractive text summarisation of academic papers”, 2020, Master's Thesis, pp. 1-95 (Year: 2020). [cited by examiner]
Fan et al “Controllable abstractive summarization”, 2017, arXiv preprint arXiv:1711.05217. Nov. 14, 2017. (Year: 2017). [cited by examiner]
Kikuchi et al, “Controlling output length in neural encoder-decoders”, 2016, arXiv preprint arXiv:1609.09552. Sep. 30, 2016. (Year: 2016). [cited by examiner]
Weiwei Hou et al., “A Token-wise CNN-based Method for Sentence Compression”, International Conference on Neural Information Processing 2020, Published Nov. 19, 2020, 12 pages. [cited by applicant]