IP Library › Granted Patent US 12,743,407
Granted Patent B2
US 12,743,407 · App. 18/315,530 · Granted Sep 22, 2026

Processing data in a data format with key-value pairs

Inventor: Ke Du (Xi'An, CN)
Assignee: International Business Machines Corporation
G06F16/2246G06F16/2365
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,743,407
App. No.
18/315,530
Granted
Sep 22, 2026
Kind
B2
Abstract

Disclosed are a computer-implemented method, a computer system and a computer program product for processing data in a data format with key-value pairs. All keys can be extracted from a plurality of data objects of the data. Duplicated keys can be removed such that only one of the same keys remains. The data in the data format can be reorganized into a format with a key portion followed by a value portion. The extracted keys can be arranged in the key portion in a predetermined order. The values of each of the plurality of data objects can be arranged in the value portion in an order corresponding to the order of the keys.

Claims (47)

1 . A computer-implemented method for processing data in a data format with key-value pairs, comprising:

extracting all keys from a plurality of data objects, wherein duplicated keys are removed such that only one of a same key remains;

reorganizing the data in the data format into a format with a key portion followed by a value portion, wherein the extracted keys are arranged in the key portion in a predetermined order, and the values of each of the plurality of data objects are arranged in the value portion in an order corresponding to the order of the keys;

grouping the values corresponding to each of the extracted keys for the plurality of data objects into a respective value group, wherein the respective value groups are arranged in respective value partitions of the value portion in the order corresponding to the order of the keys;

determining a reused part shared by the values of different data objects of the plurality of data objects, wherein the reused part of the values sharing the reused part in the value portion is replaced by a reference to the reused part, wherein the value portion comprises a reused-part partition in which the reused part is arranged in a position corresponding to a reference to the reused part;

arranging all respective value groups in a value portion in an order corresponding to an order of the keys, such that a consistent order is used for the keys and values, wherein the reused part is arranged in the reused-part partition as one of a list of reused parts shared by the values of the different data objects; and

compressing data into a reorganized data format with the key portion followed by a value portion.

2 . The computer-implemented method of claim 1 , wherein the respective value groups are arranged in the value portion in the order corresponding to the order of the keys.

3 . The computer-implemented method of claim 1 , wherein the reference to the reused part comprises the position of the reused part within the list of the reused-part partition.

4 . The computer-implemented method of claim 1 , wherein the value partitions of the value portion following the reused-part partition, wherein the reused part of the values sharing the reused part in a respective value partition is replaced by the reference to the reused part.

5 . The computer-implemented method of claim 1 , further comprising:

identifying one of the keys of the plurality of data objects of the data based on a search keyword;

determining a search range within the value portion corresponding to the identified key based on the order of the keys; and

searching for the search keyword within the search range corresponding to the identified key.

6 . The computer-implemented method of claim 1 , wherein the predetermined order for the extracted keys is an order of traversing leaf nodes of a tree structure of the data in the data format with key-value pairs.

7 . The computer-implemented method of claim 1 , further comprising: storing the data in the reorganized format.

8 . The computer-implemented method of claim 1 , further comprising: data format comprises JavaScript Object Notation (JSON).

9 . A computer system for processing data in a data format with key-value pairs, comprising:

one or more computer processors, one or more computer-readable storage media, and program instructions stored on the one or more of the computer-readable storage media for execution by at least one of the one or more processors, wherein the computer system is capable of performing a method comprising:

extracting all keys from a plurality of data objects, wherein duplicated keys are removed such that only one of a same key remains;

reorganizing the data in the data format into a format with a key portion followed by a value portion, wherein the extracted keys are arranged in the key portion in a predetermined order, and the values of each of the plurality of data objects are arranged in the value portion in an order corresponding to the order of the keys;

grouping the values corresponding to each of the extracted keys for the plurality of data objects into a respective value group, wherein the respective value groups are arranged in respective value partitions of the value portion in the order corresponding to the order of the keys;

determining a reused part shared by the values of different data objects of the plurality of data objects, wherein the reused part of the values sharing the reused part in the value portion is replaced by a reference to the reused part; determining a reused part shared by the values of different data objects of the plurality of data objects, wherein the reused part of the values sharing the reused part in the value portion is replaced by a reference to the reused part, wherein the value portion comprises a reused-part partition in which the reused part is arranged in a position corresponding to a reference to the reused part;

arranging all respective value groups in a value portion in an order corresponding to an order of the keys, such that a consistent order is used for the keys and values wherein the reused part is arranged in the reused-part partition as one of a list of reused parts shared by the values of the different data objects; and

compressing data into a reorganized data format with the key portion followed by a value portion.

10 . The computer system of claim 9 , wherein the respective value groups are arranged in the value portion in the order corresponding to the order of the keys.

11 . The computer system of claim 9 , wherein the reference to the reused part comprises the position of the reused part within the list of the reused-part partition.

12 . The computer system of claim 9 , wherein the reused part of the values sharing the reused part in a respective value partition is replaced by the reference to the reused part.

13 . The computer system of claim 9 , further comprising:

identifying one of the keys of the plurality of data objects of the data based on a search keyword;

determining a search range within the value portion corresponding to the 1 identified key based on the order of the keys; and

searching for the search keyword within the search range corresponding to the identified key.

14 . The computer system of claim 9 , wherein the predetermined order for the extracted keys is an order of traversing leaf nodes of a tree structure of the data in the data format with key-value pairs.

15 . The computer system of claim 9 , further comprising: storing the data in the reorganized format.

16 . The computer system of claim 9 , further comprising: data format comprises JavaScript Object Notation (JSON).

17 . A computer program product for processing data in a data format with key-value pairs, the computer program product comprising:

one or more computer-readable storage media and program instructions stored on the one or more computer-readable storage media, the program instructions executable by a computing system to cause the computing system to perform a method comprising:

extracting all keys from a plurality of data objects, wherein duplicated keys are removed such that only one of a same key remains;

reorganizing the data in the data format into a format with a key portion followed by a value portion, wherein the extracted keys are arranged in the key portion in a predetermined order, and the values of each of the plurality of data objects are arranged in the value portion in an order corresponding to the order of the keys;

grouping the values corresponding to each of the extracted keys for the plurality of data objects into a respective value group, wherein the respective value groups are arranged in respective value partitions of the value portion in the order corresponding to the order of the keys;

determining a reused part shared by the values of different data objects of the plurality of data objects, wherein the reused part of the values sharing the reused part in the value portion is replaced by a reference to the reused part determining a reused part shared by the values of different data objects of the plurality of data objects, wherein the reused part of the values sharing the reused part in the value portion is replaced by a reference to the reused part, wherein the value portion comprises a reused-part partition in which the reused part is arranged in a position corresponding to a reference to the reused part;

arranging all respective value groups in a value portion in an order corresponding to an order of the keys, such that a consistent order is used for the keys and values wherein the reused part is arranged in the reused-part partition as one of a list of reused parts shared by the values of the different data objects; and

compressing data into a reorganized data format with the key portion followed by a value portion.

18 . The computer program product of claim 17 , further comprising:

identifying one of the keys of the plurality of data objects of the data based on a search keyword;

determining a search range within the value portion corresponding to the identified key based on the order of the keys; and

searching for the search keyword within the search range corresponding to the identified key.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2023
From: DU, KE
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 063608/0780 →
Continuity (1)
Related Publication 20240378183A1 · Nov 14, 2024
References Cited (18)
US 8190835B1 · Yueh · 2012 [cited by examiner]
US 11146286B2 · Shetty · 2021 [cited by applicant]
US 11178212B2 · Maurer · 2021 [cited by applicant]
US 20100211572A1 · Beyer · 2010 [cited by examiner]
US 20190004726A1 · Li · 2019 [cited by examiner]
US 20220309549A1 · Xu · 2022 [cited by examiner]
US 20230282013A1 · Raad · 2023 [cited by examiner]
CN 103401562A · 2016 [cited by applicant]
CN 108156173A · 2018 [cited by applicant]
CN 109450450A · 2019 [cited by applicant]
CN 110247665A · 2019 [cited by applicant]
CN 111342933A · 2022 [cited by applicant]
CN 114665887A · 2022 [cited by applicant]
Cao et al. TDDFS: A Tier-Aware Data Deduplication-Based File System. ACM Transaction on Storage 15:1, 2019, pp. 1-26. (Year: 2019). [cited by examiner]
Data deduplication vs. data compression. 2022, pp. 1-10. https://www.lytics.com/blog/data-deduplication-vs-data-compression/. ( Year: 2022). [cited by examiner]
Storer et al. Secure data deduplication. StorageSS'08, pp. 1-10. (Year: 2008). [cited by examiner]
Buono et al., “Enhance Inter-service Communication in Supersonic K-Native REST-based Java Microservice Architectures”, Kristianstad University Sweden, Spring Semester 2021, 56 pages. [cited by applicant]
Muller, “JSON Compression”, heap.ch/blog/2016/06/19/json-compression/, Jun. 19, 2016, 7 pages. [cited by applicant]