IP Library Granted Patent US 12,339,980
Granted Patent B2
US 12,339,980 · App. 17/431,719 · Granted Jun 24, 2025

Data replacement apparatus, data replacement method, and program

Inventor: Satoshi Hasegawa (Musashino, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G06F21/6218G06F16/285
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,339,980
App. No.
17/431,719
Granted
Jun 24, 2025
Kind
B2
Abstract

A data replacement apparatus that can perform microaggregation of large-scale data at high speed using only a primary storage device of a small capacity. The data replacement apparatus includes an attribute value set retrieval unit that retrieves a grouped attribute value set into a primary storage device when a size of the grouped attribute value set is equal to or smaller than a predefined size and retrieves the grouped attribute value set into a secondary storage device when the size of the grouped attribute value set is larger than the predefined size. Further, there is a median computation unit that computes a median of the grouped attribute value set at the primary storage device or at the secondary storage device and a division determination unit that sets respective ones of the two attribute value sets formed by the division as new groups.

Claims (24)

1. A data replacement apparatus for replacing attribute values with representative values for each of groups, the data replacement apparatus comprising:

attribute value set retrieval circuitry that retrieves a grouped attribute value set into a primary storage device when a size of the grouped attribute value set is equal to or smaller than a predefined size and retrieves the grouped attribute value set into a secondary storage device when the size of the grouped attribute value set is larger than the predefined size, wherein the primary storage device is physically separate from the secondary storage device, and the secondary storage device is slower than the primary storage device;

median computation circuitry that computes a median of the grouped attribute value set at the primary storage device or at the secondary storage device;

division determination circuitry that, if a size of each of two attribute value sets which are formed by dividing the grouped attribute value set into two parts based on the median is equal to or greater than a predetermined threshold, sets respective ones of the two attribute value sets formed by the division as new groups;

a joined set generation circuitry that generates a joined set which is formed by arranging record numbers associated with the attribute values such that the attribute values in each of the groups which have converged after repeated execution of processing by the attribute value set retrieval circuitry, the median computation circuitry, and the division determination circuitry are consecutive;

a rearrangement circuitry that rearranges the attribute values in the secondary storage device based on the joined set;

a representative value replacement circuitry that sequentially executes processing for retrieving some of the rearranged attribute values from the secondary storage device into the primary storage device, and replaces the attribute values retrieved into the primary storage device with the representative values; and

a re-rearrangement circuitry that moves the representative values to the secondary storage device and rearranges them into an original order.

2. A data replacement method for replacing attribute values with representative values for each of groups, the data replacement method comprising:

retrieving a grouped attribute value set into a primary storage device when a size of the grouped attribute value set is equal to or smaller than a predefined size and retrieving the grouped attribute value set into a secondary storage device when the size of the grouped attribute value set is larger than the predefined size, wherein the primary storage device is physically separate from the secondary storage device and the secondary storage device is slower than the primary storage device;

computing a median of the grouped attribute value set at the primary storage device or at the secondary storage device;

setting respective ones of the two attribute value sets formed by the division as new groups, when a size of each of two attribute value sets which are formed by dividing the grouped attribute value set into two parts based on the median is equal to or greater than a predetermined threshold;

generating a joined set which is formed by arranging record numbers associated with the attribute values such that the attribute values in each of the groups which have converged after repeated execution of the retrieving, the computing and the setting are consecutive;

rearranging the attribute values in the secondary storage device based on the joined set;

sequentially executing processing for retrieving some of the rearranged attribute values from the secondary storage device into the primary storage device, and replacing the attribute values retrieved into the primary storage device with the representative values; and

moving the representative values to the secondary storage device and rearranging them into an original order.

3. A non-transitory computer-readable storage medium storing a program for causing a computer to perform a data replacement method for replacing attribute values with representative values for each of groups, the data replacement method comprising:

retrieving a grouped attribute value set into a primary storage device when a size of the grouped attribute value set is equal to or smaller than a predefined size and retrieving the grouped attribute value set into a secondary storage device when the size of the grouped attribute value set is larger than the predefined size, wherein the primary storage device is physically separate from the secondary storage device, and the secondary storage device is slower than the primary storage device;

computing a median of the grouped attribute value set at the primary storage device or at the secondary storage device;

setting respective ones of the two attribute value sets formed by the division as new groups, when a size of each of two attribute value sets which are formed by dividing the grouped attribute value set into two parts based on the median is equal to or greater than a predetermined threshold;

generating a joined set which is formed by arranging record numbers associated with the attribute values such that the attribute values in each of the groups which have converged after repeated execution of the retrieving, the computing and the setting are consecutive;

rearranging the attribute values in the secondary storage device based on the joined set:

sequentially executing processing for retrieving some of the rearranged attribute values from the secondary storage device into the primary storage device, and replacing the attribute values retrieved into the primary storage device with the representative values; and

moving the representative values to the secondary storage device and rearranging them into an original order.

Assignments (2)
CHANGE OF NAME Recorded Aug 20, 2025
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 072801/0812 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 18, 2021
From: HASEGAWA, SATOSHI
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 057211/0075 →
Priority Claims (1)
JP 2019-043663 · Mar 11, 2019 · national
Continuity (1)
Related Publication 20220138338A1 · May 5, 2022
References Cited (9)
US 8751499B1 · Carasso · 2014 [cited by examiner]
US 11698901B1 · Porath · 2023 [cited by examiner]
US 20090182789A1 · Sandorfi · 2009 [cited by examiner]
US 20140189858A1 · Chen et al. · 2014 [cited by applicant]
US 20160379011A1 · Koike et al. · 2016 [cited by applicant]
JP 2016115112A · 2016 [cited by applicant]
Lefevre et al., “Mondrian Multidimensional K-Anonymity”, In Proceedings of the 22nd International Conference on Data Engineering, 2006, pp. 1-11. [cited by applicant]
LeFevre et al., “Workload-Aware Anonymization Techniques for Large-Scale Datasets”, ACM Transactions on Database Systems, vol. 33, No. 3, Article 17, Publication date: Aug. 2008, pp. 17:1-17:47, total 47 pages. [cited by applicant]
Zhang et al., “MRMondrian: Scalable Multidimensional Anonymisation for Big Data Privacy Preservation”, IEEE Transactions on Big Data, IEEE, vol. 8, No. 1, Jan./Feb. 2022, pp. 125-139, total 15 pages. [cited by applicant]