IP Library Granted Patent US 12,423,327
Granted Patent B2
US 12,423,327 · App. 17/609,765 · Granted Sep 23, 2025

Information processing apparatus, information processing method and program for anonymizing data

Inventor: Yoshiyuki Mihara (Tokyo, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G06F16/285G06F16/282
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,423,327
App. No.
17/609,765
Granted
Sep 23, 2025
Kind
B2
Abstract

With respect to an information processing device which anonymizes data composed of records including one or more items through statistical processing, the information processing device includes a memory, and a processor configured to classify respective records constituting the data into one or more first sets, based on masking target items, a dictionary, and a selected hierarchy level, classify the respective records into one or more second sets with respect to a number of records belonging to each of the one or more first sets, and calculate a number of records of each of the one or more second sets and a ratio of records belonging to each of the one or more second sets to the records constituting the data, change the selected hierarchy level based on the ratio and priority set in advance, and create a statistically processed record by statistically processing records belonging to a same first set.

Claims (49)

1. An information processing device which anonymizes data composed of records including one or more items through statistical processing, the information processing device comprising:

a memory, and

a processor configured to:

determine one or more sets of data by classifying records based at least on:

masking target items for marking items of the records,

a dictionary which expresses categories of item values of at least an item of the records in a tree structure for each of the masking target items, and

a selected hierarchy level indicating a hierarchy level selected among the hierarchy levels in the tree structure of the dictionary for each of the masking target items;

calculate a number of one or more records belonging to each of the one or more sets;

determine a set to be deleted and a set to be statistically processed among the one or more sets based on the number of the records belonging to each of the one or more sets and a predetermined number, wherein a number of one or more records belonging to the set to be deleted being less than the predetermined number, and a number of one or more records belonging to the set to be statistically processed being greater than or equal to the predetermined number; and

generate anonymized data by deleting the one or more records belonging to the set to be deleted and statistically processing the one or more records belonging to the set to be statistically processed, wherein anonymizing data comprises automatic determination of amount of data pertaining to identification of an individual to be deleted or replaced,

wherein the one or more sets are determined based on categories in the selected hierarchy level, and upon selection of a hierarchy level by a user, information of corresponding items expressed in a hierarchy level lower than the selected hierarchy level is masked.

2. The information processing device according to claim 1 ,

wherein the determining of the set to be deleted includes determining one or more sets to be deleted,

wherein the processor is further configured to calculate a sum of the number of the one or more records belonging to each of the one or more sets to be deleted and a ratio of the calculated sum to a number of the records, and

wherein the processor deletes the one or more records belonging to each of the one or more sets to be deleted, in response to determining that a predetermined termination condition is satisfied, the predetermined termination condition including a condition that the ratio is less than or equal to a predetermined ratio.

3. The information processing device according to claim 2 ,

wherein the processor is further configured to change the selected hierarchy level, and

wherein the changing, the determining of the one or more sets, the calculating, the determining of the set to be deleted, and the deleting are repeatedly performed until the predetermined termination condition is satisfied.

4. The information processing device according to claim 3 ,

wherein the changing is performed based on the ratio and priority set in advance.

5. The information processing device according to claim 4 ,

wherein the priority is values set for the masking target items.

6. The information processing device according to claim 4 ,

wherein the priority is one or more types of scores calculated using a predetermined method.

7. The information processing device according to claim 6 ,

wherein the priority is a sum or a weighted sum of a plurality of scores respectively calculated using predetermined methods, or a score determined by priority set for each of the plurality of scores.

8. The information processing device according to claim 1 ,

wherein, in the determining of the one or more sets, the item values of each of the masking target items are abstracted based on the categories in the selected hierarchy level, and

wherein a first set among the one or more sets includes records having same abstracted item values.

9. The information processing device according to claim 1 ,

wherein the statistically processing of the one or more records belonging to each of the one or more sets to be statistically processed includes combining the one or more records into one record and adding a new item indicating a number of the one or more records.

10. An information processing method to be performed by a computer which anonymizes data composed of records including one or more items through statistical processing, the information processing method comprising:

determining one or more sets of data by classifying records based at least on:

masking target items for masking items of the records,

a dictionary which expresses categories of item values of at least an item of the records in hierarchy levels in a tree structure for each of the masking target items, and

a selected hierarchy level indicating a hierarchy level selected among the hierarchy levels in the tree structure of the dictionary for each of the masking target items;

calculating a number of one or more records belonging to each of the one or more sets;

determining a set to be deleted and a set to be statistically processed among the one or more sets based on the number of the records belonging to each of the one or more sets and a predetermined number, wherein a number of one or more records belonging to the set to be deleted being less than the predetermined number, and a number of one or more records belonging to the set to be statistically processed being greater than or equal to the predetermined number; and

generating anonymized data by deleting the one or more records belonging to the set to be deleted and statistically processing the one or more records belonging to the set to be statistically processed, wherein anonymizing data comprises automatic determination of amount of data pertaining to identification of an individual to be deleted or replaced,

wherein the one or more sets are determined based on categories in the selected hierarchy level, and upon selection of a hierarchy level by a user, information of corresponding items expressed in a hierarchy level lower than the selected hierarchy level is masked.

11. A non-transitory computer-readable recording medium having stored therein program instructions for causing a computer to execute operations comprising:

determining one or more sets of data by classifying records based at least on:

masking target items for masking items of the records,

a dictionary which expresses categories of item values of at least an item of the records in hierarchy levels in a tree structure for each of the masking target items, and

a selected hierarchy level indicating a hierarchy level selected among the hierarchy levels in the tree structure of the dictionary for each of the masking target items;

calculating a number of one or more records belonging to each of the one or more sets;

determining a set to be deleted and a set to be statistically processed among the one or more sets based on the number of the records belonging to each of the one or more sets and a predetermined number, a number of one or more records belonging to the set to be deleted being less than the predetermined number, and a number of one or more records belonging to the set to be statistically processed being greater than or equal to the predetermined number; and

generating anonymized data by deleting the one or more records belonging to the set to be deleted and statistically processing the one or more records belonging to the set to be statistically processed, wherein anonymizing data comprises automatic determination of amount of data pertaining to identification of an individual to be deleted or replaced,

wherein the one or more sets are determined based on categories in the selected hierarchy level, and upon selection of a hierarchy level by a user, information of corresponding items expressed in a hierarchy level lower than the selected hierarchy level is masked.

Assignments (2)
CHANGE OF NAME Recorded Oct 22, 2025
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 073184/0535 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 8, 2021
From: MIHARA, YOSHIYUKI
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 058052/0182 →
Continuity (1)
Related Publication 20220229853A1 · Jul 21, 2022
References Cited (33)
US 20040111453A1 · Harris · 2004 [cited by examiner]
US 20050240572A1 · Sung et al. · 2005 [cited by applicant]
US 20100153184A1 · Caffrey et al. · 2010 [cited by applicant]
US 20110131530A1 · Oosterholt · 2011 [cited by examiner]
US 20110179011A1 · Cardno et al. · 2011 [cited by applicant]
US 20130138698A1 · Harada et al. · 2013 [cited by applicant]
US 20140244296A1 · Linn et al. · 2014 [cited by applicant]
US 20150033356A1 · Takenouchi · 2015 [cited by applicant]
US 20150101061A1 · Jonas · 2015 [cited by examiner]
US 20150169895A1 · Gkoulalas-Divanis · 2015 [cited by examiner]
US 20160070776A1 · Furusho · 2016 [cited by examiner]
US 20160379011A1 · Koike · 2016 [cited by examiner]
US 20170061156A1 · Hamamoto et al. · 2017 [cited by applicant]
US 20170068828A1 · Nishi · 2017 [cited by examiner]
US 20170126694A1 · Nerurkar · 2017 [cited by examiner]
US 20180012039A1 · Takahashi · 2018 [cited by examiner]
US 20180046679A1 · Sharifi Sedeh et al. · 2018 [cited by applicant]
US 20190066830A1 · Kumar · 2019 [cited by applicant]
US 20190147988A1 · Sharifi Sedeh et al. · 2019 [cited by applicant]
US 20190205905A1 · Raghunathan · 2019 [cited by examiner]
US 20200012886A1 · Walters · 2020 [cited by examiner]
US 20200311106A1 · Zhao · 2020 [cited by examiner]
US 20200334219A1 · Antonatos · 2020 [cited by examiner]
US 20210279366A1 · Choudhury · 2021 [cited by examiner]
US 20210382867A1 · Shiinoki · 2021 [cited by examiner]
US 20220215129A1 · Mihara · 2022 [cited by applicant]
US 20220222369A1 · Mihara · 2022 [cited by examiner]
US 20220247574A1 · Hoshino · 2022 [cited by examiner]
JP 2015046030A · 2015 [cited by applicant]
JP 2017049693A · 2017 [cited by applicant]
WO 2013121739A1 · 2013 [cited by applicant]
WO 2020235016A1 · 2020 [cited by applicant]
Nazumi et al. (2013) “One proposal regarding improvement of efficiency in k-anonymization method”, Information Processing Society of Japan, Collection of Papers of the 75-th National Convention, 2013(1), 519-520 (Mar. 6… [cited by applicant]