IP Library Granted Patent US 12,488,135
Granted Patent B2
US 12,488,135 · App. 18/615,455 · Granted Dec 2, 2025

System and method for watermarking tabular data while obscuring underlying data for improving data integrity and security

Inventors: Vamsi Krishna Potluru (New York, NY); Tucker Richard Balch (Suwanee, GA); Manuela Veloso (New York, NY)
Assignee: JPMORGAN CHASE BANK, N.A.
G06F21/6227G06F21/64
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,135
App. No.
18/615,455
Granted
Dec 2, 2025
Kind
B2
Abstract

A method and system for watermarking a dataset generated by a source system are disclosed. The method includes acquiring the dataset, distributing data elements included in the dataset over a range, and dividing the range into multiple bins according to a scheme. The method further includes designating each of the bins as a first or second type, tagging each data element according to according to a bin type of a bin the respective data element falls into. For each data element included in a bin of the second type, selecting a new value by sampling within a nearest bin of the first type and replacing the respective data element with a replacement data element including the new value, and watermarking each of the data elements originally included in the bins of the first type and replacement data elements for generating a watermarked dataset.

Claims (71)

1 . A method for watermarking a dataset generated by a source system, the method comprising:

acquiring, by a processor, the dataset generated by the source system;

distributing, by the processor, a plurality of data elements included in the dataset over a range;

dividing, by the processor, the range into a plurality of bins according to a scheme among a plurality of schemes;

designating, by the processor, each of the plurality of bins as a first type or a second type;

tagging, by the processor, each data element among the plurality of data elements according to a bin type of a bin the respective data element falls into;

for each data element included in a bin of the second type, selecting a new value by sampling within a nearest bin of the first type and replacing the respective data element with a replacement data element including the new value; and

watermarking, by the processor, each data element originally included in the bins of the first type and replacement data elements for generating a watermarked dataset.

2 . The method according to claim 1 , wherein the data values are distributed over the range using a scheme among the plurality of schemes stored in a database.

3 . The method according to claim 2 , wherein the plurality of schemes include a fixed scheme, a random scheme, a mathematical relationship scheme and a statistical distribution scheme.

4 . The method according to claim 1 , further comprising:

selecting, by the processor, the scheme among the plurality of schemes based on the dataset generated by the source system.

5 . The method according to claim 1 , wherein a number of bins of the plurality of bins is determined based on the plurality of data elements of the dataset.

6 . The method according to claim 1 , wherein a number of bins of the plurality of bins is determined based on size of one or more bins of the plurality of bins.

7 . The method according to claim 1 , wherein, in the designating, each of the plurality of bins is randomly designated as the first type or the second type.

8 . The method according to claim 1 , wherein the new value is selected using a scheme among the plurality of schemes.

9 . The method according to claim 8 , wherein the scheme is a statistical distribution of the plurality of data elements.

10 . The method according to claim 8 , wherein the scheme is a mathematical relationship with respect to the plurality of data elements.

11 . The method according to claim 8 , wherein the scheme is a random generation.

12 . The method according to claim 1 , wherein the watermarking is performed on pairs of columns of data elements, and

wherein each of the pairs of columns of data elements include a seed column.

13 . The method according to claim 1 , further comprising:

adding, by the processor, one or more artificial data elements corresponding to bins of the second type prior to the watermarking.

14 . The method according to claim 13 , wherein the watermarking with the one or more artificial data elements generates a soft watermarked dataset.

15 . The method according to claim 1 , further comprising computing a score that the dataset is watermarked.

16 . The method according to claim 15 , wherein the score that the dataset is watermarked is calculated according to a below relationship:

z

=

2

(

m

g

-

k

m

/

2

)

/

mk

wherein

z is the score,

m g is a number of data elements that have been assigned to the bin of the first type,

k is a number of numerical columns in the dataset, and

m is a number of the data elements included in the dataset.

17 . The method according to claim 16 , further comprising:

comparing the score that the dataset is watermarked against a reference threshold.

18 . The method according to claim 17 , wherein

when the score that the dataset is watermarked is greater than the reference threshold, labeling the dataset as watermarked, and

when the score that the dataset is watermarked is less than the reference threshold, selecting a different scheme among the plurality of schemes until the score that the dataset is watermarked is calculated to be greater than the reference threshold.

19 . A system for watermarking a dataset generated by a source system, the system comprising:

a memory; and

a processor,

wherein the system is configured to perform:

acquiring the dataset generated by the source system;

distributing a plurality of data elements included in the dataset over a range;

dividing the range into a plurality of bins according to a scheme among a plurality of schemes;

designating each of the plurality of bins as a first type or a second type;

tagging each data element among the plurality of data elements according to a bin type of a bin the respective data element falls into;

for each data element included in a bin of the second type, selecting a new value by sampling within a nearest bin of the first type and replacing the respective data element with a replacement data element including the new value; and

watermarking each data element originally included in the bins of the first type and replacement data elements for generating a watermarked dataset.

20 . A non-transitory computer readable storage medium that stores a computer program for watermarking a dataset generated by a source system, the computer program, when executed by a processor, causing a system to perform a plurality of processes comprising:

acquiring the dataset generated by the source system;

distributing a plurality of data elements included in the dataset over a range;

dividing the range into a plurality of bins according to a scheme among a plurality of schemes;

designating each of the plurality of bins as a first type or a second type;

tagging each data element among the plurality of data elements according to a bin type of a bin the respective data element falls into;

for each data element included in a bin of the second type, selecting a new value by sampling within a nearest bin of the first type and replacing the respective data element with a replacement data element including the new value; and

watermarking each data element originally included in the bins of the first type and replacement data elements for generating a watermarked dataset.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 7, 2024
From: POTLURU, VAMSI KRISHNA; BALCH, TUCKER RICHARD; VELOSO, MANUELA
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 068207/0730 →
Continuity (1)
Related Publication 20250298918A1 · Sep 25, 2025
References Cited (8)
US 10373299B1 · Holub · 2019 [cited by examiner]
US 11769241B1 · Rhoads · 2023 [cited by examiner]
US 20030223584A1 · Bradley · 2003 [cited by examiner]
US 20040228502A1 · Bradley · 2004 [cited by examiner]
US 20110044494A1 · Bradley · 2011 [cited by examiner]
US 20160217547A1 · Stach · 2016 [cited by examiner]
US 20180210643A1 · Ghassabian · 2018 [cited by examiner]
US 20190287226A1 · Holub · 2019 [cited by examiner]