IP Library › Granted Patent US 12,417,506
Granted Patent B2
US 12,417,506 · App. 18/270,690 · Granted Sep 16, 2025

Generating fake documents using word embeddings to deter intellectual property theft

Inventors: Venkatramanan Subrahmanian (Evanston, IL); Dongkai Chen (Mountain View, CA); Haipeng Chen (Williamsburg, VA); Deepti Poluru (Hanover, NH); Almas Abdibayev (Hanover, NH)
Assignee: Trustees of Dartmouth College
G06Q50/184G06F16/355G06F18/22G06F40/166G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,506
App. No.
18/270,690
Granted
Sep 16, 2025
Kind
B2
Abstract

A computer-implemented method, system and computer program product for generating fake documents. A corpus of domain specific documents is built and word embeddings for each word in such documents are identified as embedding vectors. Concepts in the corpus are then clustered together by clustering the embedding vectors. A feasible candidate replacement set is generated for each concept using the clustered concepts in the corpus. After such pre-processing steps are accomplished, concepts are extracted from a document. The concept importance values are computed for these extracted concepts, in which the extracted concepts are clustered into bins based on such measurements. A joint optimization problem is solved to identify both the concepts in the document to be replaced using the clustered concepts in the bins as well as the corresponding replacement concepts obtained from the clustered concepts in the corpus. Such replacements are made to generate a fake document.

Claims (59)

1. A computer-implemented method for generating fake documents, the method comprising:

extracting concepts from a document;

computing concept importance values for said extracted concepts;

clustering said extracted concepts into bins according to said computed concept importance values;

solving a joint optimization problem to identify both concepts in said document to be replaced using said clustered concepts in said bins as well as corresponding replacement concepts obtained from clustered concepts in a corpus of domain specific documents, wherein said concepts in said corpus of domain specific documents are clustered by clustering embeddings vectors, wherein word embeddings are identified for each word in documents of said corpus of domain specific documents as said embedding vectors; and

replacing said identified concepts in said document with corresponding identified replacement concepts to generate a fake document.

2. The method as recited in claim 1 , wherein a constraint in said optimization problem requires a sum of a distance between a concept and its replacement exceeding a threshold.

3. The method as recited in claim 1 , wherein an objective function in said optimization problem minimizes a sum of term frequency, inverse document frequency values of selected concepts from a set of said bins.

4. The method as recited in claim 1 , wherein an objective function in said optimization problem uses (c, f, f′) triples to maximize a difference between whether concept c in said document is replaced in both fake files f, f′.

5. The method as recited in claim 1 , wherein an objective function in said optimization problem uses an average distance between a concept and feasible candidates in a feasible candidate replacement set, wherein said feasible candidates comprise concepts from said clustered concepts in said corpus of domain specific documents.

6. The method as recited in claim 1 , wherein an objective function in said optimization problem selects a concept to be replaced in said document by minimizing an average distance between said concept and feasible candidates in a feasible candidate replacement set, wherein said feasible candidates comprise concepts from said clustered concepts in said corpus of domain specific documents.

7. The method as recited in claim 1 , wherein an objective function in said optimization problem identifies (c, c′, f) triples to identify a concept c to be replaced with concept c′ in fake file f using integer linear programming.

8. The method as recited in claim 1 further comprising:

building said corpus of domain specific documents; and

eliminating words from said domain specific documents from being replaced using a parts of speech tagging tool.

9. The method as recited in claim 8 further comprising:

computing a concept importance value for a word representing a concept from said corpus of domain specific documents for each concept.

10. The method as recited in claim 1 further comprising:

generating a feasible candidate replacement set for each concept using said clustered concepts in said corpus of domain specific documents.

11. A computer program product for generating fake documents, the computer program product comprising one or more computer readable storage mediums having program code embodied therewith, the program code comprising programming instructions for:

extracting concepts from a document;

computing concept importance values for said extracted concepts;

clustering said extracted concepts into bins according to said computed concept importance values;

solving a joint optimization problem to identify both concepts in said document to be replaced using said clustered concepts in said bins as well as corresponding replacement concepts obtained from clustered concepts in a corpus of domain specific documents, wherein said concepts in said corpus of domain specific documents are clustered by clustering embeddings vectors, wherein word embeddings are identified for each word in documents of said corpus of domain specific documents as said embedding vectors; and

replacing said identified concepts in said document with corresponding identified replacement concepts to generate a fake document.

12. The computer program product as recited in claim 11 , wherein a constraint in said optimization problem requires a sum of a distance between a concept and its replacement exceeding a threshold.

13. The computer program product as recited in claim 11 , wherein an objective function in said optimization problem minimizes a sum of term frequency, inverse document frequency values of selected concepts from a set of said bins.

14. The computer program product as recited in claim 11 , wherein an objective function in said optimization problem uses (c, f, f′) triples to maximize a difference between whether concept c in said document is replaced in both fake files f, f′.

15. The computer program product as recited in claim 11 , wherein an objective function in said optimization problem uses an average distance between a concept and feasible candidates in a feasible candidate replacement set, wherein said feasible candidates comprise concepts from said clustered concepts in said corpus of domain specific documents.

16. The computer program product as recited in claim 11 , wherein an objective function in said optimization problem selects a concept to be replaced in said document by minimizing an average distance between said concept and feasible candidates in a feasible candidate replacement set, wherein said feasible candidates comprise concepts from said clustered concepts in said corpus of domain specific documents.

17. The computer program product as recited in claim 11 , wherein an objective function in said optimization problem identifies (c, c′, f) triples to identify a concept c to be replaced with concept c′ in fake file f using integer linear programming.

18. The computer program product as recited in claim 11 , wherein the program code further comprises the programming instructions for:

building said corpus of domain specific documents; and

eliminating words from said domain specific documents from being replaced using a parts of speech tagging tool.

19. The computer program product as recited in claim 18 , wherein the program code further comprises the programming instructions for:

computing a concept importance value for a word representing a concept from said corpus of domain specific documents for each concept.

20. The computer program product as recited in claim 11 , wherein the program code further comprises the programming instructions for:

generating a feasible candidate replacement set for each concept using said clustered concepts in said corpus of domain specific documents.

21. A system, comprising:

a memory for storing a computer program for generating fake documents; and

a processor connected to said memory, wherein said processor is configured to execute program instructions of the computer program comprising:

extracting concepts from a document;

computing concept importance values for said extracted concepts;

clustering said extracted concepts into bins according to said computed concept importance values;

solving a joint optimization problem to identify both concepts in said document to be replaced using said clustered concepts in said bins as well as corresponding replacement concepts obtained from clustered concepts in a corpus of domain specific documents, wherein said concepts in said corpus of domain specific documents are clustered by clustering embeddings vectors, wherein word embeddings are identified for each word in documents of said corpus of domain specific documents as said embedding vectors; and

replacing said identified concepts in said document with corresponding identified replacement concepts to generate a fake document.

22. The system as recited in claim 21 , wherein a constraint in said optimization problem requires a sum of a distance between a concept and its replacement exceeding a threshold.

23. The system as recited in claim 21 , wherein an objective function in said optimization problem minimizes a sum of term frequency, inverse document frequency values of selected concepts from a set of said bins.

24. The system as recited in claim 21 , wherein an objective function in said optimization problem uses (c, f, f′) triples to maximize a difference between whether concept c in said document is replaced in both fake files f, f′.

25. The system as recited in claim 21 , wherein an objective function in said optimization problem uses an average distance between a concept and feasible candidates in a feasible candidate replacement set, wherein said feasible candidates comprise concepts from said clustered concepts in said corpus of domain specific documents.

26. The system as recited in claim 21 , wherein an objective function in said optimization problem selects a concept to be replaced in said document by minimizing an average distance between said concept and feasible candidates in a feasible candidate replacement set, wherein said feasible candidates comprise concepts from said clustered concepts in said corpus of domain specific documents.

27. The system as recited in claim 21 , wherein an objective function in said optimization problem identifies (c, c′, f) triples to identify a concept c to be replaced with concept c′ in fake file f using integer linear programming.

28. The system as recited in claim 21 , wherein the program instructions of the computer program further comprise:

building said corpus of domain specific documents; and

eliminating words from said domain specific documents from being replaced using a parts of speech tagging tool.

29. The system as recited in claim 28 , wherein the program instructions of the computer program further comprise:

computing a concept importance value for a word representing a concept from said corpus of domain specific documents for each concept.

30. The system as recited in claim 21 , wherein the program instructions of the computer program further comprise:

generating a feasible candidate replacement set for each concept using said clustered concepts in said corpus of domain specific documents.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2025
From: ABDIBAYEV, ALMAS
To: TRUSTEES OF DARTMOUTH COLLEGE
Reel/Frame 071631/0127 →
GOVERNMENT INTEREST AGREEMENT Recorded Feb 21, 2025
From: DARTMOUTH COLLEGE
To: THE GOVERNMENT OF THE UNITED STATES OF AMERICA AS REPRESENTED BY THE SECRETARY OF THE NAVY
Reel/Frame 070291/0805 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2023
From: SUBRAHMANIAN, VENKATRAMANAN; CHEN, DONGKAI; CHEN, HAIPENG; POLURU, DEEPTI
To: TRUSTEES OF DARTMOUTH COLLEGE
Reel/Frame 064133/0130 →
Continuity (2)
Provisional Application 63133358 · Jan 3, 2021
Related Publication 20240095856A1 · Mar 21, 2024
References Cited (10)
US 10339423B1 · Dinerstein et al. · 2019 [cited by applicant]
US 10657259B2 · Lee · 2020 [cited by examiner]
US 20170132866A1 · Kuklinski et al. · 2017 [cited by applicant]
US 20180033020A1 · Viens et al. · 2018 [cited by applicant]
US 20210065355A1 · Atzmon et al. · 2021 [cited by applicant]
CA 3021168A1 · 2019 [cited by examiner]
WO WO2012075336A1 · 2012 [cited by examiner]
Bollegala et al., “Joint Word Representation Learning Using a Corpus and a Semantic Lexicon,” Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, 2016, available at: < https://cdn.aaai.org/ojs/10340… [cited by examiner]
International Search Report and Written Opinion for PCT/US2021/065209, mailed on Mar. 24, 2022. [cited by applicant]
Abdibayev et al. “Using Word Embeddings to Deter Intellectual Property Theft through Automated Generation of Fake Documents.” ACM Trans. Manage. Inf. Syst. 12, 2, Article 13 [online), pp. 1-22, Jan. 2021. [cited by applicant]