IP Library Granted Patent US 11,763,093
Granted Patent B2
US 11,763,093 · App. 17/245,774 · Granted Sep 19, 2023

Systems and methods for a privacy preserving text representation learning framework

Inventors: Ghazaleh Beigi (Tempe, AZ); Kai Shu (Mesa, AZ); Ruocheng Guo (Sichuan, CN); Suhang Wang (Mesa, AZ); Huan Liu (Tempe, AZ)
Assignee: Arizona Board of Regents on Behalf of Arizona State University
G06F40/30G06F18/21322G06F21/6245G06F40/211G06N3/08G06F18/21326
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,763,093
App. No.
17/245,774
Granted
Sep 19, 2023
Kind
B2
Abstract

Various embodiments of a computer-implemented system which learns textual representations while filtering out potentially personally identifying data and retaining semantic meaning within the textual representations are disclosed herein.

Claims (90)

1. A method of generating a modified latent text representation for a document, comprising:

utilizing a processor in communication with a tangible storage medium storing instructions that are executed by the processor to perform operations comprising:

generating an initial latent representation representative of text in a document;

inferring an amount of noise to be added to the initial latent representation by finding a privacy budget value that minimizes a loss between a predicted semantic label and a ground truth semantic label for the initial latent representation and maximizes a loss between a predicted private attribute label and a ground truth private attribute label for the initial latent representation;

adding the amount of noise to the initial latent text representation to generate the modified latent text representation;

optimize a set of private attribute discriminator parameters and the privacy budget value that maximize the loss between the predicted private attribute label and the ground truth private attribute label for the initial latent representation;

generating the predicted private attribute label by applying a classifier to the modified latent representation;

maximizing a loss between the predicted private attribute label and the ground truth private attribute label; and

selecting the set of private attribute discriminator classifier parameters associated with the lowest loss value between the predicted private attribute label and the ground truth private attribute label.

2. The method of claim 1 , wherein the initial latent representation is generated using an auto-encoder trained to generate the initial latent representation from the text in the document.

3. The method of claim 2 , further comprising training the autoencoder by:

generating an encoded latent representation representative of the text in the document by applying an encoder to the document;

constructing a reconstructed document including reconstructed text representative of the text in the encoded latent representation by applying a decoder to the encoded latent representation; and

identifying a plurality of autoencoder parameters that minimize a loss between the text in the document and the reconstructed text in the reconstructed document.

4. The method of claim 1 , further comprising:

optimizing a set of semantic discriminator classifier parameters and the privacy budget value that minimize the loss between the predicted semantic label and the ground truth semantic label for the initial latent representation.

5. The method of claim 4 , further comprising:

adding a first amount of noise to the initial latent representation to generate the modified latent representation.

6. The method of claim 4 , further comprising:

generating the predicted semantic label by applying a second classifier to the modified latent representation;

minimizing a loss between the predicted semantic label and the ground truth semantic label; and

identifying the set of semantic discriminator classifier parameters associated with the lowest loss value between the predicted semantic label and the ground truth semantic label.

7. The method of claim 1 , further comprising:

selecting the privacy budget value that is associated with a lowest loss value between the predicted semantic label and the ground truth semantic label and that is associated with a highest loss value between the predicted private attribute label and the ground truth private attribute label.

8. The method of claim 1 , wherein the optimization of the set of private attribute discriminator parameters is modeled as a minmax game.

9. The method of claim 1 , further comprising:

determining the amount of noise to add based on the privacy budget value by sampling a value r from a uniform distribution such that:

s

(

i

)

=

-

Δ

ϵ

×

sgn

(

r

)

ln

(

1

-

2

|

r

|

)

,

i

=

1

,

2

,

.

.

.

,

d

wherein ∈ is the privacy budget value, Δ is an L 1 -sensitivity of the initial latent representation, d is a dimension of the initial latent representation, s is a noise vector and s(i) is the i-th element for noise vector s.

10. The method of claim 1 , wherein the step of finding the privacy budget value is iteratively repeated until convergence.

11. The method of claim 10 , wherein the process of adding the amount of noise to the initial latent text representation to generate the modified latent text representation runs concurrently with finding the privacy budget value.

12. A computer system for generating a modified latent text representation for a document, comprising:

at least one processor in communication with a memory and operable for execution of a plurality of modules, the plurality of modules including:

an auto-encoder configured to generate an initial latent representation representative of text in a document;

a noise adder module configured to receive the initial latent representation and add an amount of noise to the initial latent representation to generate a modified latent text representation based on a privacy budget value;

a semantic meaning discriminator module configured to optimize a set of semantic discriminator classifier parameters and the privacy budget value such that a loss is minimized between the predicted semantic label and the ground truth semantic label for the initial latent representation; and

a private attribute discriminator module configured to optimize a set of private attribute discriminator parameters and the privacy budget value such that a loss is maximized between the predicted private attribute label and the ground truth private attribute label for the initial latent representation.

13. The computer system of claim 12 , wherein the auto-encoder module is configured to:

generate an encoded latent representation representative of the text in the document by applying an encoder to the document;

construct a reconstructed document including reconstructed text representative of the text in the encoded latent representation by applying a decoder to the encoded latent representation; and

identify a plurality of autoencoder parameters that minimize a loss between the text in the document and the reconstructed text in the reconstructed document.

14. The computer system of claim 12 , wherein the semantic meaning discriminator module is configured to:

generate the predicted semantic label by applying a first classifier to the modified latent representation;

minimize the loss between the predicted semantic label and the ground truth semantic label; and

identify the set of semantic discriminator classifier parameters associated with the lowest loss value between the predicted semantic label and the ground truth semantic label.

15. The computer system of claim 14 , wherein the first classifier is implemented using a recurrent neural network that takes the set of semantic discriminator classifier parameters and the modified latent text representation as input.

16. The computer system of claim 12 , wherein the private attribute discriminator module is configured to:

generate the predicted private attribute label by applying a second classifier to the modified latent representation;

maximize a loss between the predicted private attribute label and the ground truth private attribute label; and

select the set of private attribute discriminator classifier parameters associated with the lowest loss value between the predicted private attribute label and the ground truth private attribute label.

17. The computer system of claim 16 , wherein the second classifier is implemented using a recurrent neural network that takes the set of private attribute discriminator classifier parameters and the modified latent text representation as input.

18. The computer system of claim 16 , wherein the privacy budget value is associated with a lowest loss value between the predicted semantic label and the ground truth semantic label and a highest loss value between the predicted private attribute label and the ground truth private attribute label.

Assignments (2)
CONFIRMATORY LICENSE Recorded Feb 23, 2022
From: ARIZONA STATE UNIVERSITY-TEMPE CAMPUS
To: NATIONAL SCIENCE FOUNDATION
Reel/Frame 059221/0530 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 10, 2021
From: BEIGI, GHAZALEH; SHU, KAI; GUO, RUOCHENG; WANG, SUHANG; LIU, HUAN
To: ARIZONA BOARD OF REGENTS ON BEHALF OF ARIZONA STATE UNIVERSITY
Reel/Frame 056187/0212 →
Continuity (2)
Provisional Application 63018287 · Apr 30, 2020
Related Publication 20210342546A1 · Nov 4, 2021