IP Library Granted Patent US 8,473,451
Granted Patent B1
US 8,473,451 · App. 11/086,954 · Granted Jun 25, 2013

Preserving privacy in natural language databases

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,473,451
App. No.
11/086,954
Granted
Jun 25, 2013
Kind
B1
Abstract

An apparatus and a method for preserving privacy in natural language databases are provided. Natural language input may be received. At least one of sanitizing or anonymizing the natural language input may be performed to form a clean output. The clean output may be stored.

Claims (40)

1. A method comprising:

selecting a transcription of a spoken natural language input from a speaker, the transcription comprising sensitive information and non-sensitive information;

sanitizing the sensitive information, to form a clean transcription of the spoken natural language input, the clean transcription having sanitized text and non-sanitized text;

anonymizing the non-sanitized text in the clean transcription by modifying a first feature vector associated with the non-sanitized text, to yield a modified first feature vector, such that the modified first feature vector is the same as a second feature vector of a different document from the transcription, to yield a clean anonymous text; and

storing the clean anonymous text.

2. The method of claim 1 , wherein sanitizing the sensitive information further comprises:

finding a named entity in the spoken natural language input; and

performing, on the named entity, one of value distortion, value disassociation, and value class membership to preserve privacy in a spoken natural language database.

3. The method of claim 2 , wherein sanitizing the sensitive information further comprises:

performing, on the named entity, two of value distortion, value disassociation, and value class membership to preserve privacy in the spoken natural language database.

4. The method of claim 2 , further comprising performing value class membership by replacing a value with a generic token.

5. The method of claim 4 , wherein performing value class membership further comprises:

placing an indication of one of a gender and other information in the generic token when the value represents an identification of a person.

6. The method of claim 2 , wherein finding the named entity further comprises using an automated approach to detect a name.

7. The method of claim 1 , wherein anonymizing further comprises replacing a word in the non-sanitized text with a corresponding synonym.

8. The method of claim 1 , wherein anonymizing further comprises altering a plurality of stored documents, each stored document of the plurality of stored documents comprising the clean transcription, to change corresponding feature vectors of each of the plurality of stored documents to match.

9. The method of claim 1 , wherein sanitizing the sensitive information further comprises changing a distribution of utterances and semantically labeled data in the non-sanitized text.

10. A system comprising:

a processor; and

a storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

selecting a transcription of a spoken natural language input from a speaker, the transcription comprising sensitive information and non-sensitive information;

sanitizing the sensitive information, to form a clean transcription of the spoken natural language input, the clean transcription having sanitized text and non-sanitized text;

anonymizing the non-sanitized text in the clean transcription by modifying a first feature vector associated with the non-sanitized text, to yield a modified first feature vector, such that the modified first feature vector is the same as a second feature vector of a different document from the transcription yield a clean anonymous text; and

storing the clean anonymous text.

11. The system of claim 10 , wherein sanitizing the sensitive information further comprises:

finding a named entity in the spoken natural language input; and

performing, the named entity, one of value distortion, value disassociation, and value class membership to preserve privacy in a spoken natural language database.

12. The system of claim 11 , wherein sanitizing the sensitive information further comprises performing, on the named entity, at least two of value distortion, value disassociation, and value class membership to preserve privacy in the spoken natural language database.

13. The system of claim 11 , the storage medium having additional instructions stored which result in the operations further comprising performing value class membership by replacing a value with a generic token.

14. The system of claim 11 , the storage medium having additional instructions stored which result in the operations further comprising marking the named entity using a tag.

15. The system of claim 13 , the storage medium having additional instructions stored which result in the operations further comprising performing value class membership by placing an indication of one of a gender and other information in the generic token when the replaced value represents an identification of a person.

16. The system of claim 11 , wherein finding the named entity further comprises using an automated approach to detect names.

17. The system of claim 10 , the storage medium having additional instructions stored which result in the operations further comprising performing anonymizing by using an automated approach to detect names.

18. The system of claim 10 , the storage medium having additional instructions stored which result in the operations further comprising altering a plurality of stored documents which comprise the clean transcription to change corresponding feature vectors of each of the plurality of stored documents to match.

19. The system of claim 10 , the storage medium having additional instructions stored which result in the operations further comprising performing sanitizing by changing a distribution of utterances and semantically labeled data in the non-sanitized text.

20. A non-transitory computer-readable storage medium having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

selecting a transcription of a spoken natural language input from a speaker, the transcription comprising sensitive information and non-sensitive information;

sanitizing the sensitive information, to form a clean transcription of the spoken natural language input, the clean transcription having sanitized text and non-sanitized text;

anonymizing the non-sanitized text in the clean transcription by modifying a first feature vector associated with the non-sanitized text, to yield a modified first feature vector, such that the modified first feature vector is the same as a second feature vector of a different document from the transcription to yield a clean anonymous text; and

storing the clean anonymous text.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065533/0389 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 038275/0238 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 038275/0310 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 22, 2005
From: HAKKANI-TUR, DILEK Z.; SAYGIN, YUCEL; TANG, MIN; TUR, GOKHAN
To: AT&T CORP.
Reel/Frame 016414/0026 →