IP Library Granted Patent US 9,058,813
Granted Patent B1
US 9,058,813 · App. 13/624,755 · Granted Jun 16, 2015

Automated removal of personally identifiable information

Inventor: Scott I. Blanksteen (Issaquah, WA)
Assignee: Rawles LLC
G10L15/19
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,058,813
App. No.
13/624,755
Granted
Jun 16, 2015
Kind
B1
Abstract

A natural language system may receive user-input. The user-input may include personal or restrictable information. The natural language system may provide a dual processing system. The natural language system may store a true copy of the user-input, which may include the personal or restrictable information. The natural language system may also generate an obfuscated copy of the user-input that does not contain personal or restricted information. The true copy of the user-input may be stored in a secure storage system and may be retrieved by authorized personnel, which may include the user who provided the user-input. The obfuscated copy of the user-input may be stored in a storage system and may be employed in ongoing training of the natural language system.

Claims (72)

1. A method of processing natural language communications comprising:

receiving, via a natural language system, user-input that includes restricted information spoken by a user;

identifying the restricted information in the user-input, the identifying comprising:

determining that one or more words of the user-input match one or more words of a restricted string;

determining, at least partly in response to determining that the one or more words of the user-input match the one or more words of the restricted string, a context of the one or more words of the user-input; and

determining that the one or more words of the user-input comprise the restricted information based at least in part on the context and based at least in part on the one or more words of the user-input matching the one or more words of the restricted string;

generating obfuscated user-input by removing the restricted information from the user-input;

storing the obfuscated user-input in a datastore; and

training the natural language system with training data that includes the obfuscated user-input.

2. The method of claim 1 , wherein the restricted information includes at least one of a name of a person, a telephone number, an address, financial information, a birth date, and a personal identifier.

3. The method of claim 1 , further comprising:

training the natural language system to recognize restricted information based at least in part on one or more of speech-cadences of restricted information contained in user-inputs, structures of restricted information contained in user-inputs, and flagged information contained in user-inputs, wherein the flagged information corresponds to restricted information that is flagged by humans.

4. A method of processing natural language communications comprising:

identifying, via a natural language system, restricted information in a user-input, the identifying of the restricted information comprising:

identifying a structure of a restrictable-phrase candidate included in the user-input;

comparing the structure of the restrictable-phrase candidate to a restricted structure; and

determining that the restrictable-phrase candidate includes the restricted information based at least in part on the comparing of the structure of the restrictable-phrase candidate to the restricted structure;

generating obfuscated user-input in which the restricted information is not discernible;

providing at least a portion of the user-input to a first datastore; and

providing the obfuscated user-input to a second datastore.

5. The method of claim 4 , further comprising:

retrieving the obfuscated user-input from the second datastore; and

training the natural language system with training data that includes at least a portion the obfuscated user-input.

6. The method of claim 4 , further comprising:

retrieving the user-input from the first datastore; and

training the natural language system with training data that includes at least a portion the user-input.

7. The method of claim 6 , further comprising training the natural language system to recognize restricted information based at least in part on one or more of speech-cadences of restricted information contained in user-inputs, structures of restricted information contained in user-inputs, and flagged information contained in user-inputs, wherein the flagged information corresponds to restricted information that is flagged by humans.

8. The method of claim 4 , further comprising:

restricting access to the first datastore such that unauthorized personnel of an entity that controls the first datastore are unable to access the user-input.

9. The method of claim 8 , further comprising:

providing the user-input to a person that provided the user-input.

10. The method of claim 4 , wherein the generating obfuscated user-input includes:

omitting the restricted information from the obfuscated user-input.

11. The method of claim 4 , wherein the generating obfuscated user-input includes:

replacing the restricted information with non-restricted information.

12. The method of claim 11 , wherein the replacing the restricted information with non-restricted information includes:

determining a context, within the user-input, of the restricted information; and

determining the non-restricted information based at least in part on the context of the restricted information.

13. The method of claim 4 , wherein the user-input is a first user-input and the obfuscated user-input is a first obfuscated user-input, and further comprising:

employing a machine-learning model that has been trained with training data that includes at least one second obfuscated user-input, wherein the second obfuscated user-input corresponds to a second user-input in which restricted information of the second user-input is obfuscated.

14. The method of claim 4 , further comprising:

determining whether a restrictable string candidate is associated with a restricted string;

determining a context, within the user-input, of the restrictable string candidate in response to the restrictable string candidate matching the restricted string; and

determining that the restrictable string candidate is a restricted string based at least in part on the context of the restrictable string candidate.

15. The method of claim 4 , further comprising:

obtaining user-data that is independent of the user-input;

determining a restrictable-word candidate based at least in part on the user-data;

determining a context, within the user-input, of the restrictable-word candidate; and

determining that the restrictable-word candidate is a restricted key word based at least in part on the context of the restrictable-word candidate.

16. The method of claim 15 , wherein the user-input was provided by a user, wherein the user-data comprises data from at least one of an address book of the user, a calendar of the user, and a location of the user.

17. One or more non-transitory computer-readable storage media having computer-executable instructions thereon which, when executed by a computing device, implement a method comprising:

identifying restricted information in a user-input;

determining, from the user-input, a context of the restricted information;

determining, based at least in part on the context of the restricted information, non-restricted information to substitute for the restricted information;

generating an obfuscated user-input by replacing the restricted information in the user-input with the non-restricted information;

providing at least a portion of the user-input to a first datastore; and

providing the obfuscated user-input to a second datastore.

18. The one or more non-transitory computer-readable storage media of claim 17 , the method further comprising:

retrieving the obfuscated user-input from the second datastore; and

training a natural language system with training data that includes at least a portion the obfuscated user-input.

19. A natural language system comprising:

one or more processors; and

one or more computer-readable media storing computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform acts comprising:

identifying restricted information in a user-input;

determining, from the user-input, a context of the restricted information;

determining, based at least in part on the context of the restricted information, non-restricted information to substitute for the restricted information;

generating an obfuscated user-input by replacing the restricted information in the user-input with the non-restricted information;

providing at least a portion of the user-input to a first datastore; and

providing the obfuscated user-input to a second datastore.

20. The natural language system of claim 19 , the acts further comprising:

retrieving the obfuscated user-input from the second datastore; and

training a natural language system with training data that includes at least a portion the obfuscated user-input.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 12, 2015
From: RAWLES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 037103/0084 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 21, 2013
From: BLANKSTEEN, SCOTT I.
To: RAWLES LLC
Reel/Frame 029851/0065 →