IP Library › Granted Patent US 11,138,974
Granted Patent B2
US 11,138,974 · App. 16/444,825 · Granted Oct 5, 2021

Privacy mode based on speaker identifier

Inventor: Zhenhua Wang (Seattle, WA)
Assignee: Amazon Technologies, Inc.
G10L15/22G06F3/167G06F21/64G10L15/26G10L17/00G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,138,974
App. No.
16/444,825
Granted
Oct 5, 2021
Kind
B2
Abstract

Techniques for configuring a speech processing system with a privacy mode that is associated with the identity of a user that activated the privacy mode are described. A user may speak an indication to have the speech processing system activate a privacy mode. When such an indication is detected by the speech processing system, the speech processing system determines an identity of the user, determines a unique system identifier associated with the user, and generates a privacy mode flag. The speech processing system then associates the privacy mode flag with the user's unique system identifier. The privacy mode flag indicates to components of the speech processing system that any data related to processing of the user's utterances should not be sent to long term storage, thus causing various components of the system to delete data once the respective component is finished processing with respect to an utterance of the user.

Claims (100)

1. A system, comprising:

at least one processor; and

at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:

receive first audio data corresponding to a first utterance;

perform speech processing on the first audio data to determine the first utterance requests deactivation of a privacy mode;

determine a user identifier associated with the first utterance;

delete a first association between the user identifier and a privacy mode indicator;

receive, after deleting the first association, second audio data corresponding to a second utterance corresponding to the user identifier;

perform speech processing on the second audio data to determine natural language understanding (NLU) intent data representing the second utterance; and

store, based at least in part on the second audio data being received after deletion of the first association, the NLU intent data in long-term storage.

2. The system of claim 1 , wherein deletion of the first association further enables long-term storing of automatic speech recognition (ASR) data representing the second utterance.

3. The system of claim 1 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

prior to receiving the first audio data:

receive third audio data corresponding to a third utterance;

perform speech processing on the third audio data to determine the third utterance requests activation of the privacy mode;

determine the user identifier is associated with the third utterance; and

store the first association, wherein storing the first association prevents long-term storing of first data associated with speech processing of a fourth utterance:

received after storing the first association and prior to deletion of the first association; and

corresponding to the user identifier.

4. The system of claim 3 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

receive the first audio data from a first device; and

after storing the first association:

receive, from a second device different from the first device, fourth audio data corresponding to a fifth utterance;

perform speech processing on the fourth audio data to determine the fifth utterance corresponds to a first intent other than a second intent to deactivate the privacy mode;

determine the user identifier is associated with the fifth utterance; and

delete, based at least in part on determining the user identifier is associated with the fifth utterance, speech processing data associated with the fifth utterance.

5. The system of claim 1 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

receive the first audio data from a first device; and

delete a second association between the privacy mode indicator and a first device identifier corresponding to the first device.

6. The system of claim 5 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

delete a third association between the privacy mode indicator and a second device identifier corresponding to a second device different from the first device, the second device being associated with the user identifier.

7. The system of claim 1 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

delete a second association between the privacy mode indicator and a domain of commands.

8. A method, comprising:

receiving first audio data corresponding to a first utterance;

performing speech processing on the first audio data to determine the first utterance requests deactivation of a privacy mode;

determining a user identifier associated with the first utterance;

deleting a first association between the user identifier and a privacy mode indicator;

receiving, after deleting the first association, second audio data corresponding to a second utterance corresponding to the user identifier;

performing speech processing on the second audio data to determine natural language understanding (NLU) intent data representing the second utterance; and

storing, based at least in part on the second audio data being received after deletion of the first association, the NLU intent data in long-term storage.

9. The method of claim 8 , wherein deletion of the first association further enables long-term storing of automatic speech recognition (ASR) data representing the second utterance.

10. The method of claim 8 , further comprising:

receiving the first audio data from a first device; and

deleting a second association between the privacy mode indicator and a first device identifier corresponding to the first device.

11. The method of claim 10 , further comprising:

deleting a third association between the privacy mode indicator and a second device identifier corresponding to a second device different from the first device, the second device being associated with the user identifier.

12. The method of claim 8 , further comprising:

deleting a second association between the privacy mode indicator and a domain of commands.

13. A system, comprising:

at least one processor; and

at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:

receive first audio data corresponding to a first utterance, the first audio data corresponding to a user identifier;

perform speech processing on the first audio data to determine the first utterance corresponds to first natural language understanding (NLU) intent data;

generate output data based at least in part on the first NLU intent data;

receive, after generating the output data, second audio data corresponding to a second utterance;

perform speech processing on the second audio data to determine the second utterance requests deletion of stored data corresponding to the first utterance;

determine the second audio data corresponds to the user identifier;

delete the first NLU intent data based at least in part on the second utterance and the second audio data corresponding to the user identifier; and

delete the output data based at least in part on the second utterance and the second audio data corresponding to the user identifier.

14. The system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine automatic speech recognition (ASR) data corresponding to the first utterance; and

delete the ASR data based at least in part on the second utterance and the second audio data corresponding to the user identifier.

15. The method of claim 8 , further comprising:

prior to receiving the first audio data:

receiving second audio data corresponding to a third utterance;

performing speech processing on the second audio data to determine the third utterance requests activation of the privacy mode;

determining the user identifier is associated with the third utterance; and

store the first association, wherein storing the first association prevents long-term storing of first data associated with speech processing of a fourth subsequently received utterance:

received after storing the first association and prior to deletion of the first association; and

corresponding to the user identifier.

16. The method of claim 15 , further comprising:

receiving the first audio data from a first device; and

after storing the first association:

receiving, from a second device different from the first device, third audio data corresponding to a fifth utterance;

performing speech processing on the third audio data to determine the fifth utterance corresponds to a first intent other than a second intent to deactivate the privacy mode;

determining the user identifier is associated with the fifth utterance; and

delete, based at least in part on determining the user identifier is associated with the fifth utterance, speech processing data associated with the fifth utterance.

17. The system of claim 1 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to, prior to receiving the first audio data:

receive second audio data corresponding to a third utterance;

perform speech processing on the second audio data to determine the third utterance requests an action configured to be executed using user information;

determine the user identifier is associated with the third utterance;

determine first user information associated with the user identifier;

generate first data corresponding to a genericized representation of the first user information; and

generate output data using the first data.

18. The method of claim 8 , further comprising, prior to receiving the first audio data:

receiving second audio data corresponding to a third utterance;

performing speech processing on the second audio data to determine the third utterance requests an action configured to be executed using user information;

determining the user identifier is associated with the third utterance;

determining first user information associated with the user identifier;

generating first data corresponding to a genericized representation of the first user information; and

generating output data using the first data.

19. The system of claim 1 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to, prior to receiving the first audio data:

receive second audio data corresponding to a third utterance, the second audio data corresponding to a user identifier associated with a user name;

perform speech processing on the second audio data to determine the third utterance requests establishment of a two-way communication with a first device; and

send, to the first device, first data instructing first device to refrain from displaying the user name as part of the two-way communication.

20. The method of claim 8 , further comprising, prior to receiving the first audio data:

receiving second audio data corresponding to a third utterance, the second audio data corresponding to a user identifier associated with a user name;

performing speech processing on the second audio data to determine the third utterance requests establishment of a two-way communication with a first device; and

sending, to the first device, first data instructing first device to refrain from displaying the user name as part of the two-way communication.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 18, 2019
From: WANG, ZHENHUA
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 049507/0296 →
Continuity (2)
Continuation 15612723 · Jun 2, 2017
Related Publication 20190371328A1 · Dec 5, 2019
Cited By (2)
US 12,243,532 US 12,688,329