IP Library Granted Patent US 10,909,978
Granted Patent B2
US 10,909,978 · App. 15/635,936 · Granted Feb 2, 2021

Secure utterance storage

Inventors: William Frederick Hingle Kruse (Seattle, WA); Peter Turk (Seattle, WA); Panagiotis Thomas (Mercer Island, WA)
Assignee: Amazon Technologies, Inc.
G10L15/22G06F3/0623G06F3/0659G06F3/0673G06F21/32G06F21/60G06F21/6245G06F21/6254G10L13/00G10L13/08G10L15/02G10L15/26G10L21/003G10L25/51H04M3/42008H04M3/42221G10L2021/0135H04M2203/6009
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,909,978
App. No.
15/635,936
Granted
Feb 2, 2021
Kind
B2
Abstract

Technologies for secure storage of utterances are disclosed. A computing device captures audio of a human making a verbal utterance. The utterance is provided to a speech-to-text (STT) service that translates the utterance to text. The STT service can also identify various speaker-specific attributes in the utterance. The text and attributes are provided to a text-to-speech (TTS) service that creates speech from the text and a subset of the attributes. The speech is stored in a data store that is less secure than that required for storing the original utterance. The original utterance can then be discarded. The STT service can also translate the speech generated by the TTS service to text. The text generated by the STT service from the speech and the text generated by the STT service from the original utterance are then compared. If the text does not match, the original utterance can be retained.

Claims (55)

1. An apparatus, comprising:

at least one non-transitory computer-readable storage medium to store instructions which, in response to being performed by one or more processors, cause the apparatus to:

receive first audio data comprising a first utterance of one or more first words, the first audio data having a plurality of attributes and being authorized to be stored in a first storage location;

perform speech recognition on the first audio data to identify the one or more first words;

perform speech recognition on the first audio data to identify the plurality of attributes;

determine, from the plurality of attributes, a subset of the plurality of attributes identified from the first audio data;

generate second audio data using at least the one or more first words and the subset of the plurality of attributes, the second audio data comprising a second utterance of the one or more first words;

perform speech recognition on the second audio data to identify one or more second words;

compare the one or more first words to the one or more second words; and

if the one or more first words match the one or more second words, discard the first audio data and store the second audio data in a second storage location; and

if the one or more first words do not match the one or more second words, discard the second audio data and store the first audio data in the first storage location.

2. The apparatus of claim 1 , wherein the subset of the plurality of attributes comprises attributes that do not provide personally identifiable information associated with a speaker of the first utterance.

3. The apparatus of claim 1 , wherein the at least one non-transitory computer-readable storage medium to store further instructions which, in response to being performed by the one or more processors, cause the apparatus to provide a user interface for defining the subset of the plurality of attributes or for defining the plurality of attributes that are to be identified in the first audio data.

4. The apparatus of claim 1 , wherein the at least one non-transitory computer-readable storage medium to store further instructions which, in response to being performed by the one or more processors, cause the apparatus to provide a user interface for specifying whether the first audio data is to be stored or deleted.

5. The apparatus of claim 1 , wherein:

the plurality of attributes include first attributes that are associated with a first security level;

the subset of the first plurality of attributes include second attributes that are associated with a second security level that is less than the first security level; and

the second audio data is associated with a second plurality of attributes that include third attributes that are associated with the second security level.

6. The apparatus of claim 2 , wherein:

the at least one non-transitory computer-readable storage medium to store further instructions which, in response to being performed by the one or more processors, cause the apparatus to further generate the second audio based at least on one or more selected attributes, the one or more first words, and the subset of the plurality of attributes; and

the plurality of attributes includes the one or more selected attributes and the subset of the plurality of attributes.

7. The apparatus of claim 5 , wherein the second audio data is authorized to be stored in a second storage location that is associated with the second security level.

8. A system comprising:

one or more processors; and

one or more computer-executable instructions stored in memory and executable by the one or more processors to:

cause a speech-to-text service to perform speech recognition on first audio data to:

identify one or more first words; and

identify a plurality of attributes in the first audio data, the first audio data being authorized to be stored in a storage location;

determining, from the plurality of attributes, a subset of the plurality of attributes associated with the first audio data;

cause a text-to-speech service to generate second audio data using at least the one or more first words and the subset of the plurality of attributes;

cause speech recognition to be performed on the second audio data to identify one or more second words;

cause the one or more first words to be compared to the one or more second words; and

if the one or more first words match the one or more second words, cause the first audio data to be discarded and cause the second audio data to be stored; and

if the one or more first words do not match the one or more second words, cause the second audio data to be discarded and cause the first audio data to be stored.

9. The processor of claim 8 , wherein the subset of the plurality of attributes comprises attributes that do not provide personally identifiable information of a speaker associated with the first audio data.

10. The processor of claim 8 , wherein the one or more computer-executable instructions are further executable by the one or more processors to provide a user interface for defining the subset of the plurality of attributes.

11. The processor of claim 8 , wherein the one or more computer-executable instructions are further executable by the one or more processors to cause a user interface to be provided for specifying whether the first audio data is to be stored or deleted.

12. The processor of claim 8 , wherein the one or more computer-executable instructions are further executable by the one or more processors to receive an indication, via a user interface, of one or more selected attributes that are to be included in the subset of the first plurality of attributes.

13. The processor of claim 8 , wherein the one or more computer-executable instructions are further executable by the one or more processors to:

cause the speech-to-text service to identify, from the one or more first words, at least a word that includes personally identifiable information and one or more second words that do not include the personally identifiable information; and

wherein the text-to-speech service generates the second audio data based at least on the one or more second words and the subset of the first plurality of attributes.

14. The processor of claim 8 , wherein the storage location is associated with a first security level, and wherein the second audio data is authorized to be stored in a second storage location that is associated with a second security level that is less than the first security level.

15. A computer-implemented method, comprising:

causing speech recognition to be performed on first audio data to identify one or more first words and a plurality of attributes in the first audio data, the first audio data being authorized to be stored in a storage location;

determining, from the plurality of attributes, a subset of the first plurality of attributes associated with the first audio data;

causing second audio data to be generated using at least the one or more first words and the subset of the plurality of attributes;

causing speech recognition to be performed on the second audio data to identify one or more second words;

causing the one or more first words to be compared to the one or more second words; and

if the one or more first words match the one or more second words, causing the first audio data to be discarded and causing the second audio data to be stored; and

if the one or more first words do not match the one or more second words, causing the second audio data to be discarded and causing the first audio data to be stored.

16. The computer-implemented method of claim 15 , wherein the subset of the plurality of attributes comprises second attributes that do not provide personally identifiable information of the speaker associated with the first audio data.

17. The computer-implemented method of claim 15 , further comprising providing a user interface for defining the subset of the plurality of attributes.

18. The computer-implemented method of claim 15 , wherein the first audio data is associated with a speaker and the plurality of attributes are associated with a first voice of the speaker.

19. The computer-implemented method of claim 18 , wherein the second audio data includes a second plurality of attributes that are associated with a second voice that is different than the first voice of the speaker.

20. The computer-implemented method of claim 15 , wherein the storage location is associated with a first security level, and wherein the second audio data is authorized to be stored in a second storage location that is associated with a second security level that is less than the first security level.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 7, 2017
From: KRUSE, WILLIAM FREDERICK HINGLE; TURK, PETER; THOMAS, PANAGIOTIS
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 042928/0230 →
Continuity (1)
Related Publication 20190005952A1 · Jan 3, 2019