IP Library Granted Patent US 8,977,555
Granted Patent B2
US 8,977,555 · App. 13/723,026 · Granted Mar 10, 2015

Identification of utterance subjects

Inventors: Fred Torok (Seattle, WA); Frédéric Johan Georges Deramat (Bellevue, WA); Vikram Kumar Gundeti (Seattle, WA)
Assignee: Amazon Technologies, Inc.
G06F17/3074G06F17/30684
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,977,555
App. No.
13/723,026
Granted
Mar 10, 2015
Kind
B2
Abstract

Features are disclosed for generating markers for elements or other portions of an audio presentation so that a speech processing system may determine which portion of the audio presentation a user utterance refers to. For example, an utterance may include a pronoun with no explicit antecedent. The marker may be used to associate the utterance with the corresponding content portion for processing. The markers can be provided to a client device with a text-to-speech (“TTS”) presentation. The markers may then be provided to a speech processing system along with a user utterance captured by the client device. The speech processing system, which may include automatic speech recognition (“ASR”) modules and/or natural language understanding (“NLU”) modules, can generate hints based on the marker. The hints can be provided to the ASR and/or NLU modules in order to aid in processing the meaning or intent of a user utterance.

Claims (86)

1. A system comprising:

a computer-readable memory storing executable instructions; and

one or more processors in communication with the computer-readable memory, wherein the one or more processors are programmed by the executable instructions to at least:

generate text to be presented to a user, wherein the text comprises a sequence of items;

generate an audio presentation using the text;

associate a plurality of identifiers with the sequence of items, wherein each item of the sequence of items is associated with at least one identifier of the plurality of identifiers;

transmit, to a client device, the audio presentation and the plurality of identifiers;

receive, from the client device:

audio data comprising a user utterance; and

a first identifier of the plurality of identifiers;

perform speech recognition on the user utterance to obtain speech recognition results;

identify a first item, of the sequence of items, based at least partly on the first identifier and the speech recognition results; and

perform an action based at least partly on the first item.

2. The system of claim 1 , wherein the sequence of items comprises a sequence of reminders, a sequence of items on a task list, or a sequence of items available for purchase.

3. The system of claim 1 , wherein the first identifier and the audio data are received in a single data transmission.

4. The system of claim 1 , wherein the one or more processors are further programmed by the executable instructions to:

generate a hint using the sequence of items; and

identify the first item using the hint.

5. The system of claim 1 , wherein the one or more processors are further programmed by the executable instructions to at least:

receive, from the client device, information about a second audio presentation being presented on the client device; and

identify the first item using the information.

6. A computer-implemented method comprising:

under control of one or more computing devices configured with specific computer-executable instructions,

transmitting, to a client device:

an audio presentation comprising a first portion and a second portion, wherein the first portion corresponds to a first item and the second portion corresponds to a second item;

a first marker corresponding to the first item; and

a second marker corresponding to the second item;

receiving, from the client device:

audio data comprising a user utterance; and

marker data comprising the first marker or the second marker; and

selecting an item based at least on the marker data or the audio data, wherein the selected item comprises the first item or the second item.

7. The computer-implemented method of claim 6 , wherein the marker data comprises the second marker, the computer-implemented method further comprising:

determining an amount of time between a first time that presentation of the second portion was initiated, and a second time that the user utterance was initiated.

8. The computer-implemented method of claim 7 , where the amount of time is less than a predetermined threshold, and wherein identifying the related portion comprises identifying the first portion based at least partly on the amount of time.

9. The computer-implemented method of claim 7 , where the amount of time exceeds a predetermined threshold, and wherein identifying the related portion comprises identifying the second portion based at least partly on the amount of time.

10. The computer-implemented method of claim 6 , wherein the marker data further comprises a first presentation identifier corresponding to the audio presentation and a second presentation identifier corresponding to a second audio presentation being presented on the client device when the user utterance was initiated.

11. The computer-implemented method of claim 10 , further comprising:

determining that the user utterance relates to the audio presentation based at least partly on the marker data and the user utterance.

12. The computer-implemented method of claim 6 , wherein the audio presentation is transmitted in a first data stream, and the first marker and second marker are transmitted in one of the first data stream or a second data stream.

13. The computer-implemented method of claim 6 , wherein the first portion and the second portion are transmitted in separate transmissions.

14. The computer-implemented method of claim 6 , wherein the marker data indicates which portion of the audio presentation was being presented on the client device when the utterance was initiated.

15. The computer-implemented method of claim 6 , further comprising:

performing an action based at least partly on the selected item.

16. A non-transitory computer readable medium comprising executable code that, when executed by a processor, causes a computing device to perform a process comprising:

transmitting, to a client device:

an audio presentation comprising a first portion and a second portion, wherein the first portion corresponds to a first item and the second portion corresponds to a second item;

a first marker corresponding to the first item; and

a second marker corresponding to the second item;

receiving, from the client device:

audio data comprising a user utterance; and

marker data comprising the first marker or the second marker; and

selecting an item based at least on the marker data or the audio data, wherein the selected item comprises the first item or the second item.

17. The non-transitory computer readable medium of claim 16 , wherein the marker data and the audio data are received in a single data stream.

18. The non-transitory computer readable medium of claim 16 , wherein the marker data and the audio data are received in separate data transmissions.

19. The non-transitory computer readable medium of claim 16 , the process further comprising:

generating a hint using the first item and the second item; and

selecting the applicable item using the hint.

20. The non-transitory computer readable medium of claim 16 , the process further comprising:

generating a hint using the marker data; and

selecting the applicable item using the hint.

21. The non-transitory computer readable medium of claim 16 , the process further comprising:

receiving, from the client device, information about a second audio presentation being presented on the client device; and

selecting the applicable item using the information.

22. The non-transitory computer readable medium of claim 16 , wherein the audio presentation is transmitted in a first data stream, and the first marker and second marker are transmitted in one of the first data stream or a second data stream.

23. The non-transitory computer readable medium of claim 16 , wherein the marker data indicates which portion of the audio presentation was being presented on the client device when the utterance was initiated.

24. The non-transitory computer readable medium of claim 16 , the process further comprising:

performing an action based at least partly on the selected item.

25. A non-transitory computer readable medium comprising executable code that, when executed by a processor, causes a computing device to perform a process comprising:

receiving, from a speech processing system:

an audio presentation comprising a first portion corresponding to a first item and a second portion corresponding to a second item;

a first marker corresponding to the first item; and

a second marker corresponding to the second item;

presenting the audio presentation; and

transmitting, to the speech processing system:

audio data received via an audio input component of the computing device; and

marker data comprising at least one of the first marker or the second marker.

26. The non-transitory computer readable medium of claim 25 , wherein the audio presentation is received in a first data stream, and the first marker and second marker are received in one of the first data stream or a second data stream.

27. The non-transitory computer readable medium of claim 25 , wherein the marker data and the audio data are transmitted in a single data stream.

28. The non-transitory computer readable medium of claim 25 , wherein the marker data and the audio data are transmitted in separate data transmissions.

29. The non-transitory computer readable medium of claim 25 , the process further comprising:

initiating, substantially concurrently with presenting the audio presentation, a data stream to the speech processing system, the data stream comprising the audio data.

30. The non-transitory computer readable medium of claim 25 , wherein the marker data comprises the first marker, and wherein the marker data is transmitted substantially concurrently with presentation of the first item.

31. The non-transitory computer readable medium of claim 25 , the process further comprising:

presenting a second audio presentation substantially concurrently with presentation of the audio presentation; and

transmitting, to the speech processing system, a first presentation identifier corresponding to the audio presentation and a second presentation identifier corresponding to a second audio presentation.

32. The non-transitory computer readable medium of claim 25 , wherein the first marker comprises a first identifier and the second marker comprises a second identifier.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2015
From: TOROK, FRED; DERAMAT, FREDERIC JOHAN GEORGES; GUNDETI, VIKRAM
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 034848/0789 →
Continuity (1)
Related Publication 20140180697A1 · Jun 26, 2014