IP Library › Granted Patent US 11,657,812
Granted Patent B2
US 11,657,812 · App. 17/030,445 · Granted May 23, 2023

Message playback using a shared device

Inventors: Christo Frank Devaraj (Seattle, WA); Brian Oliver (Seattle, WA); Sumedha Arvind Kshirsagar (Seattle, WA); Gregory Michael Hart (Mercer Island, WA); Ran Mokady (Seattle, WA)
Assignee: Amazon Technologies, Inc.
G10L15/22G06F3/167G06F16/685G10L13/08G10L15/02G10L15/1815G10L15/26G10L17/06G10L17/22H04L51/226H04L67/63G10L17/00G10L2015/223H04L51/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,657,812
App. No.
17/030,445
Granted
May 23, 2023
Kind
B2
Abstract

Methods and systems for providing message playback using a shared electronic device is described herein. In response to receiving a request to output messages, a speech-processing system may determine a group account associated with a requesting device, and may determine messages stored by a message data store for the group account. Speaker identification processing may also be performed to determine a speaker of the request. A user account associated with the speaker, and messages stored for the user account, may be determined. A summary response indicating the user account's messages and the group account's message may then be generated such that the user account messages are identified prior to the group account's messages. The messages may then be analyzed to determine an appropriate voice user interface for the requester such that the playback of the messages using a shared electronic device is more natural and conversational.

Claims (74)

1. A computer-implemented method, comprising:

receiving audio data corresponding to a first natural language request to send a message;

processing the audio data to determine language processing data;

processing the audio data to determine at least one vocal characteristic of the first natural language request;

based on the at least one vocal characteristic, determining an importance of the message;

using the language processing data and the importance of the message, determining second data representing how the message is to be sent;

determining a first messaging account identifier corresponding to a recipient of the message; and

causing the message to be sent based at least in part on the first messaging account identifier and the second data.

2. The computer-implemented method of claim 1 , wherein determining the importance of the message is further based on

the language processing data.

3. The computer-implemented method of claim 1 , wherein determining the second data comprises:

determining the language processing data represents a name;

determining the name corresponds to a first identifier; and

including the first identifier in the second data.

4. The computer-implemented method of claim 3 , further comprising:

determining a first portion of the first natural language request corresponds to a group corresponding to a second identifier;

determining a second portion of the first natural language request corresponds to a message payload;

determining, using the language processing data, that the second portion includes the name; and

in response to the second portion including the name, including the first identifier in the second data and omitting the second identifier from the second data.

5. The computer-implemented method of claim 1 , further comprising:

performing speaker identification using the audio data to determine a speaker identifier; and

including the speaker identifier in the second data.

6. The computer-implemented method of claim 5 , further comprising:

receiving the audio data from a device associated with a group account; and

based at least on the speaker identification, associating the speaker identifier with a sender identifier of the message rather than the group account.

7. The computer-implemented method of claim 1 , further comprising:

receiving the audio data from a device associated with a group account;

failing to identify a user who originated the first natural language request; and

based at least on failing to identify the user who originated the first natural language request, associating the group account with a sender identifier of the message.

8. The computer-implemented method of claim 1 , further comprising:

performing automatic speech recognition (ASR) using the audio data to determine ASR data; and

performing natural language understanding (NLU) using the ASR data to determine NLU data,

wherein the language processing data includes at least one of the ASR data or the NLU data.

9. The computer-implemented method of claim 8 , further comprising:

determining the importance of the message further based on the ASR data.

10. The computer-implemented method of claim 1 , wherein the at least one vocal characteristic comprises at least one of an inflection or a volume.

11. A system comprising:

at least one processor; and

at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:

receive audio data corresponding to a first natural language request to send a message;

process the audio data to determine language processing data;

process the audio data to determine at least one vocal characteristic of the first natural language request;

based on the at least one vocal characteristic, determining an importance of the message;

using the language processing data and the importance of the message, determining second data representing how the message is to be sent;

determine a first messaging account identifier corresponding to a recipient of the message; and

cause the message to be sent based at least in part on the first messaging account identifier and the second data.

12. The system of claim 11 , wherein the instructions that cause the system to determine the importance of the message comprise instructions that, when executed by the at least one processor, further cause the system to:

determine the importance of the message further based on the language processing data.

13. The system of claim 11 , wherein the instructions that cause the system to determine the second data comprise instructions that, when executed by the at least one processor, further cause the system to:

determine the language processing data represents a name;

determine the name corresponds to a first identifier; and

include the first identifier in the second data.

14. The system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine a first portion of the first natural language request corresponds to a group corresponding to a second identifier;

determine a second portion of the first natural language request corresponds to a message payload;

determine, using the language processing data, that the second portion includes the name; and

in response to the second portion including the name, include the first identifier in the second data and omit the second identifier from the second data.

15. The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

perform speaker identification using the audio data to determine a speaker identifier; and

include the speaker identifier in the second data.

16. The system of claim 15 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

receive the audio data from a device associated with a group account; and

based at least on the speaker identification, associate the speaker identifier with a sender identifier of the message rather than the group account.

17. The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

receive the audio data from a device associated with a group account;

fail to identify a user who originated the first natural language request; and

based at least on failing to identify the user who originated the first natural language request, associate the group account with a sender identifier of the message.

18. The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

perform automatic speech recognition (ASR) using the audio data to determine ASR data; and

perform natural language understanding (NLU) using the ASR data to determine NLU data,

wherein the language processing data includes at least one of the ASR data or the NLU data.

19. The system of claim 18 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine the importance of the message based on the ASR data.

20. The system of claim 11 , wherein the at least one vocal characteristic comprises at least one of an inflection or a volume.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 24, 2020
From: DEVARAJ, CHRISTO FRANK; OLIVER, BRIAN; KSHIRSAGAR, SUMEDHA ARVIND; HART, GREGORY MICHAEL; MOKADY, RAN
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 053867/0477 →
Continuity (3)
Continuation 16251901 · Jan 18, 2019
Continuation 15392810 · Dec 28, 2016
Related Publication 20210110823A1 · Apr 15, 2021