IP Library Granted Patent US 11,081,099
Granted Patent B2
US 11,081,099 · App. 16/722,942 · Granted Aug 3, 2021

Automated speech pronunciation attribution

Inventors: Justin Lewis (South San Francisco, CA); Lisa Takehana (San Bruno, CA)
Assignee: GOOGLE LLC
G10L13/02G06F3/167G10L13/00G10L15/02G10L15/22G10L25/51H04L67/18H04L67/24H04L67/306
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,081,099
App. No.
16/722,942
Granted
Aug 3, 2021
Kind
B2
Abstract

Methods, systems, and apparatus for determining candidate user profiles as being associated with a shared device, and identifying, from the candidate user profiles, candidate pronunciation attributes associated with at least one of the candidate user profiles determined to be associated with the shared device. The methods, systems, and apparatus are also for receiving, at the shared device, a spoken utterance; determining a received pronunciation attribute based on received audio data corresponding to the spoken utterance; comparing the received pronunciation attribute to at least one of the candidate pronunciation attributes; and selecting a particular pronunciation attribute from the candidate pronunciation attributes based on a result of the comparison of the received pronunciation attribute to at least one of the candidate pronunciation attributes. With the methods, systems, and apparatus, the particular pronunciation attribute, selected from the candidate pronunciation attributes, is provided for outputting audio associated with the spoken utterance.

Claims (75)

1. A method implemented by one or more processors, the method comprising:

receiving, at a shared digital assistant device, a spoken utterance of a user;

determining that the spoken utterance matches a plurality of candidate user profiles;

determining that the spoken utterance corresponds to a command associated with an action to be performed by an assistant of the shared digital assistant device;

determining whether the action to be performed by the assistant must be attributed to a specific user profile; and

when it is determined that the action to be performed by the assistant must be attributed to a specific user profile:

selecting a particular user profile of the candidate user profiles, wherein selecting the particular user profile comprises:

providing, at a user interface of the shared digital assistant device, a question related to identifying information;

receiving, at the shared digital assistant device, user input responding to the question;

comparing the user input responding to the question to corresponding identifying information for at least one of the plurality of candidate user profiles;

identifying, based on the comparing, the particular user profile, of the plurality of candidate user profiles, as the specific user profile; and

subsequent to identifying the particular user profile:

attributing the action to the particular user profile;

performing the action associated with the command corresponding to the spoken utterance of the user; and

providing, at the user interface of the shared digital assistant device, first audio output related to the command, the action, or the attribution; and

when it is determined that the action to be performed by the assistant need not be attributed to a specific user profile:

providing, at the user interface of the shared digital assistant device, second audio output related to the action or the command.

2. The method of claim 1 , wherein each candidate user profile is associated with corresponding pronunciation attributes.

3. The method of claim 2 , wherein the first audio output includes one or more of the corresponding pronunciation attributes associated with the particular user profile.

4. The method of claim 2 , wherein the comparing further comprises comparing the user input responding to the question to the corresponding pronunciation attributes associated with at least one of the plurality of candidate user profiles.

5. The method of claim 2 , wherein determining that the spoken utterance matches a plurality of candidate user profiles comprises:

determining one or more pronunciation attributes of the spoken utterance;

comparing the one or more pronunciation attributes of the spoken utterance to corresponding pronunciation attributes associated with a plurality of user profiles; and

identifying, based on the comparing, the plurality of candidate user profiles of the plurality of user profiles.

6. The method of claim 1 , wherein the identifying information includes a phone number.

7. A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving, at a shared digital assistant device, a spoken utterance of a user;

determining that the spoken utterance matches a plurality of candidate user profiles;

determining that the spoken utterance corresponds to a command associated with an action to be performed by an assistant of the shared digital assistant device;

determining whether the action to be performed by the assistant must be attributed to a specific user profile; and

when it is determined that the action to be performed by the assistant must be attributed to a specific user profile:

selecting a particular user profile of the candidate user profiles, wherein selecting the particular user profile comprises:

providing, at a user interface of the shared digital assistant device, a question related to identifying information;

receiving, at the shared digital assistant device, user input responding to the question;

comparing the user input responding to the question to corresponding identifying information for at least one of the plurality of candidate user profiles;

identifying, based on the comparing, the particular user profile, of the plurality of candidate user profiles, as the specific user profile; and

subsequent to identifying the particular user profile:

attributing the action to the particular user profile;

performing the action associated with the command corresponding to the spoken utterance of the user; and

providing, at the user interface of the shared digital assistant device, first audio output related to the command, the action, or the attribution; and

when it is determined that the action to be performed by the assistant need not be attributed to a specific user profile in a database:

providing, at the user interface of the shared digital assistant device, second audio output related to the command or the action.

8. The system of claim 7 , wherein each candidate user profile is associated with corresponding pronunciation attributes.

9. The system of claim 8 , wherein the first audio output includes one or more of the corresponding pronunciation attributes associated with the particular user profile.

10. The system of claim 8 , wherein the comparing further comprises comparing the user input responding to the question to the corresponding pronunciation attributes associated with at least one of the plurality of candidate user profiles.

11. The system of claim 8 , wherein determining that the spoken utterance matches a plurality of candidate user profiles comprises:

determining one or more pronunciation attributes of the spoken utterance;

comparing the one or more pronunciation attributes of the spoken utterance to corresponding pronunciation attributes associated with a plurality of user profiles; and

identifying, based on the comparing, the plurality of candidate user profiles of the plurality of user profiles.

12. The system of claim 7 , wherein the identifying information includes a phone number.

13. A computer-readable storage device storing instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

receiving, at a shared digital assistant device, a spoken utterance of a user;

determining that the spoken utterance matches a plurality of candidate user profiles;

determining that the spoken utterance corresponds to a command associated with an action to be performed by an assistant of the shared digital assistant device;

determining whether the action to be performed by the assistant must be attributed to a specific user profile; and

when it is determined that the action to be performed by the assistant must be attributed to a specific user profile:

selecting a particular user profile of the candidate user profiles, wherein selecting the particular user profile comprises:

providing, at a user interface of the shared digital assistant device, a question related to identifying information;

receiving, at the shared digital assistant device, user input responding to the question;

comparing the user input responding to the question to corresponding identifying information for at least one of the plurality of candidate user profiles;

identifying, based on the comparing, the particular user profile, of the plurality of candidate user profiles, as the specific user profile; and

subsequent to identifying the particular user profile:

attributing the action to the particular user profile;

performing the action associated with the command corresponding to the spoken utterance of the user; and

providing, at the user interface of the shared digital assistant device, first audio output related to the command, the action, or the attribution; and

when it is determined that the action to be performed by the assistant need not be attributed to a specific user profile in a database:

providing, at the user interface of the shared digital assistant device, second audio output related to the command or the action.

14. The computer-readable storage device of claim 13 , wherein each candidate user profile is associated with corresponding pronunciation attributes.

15. The computer-readable storage device of claim 14 , wherein the first audio output includes one or more of the corresponding pronunciation attributes associated with the particular user profile.

16. The computer-readable storage device of claim 14 , wherein the comparing further comprises comparing the user input responding to the question to the corresponding pronunciation attributes associated with at least one of the plurality of candidate user profiles.

17. The computer-readable storage device of claim 14 , wherein determining that the spoken utterance matches a plurality of candidate user profiles comprises:

determining one or more pronunciation attributes of the spoken utterance;

comparing the one or more pronunciation attributes of the spoken utterance to corresponding pronunciation attributes associated with a plurality of user profiles; and

identifying, based on the comparing, the plurality of candidate user profiles of the plurality of user profiles.

18. The computer-readable storage device of claim 13 , wherein the identifying information includes a phone number.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 22, 2020
From: LEWIS, JUSTIN; TAKEHANA, LISA
To: GOOGLE INC.
Reel/Frame 052735/0688 →
CHANGE OF NAME Recorded May 22, 2020
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 052743/0483 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2019
From: LEWIS, JUSTIN; TAKEHANA, LISA
To: GOOGLE INC.
Reel/Frame 051346/0094 →
ENTITY CONVERSION Recorded Dec 20, 2019
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 051395/0953 →
Continuity (3)
Continuation 15995380 · Jun 1, 2018
Continuation 15394104 · Dec 29, 2016
Related Publication 20200243063A1 · Jul 30, 2020
Cited By (1)
US 12,614,542