IP Library › Granted Patent US 10,679,614
Granted Patent B2
US 10,679,614 · App. 16/393,785 · Granted Jun 9, 2020

Systems and method to resolve audio-based requests in a networked environment

Inventors: Pedro Gonnet Anders (Zurich, CH); Victor Carbune (Winterthur, CH); Daniel Keysers (Stallikon, CH); Thomas Deselaers (Zurich, CH); Sandro Feuz (Zurich, CH)
Assignee: GOOGLE LLC
G10L15/1815G10L15/19G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,679,614
App. No.
16/393,785
Granted
Jun 9, 2020
Kind
B2
Abstract

Techniques are described herein for enabling an automated assistant to adjust its behavior depending on a detected vocabulary level or other vocal characteristics of an input utterance provided to an automated assistant. The estimated vocabulary level or other vocal characteristics may be used to influence various aspects of a data processing pipeline employed by the automated assistant. In some implementations, one or more tolerance thresholds associated with, for example, grammatical tolerances or vocabulary tolerances, may be adjusted based on the estimated vocabulary level or vocal characteristics of the input utterance.

Claims (53)

1. A system to resolve requests in an audio-based networked system, comprising a computing device comprising one or more processors and a memory, the one or more processors to execute:

a proficiency detector to:

receive a vocal utterance captured at a client device;

determine a vocal characteristic of the vocal utterance captured at the client device;

a speech-to-text module to select a query understanding model from a plurality of candidate query understanding models based on the vocal characteristic;

an intent matcher to determine an intent of the vocal utterance using the query understanding model;

a fulfillment module to select a content item based on the intent and one or more keywords parsed from the vocal utterance; and

an interface to transmit the content item to the client device.

2. The system of claim 1 , comprising:

a content selector component to select a digital component based on the one or more keywords parsed from the vocal utterance; and

the interface to transmit the digital component to the client device.

3. The system of claim 1 , wherein each of the plurality of candidate query understanding models include a different grammatical tolerance.

4. The system of claim 1 , comprising:

the proficiency detector to determine a vocabulary level of the vocal utterance; and

the fulfillment module to select the content item based on the vocabulary level of the vocal utterance.

5. The system of claim 1 , comprising:

the proficiency detector to determine a vocabulary level of the vocal utterance; and

the speech-to-text module to select the query understanding model from the plurality of candidate query understanding models based on the vocabulary level of the vocal utterance.

6. The system of claim 1 , wherein the vocal characteristic comprises at least one of phonemes, a pitch, frequency components, or a cadence of the vocal utterance.

7. The system of claim 1 , comprising:

the fulfillment module to select the content item based to match a vocabulary level of the vocal utterance.

8. The system of claim 1 , comprising:

a text-to-speech module to select a voice synthesis model based on the vocal characteristic of the vocal utterance, the voice synthesis model to render the content item at the client device.

9. The system of claim 1 , comprising:

an invocation module to set an invocation threshold to invoke processing of vocal utterances based on the vocal characteristic.

10. The system of claim 1 , comprising:

a natural language processor to increase a tolerance to at least one of grammatical, vocabulary, or pronunciation errors based on the vocal characteristic.

11. A method implemented using one or more processors, comprising:

receiving, by a proficiency detector executed by one or more processors, a vocal utterance captured at a client device;

determining, by the proficiency detector executed by the one or more processors, a vocal characteristic based on the vocal utterance captured at the client device;

selecting, by a speech-to-text module executed by the one or more processors, a query understanding model from a plurality of candidate query understanding models based on the vocal characteristic;

determining, by an intent matcher executed by the one or more processors, an intent of the vocal utterance using the query understanding model;

selecting, by a fulfillment module executed by the one or more processors, a content item based on the intent and one or more keywords parsed from the vocal utterance; and

transmitting, by the one or more processors, the content item to the client device.

12. The method of claim 11 , comprising:

selecting, by a content selector component executed by the one or more processors, a digital component, based on the one or more keywords parsed from the vocal utterance; and

transmitting, by the one or more processors, the digital component to the client device.

13. The method of claim 11 , wherein each of the plurality of candidate query understanding models include a different grammatical tolerance.

14. The method of claim 11 , comprising:

determining a vocabulary level of the vocal utterance; and

selecting the content item based on the vocabulary level of the vocal utterance.

15. The method of claim 11 , comprising:

determining a vocabulary level of the vocal utterance; and

selecting the query understanding model from the plurality of candidate query understanding models based on the vocabulary level of the vocal utterance.

16. The method of claim 11 , wherein the vocal characteristic comprises at least one of phonemes, a pitch, frequency components, or a cadence of the vocal utterance.

17. The method of claim 11 , comprising:

selecting, by the fulfillment module executed by the one or more processors, the content item based to match the vocal characteristic.

18. The method of claim 11 , comprising:

selecting a voice synthesis model based on the vocal characteristic to render the content item at the client device.

19. The method of claim 11 , comprising:

setting, by an invocation module executed by the one or more processors and based on the vocal characteristic, an invocation threshold to invoke processing of vocal utterances.

20. The method of claim 11 , comprising:

increasing, by a natural language processor executed by the one or more processors, a tolerance to at least one of grammatical, vocabulary, or pronunciation errors based on the vocal characteristic.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 6, 2019
From: ANDERS, PEDRO GONNET; CARBUNE, VICTOR; KEYSERS, DANIEL; DESELAERS, THOMAS; FEUZ, SANDRO
To: GOOGLE LLC
Reel/Frame 049978/0507 →
Continuity (2)
Continuation In Part 15954174 · Apr 16, 2018
Related Publication 20190348030A1 · Nov 14, 2019