IP Library Granted Patent US 12,731,588
Granted Patent B2
US 12,731,588 · App. 17/310,175 · Granted Sep 8, 2026

Voice query QoS based on client-computed content metadata

Inventors: Matthew Sharifi (Kilchberg, CH); Aleksandar Kracun (New York, NY)
Assignee: Google LLC
G10L15/30G06F16/63G10L15/08G10L15/22H04L67/568G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,588
App. No.
17/310,175
Granted
Sep 8, 2026
Kind
B2
Abstract

A method includes receiving an automated speech recognition (ASR) request from a user device that includes a speech input captured by the user device and content metadata associated with the speech input. The content metadata is generated by the user device. The method also includes determining a priority score for the ASR request based on the content metadata associated with the speech input and caching the ASR request in a pre-processing backlog of pending ASR requests each having a corresponding priority score. The pending ASR requests in the pre-processing backlog are ranked in order of the priority scores. The method also includes providing, from the pre-processing backlog, one or more of the pending ASR requests to a backend-side ASR module, wherein pending ASR requests associated with higher priority scores are processed before pending ASR requests associated with lower priority scores.

Claims (86)

1 . A method comprising:

receiving, at data processing hardware of a query processing backend, an automated speech recognition (ASR) request from a user device, the ASR request comprising:

a speech input captured by the user device that includes a voice query; and

content metadata associated with the speech input, the content metadata collected from two or more of a plurality of sources and generated by the user device, the content metadata indicating a likelihood that the corresponding ASR request corresponds to speech output from a real user in proximity to the user device instead of broadcast speech spoken by a human but that emanates from a non-human source;

determining, by the data processing hardware, a priority score for the ASR request based on the content metadata associated with the speech input;

caching, by the data processing hardware, the ASR request in a pre-processing backlog of pending ASR requests each having a corresponding priority score, the pending ASR requests in the pre-processing backlog ranked in order of the priority scores;

in response to caching one or more new ASR requests in the pre-processing backlog of pending ASR requests, re-ranking, by the data processing hardware, all of the pending ASR requests in the pre-processing backlog in order of the priority scores;

providing, by the data processing hardware from the pre-processing backlog, one or more of the pending ASR requests to a backend-side ASR module based on processing availability of the backend-side ASR module, wherein pending ASR requests associated with higher priority scores are processed by the backend-side ASR module before pending ASR requests associated with lower priority scores; and

transmitting, from the data processing hardware, on-device processing instructions to the user device, the on-device processing instructions instructing the user device to transcribe at least a portion of any new speech inputs captured by the user device using a local ASR module residing on the user device when the user device determines the query processing backend is overloaded.

2 . The method of claim 1 , wherein the backend-side ASR module is configured to, in response to receiving each pending ASR request from the pre-processing backlog of pending ASR requests, process the pending ASR request to generate an ASR result for a corresponding speech input associated with the pending ASR request.

3 . The method of claim 1 , further comprising rejecting, by the data processing hardware, any pending ASR requests residing in the pre-processing backlog for a period of time that satisfies a timeout threshold from being processed by the backend-side ASR module.

4 . The method of claim 1 , further comprising, in response to receiving a new ASR request having a respective priority score less than a priority score threshold, rejecting, by the data processing hardware, the new ASR request from being processed by the backend-side ASR module.

5 . The method of claim 1 , wherein the content metadata associated with the speech input further represents a likelihood that the corresponding ASR request will be successfully processed by the backend-side ASR module.

6 . The method of claim 1 , wherein the content metadata associated with the speech input further represents a likelihood that processing of the corresponding ASR request will have an impact on a user associated with the user device.

7 . The method of claim 1 , wherein the content metadata associated with the speech input and generated by the user device comprises at least one of:

a login indicator indicating whether or not a user associated with the user device is logged in to the user device;

a speaker-identification score for the speech input indicating a likelihood that the speech input matches a speaker profile associated with the user device;

a broadcasted-speech score for the speech input indicating a likelihood that the speech input corresponds to broadcasted or synthesized speech output from a non-human source;

a hotword confidence score indicating a likelihood that one or more terms preceding the voice query in the speech input corresponds to a predefined hotword;

an activity indicator indicating whether or not a multi-turn-interaction is in progress between the user device and the query processing backend;

an audio signal score of the speech input;

a spatial-localization score indicating a distance and position of a user relative to the user device;

a transcription of the speech input generated by the local ASR module residing on the user device;

a user device behavior signal indicating a current behavior of the user device; or

an environmental condition signal indicating current environmental conditions relative to the user device.

8 . The method of claim 1 , wherein the user device is configured to, in response to detecting a hotword that precedes the voice query in a spoken utterance:

capture the speech input comprising the voice query;

generate the content metadata associated with the speech input; and

transmit the corresponding ASR request to the data processing hardware.

9 . The method of claim 8 , wherein the speech input further comprises the hotword.

10 . The method of claim 1 , wherein the user device is configured to determine the query processing backend is overloaded by at least one of:

obtaining historical data associated with previous ASR requests communicated by the user device to the data processing hardware;

receiving, from the data processing hardware, a schedule of past and/or predicted overload conditions at the query processing backend; or

receiving an overload condition status notification from the data processing hardware on the fly indicating a present overload condition at the processing backend.

11 . The method of claim 1 , wherein the on-device processing instructions further instruct the user device to:

transcribe a new speech input using the local ASR module residing on the device to generate a transcription;

interpret the transcription of the new speech input to determine a voice query corresponding to the new speech input;

determine whether the user device can execute an action associated with the voice query corresponding to the new speech input; and

transmit the transcription of the speech input to the query processing backend when the user device is unable to execute the action associated with the voice query.

12 . The method of claim 1 , wherein the on-device processing instructions comprise one or more thresholds that corresponding portions of the content metadata must satisfy in order for the user device to transmit the ASR request to the query processing backend.

13 . The method of claim 12 , wherein the on-device processing instructions further instruct the user device to drop the ASR request when at least one of the thresholds are dissatisfied.

14 . The method of claim 1 , wherein the user device is configured to determine the query processing backend is overloaded by obtaining historical data that indicates past instances of overload conditions at the query processing backend.

15 . A system comprising:

data processing hardware of a query processing backend; and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving an automated speech recognition (ASR) request from a user device, the ASR request comprising:

a speech input captured by the user device that includes a voice query; and

content metadata associated with the speech input, the content metadata collected from two or more of a plurality of sources and generated by the user device and indicating a likelihood that the corresponding ASR request corresponds to speech output from a real user in proximity to the user device instead of broadcast speech spoken by a human but that emanates from a non-human source;

determining a priority score for the ASR request based on the content metadata associated with the speech input;

caching the ASR request in a pre-processing backlog of pending ASR requests each having a corresponding priority score, the pending ASR requests in the pre-processing backlog ranked in order of the priority scores;

in response to caching one or more new ASR requests in the pre-processing backlog of pending ASR requests, re-ranking, by the data processing hardware, all of the pending ASR requests in the pre-processing backlog in order of the priority scores;

providing, from the pre-processing backlog, one or more of the pending ASR requests to a backend-side ASR module based on processing availability of the backend-side ASR module, wherein pending ASR requests associated with higher priority scores are processed by the backend-side ASR module before pending ASR requests associated with lower priority scores; and

transmitting, from the data processing hardware, on-device processing instructions to the user device, the on-device processing instructions instructing the user device to transcribe at least a portion of any new speech inputs captured by the user device using a local ASR module residing on the user device when the user device determines the query processing backend is overloaded.

16 . The system of claim 15 , wherein the backend-side ASR module is configured to, in response to receiving each pending ASR request from the pre-processing backlog of pending ASR requests, process the pending ASR request to generate an ASR result for a corresponding speech input associated with the pending ASR request.

17 . The system of claim 15 , wherein the operations further comprise rejecting any pending ASR requests residing in the pre-processing backlog for a period of time that satisfies a timeout threshold from being processed by the backend-side ASR module.

18 . The system of claim 15 , wherein the operations further comprise, in response to receiving a new ASR request having a respective priority score less than a priority score threshold, rejecting the new ASR request from being processed by the backend-side ASR module.

19 . The system of claim 15 , wherein the content metadata associated with the speech input further represents a likelihood that the corresponding ASR request will be successfully processed by the backend-side ASR module.

20 . The system of claim 15 , wherein the content metadata associated with the speech input further represents a likelihood that processing of the corresponding ASR request will have an impact on a user associated with the user device.

21 . The system of claim 15 , wherein the content metadata associated with the speech input and generated by the user device comprises at least one of:

a login indicator indicating whether or not a user associated with the user device is logged in to the user device;

a speaker-identification score for the speech input indicating a likelihood that the speech input matches a speaker profile associated with the user device;

a broadcasted-speech score for the speech input indicating a likelihood that the speech input corresponds to broadcasted or synthesized speech output from a non-human source;

a hotword confidence score indicating a likelihood that one or more terms preceding the voice query in the speech input corresponds to a predefined hotword;

an activity indicator indicating whether or not a multi-turn-interaction is in progress between the user device and the query processing backend;

an audio signal score of the speech input;

a spatial-localization score indicating a distance and position of a user relative to the user device;

a transcription of the speech input generated by the local ASR module residing on the user device;

a user device behavior signal indicating a current behavior of the user device; or

an environmental condition signal indicating current environmental conditions relative to the user device.

22 . The system of claim 15 , wherein the user device is configured to, in response to detecting a hotword that precedes the voice query in a spoken utterance:

capture the speech input comprising the voice query;

generate the content metadata associated with the speech input; and

transmit the corresponding ASR request to the data processing hardware.

23 . The system of claim 22 , wherein the speech input further comprises the hotword.

24 . The system of claim 15 , wherein the user device is configured to determine the query processing backend is overloaded by at least one of:

obtaining historical data associated with previous ASR requests communicated by the user device to the data processing hardware;

receiving, from the data processing hardware, a schedule of past and/or predicted overload conditions at the query processing backend; or

receiving an overload condition status notification from the data processing hardware on the fly indicating a present overload condition at the processing backend.

25 . The system of claim 15 , wherein the on-device processing instructions further instruct the user device to:

transcribe a new speech input using the local ASR module residing on the device to generate a transcription;

interpret the transcription of the new speech input to determine a voice query corresponding to the new speech input;

determine whether the user device can execute an action associated with the voice query corresponding to the new speech input; and

transmit the transcription of the speech input to the query processing backend when the user device is unable to execute the action associated with the voice query.

26 . The system of claim 15 , wherein the on-device processing instructions comprise one or more thresholds that corresponding portions of the content metadata must satisfy in order for the user device to transmit the ASR request to the query processing backend.

27 . The system of claim 26 , wherein the on-device processing instructions further instruct the user device to drop the ASR request when at least one of the thresholds are dissatisfied.

28 . The system of claim 15 , wherein the user device is configured to determine the query processing backend is overloaded by obtaining historical data that indicates past instances of overload conditions at the query processing backend.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 24, 2022
From: SHARIFI, MATTHEW; KRACUN, ALEKSANDAR
To: GOOGLE LLC
Reel/Frame 058735/0959 →
Continuity (1)
Related Publication 20220093104A1 · Mar 24, 2022
References Cited (38)
US 6366882B1 · Bijl · 2002 [cited by examiner]
US 9401140B1 · Weber et al. · 2016 [cited by applicant]
US 10504504B1 · Meltzner · 2019 [cited by examiner]
US 10699706B1 · Jayavel · 2020 [cited by examiner]
US 20090052636A1 · Webb et al. · 2009 [cited by applicant]
US 20100010823A1 · Scipioni et al. · 2010 [cited by applicant]
US 20120173477A1 · Coutts · 2012 [cited by examiner]
US 20130151250A1 · VanBlon · 2013 [cited by applicant]
US 20140214429A1 · Pantel · 2014 [cited by applicant]
US 20150331490A1 · Yamada · 2015 [cited by examiner]
US 20150348548A1 · Piernot et al. · 2015 [cited by applicant]
US 20160005394A1 · Hiroe · 2016 [cited by examiner]
US 20160026627A1 · Johnston et al. · 2016 [cited by applicant]
US 20160098992A1 · Renard et al. · 2016 [cited by applicant]
US 20160240194A1 · Lee et al. · 2016 [cited by applicant]
US 20160247520A1 · Kikugawa · 2016 [cited by examiner]
US 20170083285A1 · Meyers et al. · 2017 [cited by applicant]
US 20170243577A1 · Wingate · 2017 [cited by applicant]
US 20180308472A1 · Lopez Moreno et al. · 2018 [cited by applicant]
US 20180330728A1 · Gruenstein et al. · 2018 [cited by applicant]
US 20190139541A1 · Andersen et al. · 2019 [cited by applicant]
US 20190213278A1 · Min · 2019 [cited by examiner]
US 20190214002A1 · Park · 2019 [cited by examiner]
US 20190268465A1 · Broidy · 2019 [cited by examiner]
US 20200034551A1 · Cantrell · 2020 [cited by examiner]
US 20200184959A1 · Yasa et al. · 2020 [cited by applicant]
US 20200184966A1 · Yavagal · 2020 [cited by applicant]
US 20200184967A1 · Gupta et al. · 2020 [cited by applicant]
US 20200193982A1 · Kim · 2020 [cited by applicant]
US 20200243085A1 · Zhou · 2020 [cited by examiner]
US 20200243094A1 · Thomson · 2020 [cited by examiner]
CN 109785845B · 2021 [cited by examiner]
JP 2016004270A · 2016 [cited by applicant]
JP 2022519648A · 2022 [cited by applicant]
International Search Report and Written Opinion for the PCT Application No. PCT/US2019/016882 dated Oct. 10, 2019. [cited by applicant]
1 Intellectual Property India, Examination Report for application 202127030937, dated Mar. 7, 2022. [cited by applicant]
European Search Report for the related Application No. 22215558.2, dated Apr. 4, 2023, 10 pages. [cited by applicant]
Office Action issued in related Japanese Patent Application No. 2024-062090, dated May 27, 2025. [cited by applicant]