Voice query QoS based on client-computed content metadata
A method includes receiving an automated speech recognition (ASR) request from a user device that includes a speech input captured by the user device and content metadata associated with the speech input. The content metadata is generated by the user device. The method also includes determining a priority score for the ASR request based on the content metadata associated with the speech input and caching the ASR request in a pre-processing backlog of pending ASR requests each having a corresponding priority score. The pending ASR requests in the pre-processing backlog are ranked in order of the priority scores. The method also includes providing, from the pre-processing backlog, one or more of the pending ASR requests to a backend-side ASR module, wherein pending ASR requests associated with higher priority scores are processed before pending ASR requests associated with lower priority scores.
1 . A method comprising:
receiving, at data processing hardware of a query processing backend, an automated speech recognition (ASR) request from a user device, the ASR request comprising:
a speech input captured by the user device that includes a voice query; and
content metadata associated with the speech input, the content metadata collected from two or more of a plurality of sources and generated by the user device, the content metadata indicating a likelihood that the corresponding ASR request corresponds to speech output from a real user in proximity to the user device instead of broadcast speech spoken by a human but that emanates from a non-human source;
determining, by the data processing hardware, a priority score for the ASR request based on the content metadata associated with the speech input;
caching, by the data processing hardware, the ASR request in a pre-processing backlog of pending ASR requests each having a corresponding priority score, the pending ASR requests in the pre-processing backlog ranked in order of the priority scores;
in response to caching one or more new ASR requests in the pre-processing backlog of pending ASR requests, re-ranking, by the data processing hardware, all of the pending ASR requests in the pre-processing backlog in order of the priority scores;
providing, by the data processing hardware from the pre-processing backlog, one or more of the pending ASR requests to a backend-side ASR module based on processing availability of the backend-side ASR module, wherein pending ASR requests associated with higher priority scores are processed by the backend-side ASR module before pending ASR requests associated with lower priority scores; and
transmitting, from the data processing hardware, on-device processing instructions to the user device, the on-device processing instructions instructing the user device to transcribe at least a portion of any new speech inputs captured by the user device using a local ASR module residing on the user device when the user device determines the query processing backend is overloaded.
2 . The method of claim 1 , wherein the backend-side ASR module is configured to, in response to receiving each pending ASR request from the pre-processing backlog of pending ASR requests, process the pending ASR request to generate an ASR result for a corresponding speech input associated with the pending ASR request.
3 . The method of claim 1 , further comprising rejecting, by the data processing hardware, any pending ASR requests residing in the pre-processing backlog for a period of time that satisfies a timeout threshold from being processed by the backend-side ASR module.
4 . The method of claim 1 , further comprising, in response to receiving a new ASR request having a respective priority score less than a priority score threshold, rejecting, by the data processing hardware, the new ASR request from being processed by the backend-side ASR module.
5 . The method of claim 1 , wherein the content metadata associated with the speech input further represents a likelihood that the corresponding ASR request will be successfully processed by the backend-side ASR module.
6 . The method of claim 1 , wherein the content metadata associated with the speech input further represents a likelihood that processing of the corresponding ASR request will have an impact on a user associated with the user device.
7 . The method of claim 1 , wherein the content metadata associated with the speech input and generated by the user device comprises at least one of:
a login indicator indicating whether or not a user associated with the user device is logged in to the user device;
a speaker-identification score for the speech input indicating a likelihood that the speech input matches a speaker profile associated with the user device;
a broadcasted-speech score for the speech input indicating a likelihood that the speech input corresponds to broadcasted or synthesized speech output from a non-human source;
a hotword confidence score indicating a likelihood that one or more terms preceding the voice query in the speech input corresponds to a predefined hotword;
an activity indicator indicating whether or not a multi-turn-interaction is in progress between the user device and the query processing backend;
an audio signal score of the speech input;
a spatial-localization score indicating a distance and position of a user relative to the user device;
a transcription of the speech input generated by the local ASR module residing on the user device;
a user device behavior signal indicating a current behavior of the user device; or
an environmental condition signal indicating current environmental conditions relative to the user device.
8 . The method of claim 1 , wherein the user device is configured to, in response to detecting a hotword that precedes the voice query in a spoken utterance:
capture the speech input comprising the voice query;
generate the content metadata associated with the speech input; and
transmit the corresponding ASR request to the data processing hardware.
9 . The method of claim 8 , wherein the speech input further comprises the hotword.
10 . The method of claim 1 , wherein the user device is configured to determine the query processing backend is overloaded by at least one of:
obtaining historical data associated with previous ASR requests communicated by the user device to the data processing hardware;
receiving, from the data processing hardware, a schedule of past and/or predicted overload conditions at the query processing backend; or
receiving an overload condition status notification from the data processing hardware on the fly indicating a present overload condition at the processing backend.
11 . The method of claim 1 , wherein the on-device processing instructions further instruct the user device to:
transcribe a new speech input using the local ASR module residing on the device to generate a transcription;
interpret the transcription of the new speech input to determine a voice query corresponding to the new speech input;
determine whether the user device can execute an action associated with the voice query corresponding to the new speech input; and
transmit the transcription of the speech input to the query processing backend when the user device is unable to execute the action associated with the voice query.
12 . The method of claim 1 , wherein the on-device processing instructions comprise one or more thresholds that corresponding portions of the content metadata must satisfy in order for the user device to transmit the ASR request to the query processing backend.
13 . The method of claim 12 , wherein the on-device processing instructions further instruct the user device to drop the ASR request when at least one of the thresholds are dissatisfied.
14 . The method of claim 1 , wherein the user device is configured to determine the query processing backend is overloaded by obtaining historical data that indicates past instances of overload conditions at the query processing backend.
15 . A system comprising:
data processing hardware of a query processing backend; and
memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:
receiving an automated speech recognition (ASR) request from a user device, the ASR request comprising:
a speech input captured by the user device that includes a voice query; and
content metadata associated with the speech input, the content metadata collected from two or more of a plurality of sources and generated by the user device and indicating a likelihood that the corresponding ASR request corresponds to speech output from a real user in proximity to the user device instead of broadcast speech spoken by a human but that emanates from a non-human source;
determining a priority score for the ASR request based on the content metadata associated with the speech input;
caching the ASR request in a pre-processing backlog of pending ASR requests each having a corresponding priority score, the pending ASR requests in the pre-processing backlog ranked in order of the priority scores;
in response to caching one or more new ASR requests in the pre-processing backlog of pending ASR requests, re-ranking, by the data processing hardware, all of the pending ASR requests in the pre-processing backlog in order of the priority scores;
providing, from the pre-processing backlog, one or more of the pending ASR requests to a backend-side ASR module based on processing availability of the backend-side ASR module, wherein pending ASR requests associated with higher priority scores are processed by the backend-side ASR module before pending ASR requests associated with lower priority scores; and
transmitting, from the data processing hardware, on-device processing instructions to the user device, the on-device processing instructions instructing the user device to transcribe at least a portion of any new speech inputs captured by the user device using a local ASR module residing on the user device when the user device determines the query processing backend is overloaded.
16 . The system of claim 15 , wherein the backend-side ASR module is configured to, in response to receiving each pending ASR request from the pre-processing backlog of pending ASR requests, process the pending ASR request to generate an ASR result for a corresponding speech input associated with the pending ASR request.
17 . The system of claim 15 , wherein the operations further comprise rejecting any pending ASR requests residing in the pre-processing backlog for a period of time that satisfies a timeout threshold from being processed by the backend-side ASR module.
18 . The system of claim 15 , wherein the operations further comprise, in response to receiving a new ASR request having a respective priority score less than a priority score threshold, rejecting the new ASR request from being processed by the backend-side ASR module.
19 . The system of claim 15 , wherein the content metadata associated with the speech input further represents a likelihood that the corresponding ASR request will be successfully processed by the backend-side ASR module.
20 . The system of claim 15 , wherein the content metadata associated with the speech input further represents a likelihood that processing of the corresponding ASR request will have an impact on a user associated with the user device.
21 . The system of claim 15 , wherein the content metadata associated with the speech input and generated by the user device comprises at least one of:
a login indicator indicating whether or not a user associated with the user device is logged in to the user device;
a speaker-identification score for the speech input indicating a likelihood that the speech input matches a speaker profile associated with the user device;
a broadcasted-speech score for the speech input indicating a likelihood that the speech input corresponds to broadcasted or synthesized speech output from a non-human source;
a hotword confidence score indicating a likelihood that one or more terms preceding the voice query in the speech input corresponds to a predefined hotword;
an activity indicator indicating whether or not a multi-turn-interaction is in progress between the user device and the query processing backend;
an audio signal score of the speech input;
a spatial-localization score indicating a distance and position of a user relative to the user device;
a transcription of the speech input generated by the local ASR module residing on the user device;
a user device behavior signal indicating a current behavior of the user device; or
an environmental condition signal indicating current environmental conditions relative to the user device.
22 . The system of claim 15 , wherein the user device is configured to, in response to detecting a hotword that precedes the voice query in a spoken utterance:
capture the speech input comprising the voice query;
generate the content metadata associated with the speech input; and
transmit the corresponding ASR request to the data processing hardware.
23 . The system of claim 22 , wherein the speech input further comprises the hotword.
24 . The system of claim 15 , wherein the user device is configured to determine the query processing backend is overloaded by at least one of:
obtaining historical data associated with previous ASR requests communicated by the user device to the data processing hardware;
receiving, from the data processing hardware, a schedule of past and/or predicted overload conditions at the query processing backend; or
receiving an overload condition status notification from the data processing hardware on the fly indicating a present overload condition at the processing backend.
25 . The system of claim 15 , wherein the on-device processing instructions further instruct the user device to:
transcribe a new speech input using the local ASR module residing on the device to generate a transcription;
interpret the transcription of the new speech input to determine a voice query corresponding to the new speech input;
determine whether the user device can execute an action associated with the voice query corresponding to the new speech input; and
transmit the transcription of the speech input to the query processing backend when the user device is unable to execute the action associated with the voice query.
26 . The system of claim 15 , wherein the on-device processing instructions comprise one or more thresholds that corresponding portions of the content metadata must satisfy in order for the user device to transmit the ASR request to the query processing backend.
27 . The system of claim 26 , wherein the on-device processing instructions further instruct the user device to drop the ASR request when at least one of the thresholds are dissatisfied.
28 . The system of claim 15 , wherein the user device is configured to determine the query processing backend is overloaded by obtaining historical data that indicates past instances of overload conditions at the query processing backend.