IP Library Granted Patent US 9,679,560
Granted Patent B2
US 9,679,560 · App. 14/770,371 · Granted Jun 13, 2017

Server-side ASR adaptation to speaker, device and noise condition via non-ASR audio transmission

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,679,560
App. No.
14/770,371
Granted
Jun 13, 2017
Kind
B2
Abstract

A mobile device is adapted for automatic speech recognition (ASR). A user interface for interaction with a user includes an input microphone for obtaining speech inputs from the user for automatic speech recognition, and an output interface for system output to the user based on ASR results that correspond to the speech input. A local controller obtains a sample of non-ASR audio from the input microphone for ASR-adaptation to channel-specific ASR characteristics, and then provides a representation of the non-ASR audio to a remote ASR server for server-side adaptation to the channel-specific ASR characteristics, and then provides a representation of an unknown ASR speech input from the input microphone to the remote ASR server for determining ASR results corresponding to the unknown ASR speech input, and then provides the system output to the output interface.

Claims (45)

1. A mobile device adapted for automatic speech recognition (ASR) and employing at least one hardware implemented computer processor, the mobile device comprising:

an input microphone for obtaining speech inputs from a user for automatic speech recognition;

an output interface for providing a system output to the user; and

a local controller configured to:

obtain a sample comprising non-ASR audio from the input microphone,

provide a representation of the non-ASR audio to a remote ASR server for server-side adaptation to channel-specific ASR characteristics,

obtain a sample comprising an unknown ASR speech input from the input microphone,

provide a representation of the unknown ASR speech input to the remote ASR server,

receive, from the remote ASR server, ASR results corresponding to the unknown ASR speech input, and

provide, based on the ASR results corresponding to the unknown ASR speech input, the system output to the output interface.

2. The mobile device according to claim 1 , wherein the non-ASR audio comprises audio sampled by the input microphone during a rolling sample window before the unknown ASR speech input.

3. The mobile device according to claim 1 , wherein the non-ASR audio comprises non-ASR speech audio sampled by the input microphone before the unknown ASR speech input.

4. The mobile device according to claim 3 , wherein the non-ASR speech audio comprises speech data sampled from data windows having a length of one second or less.

5. The mobile device according to claim 1 , wherein the representation of the non-ASR audio comprises pre-processed ASR adaptation data produced by the mobile device from the sample comprising non-ASR audio.

6. The mobile device according to claim 5 , wherein the ASR adaptation data comprises at least one of:

background noise model data, and

ASR acoustic model adaptation data.

7. The mobile device according to claim 1 , wherein the representation of the non-ASR audio is limited to speech feature data.

8. A method comprising:

obtaining, by an input microphone on a mobile device, a sample comprising non-automatic speech recognition (ASR) audio;

transmitting, by the mobile device and to a server, a representation of the non-ASR audio for server-side adaptation to channel-specific ASR characteristics;

receiving, by the input microphone on the mobile device, a sample comprising unknown ASR speech input;

transmitting, by the mobile device and to the server, a representation of the unknown ASR speech input;

receiving, from the server, ASR results corresponding to the unknown ASR speech input; and

outputting, by the mobile device, the ASR results.

9. The method according to claim 8 , wherein the non-ASR audio comprises audio sampled by the input microphone during a rolling sample window before the unknown ASR speech input.

10. The method according to claim 8 , wherein the non-ASR audio comprises non-ASR speech audio sampled by the input microphone before the unknown ASR speech input.

11. The method according to claim 10 , wherein the non-ASR speech audio comprises speech data sampled from data windows having a length of one second or less.

12. The method according to claim 8 , wherein the representation of the non-ASR audio comprises pre-processed ASR adaptation data produced by the mobile device from the sample comprising non-ASR audio.

13. The method according to claim 12 , wherein the ASR adaptation data comprises at least one of:

background noise model data, and

ASR acoustic model adaptation data.

14. The method according to claim 8 , wherein the representation of the non-ASR audio is limited to speech feature data.

15. A non-transitory computer-readable medium having computer-executable program instructions stored thereon that, when executed by a processor, cause the processor to:

obtain, using a microphone, a sample comprising non-automatic speech recognition (ASR) audio;

transmit, to a server, a representation of the non-ASR audio for server-side adaptation to channel-specific ASR characteristics;

obtain, using the microphone, a sample comprising unknown ASR speech input;

transmit, to the server, a representation of the unknown ASR speech input;

receive, from the server, ASR results corresponding to the unknown ASR speech input; and

the ASR results.

16. The non-transitory computer-readable medium according to claim 15 , wherein the non-ASR audio comprises audio sampled during a rolling sample window before the unknown ASR speech input.

17. The non-transitory computer-readable medium according to claim 15 , wherein the non-ASR audio comprises non-ASR speech audio sampled by the microphone before the unknown ASR speech input.

18. The non-transitory computer-readable medium according to claim 17 , wherein the non-ASR speech audio comprises speech data sampled from data windows having a length of one second or less.

19. The non-transitory computer-readable medium according to claim 15 , wherein the representation of the non-ASR audio comprises pre-processed ASR adaptation data produced from the sample comprising non-ASR audio.

20. The non-transitory computer-readable medium according to claim 15 , wherein the representation of the non-ASR audio is limited to speech feature data.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065533/0389 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 14, 2015
From: WILLETT, DANIEL; DAHAN, JEAN-GUY E.; GANONG, WILLIAM F., III; WU, JIANXIONG
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 036552/0813 →