IP Library Granted Patent US 9,792,906
Granted Patent B2
US 9,792,906 · App. 15/171,177 · Granted Oct 17, 2017

Method and apparatus for identifying acoustic background environments based on time and speed to enhance automatic speech recognition

Inventor: Mazin Gilbert (Warren, NJ)
Assignee: Nuance Communications, Inc.
G10L15/20G10L15/07G10L15/26G10L15/30G10L21/0216G10L15/065
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,792,906
App. No.
15/171,177
Granted
Oct 17, 2017
Kind
B2
Abstract

Disclosed are systems, methods, and computer readable media for identifying an acoustic environment of a caller. The method embodiment comprises analyzing acoustic features of a received audio signal from a caller, receiving meta-data information based on a previously recorded time and speed of the caller, classifying a background environment of the caller based on the analyzed acoustic features and the meta-data, selecting an acoustic model matched to the classified background environment from a plurality of acoustic models, and performing speech recognition as the received audio signal using the selected acoustic model.

Claims (38)

1. A method comprising:

identifying a speed of a caller in a communication session;

classifying a background environment of the caller based at least in part on the speed of the caller, to yield a background environment classification;

selecting an acoustic model matched to the background environment classification from a plurality of acoustic models; and

performing speech recognition on a received audio signal from the caller using the acoustic model.

2. The method of claim 1 , where meta-data associated with the received audio signal comprises one of global positioning system coordinates, elevation, automatic number identification information, computing device identification number (comprised of an internet protocol address or MAC address), uniform resource locator address, individual environmental habits, personal profile information, time, and rate of movement.

3. The method of claim 2 , wherein the meta-data comprises personal information associated with the caller and comprises probabilities that the caller is in a particular background environment.

4. The method of claim 1 , wherein the background environment classification comprises one of office, airport, street, vehicle, train and home.

5. The method of claim 4 , wherein the background environment is classified based on two levels comprising a first level from the listing of background environments and a second, finer, level based on specific geographic location.

6. The method comprising claim 1 , wherein classifying the background of the caller is further based at least in part on acoustic features of the received audio signal.

7. The method of claim 1 , wherein the acoustic features comprise one of estimates of background energy, signal-to-noise ratio, and spectral characteristics of the background environment.

8. The method of claim 1 , wherein speech recognition is applied to provide speech transcription of the received audio signal.

9. The method of claim 1 , further comprising:

classifying a first background environment in a call and thereafter classifying a second background environment; and

transitioning from a first acoustic model associated with the first background environment to a second acoustic model associated with the second background environment by:

starting the second acoustic model at an initial state similar to an ending state of the first acoustic model when the first acoustic model and the second acoustic model have similar structure; and

applying a morphing algorithm to the transition from the first acoustic model to the second acoustic model if the first acoustic model and the second acoustic model have dissimilar structures.

10. The method comprising claim 1 , wherein classifying the background of the caller is further based at least in part on acoustic features of the received audio signal.

11. The method comprising claim 1 , wherein classifying the background of the caller is further based at least in part on acoustic features of the received audio signal.

12. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

identifying a speed of a caller in a communication session;

classifying a background environment of the caller based at least in part on the speed of the caller, to yield a background environment classification;

selecting an acoustic model matched to the background environment classification from a plurality of acoustic models; and

performing speech recognition on a received audio signal from the caller using the acoustic model.

13. The system of claim 12 , where meta-data associated with the received audio signal comprises one of global positioning system coordinates, elevation, automatic number identification information, computing device identification number (comprised of an internet protocol address or MAC address), uniform resource locator address, individual environmental habits, personal profile information, time, and rate of movement.

14. The system of claim 13 , wherein the meta-data comprises personal information associated with the caller and comprises probabilities that the caller is in a particular background environment.

15. The system of claim 12 , wherein the background environment classification comprises one of office, airport, street, vehicle, train and home.

16. The system of claim 15 , wherein the background environment is classified based on two levels comprising a first level from the listing of background environments and a second, finer, level based on specific geographic location.

17. The system of claim 12 , wherein the acoustic features comprise one of estimates of background energy, signal-to-noise ratio, and spectral characteristics of the background environment.

18. The system of claim 12 , wherein speech recognition is applied to provide speech transcription of the received audio signal.

19. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

identifying a speed of a caller in a communication session;

classifying a background environment of the caller based at least in part on the speed of the caller, to yield a background environment classification;

selecting an acoustic model matched to the background environment classification from a plurality of acoustic models; and

performing speech recognition on a received audio signal from the caller using the acoustic model.

20. The computer-readable storage device of claim 19 , wherein the acoustic features comprise one of estimates of background energy, signal-to-noise ratio, and spectral characteristics of the background environment.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065552/0934 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 039511/0604 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 039511/0643 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2016
From: GILBERT, MAZIN
To: AT&T CORP.
Reel/Frame 039785/0844 →
Continuity (3)
Continuation 14312116 · Jun 23, 2014
Continuation 11754814 · May 29, 2007
Related Publication 20160275948A1 · Sep 22, 2016