IP Library Granted Patent US 7,624,015
Granted Patent B1
US 7,624,015 · App. 11/276,502 · Granted Nov 24, 2009

Recognizing the numeric language in natural spoken dialogue

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,624,015
App. No.
11/276,502
Granted
Nov 24, 2009
Kind
B1
Abstract

A system and a method are provided. A speech recognition processor receives unconstrained input speech and outputs a string of words. The speech recognition processor is based on a numeric language that represents a subset of a vocabulary. The subset includes a set of words identified as being for interpreting and understanding number strings. A numeric understanding processor contains classes of rules for converting the string of words into a sequence of digits. The speech recognition processor utilizes an acoustic model database. A validation database stores a set of valid sequences of digits. A string validation processor outputs validity information based on a comparison of a sequence of digits output by the numeric understanding processor with valid sequences of digits in the validation database.

Claims (37)

1. A system, comprising:

a speech recognition processor that receives unconstrained input speech and outputs a string of words, the speech recognition processor being based on a numeric language that represents a subset of a vocabulary, the subset including a set of words identified as being relevant for interpreting and understanding number strings;

a numeric understanding processor containing classes of rules for converting the string of words into a sequence of digits;

an acoustic model database utilized by the speech recognition processor, the acoustic model database organized with two sets of subword units, a first set of hidden Markov models that characterize acoustic features of numeric words and a second set that characterizes acoustic features of remaining vocabulary words in the acoustic model database, and when each word in the first set is modeled by three segments comprising a head, a body and a tail, and wherein each word has one body and a plurality of heads and tails, and wherein the three segment structure only applies to single digit words;

a validation database that stores a set of valid sequences of digits; and

a string validation processor that outputs validity information based on a comparison of a sequence of digits output by the numeric understanding processor with valid sequences of digits in the validation database.

2. The system of claim 1 , wherein the valid sequences of digits in the validation database include valid telephone numbers.

3. The system of claim 1 , wherein the valid sequences of digits in the validation database include valid credit card numbers.

4. The system of claim 1 , further comprising:

a set of filler models that characterizes out-of-vocabulary features.

5. The system of claim 1 , further comprising:

an utterance verification processor that identifies out-of-vocabulary utterances and utterances that are poorly recognized.

6. The system of claim 1 , further comprising:

a dialogue manager processor that initiates an action based on the validity information.

7. The system of claim 1 , further comprising:

a language model database that stores data describing a structure and a sequence of words and phrases.

8. The system of claim 1 , wherein the classes of rules include a restarts rule.

9. A speech recognition method comprising:

performing a speech recognition process on a received speech signal to produce a string of words, the speech recognition process being based on a numeric language that represents a subset of a vocabulary, the subset of the vocabulary including a set of words identified as being relevant for interpreting and understanding number strings and acoustic model database being used by the speech recognition process, the acoustic model database organized with two sets of subword units, a first set of hidden Markov models that characterize acoustic features of numeric words and a second set that characterizes acoustic features of remaining vocabulary words in the acoustic model database, and when each word in the first set is modeled by three segments comprising a head, a body and a tail, and wherein each word has one body and a plurality of heads and tails and wherein the three segment structure only applies to single digit words;

converting, based on a set of rules, the string of words into a sequence of digits; and

producing validity information based on a comparison of the sequence of digits with valid sequences of digits in a database.

10. The speech recognition method of claim 9 , wherein the valid sequences of digits in the database include valid telephone numbers.

11. The speech recognition method of claim 9 , wherein the valid sequences of digits in the database include valid credit card numbers.

12. The speech recognition method of claim 9 , further comprising initiating an action based on the validity information.

13. The speech recognition method of claim 9 , wherein performing a speech recognition process on a received speech signal to produce a string of words, further comprises:

using a language model database that stores data describing a structure and a sequence of words and phrases.

14. The speech recognition method of claim 9 , wherein the set of rules include a restarts rule.

15. A system, comprising:

means for receiving unconstrained input speech and for outputting a string of words, the means for receiving unconstrained input speech and for outputting a string of words being based on a numeric language that represents a subset of a vocabulary, the subset including a set of words identified as being relevant for interpreting and understanding number strings and acoustic model database being used by the speech recognition process the acoustic model database organized with two sets of subword units, a first set of hidden Markov models that characterize acoustic features of numeric words and a second set that characterizes acoustic features of remaining vocabulary words in the acoustic model database, and when each word in the first set is modeled by three segments comprising a head, a body and a tail, and wherein each word has one body and a plurality of heads and tails and wherein the three segment structure only applies to single digit words;

means for converting the string of words into a sequence of digits, the means for converting the string of words into a sequence of digits including classes of rules for converting the string of words into the sequence of digits;

a validation database that stores a set of valid sequences of digits; and

means for outputting validity information based on a comparison of a sequence of digits, output by the means for converting the string of words into a sequence of digits, with valid sequences of digits in the validation database.

16. The system of claim 15 , wherein the valid sequences of digits in the validation database include valid telephone numbers.

17. The system of claim 15 , wherein the valid sequences of digits in the validation database include valid credit card numbers.

18. The system of claim 15 , further comprising means for initiating an action based on the validity information.

19. The system of claim 15 , wherein the means for receiving unconstrained input speech and for outputting a string of words further comprises:

means for using a language model database that stores data describing a structure and a sequence of words and phrases.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041498/0316 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 12, 2017
From: BUNTSCHUH, BRUCE MELVIN
To: AT&T CORP.
Reel/Frame 040956/0358 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 038529/0164 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 038529/0240 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2016
From: RAHIM, MAZIN G.; RICCARDI, GIUSEPPE; WRIGHT, JEREMY HUNTLEY; GORIN, ALLEN LOUIS
To: AT&T CORP.
Reel/Frame 038292/0011 →