IP Library Granted Patent US 7,389,229
Granted Patent B2
US 7,389,229 · App. 10/685,410 · Granted Jun 17, 2008

Unified clustering tree

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,389,229
App. No.
10/685,410
Granted
Jun 17, 2008
Kind
B2
Abstract

A unified clustering tree ( 500 ) generates phoneme clusters based on an input sequence of phonemes. The number of possible clusters is significantly less than the number of possible combinations of input phonemes. Nodes ( 510, 511 ) in the unified clustering tree are arranged into levels such that the clustering tree generates clusters for multiple speech recognition models. Models that correspond to higher levels in the unified clustering tree are coarse models relative to more fine-grain models at lower levels of the clustering tree.

Claims (32)

1. A speech recognition system comprising:

a clustering tree configured to classify a series of sounds into predefined clusters based on one of the sounds and on a predetermined number of neighboring sounds that surround the one of the sounds, where the clustering tree comprises:

a first level with a first hierarchical arrangement of decision nodes in which the decision nodes of the first hierarchical arrangement are associated with a first group of questions relating to the series of sounds,

a second level with a second hierarchical arrangement of decision nodes in which the decision nodes of the second hierarchical arrangement are associated with a second group of questions relating to the series of sounds, the second group of questions discriminating at a finer level of granularity within the series of sounds than the first group of questions, and

a third level with a third hierarchical arrangement of decision nodes in which the decision nodes of the third hierarchical arrangement are associated with a third group of questions discriminating at a finer level of granularity within the series of sounds than the second group of questions; and

a plurality of speech recognition models trained to recognize speech based on the predefined clusters, the plurality of speech recognition models comprising:

a first model associated with the first level and including a triphone non-crossword speech recognition model,

a second model associated with the second level and including a quinphone non-crossword speech recognition model, and

a third model associated with the third level and including a quinphone crossword speech recognition model.

2. The system of claim 1 , wherein the clustering tree is formed by freezing building of the first level of the clustering tree before building the second level of the clustering tree.

3. The system of claim 2 , wherein the clustering tree is further formed by freezing building of the first level of the clustering tree when an entropy level of the first level of the clustering tree is below a predetermined threshold.

4. The system of claim 1 wherein the clustering tree is further formed by freezing building of the second level of the clustering tree before building the third level of the clustering tree.

5. The system of claim 4 , wherein the clustering tree is further formed by freezing building of the second level of the clustering tree when an entropy level of the second level of the clustering tree is below a predetermined threshold.

6. The system of claim 1 , wherein the clustering tree is further built to include terminal nodes that assign each of the groups of sound into one of the sound clusters.

7. The system of claim 1 , wherein the first group of questions includes questions that relate to the series of sounds as a sound being modeled and one context sound before and after the sound being modeled.

8. The system of claim 7 , wherein the second group of questions includes questions that relate to the series of sounds as the sound being modeled and two context sounds before and after the sound being modeled.

9. The system of claim 1 wherein higher ones of the hierarchical levels include nodes that correspond to more general questions than questions corresponding to nodes at lower ones of the hierarchical levels.

10. The system of claim 1 , wherein the sounds are represented by phonemes.

11. The system of claim 1 , wherein the clustering tree comprises:

decision nodes associated with questions that relate to the series of sounds, and

terminal nodes that define a sound cluster to which the series of sounds belong.

12. The system of claim 11 , wherein the decision nodes and the terminal nodes are defined hierarchically relative to one another.

13. The system of claim 12 , wherein the decision nodes correspond to lower ones of the levels in the hierarchically defined nodes are associated with more detailed questions than decision nodes corresponding to higher ones of the levels in the hierarchically defined nodes.

14. A device comprising:

means for classifying a series of sounds into predefined clusters using a clustering tree and based on one of the sounds and a predetermined number of neighboring sounds that surround the one of the sounds, where the clustering tree includes:

a first level with a first hierarchical arrangement of decision nodes in which the decision nodes of the first hierarchical arrangement are associated with a first group of questions relating to the series of sounds;

a second level with a second hierarchical arrangement of decision nodes in which the decision nodes of the second hierarchical arrangement are associated with a second group of questions relating to the series of sounds, the second group of questions discriminating at a finer level of granularity within the series of sounds than the first group of questions; and

a third level with a third hierarchical arrangement of decision nodes in which the decision nodes of the third hierarchical arrangement are associated with a third group of questions discriminating at a finer level of granularity within the series of sounds than the second group of questions; and

means for training a plurality of speech recognition models to recognize speech based on the predefined clusters, the speech recognition models including:

a first model associated with the first level and including a triphone non-crossword speech recognition model,

a second model associated with the second level and including a quinphone non-crossword speech recognition model, and

a third model associated with the third level and including a quinphone crossword speech recognition model.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2010
From: BBN TECHNOLOGIES CORP.
To: RAMP HOLDINGS, INC. (F/K/A EVERYZING, INC.)
Reel/Frame 023973/0141 →
RELEASE OF SECURITY INTEREST Recorded Oct 27, 2009
From: BANK OF AMERICA, N.A. (SUCCESSOR BY MERGER TO FLEET NATIONAL BANK)
To: BBN TECHNOLOGIES CORP. (AS SUCCESSOR BY MERGER TO BBNT SOLUTIONS LLC)
Reel/Frame 023427/0436 →
MERGER Recorded Mar 2, 2006
From: BBNT SOLUTIONS LLC
To: BBN TECHNOLOGIES CORP.
Reel/Frame 017274/0318 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2005
From: BILLA, JAYADEV; KIECZA, DANIEL; KUBALA, FRANCIS G.
To: BBNT SOLUTIONS LLC
Reel/Frame 015728/0955 →
PATENT & TRADEMARK SECURITY AGREEMENT Recorded May 12, 2004
From: BBNT SOLUTIONS LLC
To: FLEET NATIONAL BANK, AS AGENT
Reel/Frame 014624/0196 →