IP Library Granted Patent US 7,302,393
Granted Patent B2
US 7,302,393 · App. 10/539,454 · Granted Nov 27, 2007

Sensor based approach recognizer selection, adaptation and combination

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,302,393
App. No.
10/539,454
Granted
Nov 27, 2007
Kind
B2
Abstract

A method and respective system for operating a speech recognition system, in which a plurality of recognizer programs are accessible to be activated for speech recognition, and are combined on a per need basis in order to efficiently improve the results of speech recognition done by a single recognizer. In order to adapt such system to the dynamically changing acoustic conditions of various operating environments and to the particular requirements of running in embedded systems having only a limited computing power available, it is proposed to a) collect selection base data characterizing speech recognition boundary conditions, e.g. the speaker person and the environmental noise, etc., with sensor means, and b) using program-controlled arbiter means for evaluating the collected data, e.g., a decision engine including software mechanism and a physical sensor, to select the best suited recognizer or a combination thereof out of the plurality of available recognizers.

Claims (50)

1. A method for operating a speech recognition system, in which a program-controlled recognizer performs the steps of:

dissecting a speech signal into frames and computing any kind of feature vector for each frame;

labeling frames by characters or groups of them yielding a plurality of labels per phoneme; and

decoding said labels according a predetermined acoustic model to construct one or more words or fragments of a word,

wherein a plurality of recognizers are accessible to be activated for speech recognition, and are combined in order to balance the results of speech recognition done by a single recognizer, the method further comprising:

a) collecting selection base data characterizing speech recognition boundary conditions with sensor means;

b) using program-controlled arbiter means for evaluating the collected data; and

c) selecting the best suited recognizer or a combination thereof out of the plurality of available recognizers according to said evaluation.

2. The method according to claim 1 , in which said sensor means is one or more of:

a decision logic, including software program, physical sensors or a combination thereof.

3. The method according to claim 1 , further comprising the steps of:

a) processing a physical sensor output in a decision logic implementing one or more of statistical tests, decision trees, and fuzzy membership functions; and

b) returning from said process a confidence value to be used in the sensor select/combine decision.

4. The method according to claim 1 , in which selection base data which have led to a recognizer select decision, is stored in a database for a repeated fast access thereof in order to obtain a fast selection of recognizers.

5. The method according to claim 1 , further comprising the step of:

selecting the number and/or combination of recognizers dependent of the current processor load.

6. The method according to claim 1 , further comprising the step of:

storing the mapping rule how one acoustic model is transformed to another one, instead of storing a plurality of models themselves.

7. A computer system comprising:

means for dissecting a speech signal into frames and computing any kind of feature vector for each frame;

means for labeling frames by characters or groups of them yielding a plurality of labels per phoneme;

means for decoding said labels according a predetermined acoustic model to construct one or more words or fragments of a word,

wherein a plurality of recognizers are accessible to be activated for speech recognition, and are combined in order to balance the results of speech recognition done by a single recognizer, the computer system carrying out the method of:

a) collecting selection base data characterizing speech recognition boundary conditions with sensor means;

b) using program-controlled arbiter means for evaluating the collected data; and

c) selecting the best suited recognizer or a combination thereof out of the plurality of available recognizers according to said evaluation.

8. The system according to claim 7 , in which said sensor means is one or more of:

a decision logic, a software program, physical sensors or a combination thereof.

9. A computer program for execution in a data processing system comprising computer program code portions for performing respective steps:

dissecting a speech signal into frames and computing any kind of feature vector for each frame;

labeling frames by characters or groups of them yielding a plurality of labels per phoneme; and

decoding said labels according a predetermined acoustic model to construct one or more words or fragments of a word,

wherein a plurality of recognizers are accessible to be activated for speech recognition, and are combined in order to balance the results of speech recognition done by a single recognizer, the method further comprising:

a) collecting selection base data characterizing speech recognition boundary conditions with sensor means;

b) using program-controlled arbiter means for evaluating the collected data; and

c) selecting the best suited recognizer or a combination thereof out of the plurality of available recognizers according to said evaluation,

when said computer program code portions are executed on a computer.

10. The computer program according to claim 9 , in which said sensor means is one or more of:

a decision logic, a software program, physical sensors or a combination thereof.

11. A computer program product stored on a computer usable medium comprising computer readable program means for causing a computer to perform the steps of:

dissecting a speech signal into frames and computing any kind of feature vector for each frame;

labeling frames by characters or groups of them yielding a plurality of labels per phoneme; and

decoding said labels according a predetermined acoustic model to construct one or more words or fragments of a word,

wherein a plurality of recognizers are accessible to be activated for speech recognition, and are combined in order to balance the results of speech recognition done by a single recognizer, the method further comprising:

a) collecting selection base data characterizing speech recognition boundary conditions with sensor means;

b) using program-controlled arbiter means for evaluating the collected data; and

c) selecting the best suited recognizer or a combination thereof out of the plurality of available recognizers according to said evaluation,

when said computer program product is executed on a computer.

12. The computer program product according to claim 11 , in which said sensor means is one or more of:

a decision logic, a software program, physical sensors or a combination thereof.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 10, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065536/0574 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065531/0665 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022354/0566 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 18, 2005
From: FISCHER, VOLKER; KUNZMANN, SIEGFRIED
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 016417/0133 →