IP Library Granted Patent US 9,524,718
Granted Patent B2
US 9,524,718 · App. 14/391,200 · Granted Dec 20, 2016

Speech recognition server integration device that is an intermediate module to relay between a terminal module and speech recognition server and speech recognition server integration method

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,524,718
App. No.
14/391,200
Granted
Dec 20, 2016
Kind
B2
Abstract

The speech recognition result through the general-purpose server and that through the specialized speech recognition server are integrated in an optimum manner, thereby, a speech recognition function least in errors in the end being provided. The specialized speech recognition server 108 is constructed with the words contained in the user dictionary data in use as well as the performance of the general-purpose speech recognition server 106 is preliminarily evaluated with such user dictionary data. Based on such evaluation result, information related to which recognition results through the specialized and general-purpose speech recognition servers are adopted and to how the adopted recognition results are weighted to obtain an optimum recognition result is preliminarily retained in the form of a database. Upon executing recognition, an optimum recognition result is obtained by comparing the recognition results through the specialized and general-purpose servers with the parameter for recognition result integration 118.

Claims (47)

1. A speech recognition server integration device, which is an intermediate module to relay between a terminal module by which a user performs operations by way of speech and a speech recognition server to recognize speech data and to feed back its recognition result, the speech recognition server integration device being configured to:

learn and preserve a parameter for recognition result integration based on words registered by the user or a list of words frequently used by the user;

receive speech data spoken by the user for speech recognition from the terminal module;

transmit the received speech data to a general-purpose speech recognition server and a specialized speech recognition server;

receive recognition results of the speech data through the general-purpose speech recognition server and the specialized speech recognition server;

compare the recognition results through the general-purpose speech recognition server and the specialized speech recognition server with the preserved parameter for recognition result integration and to select an optimum recognition result;

transmit the selected recognition result to the terminal module;

receive the words registered by the user or the list of words frequently used by the user from the terminal module;

make synthesized speech based on the received words;

transmit the made synthesized speech to the general-purpose speech recognition server and the specialized speech recognition server; and

receive the recognition results of the synthesized speech through the general-purpose speech recognition server and the specialized speech recognition server,

wherein the speech recognition server integration device concurrently analyses the words from which the synthesized speech derives and the recognition results to learn and preserve the parameter for recognition result integration.

2. The speech recognition server integration device according to claim 1 , wherein the speech recognition server integration device is further configured to:

receive the words registered by the user or the list of words frequently used by the user from the terminal module;

receive a list of words for recognition from the general-purpose speech recognition server; and

compare the list of words for recognition with the list of words received from the terminal module and to estimate similarity between the lists,

wherein the speech recognition server integration device preserves an estimation result as the parameter for recognition result integration.

3. The speech recognition server integration device according to claim 1 ,

wherein the specialized speech recognition server makes a list of words for recognition target based on the words registered by the user or the list of words frequently used by the user and is capable of recognizing words contained in the list of words for recognition target with high precision.

4. The speech recognition server integration device according to claim 1 ,

wherein the specialized speech recognition server is incorporated in the speech recognition server integration device or the terminal module as a specialized speech recognition module.

5. The speech recognition server integration device according to claim 1 ,

wherein the parameter for recognition result integration accumulates accuracies and errors of the recognition results of the words registered or frequently used by the user through the speech recognition servers,

wherein the speech recognition server integration device extracts the recognition results of a certain word through the speech recognition servers from the parameter for recognition result integration based on the recognition result through the specialized speech recognition server and extracts only the recognition results through the speech recognition servers whose recognition results are accurate to select an optimum recognition result based on the extracted recognition results.

6. The speech recognition server integration device according to claim 1 ,

wherein the parameter for recognition result integration accumulates accuracies and errors of the recognition results of the words registered and frequently used by the user through the speech recognition servers and values denoting confidence measures of the recognition results of the individual words through the speech recognition servers,

wherein the speech recognition server integration device extracts the recognition results and the confidence measures of the recognition results of a certain word through the speech recognition servers from the parameter for recognition result integration based on the recognition result through the specialized speech recognition server and extracts only the recognition results and the confidence measures of the recognition results through the speech recognition servers whose extracted recognition results are accurate to integrate the extracted recognition results by weighting with the confidence measure on the extracted recognition results.

7. The speech recognition server integration device according to claim 1 ,

wherein the parameter for recognition result integration measures time required for the speech recognition servers recognizing the words registered or frequently used by the user and accumulates the measured time values,

wherein the speech recognition server integration device extracts time required for the speech recognition servers recognizing a certain word from the parameter for recognition result integration based on the recognition result through the specialized speech recognition server and obtains an allowable upper value of the time required for the speech recognition servers recognizing a certain word, which value is determined depending on an application in use, and extracts only the recognition results through the speech recognition servers whose time required for recognizing a certain word goes below the allowable upper value to select an optimum recognition result based on the extracted recognition results.

8. The speech recognition server integration device according to claim 1 ,

wherein the parameter for recognition result integration accumulates accuracies and errors of the recognition results of the words registered or frequently used by the user and one misrecognition result or a plurality of misrecognition results through the speech recognition servers,

wherein the speech recognition server integration device extracts the accuracies and errors as well as misrecognition results of the recognition results of a certain word through the speech recognition servers from the parameter for recognition result integration based on the recognition result through the specialized speech recognition server and compares the extracted misrecognition result with a recognition result upon the recognition being executed when the extracted recognition result is an error and makes the recognition result effective only when it is determined that the comparison results are the same to select an optimum recognition result based on the recognition result made effective.

9. A speech recognition server integration method comprising:

a step of learning and preserving a parameter for recognition result integration based on words registered by a user or a list of words frequently used by the user;

a step to transmit data of a speech spoken by the user for speech recognition to a general-purpose speech recognition server and a specialized speech recognition server;

a step to receive recognition results of the speech data through the general-purpose speech recognition server and the specialized speech recognition server;

a step to compare the recognition result through the general-purpose speech recognition server and the recognition result through the specialized speech recognition server with the parameter for recognition result integration to select an optimum speech recognition result;

a step of making a synthesized speech based on the words registered or frequently used by the user;

a step of transmitting the made synthesized speech to the general-purpose speech recognition server and the specialized speech recognition server; and

a step of receiving the recognition results of the synthesized speech through the general-purpose speech recognition server and the specialized speech recognition server,

wherein in the step of learning and preserving the parameter for recognition result integration the words from which the synthesized speech derives and the recognition results are concurrently analyzed to make the parameter for recognition result integration learnt and preserved.

10. The speech recognition server integration method according to claim 9 , further comprising:

a step of acquiring the words registered by the user or the list of words frequently used by the user;

a step of receiving a list of words for recognition from the general-purpose speech recognition servers; and

a step of comparing the list of words for recognition with the words registered by the user or the list of words frequently used by the user and estimating similarity between them,

wherein in the step of learning and preserving the parameter for recognition result integration the estimation result is preserved as the parameter for recognition result integration.

Assignments (4)
CHANGE OF NAME Recorded Jan 7, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 069839/0593 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2022
From: FAURECIA CLARION ELECTRONICS CO., LTD.
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 060368/0973 →
CHANGE OF NAME Recorded Apr 5, 2022
From: CLARION CO., LTD.
To: FAURECIA CLARION ELECTRONICS CO., LTD.
Reel/Frame 059501/0917 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2014
From: OBUCHI, YASUNARI; HOMMA, TAKESHI
To: CLARION CO., LTD.
Reel/Frame 034021/0331 →