IP Library Granted Patent US 7,363,228
Granted Patent B2
US 7,363,228 · App. 10/666,956 · Granted Apr 22, 2008

Speech recognition system and method

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,363,228
App. No.
10/666,956
Granted
Apr 22, 2008
Kind
B2
Abstract

A computer system and method is disclosed that includes a telephony server that receives a spoken dialing command, sends the command to a speech recognition server, and dials a command based on the result. A computer system and method is disclosed that improves audio message delivery reliability. A computer system and method is disclosed that improves audio message manipulation. A computer system and method is disclosed that manages memory when audio messages are received. A system and method is disclosed that supports multiple speech recognition engines.

Claims (47)

1. A method comprising:

detecting a phone in an off-hook state;

retrieving with a telephony server information associated with a user assigned to the phone;

generating a custom input grammar with the telephony server using the information;

generating a dial-tone with the telephony server;

receiving with the telephony server a command spoken into the phone;

processing the spoken command with the telephony server to locate a corresponding entry in the custom input grammar; and

executing a command operation associated with the corresponding entry.

2. The method of claim 1 , wherein the custom input grammar is not generated until an identification of a person who spoke the command is performed, and wherein the custom input grammar is then generated based on the particular profile of the person.

3. The method of claim 1 , wherein said processing comprises:

sending the spoken command to a speech recognition server, said speech recognition server processing the spoken command, locating the corresponding entry in the custom input grammar, and returning the corresponding entry to the telephony server.

4. The method of claim 3 , wherein the speech recognition server verifies the identity of a person that spoke the command to ensure the person is authorized to access the custom input grammar before locating the corresponding entry in the custom input grammar.

5. The method of claim 1 , wherein the custom input grammar is generated from a text-based contacts database associated with the user assigned to the phone.

6. The method of claim 5 , wherein the spoken command is a name of a person in the text-based contacts database associated with the user assigned to the phone.

7. The method of claim 1 , wherein the command is spoken into the phone by a person other than the user assigned to the phone.

8. The method of claim 1 , wherein said generating the custom input grammar is only performed if the custom input grammar does not already exist for the user associated with the phone or if the custom grammar exists but needs updated due to modifications in an underlying data source.

9. The method of claim 1 , wherein the dial-tone is cancelled when the telephony application processor begins receiving the command spoken into the phone.

10. The method of claim 1 , wherein said processing the spoken command with the telephony application processor comprises:

sending a recognition request to a speech recognition server;

receiving a probing request from the speech recognition server;

sending a UDP probe response message to a probing port number of the speech recognition server;

sending the spoken command to the speech recognition server, said speech recognition server determining a translated result based on the custom input grammar; and

receiving the translated result from the speech recognition server.

11. A system comprising:

a speech recognition server; and

a telephony application server coupled to the speech recognition server over a network, the telephony application server being operative to detect a phone in an off-hook state, retrieve information associated with a user assigned to the phone, generate a custom input grammar using the information, generate a dial-tone, receive a command spoken into the phone, send the spoken command to the speech recognition server, receive a corresponding entry based on the custom input grammar from the speech recognition server and execute a command operation associated with the corresponding entry.

12. The system of claim 11 , wherein the speech recognition server is operative to support a plurality of speech recognition engines.

13. The system of claim 11 , wherein the speech recognition server is operative to send a port number of a probing endpoint to the telephony application server, send a probing request to the telephony application server, and receive from the telephony application server a UDP probe response message at the port number.

14. A system comprising:

multiple speech recognition engines residing on one or more speech recognition servers; and

a telephony server having a telephony application processor operable to translate vendor-neutral interfaces to and from a specific syntax requires by each of the multiple recognition engines.

15. The system of claim 14 , wherein the telephony application processor is operable to perform speaker identification and verification as part of a recognition operation.

16. The system of claim 14 , wherein the telephony application processor is operable to send recognition requests to at least two of the multiple speech recognition engines at the same time.

17. A method, comprising:

offering a telephony application interface routine including a voice recognition interface operable with multiple speech recognition engines;

providing the telephony application interface to a first customer having a pre-established grammar for a first one of the speech recognition engines;

the first customer operating the telephony application interface with the pre-established grammar of the first one of the speech recognition engines;

providing the telephony application interface to a second customer having a second one of the speech recognition engines; and

the second customer operating the telephony application interface with the second one of the speech recognition engines.

18. A method comprising:

detecting a user being connected to a telephony server;

identifying the user;

retrieving information associated with the user;

generating a custom input grammar using the information;

receiving with the telephony server a command spoken by the user;

processing the spoken command to locate a corresponding entry in the custom input grammar; and

executing a command operation associated with the corresponding entry.

Assignments (6)
NOTICE OF SUCCESSION OF SECURITY INTERESTS AT REEL/FRAME 040815/0001 Recorded Feb 3, 2025
From: BANK OF AMERICA, N.A., AS RESIGNING AGENT
To: GOLDMAN SACHS BANK USA, AS SUCCESSOR AGENT
Reel/Frame 070498/0001 →
CHANGE OF NAME Recorded Jun 7, 2024
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: GENESYS CLOUD SERVICES, INC.
Reel/Frame 067651/0783 →
MERGER Recorded Jul 1, 2018
From: INTERACTIVE INTELLIGENCE GROUP, INC.
To: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
Reel/Frame 046463/0839 →
SECURITY AGREEMENT Recorded Dec 5, 2016
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC., AS GRANTOR; ECHOPASS CORPORATION; INTERACTIVE INTELLIGENCE GROUP, INC.; BAY BRIDGE DECISION TECHNOLOGIES, INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 040815/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2016
From: INTERACTIVE INTELLIGENCE, INC.
To: INTERACTIVE INTELLIGENCE GROUP, INC.
Reel/Frame 040647/0285 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2004
From: WYSS, FELIX I.; O'CONNOR, KEVIN
To: INTERACTIVE INTELLIGENCE, INC.
Reel/Frame 014255/0734 →