IP Library Granted Patent US 7,689,415
Granted Patent B1
US 7,689,415 · App. 11/925,558 · Granted Mar 30, 2010

Real-time speech recognition over the internet

Assignee: GlobalEnglish Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,689,415
App. No.
11/925,558
Granted
Mar 30, 2010
Kind
B1
Abstract

Methods and systems for handling speech recognition processing in effectively real-time, via the Internet, in order that users do not experience noticeable delays from the start of an exercise until they receive responsive feedback. A user uses a client to access the Internet and a server supporting speech recognition processing, e.g., for language learning activities. The user inputs speech to the client, which transmits the user speech to the server in approximate real-time. The server evaluates the user speech in context of the current speech recognition exercise being executed, and provides responsive feedback to the client, again, in approximate real-time, with minimum latency delays. The client upon receiving responsive feedback from the server, displays, or otherwise provides, the feedback to the user.

Claims (43)

1. A system supporting speech recognition comprising:

two or more clients, each client comprising the capability to receive audio speech from a user, store the audio speech in one or more buffers, each buffer comprising a portion of the received audio speech, encode a buffer of the received audio speech before all of the audio speech is received, package the encoded buffer to receive audio speech into one or more packets to be transmitted over the Internet before all of the audio speech is received, and transmit a packet of encoded audio speech over the Internet before all of the audio speech is received; and

a server, the server comprising the capability to receive packets of encoded audio speech from at least two clients, decode each of the packets of audio speech and store the resultant raw speech into one or more buffers for the respective client, and evaluate the resultant raw speech received from each of the at least two clients,

wherein the server further comprises the capability to transmit a response to a client, the response a result of the server's evaluation of the resultant raw speech received from the client,

a client of the two or more clients further comprises the capability to receive the response from the server,

the response is in text format, and a client of the two or more clients comprises a text-to-speech engine which converts the text format response to audio data, and an audio output device that the client uses to output the audio data to a user, and

a level of processing used in the evaluation of the resultant raw speech received from a client is alterable based on a parameter communicated between the client and the server.

2. The system of claim 1 wherein the parameter is communicated to the server by URL.

3. The system of claim 1 wherein a software application, executing on a client, receives the audio speech from the user, and the software application of the client converts the text format response to audio data.

4. A system supporting speech recognition comprising:

two or more clients, each client comprising the capability to receive audio speech from a user, store the audio speech in one or more buffers, each buffer comprising a portion of the received audio speech, encode a buffer of the received audio speech before all of the audio speech is received, package the encoded buffer to receive audio speech into one or more packets to be transmitted over the Internet before all of the audio speech is received, and transmit a packet of encoded audio speech over the Internet before all of the audio speech is received; and

a server, the server comprising the capability to receive packets of encoded audio speech from at least two clients, decode each of the packets of audio speech and store the resultant raw speech into one or more buffers for the respective client, and evaluate the resultant raw speech received from each of the at least two clients,

wherein the server further comprises the capability to transmit a response to a client, the response a result of the server's evaluation of the resultant raw speech received from the client,

a client of the two or more clients further comprises the capability to receive the response from the server,

the response is in text format, and a client of the two or more clients comprises a text-to-speech engine which converts the text format response to audio data, and an audio output device that the client uses to output the audio data to a user, and

a processing time used to evaluate the resultant raw speech will vary based on a value communicated to the server from a client.

5. The system of claim 4 wherein the value communicated to the server from a client is a user objective, a user selects the user objective at a client, the client transmits the user objective to the server, and the server evaluates the resultant raw speech received from the client based on the user objective.

6. The system of claim 4 wherein the value communicated to the server from a client is a pronunciation accuracy objective, a user selects the pronunciation accuracy objective at a client, the client transmits the pronunciation accuracy objective to the server, and the server evaluates the resultant raw speech received from the client based on the pronunciation accuracy objective.

7. The system of claim 4 wherein the encoded audio speech is in a compressed format.

8. The system of claim 4 wherein before the client receives audio speech from a user, the server transmits a file to a client, the client presents the file in at least one of an audio or visual format to the user.

9. The system of claim 4 wherein the client shows the text format response on a display of the client.

10. A system supporting speech recognition comprising:

two or more clients, each client comprising the capability to receive audio speech from a user, store the audio speech in one or more buffers, each buffer comprising a portion of the received audio speech, encode a buffer of the received audio speech before all of the audio speech is received, package the encoded buffer to receive audio speech into one or more packets to be transmitted over the Internet before all of the audio speech is received, and transmit a packet of encoded audio speech over the Internet before all of the audio speech is received; and

a server, the server comprising the capability to receive packets of encoded audio speech from at least two clients, decode each of the packets of audio speech and store the resultant raw speech into one or more buffers for the respective client, and evaluate the resultant raw speech received from each of the at least two clients,

wherein the server further comprises two or more stored text format files, and the server selects a stored text format file to transmit to a client of the two or more clients as a result of the server's evaluation of the resultant raw speech received from the client,

the server further comprises the capability to partition a stored text format file into two or more packets for the transmission over the Internet, and to transmit each packet over the Internet to a client,

a client further comprises an audio output device, and the capability to receive the packets of text format, convert the packets of text format to audio data and play the audio data to a user, and

a level of processing used in the evaluation of the resultant raw speech received from a client is alterable based on a parameter communicated between the client and the server.

11. The system of claim 10 wherein the parameter is communicated to the server by URL.

12. The system of claim 10 wherein the client shows the received text format file on a display of the client.

13. The system of claim 10 wherein a software application, executing on a client, receives the audio speech from the user, and the software application converts the packets of text format to audio data.

14. The system of claim 10 wherein the two or more stored text format files are different from each other.

15. A system supporting speech recognition comprising:

two or more clients, each client comprising the capability to receive audio speech from a user, store the audio speech in one or more buffers, each buffer comprising a portion of the received audio speech, encode a buffer of the received audio speech before all of the audio speech is received, package the encoded buffer to receive audio speech into one or more packets to be transmitted over the Internet before all of the audio speech is received, and transmit a packet of encoded audio speech over the Internet before all of the audio speech is received; and

a server, the server comprising the capability to receive packets of encoded audio speech from at least two clients, decode each of the packets of audio speech and store the resultant raw speech into one or more buffers for the respective client, and evaluate the resultant raw speech received from each of the at least two clients,

wherein the server further comprises two or more stored text format files, and the server selects a stored text format file to transmit to a client of the two or more clients as a result of the server's evaluation of the resultant raw speech received from the client,

the server further comprises the capability to partition a stored text format file into two or more packets for the transmission over the Internet, and to transmit each packet over the Internet to a client,

a client further comprises an audio output device, and the capability to receive the packets of text format, convert the packets of text format to audio data and play the audio data to a user, and

a processing time used to evaluate the resultant raw speech will vary based on a value communicated to the server from a client.

16. The system of claim 15 wherein the value communicated to the server from a client is a user objective, a user selects the user objective at a client, the client transmits the user objective to the server, and the server evaluates the resultant raw speech received from the client based on the user objective.

17. The system of claim 15 wherein the value communicated to the server from a client is a pronunciation accuracy objective, a user selects the pronunciation accuracy objective at a client, the client transmits the pronunciation accuracy objective to the server, and the server evaluates the resultant raw speech received from the client based on the pronunciation accuracy objective.

18. The system of claim 15 wherein the encoded audio speech is in a compressed format.

19. The system of claim 15 wherein before the client receives audio speech from a user, the server transmits a file to a client, the client presents the file in at least one of an audio or visual format to the user.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 4, 2017
From: JOCHUMSON, CHRISTOPHER SCOTT
To: GLOBALENGLISH CORPORATION
Reel/Frame 042246/0377 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2016
From: PEARSON ENGLISH CORPORATION
To: PEARSON EDUCATION, INC.
Reel/Frame 040214/0022 →
CHANGE OF NAME Recorded Sep 21, 2016
From: GLOBALENGLISH CORPORATION
To: PEARSON ENGLISH CORPORATION
Reel/Frame 039817/0925 →
Continuity (3)
Division 1071111400 · Aug 24, 2004
Continuation 1019939500 · Jul 19, 2002
Continuation 0941204300 · Oct 4, 1999