IP Library Granted Patent US 7,330,815
Granted Patent B1
US 7,330,815 · App. 10/711,114 · Granted Feb 12, 2008

Method and system for network-based speech recognition

Assignee: GlobalEnglish Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,330,815
App. No.
10/711,114
Granted
Feb 12, 2008
Kind
B1
Abstract

Methods and systems for handling speech recognition processing in effectively real-time, via the internet, in order that users do not experience noticeable delays from the start of an exercise until they receive responsive feedback. A user uses a client to access the internet and a server supporting speech recognition processing, e.g., for language learning activities. The user inputs speech to the client, which transmits the user speech to the server in approximate real-time. The server evaluates the user speech in context of the current speech recognition exercise being executed, and provides responsive feedback to the client, again, in approximate real-time, with minimum latency delays. The client upon receiving responsive feedback from the server, displays, or otherwise provides, the feedback to the user.

Claims (45)

1. A system supporting speech recognition comprising:

two or more clients, each client comprising the capability to receive audio speech from a user, store the audio speech in one or more buffers, each buffer comprising a portion of the received audio speech, encode a buffer of the received audio speech before all of the audio speech is received, package the encoded buffer to receive audio speech into one or more packets to be transmitted over the internet before all of the audio speech is received, and transmit a packet of encoded audio speech over the internet before all of the audio speech is received; and

a server, the server comprising the capability to receive packets of encoded audio speech from at least two clients, decode each of the packets of audio speech and store the resultant raw speech into one or more buffers for the respective client, and a processing time used to evaluate the resultant raw speech will vary based on a value communicated to the server from each respective client.

2. The system of claim 1 wherein the server further comprises the capability to transmit a response to a client, the response a result of the server's evaluation of the resultant raw speech received from the client, and

a client of the two or more clients further comprises the capability to receive the response from the server.

3. The system of claim 2 wherein the response is a text response, and a client of the two or more clients comprises a screen on which the client displays the text response.

4. The system of claim 1 wherein the one or more buffers comprise a linked list of buffers.

5. The system of claim 1 wherein a user selects a user objective at a client, the client transmits the user objective to the server, and the server evaluates the resultant raw speech received from the client based on the user objective.

6. The system of claim 5 wherein the user objective comprises pronunciation accuracy.

7. The system of claim 5 wherein the user objective comprises grammar.

8. The system of claim 1 wherein the encoded audio speech is in a compressed format.

9. The system of claim 1 wherein the value is communicated by URL.

10. A system supporting speech recognition comprising:

two or more clients, each client comprising the capability to receive audio speech from a user, store the audio speech in one or more buffers in a raw uncompressed audio format each buffer comprising a portion of the received audio speech encode a buffer of the received audio speech before all of the audio speech is received package the encoded buffer to receive audio speech into one or more packets to be transmitted over the Internet before all of the audio speech is received, and transmit a packet of encoded audio speech over the Internet before all of the audio speech is received; and

a server, the server comprising the capability to receive packets of encoded audio speech from at least two clients decode each of the packets of audio speech and store the resultant raw speech into one or more buffers for the respective client and evaluate the resultant raw speech received from each of the at least two clients,

wherein the server further comprises two or more stored text format files, and the server selects a stored text format file to transmit to a client of the two or more clients as a result of the server's evaluation of the resultant raw speech received from the client, and the server adjusts a processing time used to evaluate the resultant raw speech based on a value in a URL connection between the client and the server.

11. The system of claim 10 wherein the server further comprises the capability to partition a stored text format file into two or more packets for the transmission over the Internet, and to transmit each packet over the Internet to a client.

12. A system supporting speech recognition comprising:

two or more clients, each client comprising the capability to receive audio speech from a user, store the audio speech in one or more buffers organized as a linked list in a raw uncompressed audio format, each buffer comprising a portion of the received audio speech, encode a buffer of the received audio speech before all of the audio speech is received, package the encoded buffer to receive audio speech into one or more packets to be transmitted over a network before all of the audio speech is received, and transmit a packet of encoded audio speech over the network before all of the audio speech is received; and

a server, the server comprising the capability to receive packets of encoded audio speech from at least two clients, decode each of the packets of audio speech and store the resultant raw speech into one or more buffers for the respective client, and evaluate the resultant raw speech received from each of the at least two clients, wherein a level of processing used in the evaluation of the resultant raw speech received from each of the at least two clients is alterable based on a value communicated between the clients and the server.

13. The system of claim 12 wherein the encoded audio speech is in a compressed format.

14. The system of claim 12 wherein the server further comprises the capability to transmit a response to a client, the response a result of the server's evaluation of the resultant raw speech received from the client, and

a client of the two or more clients further comprises the capability to receive the response from the server.

15. The system of claim 14 wherein the response is a text response, and a client of the two or more clients comprises a screen on which the client displays the text response.

16. The system of claim 12 wherein a user selects a user objective at a client, the client transmits the user objective to the server, and the server evaluates the resultant raw speech received from the client based on the user objective.

17. The system of claim 16 wherein the user objective comprises pronunciation accuracy.

18. The system of claim 16 wherein the user objective comprises grammar.

19. The system of claim 12 wherein the one or more buffers comprise a linked list of buffers.

20. The system of claim 12 wherein the server further comprises the capability to partition a stored text format file into two or more packets for the transmission over the network, and to transmit each packet over the network to a client.

21. The system of claim 12 wherein the value is communicated by URL.

22. A system supporting speech recognition comprising:

two or more clients, each client comprising the capability to receive audio speech from a user, store the audio speech in one or more buffers in a raw uncompressed audio format, each buffer comprising a portion of the received audio speech, encode a buffer of the received audio speech before all of the audio speech is received, package the encoded buffer to receive audio speech into one or more packets to be transmitted over the internet before all of the audio speech is received, and transmit a packet of encoded audio speech over the internet before all of the audio speech is received; and

a server, the server comprising the cap ability to receive packets of encoded audio speech from at least two clients, decode each of the packets of audio speech and store the resultant raw speech into one or more buffers for the respective client, and evaluate the resultant raw speech received from each of the at least two clients,

wherein a user selects a user objective at a client, the client transmits the user objective to the server, and the server evaluates the resultant raw speech received from the client based on the user objective and a value communicated to the server by URL.

23. A system supporting speech recognition comprising:

two or more clients, each client comprising the capability to receive audio speech from a user, store the audio speech in one or more buffers in a raw uncompressed audio format, each buffer comprising a portion of the received audio speech, encode a buffer of the received audio speech before all of the audio speech is received package the encoded buffer to receive audio speech into one or more packets to be transmitted over the internet before all of the audio speech is received, and transmit a packet of encoded audio speech over the internet before all of the audio speech is received; and

a server, the server comprising the capability to receive packets of encoded audio speech from at least two clients, decode each of the packets of audio speech and store the resultant raw speech into one or more buffers for the respective client and evaluate the resultant raw speech received from each of the at least two clients,

wherein before the client receives audio speech from a user, the server transmits a file to a client, the client presents the file in at least one of an audio or visual format to the user, and the server evaluates the resultant raw speech received from the client in connection with the file transmitted from the server to the client and a processing time used to evaluate the resultant raw speech will vary based on a value communicated to the server from the client.

24. A system supporting speech recognition comprising:

two or more clients, each client comprising the capability to receive audio speech from a user, store the audio speech in one or more buffers in a raw uncompressed audio format, each buffer comprising a portion of the received audio speech, encode a buffer of the received audio speech before all of the audio speech is received package the encoded buffer to receive audio speech into one or more packets to be transmitted over the internet before all of the audio speech is received, and transmit a packet of encoded audio speech over the internet before all of the audio speech is received; and

a server, the server comprising the capability to receive packets of encoded audio speech from at least two clients decode each of the packets of audio speech and store the resultant raw speech into one or more buffers for the respective client, and evaluate the resultant raw speech received from each of the at least two clients,

wherein the server transmits a first file to a client, the client presents the first file in at least one of an audio or visual format to the user, after presenting the first file to the user, the client receives audio speech from the user, and

the server evaluates the resultant raw speech received from the client in connection with the first file transmitted from the server to the client and a processing time used by the server to evaluate the resultant raw speech is alterable based on a value communicated from the client to the server.

25. The system of claim 24 wherein the server transmits a second file to the client, the client presents the second file in at least one of an audio or visual format to the user, after presenting the second file to the user, the client receives audio speech from the user, and

the server evaluates the resultant raw speech received from the client in connection with the second file transmitted from the server to the client.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2016
From: PEARSON ENGLISH CORPORATION
To: PEARSON EDUCATION, INC.
Reel/Frame 040214/0022 →
CHANGE OF NAME Recorded Sep 21, 2016
From: GLOBALENGLISH CORPORATION
To: PEARSON ENGLISH CORPORATION
Reel/Frame 039817/0925 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 24, 2004
From: JOCHUMSON, CHRISTOPHER S
To: GLOBALENGLISH CORPORATION
Reel/Frame 015029/0233 →
Continuity (2)
Continuation 1019939500 · Jul 19, 2002
Continuation 0941204300 · Oct 4, 1999