IP Library Granted Patent US 9,947,321
Granted Patent B1
US 9,947,321 · App. 15/438,096 · Granted Apr 17, 2018

Real-time interactive voice recognition and response over the internet

Inventor: Christopher S. Jochumson (Novato, CA)
Assignee: PEARSON EDUCATION, INC.
G10L15/30G10L13/08G10L15/22G10L19/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,947,321
App. No.
15/438,096
Granted
Apr 17, 2018
Kind
B1
Abstract

Methods and systems for handling speech recognition processing in effectively real-time, via the Internet, in order that users do not experience noticeable delays from the start until they receive responsive feedback. A user uses a client to access the Internet and a server supporting speech recognition processing. The user inputs speech to the client, which transmits the user speech to the server in approximate real-time. The server evaluates the user speech, and provides responsive feedback to the client, again, in approximate real-time, with minimum latency delays. The client upon receiving responsive feedback from the server, displays, or otherwise provides, the feedback to the user.

Claims (43)

1. A method comprising:

receiving, by a server from a client device, a data transmission including one or more packets encoding voice data, a first value, and a question count identifier;

processing, by the server, the voice data to generate a response to the voice data, wherein the server processes the voice data for a period of time at least partially determined by the first value; and

transmitting, by the server in response to the data transmission, the response to the voice data to the client device, wherein the response is determined by:

identifying a portion of an audio file using the question count identifier, and

transmitting the portion of the audio file to the client device.

2. The method of claim 1 , wherein the first value includes an accuracy value and the server is configured to process the voice data for an amount of processor time determined by the accuracy value.

3. The method of claim 1 , wherein the first value includes a threshold value and the server is configured to process the voice data by applying a threshold level of recognition to a spoken word in the voice data, wherein the threshold level of recognition is determined by the threshold value.

4. The method of claim 1 , wherein the response includes text data and the client device is configured to process the text data into audio data.

5. The method of claim 1 , wherein transmitting the response to the voice data to the client device over the Internet includes causing a display of visual information on a screen of the client device.

6. The method of claim 1 , wherein transmitting the response to the voice data to the client device over the Internet includes causing an output of audio information on an audio output device of the client device.

7. The method of claim 1 , wherein the response is communicated from server to a client over the Internet, and causes a speech output on an audio output device of the client device.

8. The method of claim 1 , wherein transmitting the response to the voice data to the client device over the Internet includes causing an output of audio and visual information by the client device.

9. A method, comprising:

transmitting, by a client device to a server, a data transmission including one or more packets encoding voice data and a first value;

receiving, by the client device from the server, a response to the voice data, wherein the voice data has been processed by the server for a period of time at least partially determined by the first value;

transmitting, to the server and after receiving the response to the voice data, a request for a correct speech file; and

receiving, from the server, an audio file that is the correct speech file.

10. The method of claim 9 , wherein the first value includes an accuracy value and the server is configured to process the voice data for an amount of processor time determined by the accuracy value.

11. The method of claim 9 , wherein the first value includes a threshold value and the server is configured to process the voice data by applying a threshold level of recognition to a spoken word in the voice data, wherein the threshold level of recognition is determined by the threshold value.

12. The method of claim 9 , wherein the response includes text data and including processing the text data into audio data.

13. The method of claim 9 , including:

processing the response to the voice data into visual information; and

displaying the visual information on a screen of the client device.

14. The method of claim 9 , including:

processing the response to the voice data into audio information; and

outputting the audio information on an audio output device of the client device.

15. The method of claim 9 , including:

processing the response to the voice data into speech output; and

outputting the speech output on an audio output device of the client device.

16. A system, comprising:

a server including a processor configured to:

receive a data transmission from a client device, the data transmission including one or more packets encoding voice data, a first value, and a question count identifier;

process the voice data to generate a response to the voice data, wherein the voice data is processed for a period of time at least partially determined by the first value; and

transmit, in response to the data transmission, the response to the voice data to the client device, wherein the response is determined by:

identifying a portion of an audio file using the question count identifier, and

transmitting the portion of the audio file to the client device.

17. The system of claim 16 , wherein the first value includes an accuracy value and the processor is configured to process the voice data for an amount of processor time determined by the accuracy value.

18. The system of claim 16 , wherein the first value includes a threshold value and the processor is configured to process the voice data by applying a threshold level of recognition to a spoken word in the voice data, wherein the threshold level of recognition is determined by the threshold value.

19. The system of claim 16 , wherein the processor is configured to:

execute one or more speech processing threads; and

input the voice data into the one or more speech processing threads for processing.

20. The system of claim 16 , including a storage device configured to store a plurality of response data files and wherein processing the voice data to generate a response to the voice data includes identifying a first response data file in the plurality of response data files using the voice data and the response to the voice data includes at least a portion of the first response data file.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 21, 2017
From: JOCHUMSON, CHRISTOPHER S.
To: GLOBALENGLISH CORPORATION
Reel/Frame 041320/0826 →
CHANGE OF NAME Recorded Feb 21, 2017
From: GLOBALENGLISH CORPORATION
To: PEARSON ENGLISH CORPORATION
Reel/Frame 041322/0799 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 21, 2017
From: PEARSON ENGLISH CORPORATION
To: PEARSON EDUCATION, INC.
Reel/Frame 041322/0975 →
Continuity (7)
Continuation 14828399 · Aug 17, 2015
Continuation 13407611 · Feb 28, 2012
Continuation 12942834 · Nov 9, 2010
Continuation 11925584 · Oct 26, 2007
Division 10711114 · Aug 24, 2004
Continuation 10199395 · Jul 19, 2002
Continuation 09412043 · Oct 4, 1999