IP Library Granted Patent US 7,382,770
Granted Patent B2
US 7,382,770 · App. 10/374,262 · Granted Jun 3, 2008

Multi-modal content and automatic speech recognition in wireless telecommunication systems

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,382,770
App. No.
10/374,262
Granted
Jun 3, 2008
Kind
B2
Abstract

A communication architecture for delivery of grammar and speech related information such as text-to-speech (TTS) data to a speech recognition server operating with a wireless telecommunication system for use with automatic speech recognition and interactive voice-based applications. In the invention, a mobile client retrieves a Web page containing multi-modal content hosted on a origin server via WAP gateway. The content may include a grammar file and/or TTS strings embedded in the content or reference URL(s) pointing to their storage locations. The client then sends the grammar and/or TTS strings to a speech recognition server via a wireless packet streaming protocol channel. When URL(s) are received by the client and sent to the SRS, the grammar file and/or TTS strings are obtained via a high speed HTTP connection. The speech processing results and the synthesized speech are returned to the client over the established wireless UDP connection.

Claims (39)

1. A method comprising:

sending a request for a Web page from a mobile client to a gateway, wherein the mobile client is in wireless communication with the gateway;

retrieving the Web page from an origin server to the gateway;

returning the Web page to the client;

determining whether the Web page contains multi-modal components;

sending a multi-modal components from the client to a speech recognition server using a direct wireless packet streaming protocol connection, wherein the speech recognition server includes a speech recognizer and a text-to-speech synthesizer;

obtaining a grammar file or text-to-speech markup strings by the speech recognition server from a remotely located server using an established hypertext transport protocol network connection from at least one universal resource locator references sent from the client;

loading the received grammars in a speech recognizer for performing speech recognition and text-to-speech markup strings into the speech synthesizer for producing synthesized speech; and

returning speech recognition results from the speech recognizer and produced synthesized speech to the client over said wireless packet streaming protocol connection.

2. The method according to claim 1 , wherein said wireless telecommunication system operates in accordance with Wireless Application Protocol.

3. The method according to claim 1 , wherein the multi-modal components include grammar, text-to-speech markup strings, pre-recorded audio, video, or music markup, or universal resource locator references of any of those mentioned.

4. The method according to claim 3 , wherein the grammar and text-to-speech markup strings are embedded in the Web page.

5. The method according to claim 1 , wherein the wireless packet streaming protocol connection is a wireless user datagram protocol connection.

6. An apparatus comprising:

a processor configured to control operations of the apparatus; and

memory storing executable instructions that, when executed by the processor, cause the apparatus to perform:

interfacing with a proxy gateway via a data protocol standard to retrieve a Web page located on an origin server;

extracting multi-modal components from said Web page for transmission to a speech recognition server;

generating speech parameters for use with said speech recognition sewer;

establishing a direct wireless packet streaming protocol connection for wireless communication with said speech recognition server;

sending the multi-modal components to the speech recognition server using the established direct wireless packet streaming protocol connection; and

receiving speech recognition results and produced synthesized speech from the speech recognition server via said wireless packet streaming protocol connection.

7. The apparatus according to claim 6 , wherein the data protocol standard is Wireless Application Protocol.

8. The apparatus according to claim 6 , wherein said multi-modal components includes any one of grammar, text-to-speech markup strings, pre-recorded audio, video, or music markup, or URL references of any of those mentioned.

9. The apparatus according to claim 6 , wherein the generated speech parameters in the client are used together with a distributed speech recognition system comprising a remote speech recognition server.

10. The apparatus according to claim 6 , wherein the packet streaming protocol connection is a wireless user datagram protocol connection.

11. An apparatus, comprising:

a processor configured to control operations of the apparatus: and

memory storing executable instructions that, when executed by the processor, cause the apparatus to perform:

establishing a direct wireless packet streaming protocol connection with a mobile client;

receiving multi-modal components from the client via the established direct wireless packet streaming protocol connection;

obtaining a grammar file or text-to-speech markup strings from a remotely located server using an established hypertext transport protocol network connection from at least one universal resource locator reference sent from the mobile client;

loading the received grammars in a speech recognizer for performing speech recognition and text-to-speech markup strings into the speech synthesizer for producing synthesized speech; and

returning speech recognition results from the speech recognizer and produced synthesized speech to the mobile client over the said wireless packet streaming protocol connection.

12. The apparatus according to claim 11 , wherein the wireless packet streaming protocol connection is a wireless user datagram protocol connection.

13. The apparatus according to claim 12 , wherein the mobile client and apparatus each possesses a user datagram protocol port and associated hardware and software to facilitate communication via a wireless user datagram protocol connection.

14. The apparatus according to claim 11 , wherein the speech recognition server further comprises:

a speech recognizer, a text-to-speech processor, and security hardware and software for ensuring the secure transfer of communications data.

15. The apparatus according to claim 11 , wherein the hypertext transport protocol network connection is a high speed Internet connection.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 9, 2015
From: NOKIA CORPORATION
To: NOKIA TECHNOLOGIES OY
Reel/Frame 035601/0919 →