IP Library Granted Patent US 8,750,463
Granted Patent B2
US 8,750,463 · App. 11/931,038 · Granted Jun 10, 2014

Mass-scale, user-independent, device-independent voice messaging system

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,750,463
App. No.
11/931,038
Granted
Jun 10, 2014
Kind
B2
Abstract

A mass-scale, user-independent, device-independent, voice messaging system that converts unstructured voice messages into text for display on a screen is disclosed. The system comprises (i) computer implemented sub-systems and also (ii) a network connection to human operators providing transcription and quality control; the system being adapted to optimise the effectiveness of the human operators by further comprising 3 core sub-systems, namely (i) a pre-processing front end that determines an appropriate conversion strategy; (ii) one or more conversion resources; and (iii) a quality control sub-system.

Claims (20)

1. A voice messaging system for converting an audio voice message from a caller into text, the voice messaging system comprising:

a plurality of conversion resources for converting the audio voice message into the text for an intended recipient, the plurality of conversion resources comprising:

at least one automatic speech recognition (ASR) system to automatically recognize at least some of the audio voice message and generate a plurality of candidate word or phrase sequences; and

a computer implemented boundary selection sub-system adapted to process the audio voice message to determine a type of at least one portion of the audio voice message as being a greeting portion, a message body portion and/or a tail portion of the audio voice message received from the caller and to select, based on a selected optimal conversion strategy, one or more different pieces of ASR for recognizing the at least one portion based on the type determined.

2. The system of claim 1 in which different conversion strategies are applied to each portion for which a type is determined, the applied strategy being optimal for converting the portion of the audio voice message to text based on whether the respective portion is determined to be a greeting portion, a message body portion or a tail portion.

3. The system of claim 1 in which the greeting portion, the message body portion and the tail portion of the audio voice message have different quality requirements and a quality assessment sub-system applies the respective standard to each portion for which the type is determined.

4. The system of claim 1 in which the type of the at least one portion is determined at least in part by detecting a boundary between portions of the message corresponding to the greeting portion, the message body portion and/or the tail portion.

5. The system of claim 4 in which the boundary between the greeting portion, the message body portion and/or the tail portion of the audio voice message are detected or inferred at regions in the message where the speech density alters.

6. The system of claim 4 in which the boundary between the greeting portion, the message body portion and/or the tail portion of the audio voice message are detected or inferred at a pause in the message.

7. The system of claim 4 in which the boundary between the greeting portion, the message body portion and/or the tail portion of the audio voice message are detected or inferred as arising at a pre-defined proportion of the message.

8. The system of claim 4 in which a boundary of a greeting portion is inferred at 15% or less of the entire message length.

9. The system of claim 1 in which the audio voice message is intended for a mobile telephone and the audio voice message is converted to text and sent to the mobile telephone.

10. The system of claim 1 in which the audio voice message is intended for an instant messaging service and the audio voice message is converted to text and sent to the instant messaging service for display on a screen.

11. The system of claim 1 in which the audio voice message is intended for a web service and the audio voice message is converted to text and sent to a server for display as part of the web service.

12. The system of claim 1 in which the audio voice message is converted to text format and sent as a text message.

13. The system of claim 1 in which the audio voice message is converted to text format and sent as an email message.

14. The system of claim 1 in which the audio voice message is converted to text format and sent as a note or memo, by email or text, to an originator of the message.

15. The system of claim 1 further comprising a mobile telephone network.

16. The system of claim 1 further comprising a mobile telephone for displaying the text converted from the audio voice message.

17. The system of claim 1 further comprising a computer display screen for displaying the text converted from the audio voice message.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 24, 2013
From: SPINVOX LIMITED
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 031266/0720 →
CORRECTIVE ASSIGNMENT TO CORRECT THE NUMBER US0884215 BY REMOVING THIS NUMBER FROM THE SECURITY AGREEMENT PREVIOUSLY RECORDED ON REEL 022923 FRAME 0447. ASSIGNOR(S) HEREBY CONFIRMS THE SECURITY AGREEMENT. Recorded Aug 10, 2010
From: SPINVOX LIMITED
To: TISBURY MASTER FUND LIMITED
Reel/Frame 024819/0229 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 24, 2009
From: DOULTON, DANIEL MICHAEL
To: SPINVOX LIMITED
Reel/Frame 023700/0134 →
SECURITY AGREEMENT Recorded Jul 7, 2009
From: SPINVOX LIMITED
To: TISBURY MASTER FUND LIMITED
Reel/Frame 022923/0447 →