IP Library Granted Patent US 8,903,053
Granted Patent B2
US 8,903,053 · App. 11/673,746 · Granted Dec 2, 2014

Mass-scale, user-independent, device-independent voice messaging system

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,903,053
App. No.
11/673,746
Granted
Dec 2, 2014
Kind
B2
Abstract

A mass-scale, user-independent, device-independent, voice messaging system that converts unstructured voice messages into text for display on a screen is disclosed. The system comprises (i) computer implemented sub-systems and also (ii) a network connection to human operators providing transcription and quality control; the system being adapted to optimize the effectiveness of the human operators by further comprising 3 core sub-systems, namely (i) a pre-processing front end that determines an appropriate conversion strategy; (ii) one or more conversion resources; and (iii) a quality control sub-system.

Claims (25)

1. A voice messaging system for converting an audio voice message from a caller to text, the voice messaging system comprising:

a plurality of conversion resources for converting the audio voice message into the text for an intended recipient, the plurality of conversion resources comprising:

a network connection to a transcription service that uses at least one human operator to assist in converting the audio voice message into the text;

at least one automatic speech recognition (ASR) system to automatically recognize at least some of the audio voice message to assist the at least one human operator in converting the audio voice message into the text;

a pre-processing front end configured to:

receive the audio voice message from the caller;

optimize the quality of the audio voice message;

determine a confidence level associated with converting the audio voice message; and

determine an appropriate conversion strategy by selecting a particular conversion resource to process the audio voice message based on the confidence level, wherein determining the appropriate conversion strategy based on the confidence level comprises,

if the confidence level is low, flagging the audio voice message as unconvertible and sending a notice that an unconvertible message was received; and

a text output device configured to output the text.

2. The system of claim 1 wherein the pre-processing front end optimizes the quality of the audio voice message by performing one or more of the following functions: removing noise, cleaning up known defects, normalizing volume, and removing sections of the audio voice message that do not contain voice content.

3. The system of claim 1 , wherein the at least one human operator performs quality assurance testing on the converted text and provides feedback to at least one of the pre-processing front-end and the at least one ASR system.

4. The system of claim 1 , wherein:

the audio voice message is a voicemail intended for a mobile telephone; and

the text output device sends the text to the mobile telephone for display on a screen of the mobile telephone.

5. The system of claim 1 , wherein:

the audio voice message is intended for an instant messaging service; and

the text output device sends the text to the instant messaging service for display on a screen of a device of the intended recipient of the audio voice message.

6. The system of claim 1 , wherein the pre-processing front end determines the appropriate conversion strategy by determining a language being used by the caller in the audio voice message based on at least one of a location of the user and a call history associated with the caller.

7. The system of claim 1 , wherein the pre-processing front end determines the appropriate conversion strategy by selecting different conversion resources to convert different portions of the audio voice message.

8. The system of claim 1 , wherein the pre-processing front end determines the appropriate conversion strategy based on the confidence level by:

if the confidence level is high, sending the text recognized by the at least one ASR system to the at least one human operator for quality assurance testing.

9. The system of claim 8 , wherein the pre-processing front end determines the appropriate conversion strategy based on the confidence level by:

if the confidence level is neither high nor low, sending the text recognized by the at least one ASR system to the at least one human operator for checking and correction.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 24, 2013
From: SPINVOX LIMITED
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 031266/0720 →
SECURITY AGREEMENT Recorded Jul 7, 2009
From: SPINVOX LIMITED
To: TISBURY MASTER FUND LIMITED
Reel/Frame 022923/0447 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2007
From: DOULTON, DANIEL MICHAEL
To: SPINVOX LIMITED
Reel/Frame 019015/0163 →