IP Library Granted Patent US 8,265,931
Granted Patent B2
US 8,265,931 · App. 12/200,292 · Granted Sep 11, 2012

Method and device for providing speech-to-text encoding and telephony service

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,265,931
App. No.
12/200,292
Granted
Sep 11, 2012
Kind
B2
Abstract

A machine-readable medium and a network device are provided for speech-to-text translation. Speech packets are received at a broadband telephony interface and stored in a buffer. The speech packets are processed and textual representations thereof are displayed as words on a display device. Speech processing is activated and deactivated in response to a command from a subscriber.

Claims (32)

1. A network device comprising:

a network interface that enables communication between the network device and a subscriber terminal;

a display interface configured to communicate with a visual display device to display textual information;

a telephone interface configured to communicate with a telephone to convey voice information of a user;

at least one processor configured to transmit speech information to be decoded and displayed as text on a visual display device of the subscriber terminal during receipt of the speech information; and

a module configured to activate and deactivate speech recognition based on communicated data from the subscriber terminal from a detector configured to respond to subscriber inputs in the form of at least one DTMF.

2. The network device of claim 1 , wherein the detector includes a dual-tone multiple frequency detector.

3. The network device of claim 1 , wherein the at least one processor is configured to identify a caller based on speech segments stored in a database.

4. The network device of claim 1 , wherein the at least one processor is further configured to identify at least one of a gender of a caller, soft-spoken words, hard-spoken words, shouting, laughter, and human expression.

5. A network device comprising:

means for transmitting, via a broadband telephony interface, speech packets to a subscriber terminal such that the speech packets can be processed and displayed as textual representations thereof on a display device of the subscriber terminal; and

means for responding to a command in the form of at least one DTMF tone from the subscriber terminal to activate and deactivate speech processing.

6. The network device of claim 5 , further comprising:

means for storing speech patterns in a database, and

means for analyzing and comparing incoming speech obtained by processing the speech packets with speech patterns stored in the database in order to provide speaker identification capability.

7. The network device of claim 6 , further comprising:

means for transmitting to the subscriber terminal an indication of a speaker identity of a speaker associated with ones of the textual representations.

8. The network device of claim 5 , further comprising:

means for analyzing characteristics of incoming speech obtained by processing the speech packets and inserting punctuation in the textual representations in response to analyzing the characteristics of incoming speech.

9. The network device of claim 8 , wherein the characteristics comprise at least one of changes in tone, volume, and inflection.

10. A non-transitory computer-readable medium having stored therein instructions which, when executed by a processor, cause the processor to perform a method comprising:

receiving, at a broadband telephony interface, speech packets;

responding to a command in the form of at least one DTMF tone, from a subscriber, to activate and deactivate speech processing; and

processing, based on the command, the speech packets to display textual representations thereof as words on a display device.

11. The non-transitory computer-readable medium of claim 10 , the instructions which, when executed by the processor, cause the processor to perform a method further comprising:

storing speech patterns in a database, and

analyzing and comparing incoming speech obtained by processing the speech packets with speech patterns stored in the database in order to provide speaker identification capability.

12. The non-transitory computer-readable medium of claim 11 , the instructions which, when executed by the processor, cause the processor to perform a method further comprising:

displaying an indication of a speaker identity of a speaker associated with ones of the textual representations.

13. The non-transitory computer-readable medium of claim 10 , the instructions which, when executed by the processor, cause the processor to perform a method further comprising:

analyzing characteristics of incoming speech obtained by processing the speech packets and inserting punctuation in the textual representations in response to analyzing the characteristics.

14. The non-transitory computer-readable medium of claim 13 , wherein the characteristics comprise at least one of changes in tone, volume, and inflection.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 038529/0164 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 038529/0240 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 30, 2016
From: CALDWELL, CHARLES DAVID; HARLOW, JOHN BRUCE; SAYKO, ROBERT J.; SHAYE, NORMAN
To: AT&T CORP.
Reel/Frame 038133/0833 →