IP Library Granted Patent US 9,324,317
Granted Patent B2
US 9,324,317 · App. 14/481,326 · Granted Apr 26, 2016

System and method for synthetically generated speech describing media content

Inventors: Linda Roberts (Boynton Beach, FL); Hong Thi Nguyen (Atlanta, GA); Horst J. Schroeter (New Providence, NJ)
Assignee: AT&T Intellectual Property I, L.P.
G10L13/033G06F3/017G06F3/04842G06F3/167G10L13/00G10L13/043G10L13/086G10L2013/083
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,324,317
App. No.
14/481,326
Granted
Apr 26, 2016
Kind
B2
Abstract

Disclosed herein are systems, methods, and computer readable-media for providing an automatic synthetically generated voice describing media content, the method comprising receiving one or more pieces of metadata for a primary media content, selecting at least one piece of metadata for output, and outputting the at least one piece of metadata as synthetically generated speech with the primary media content. Other aspects of the invention involve alternative output, output speech simultaneously with the primary media content, output speech during gaps in the primary media content, translate metadata in foreign language, tailor voice, accent, and language to match the metadata and/or primary media content. A user may control output via a user interface or output may be customized based on preferences in a user profile.

Claims (37)

1. A method comprising:

receiving a gesture from a user during a presentation of media content, wherein the gesture comprises a metadata request associated with the media content;

selecting a piece of metadata for output, to yield selected metadata, the selected metadata being responsive to the metadata request regarding the primary media content; and

outputting the selected metadata as synthetically generated speech, the synthetically generated speech having an accent selected from a plurality of accents based on the selected metadata.

2. The method of claim 1 , wherein the synthetically generated speech is output during the presentation of the media content.

3. The method of claim 1 , further comprising analyzing the media content to determine tone and prosody which indicate the accent.

4. The method of claim 1 , wherein the synthetically generated speech is output during gaps in the presentation of the media content.

5. The method of claim 1 , further comprising:

determining the metadata is in a foreign language, where the accent corresponds to the foreign language; and

translating the metadata to another language from the foreign language before output.

6. The method of claim 1 , wherein the gesture is accompanied by an oral command.

7. The method of claim 1 , wherein the metadata is output via a distinct output from that of the presentation of media content.

8. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

receiving a gesture from a user during a presentation of media content, wherein the gesture comprises a metadata request associated with the media content;

selecting a piece of metadata for output, to yield selected metadata, the selected metadata being responsive to the metadata request regarding the primary media content; and

outputting the selected metadata as synthetically generated speech, the synthetically generated speech having an accent selected from a plurality of accents based on the selected metadata.

9. The system of claim 8 , wherein the synthetically generated speech is output during the presentation of the media content.

10. The system of claim 8 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, result in operations comprising analyzing the media content to determine tone and prosody which indicate the accent.

11. The system of claim 8 , wherein the synthetically generated speech is output during gaps in the presentation of the media content.

12. The system of claim 8 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, result in operations comprising:

determining the metadata is in a foreign language, where the accent corresponds to the foreign language; and

translating the metadata to another language from the foreign language before output.

13. The system of claim 8 , wherein the gesture is accompanied by an oral command.

14. The system of claim 8 , wherein the metadata is output via a distinct output from that of the presentation of media content.

15. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

receiving a gesture from a user during a presentation of media content, wherein the gesture comprises a metadata request associated with the media content;

selecting a piece of metadata for output, to yield selected metadata, the selected metadata being responsive to the metadata request regarding the primary media content; and

outputting the selected metadata as synthetically generated speech, the synthetically generated speech having an accent selected from a plurality of accents based on the selected metadata.

16. The computer-readable storage device of claim 15 , wherein the synthetically generated speech is output during the presentation of the media content.

17. The computer-readable storage device of claim 15 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, result in operations comprising analyzing the media content to determine tone and prosody which indicate the accent.

18. The computer-readable storage device of claim 15 , wherein the synthetically generated speech is output during gaps in the presentation of the media content.

19. The computer-readable storage device of claim 15 , having additional instructions stored which, when executed by the computing device, result in operations comprising:

determining the metadata is in a foreign language, where the accent corresponds to the foreign language; and

translating the metadata to another language from the foreign language before output.

20. The computer-readable storage device of claim 15 , wherein the gesture is accompanied by an oral command.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 23, 2015
From: AT&T LABS, INC.
To: AT&T INTELLECTUAL PROPERTY I, L.P.
Reel/Frame 037121/0460 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 21, 2015
From: ROBERTS, LINDA; NGUYEN, HONG THI; SCHROETER, HORST J.
To: AT&T LABS, INC.
Reel/Frame 036846/0170 →
Continuity (2)
Continuation 12134714 · Jun 6, 2008
Related Publication 20140379350A1 · Dec 25, 2014