IP Library Granted Patent US 11,238,899
Granted Patent B1
US 11,238,899 · App. 16/739,918 · Granted Feb 1, 2022

Efficient audio description systems and methods

Inventors: Joshua Miller (Charlestown, MA); Christopher S. Antunes (Boston, MA); Lily Megan Berthold-Bond (Cambridge, MA); Andrew H. Schwartz (Roslindale, MA); Kelly J. Savietta (Madison, WI); Lanya Lee Butler (Somerville, MA); Sharon Lee Tomasulo (Medford, MA); Jeremy E. Barron (Boston, MA); Christopher E. Johnson (Belmont, MA); Roger S. Zimmerman (Boston, MA)
Assignee: 3Play Media Inc.
G11B27/036G10L15/26G10L21/055G11B27/10G11B27/28H04N5/9305H04N9/8715
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,238,899
App. No.
16/739,918
Granted
Feb 1, 2022
Kind
B1
Abstract

A computer system configured to generate an audio description of a media file is provided. The system includes a display, a memory, and a processor coupled to the display and the memory. The memory stores a media file, including video data that is accessible via a time index and audio data synchronized with the video data by the time index and a transcript of the audio data, including transcription data synchronized with the video data via the time index. The processor is configured to render, via the display, images from portions of the video data; render text from portions of the transcription data in synchrony with the images; receive input identifying a point within the time index; receive input specifying audio description data to associate with the point; store, in the memory, the audio description data; and store an association between the audio description data and the point.

Claims (89)

1. A computer system configured to generate an audio description of a media file, the computer system comprising:

a display;

a memory storing

a media file comprising video data accessible via a time index and audio data synchronized with the video data via the time index; and

a transcript of the audio data comprising transcription data synchronized with the video data via the time index; and

at least one processor coupled to the display and the memory and configured to render, via the display, one or more images from portions of the video data;

render, via the display, text from portions of the transcription data in synchrony with the one or more images;

render, via the display, at least one cell within the text, the at least one cell being associated with at least one point within the time index;

receive input identifying the at least one point within the time index;

receive input specifying audio description text to associate with the at least one point;

synthesize the audio description text to generate audio description data comprising audio data having at least one renderable duration;

store, in the memory, the audio description data; and

store, in the memory, an association between the audio description data and the at least one point.

2. The computer system of claim 1 , wherein the at least one processor is further configured to extend the media file, at one or more locations accessible via the at least one point, by the at least one renderable duration.

3. The computer system of claim 1 , wherein the at least one processor is further configured to generate a new media file that includes the audio description data synchronized with the video data according to the time index.

4. The computer system of claim 1 , wherein the at least one processor is further configured to generate a new media file that includes the video data, the audio data, and the audio description data.

5. The computer system of claim 1 , wherein the at least one processor is further configured to:

adjust a volume of at least one portion of the audio data to generate adjusted audio data; and

generate a new media file that comprises the adjusted audio data.

6. The computer system of claim 1 , wherein the at least one processor is configured to receive input identifying the at least one point via selection of an area within the text.

7. The computer system of claim 1 , wherein the at least one processor is further configured to receive input specifying a synthetic speaking style for the audio description data and to synthesize comprises to synthesize the audio description text using the synthetic speaking style.

8. The computer system of claim 1 , further comprising a keyboard coupled to the at least one processor, wherein the at least one processor is configured to receive input specifying the audio description text via the keyboard.

9. The computer system of claim 1 , wherein the at least one processor is further configured to render additional text from additional portions of the transcription data adjacent to the portions of the transcription data.

10. The computer system of claim 9 , wherein the at least one processor is further configured to:

identify a plurality of points within the time index that identify a plurality of portions of the audio data that each have one or more attributes that meet one or more predefined criteria; and

display a plurality of indications representing the plurality of points within the text and the additional text.

11. The computer system of claim 10 , wherein the at least one processor is configured to identify the plurality of points at least in part by accessing one or more of the transcription data and the audio data.

12. The computer system of claim 10 , wherein the one or more attributes comprise a duration and the one or more predefined criteria specify that the duration be at least a predefined threshold value.

13. The computer system of claim 10 , wherein the one or more attributes comprise a volume and the one or more predefined criteria specify that the volume not exceed a predefined threshold value.

14. The computer system of claim 10 , wherein the one or more attributes comprise a volume over a range of frequencies and the one or more predefined criteria specify that the volume over the range of frequencies not transgress one or more predefined threshold values.

15. The computer system of claim 1 , wherein the at least one processor is further configured to:

setup an audio description job associated with the media file;

configure the audio description job as either a standard job or an extended job; and

determine a pay rate for the audio description job.

16. The computer system of claim 1 , wherein the at least one point is suitable for overlaying the audio description data.

17. The computer system of claim 16 , wherein the at least one processor is further configured to determine that the at least one point is suitable for overlaying the audio description data.

18. The computer system of claim 16 , wherein:

the at least one cell is associated with a time interval containing the at least one point; and

the time interval is substantially equal to or longer than the at least one renderable duration of the audio data.

19. A method for generating an audio description of a media file using a computer system comprising a display and memory coupled to the display, the method comprising:

storing, by the computer system, a media file comprising video data accessible via a time index and audio data synchronized with the video data via the time index;

storing a transcript of the audio data comprising transcription data synchronized with the video data via the time index;

rendering, via the display, one or more images from portions of the video data;

rendering, via the display, text from portions of the transcription data in synchrony with the one or more images;

rendering, via the display, at least one cell within the text, the at least one cell being associated with at least one point within the time index;

receiving input identifying that least one point within the time index;

receiving input specifying audio description text to associate with the at least one point;

synthesizing the audio description text to generate audio description data comprising audio data having at least one renderable duration;

storing the audio description data; and

storing an association between the audio description data and the at least one point.

20. The method according to claim 19 , wherein the method further comprises extending the media file, at one or more locations accessible via the at least one point, by the at least one renderable duration.

21. The method according to claim 19 , further comprising generating a new media file that includes the audio description data synchronized with the video data according to the time index.

22. The method according to claim 19 , further comprising generating a new media file that includes the video data, the audio data, and the audio description data.

23. The method according to claim 19 , further comprising:

generating adjusted audio data by adjusting a volume of the audio data; and

generating a new media file that comprises the adjusted audio data.

24. The method according to claim 19 , further comprising receiving input specifying a synthetic speaking style for the audio description data, wherein synthesizing comprises synthesizing the audio description text using the synthetic speaking style.

25. The method according to claim 24 , further comprising:

identifying a plurality of points within the time index that identify a plurality of portions of the audio data that each have one or more attributes that meet one or more predefined criteria; and

displaying a plurality of indications representing the plurality of points within the text.

26. The method according to claim 19 , further comprising:

setting up an audio description job associated with the media file;

configuring the audio description job as either a standard job or an extended job; and

determining a pay rate for the audio description job.

27. A non-transitory computer readable medium storing computer-executable sequences of instructions to generate an audio description of a media file via a computer system, the sequences of instructions comprising instructions to:

store, in a memory, a media file comprising video data accessible via a time index and audio data synchronized with the video data via the time index;

store, in the memory, a transcript of the audio data comprising transcription data synchronized with the video data via the time index;

render, via a display, one or more images from portions of the video data;

render, via the display, text from portions of the transcription data in synchrony with the one or more images;

render, via the display, at least one cell within the text, the at least one cell being associated with at least one point within the time index;

receive input identifying the at least one point within the time index;

receive input specifying audio description text to associate with the at least one point;

synthesize audio description text to generate audio description data comprising audio data having at least one renderable duration;

store, in the memory, the audio description data; and

store, in the memory, an association between the audio description data and the at least one point.

28. The computer readable medium according to claim 27 , wherein the sequences of instructions further comprise instructions to extend the media file, at one or more locations accessible via the at least one point, by the at least one renderable duration.

29. The computer readable medium according to claim 27 , wherein the sequences of instructions further comprise instructions to render additional text from additional portions of the transcription data adjacent to the portions of the transcription data.

30. A transcription system configured to generate audio description snippets and a snippet manifest from a source media file, the transcription system comprising:

a time-coded transcript of the source media file; and

a synthesized audio video interface configured to

display the source media file,

display the time-coded transcript,

receive input identifying a selected time location within the source media file,

receive input specifying audio description text to associate with the selected time location, the audio description text having at least one text characteristic;

generate an estimated duration of the audio description text using at least one of the at least one text characteristics;

display the estimated duration of the audio description text;

generate an audio snippet from the audio description text;

store the audio snippet as a file; and

store, in the snippet manifest, the audio snippet, the audio description text, the selected time location in the source media file, and a duration of the snippet.

Assignments (2)
SECURITY INTEREST Recorded Feb 9, 2022
From: 3PLAY MEDIA, INC.
To: ABACUS FINANCE GROUP, LLC, AS ADMINISTRATIVE AGENT
Reel/Frame 058934/0816 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2020
From: MILLER, JOSHUA; ANTUNES, CHRISTOPHER S.; BERTHOLD-BOND, LILY MEGAN; SCHWARTZ, ANDREW H.; SAVIETTA, KELLY; BUTLER, LANYA LEE; TOMASULO, SHARON LEE; BARON, JEREMY; JOHNSON, CHRISTOPHER E.; ZIMMERMAN, ROGER
To: 3PLAY MEDIA, INC.
Reel/Frame 051994/0439 →
Continuity (2)
Continuation 16007149 · Jun 13, 2018
Provisional Application 62518911 · Jun 13, 2017
Cited By (2)
US 12,423,506 US 12,499,068