IP Library Granted Patent US 9,940,947
Granted Patent B2
US 9,940,947 · App. 15/371,776 · Granted Apr 10, 2018

Automatic rate control for improved audio time scaling

Inventor: Robert Watts (Gilroy, CA)
Assignee: TiVo Solutions Inc.
G10L21/043G06F3/165G10L15/02G10L19/167G10L2015/027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,940,947
App. No.
15/371,776
Granted
Apr 10, 2018
Kind
B2
Abstract

Input media data with an input playing speed is received and divided into input media data subsets. A first rate of audio utterance is determined for a first input media data subset in the media data subsets. A second different rate of audio utterance is determined for a second input media data subset in the media data subsets. Audio output media data is generated with an output playing speed at which audio utterance in the audio output media data is played at a preferred rate of audio utterance. The audio output media data comprises (a) a first output audio media data subset generated based on the preferred rate, the first rate, and the first input media data subset and (b) a second output audio media data subset generated based on the preferred rate, the second rate, and the second input media data subset.

Claims (36)

1. A method comprising:

performing an acoustic analysis on input media data having an input normal playback speed, wherein the input media data comprises a plurality of input media data subsets each having the same input normal playback speed;

based on results of the acoustic analysis, determining a first rate of audio utterance for a first input media data subset in the plurality of media data subsets;

based on the results of the acoustic analysis, determining a second different rate of audio utterance for a second input media data subset in the plurality of media data subsets;

generating, at a remote server, based on the first rate of audio utterance and the second rate of audio utterance, first and second output media data subsets each corresponding to a different output playing speed and time information that links between playing times of the first and second output media data subsets and playing times of the corresponding input media data; and

sending, from the server to a content client device, the input media data and a data file separate from the input media data, the data file comprising the first and second output media data subsets and the time information.

2. The method as recited in claim 1 , wherein the data file is an index file.

3. The method as recited in claim 1 , wherein the first rate of audio utterance and the second rate of audio utterance are stored in a data store accessible to the content client device.

4. The method as recited in claim 1 , wherein the input media data is transmitted to the content client device in a media stream, wherein the data file comprises one or more tags embedded in the media stream, and wherein the one or more tags are generated based at least in part on the first rate of audio utterance and the second rate of audio utterance.

5. The method as recited in claim 4 , wherein the media stream comprises input video data as well as the input audio data.

6. The method as recited in claim 1 , wherein at least one of the first rate of audio utterance or the second rate of audio utterance is one of a rate of audio utterance for sentences, a rate of audio utterance for words, or a rate of audio utterance for syllables.

7. The method as recited in claim 1 , wherein the acoustic analysis is performed in real time, near real time, or at a rate faster than real time rendering of the input media data with the content client device.

8. A non-transitory computer readable storage medium comprising instructions, which when executed by one or more processors cause performance of steps of:

performing an acoustic analysis on input media data having an input normal playback speed, wherein the input media data comprises a plurality of input media data subsets each having the same input normal playback speed;

based on results of the acoustic analysis, determining a first rate of audio utterance for a first input media data subset in the plurality of media data subsets;

based on the results of the acoustic analysis, determining a second different rate of audio utterance for a second input media data subset in the plurality of media data subsets;

generating, at a remote server, based on the first rate of audio utterance and the second rate of audio utterance, first and second output media data subsets each corresponding to a different output playing speed and time information that links between playing times of the first and second output media data subsets and playing times of the corresponding input media data; and

sending, from the server to a content client device, the input media data and a data file separate from the input media data, the data file comprising the first and second output media data subsets and the time information.

9. The medium as recited in claim 8 , wherein the data file is an index file.

10. The medium as recited in claim 8 , wherein the first rate of audio utterance and the second rate of audio utterance are stored in a data store accessible to the content client device.

11. The medium as recited in claim 8 , wherein the input media data is transmitted to the content client device in a media stream, wherein the data file comprises one or more tags embedded in the media stream, and wherein the one or more tags are generated based at least in part on the first rate of audio utterance and the second rate of audio utterance.

12. The medium as recited in claim 11 , wherein the media stream comprises input video data as well as the input audio data.

13. The medium as recited in claim 8 , wherein at least one of the first rate of audio utterance or the second rate of audio utterance is one of a rate of audio utterance for sentences, a rate of audio utterance for words, or a rate of audio utterance for syllables.

14. The medium as recited in claim 8 , wherein the acoustic analysis is performed in real time, near real time, or at a rate faster than real time rendering of the input media data with the content client device.

15. An apparatus comprising:

a subsystem, implemented at least partially in hardware, that performs an acoustic analysis on input media data having an input normal playback speed, wherein the input media data comprises a plurality of input media data subsets each having the same input normal playback speed;

a subsystem, implemented at least partially in hardware, that, based on results of the acoustic analysis, determines a first rate of audio utterance for a first input media data subset in the plurality of media data subsets;

a subsystem, implemented at least partially in hardware, that, based on the results of the acoustic analysis, determines a second different rate of audio utterance for a second input media data subset in the plurality of media data subsets;

a subsystem, implemented at least partially in hardware, that generates, at a remote server, based on the first rate of audio utterance and the second rate of audio utterance, first and second output media data subsets each corresponding to a different output playing speed and time information that links between playing times of the first and second output media data subsets and playing times of the corresponding input media data; and

a subsystem, implemented at least partially in hardware, that sends, from the server to a content client device, the input media data and a data file separate from the input media data, the data file comprising the first and second output media data subsets and the time information.

16. The apparatus as recited in claim 15 , wherein the data file is an index file.

17. The apparatus as recited in claim 15 , wherein the first rate of audio utterance and the second rate of audio utterance are stored in a data store accessible to the content client device.

18. The apparatus as recited in claim 15 , wherein the input media data is transmitted to the content client device in a media stream, wherein the data file comprises one or more tags embedded in the media stream, and wherein the one or more tags are generated based at least in part on the first rate of audio utterance and the second rate of audio utterance.

19. The apparatus as recited in claim 18 , wherein the media stream comprises input video data as well as the input audio data.

20. The apparatus as recited in claim 15 , wherein at least one of the first rate of audio utterance or the second rate of audio utterance is one of a rate of audio utterance for sentences, a rate of audio utterance for words, or a rate of audio utterance for syllables.

21. The apparatus as recited in claim 15 , wherein the acoustic analysis is performed in real time, near real time, or at a rate faster than real time rendering of the input media data with the content client device.

Assignments (8)
CHANGE OF NAME Recorded Sep 27, 2024
From: TIVO SOLUTIONS INC.
To: ADEIA MEDIA SOLUTIONS INC.
Reel/Frame 069067/0489 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: HPS INVESTMENT PARTNERS, LLC
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053458/0749 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053481/0790 →
SECURITY INTEREST Recorded Jun 1, 2020
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS INC.; VEVEO, INC.; INVENSAS CORPORATION; INVENSAS BONDING TECHNOLOGIES, INC.; TESSERA, INC.; TESSERA ADVANCED TECHNOLOGIES, INC.; DTS, INC.; PHORUS, INC.; IBIQUITY DIGITAL CORPORATION
To: BANK OF AMERICA, N.A.
Reel/Frame 053468/0001 →
PATENT SECURITY AGREEMENT Recorded Nov 25, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 051110/0006 →
SECURITY INTEREST Recorded Nov 22, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: HPS INVESTMENT PARTNERS, LLC, AS COLLATERAL AGENT
Reel/Frame 051143/0468 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2018
From: WATTS, ROBERT
To: TIVO INC.
Reel/Frame 045188/0302 →
CHANGE OF NAME Recorded Mar 23, 2017
From: TIVO INC.
To: TIVO SOLUTIONS INC.
Reel/Frame 041700/0424 →
Continuity (3)
Continuation 14203391 · Mar 10, 2014
Provisional Application 61777940 · Mar 12, 2013
Related Publication 20170092291A1 · Mar 30, 2017