IP Library Granted Patent US 9,569,167
Granted Patent B2
US 9,569,167 · App. 14/203,391 · Granted Feb 14, 2017

Automatic rate control for improved audio time scaling

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,569,167
App. No.
14/203,391
Granted
Feb 14, 2017
Kind
B2
Abstract

Input media data with an input playing speed is received and divided into input media data subsets. A first rate of audio utterance is determined for a first input media data subset in the media data subsets. A second different rate of audio utterance is determined for a second input media data subset in the media data subsets. Audio output media data is generated with an output playing speed at which audio utterance in the audio output media data is played at a preferred rate of audio utterance. The audio output media data comprises (a) a first output audio media data subset generated based on the preferred rate, the first rate, and the first input media data subset and (b) a second output audio media data subset generated based on the preferred rate, the second rate, and the second input media data subset.

Claims (51)

1. A method comprising:

dividing input media data having an input normal playback speed into a plurality of input media data subsets each having the same input normal playback speed;

determining a first rate of audio utterance for a first input media data subset in the plurality of media data subsets;

determining a second different rate of audio utterance for a second input media data subset in the plurality of media data subsets;

receiving a preferred rate of audio utterance;

generating, from the plurality of input media data subsets each having the same input normal playback speed, audio output media data comprising a plurality of output media data subsets having at least two different output normal playback speeds but the same preferred rate of audio utterance; the plurality of output media data subsets comprising (a) a first output audio media data subset generated based on the preferred rate of audio utterance, the first rate of audio utterance, and the first input media data subset, and (b) a second output audio media data subset generated based on the preferred rate of audio utterance, the second rate of audio utterance, and the second input media data subset.

2. The method as recited in claim 1 , wherein the at least two output normal playback speeds vary relative to different rates of audio utterance of the plurality of input media data subsets played at the input normal playback speed.

3. The method as recited in claim 1 , further comprising storing the first rate of audio utterance and the second rate of audio utterance in a data store.

4. The method as recited in claim 1 , further comprising embedding one or more tags in a media stream, wherein the one or more tags are generated based on the first rate of audio utterance and the second rate of audio utterance.

5. The method as recited in claim 1 , wherein at least one of the first rate of audio utterance or the second rate of audio utterance is one of a rate of audio utterance for sentences, a rate of audio utterance for words, or a rate of audio utterance for syllables.

6. The method as recited in claim 1 , wherein at least one of the first rate of audio utterance or the second rate of audio utterance is derived from one or more tags embedded in a media stream.

7. A method comprising:

receiving input media data for playing at an input normal playback speed, the input media data comprising a plurality of input media data subsets each having the same input normal playback speed;

receiving a first rate of audio utterance for a first input media data subset in the plurality of media data subsets;

receiving a second different rate of audio utterance for a second input media data subset in the plurality of media data subsets;

receiving a preferred rate of audio utterance;

based at least in part on the preferred rate of audio utterance, the first rate of audio utterance and the second rate of audio utterance, generating output media data comprising a plurality of output media data subsets having at least two different output normal playback speeds but the same preferred rate of audio utterance; the plurality of output media data subsets comprising (a) a first output audio media data subset generated based on the preferred rate of audio utterance, the first rate of audio utterance, and the first input media data subset, and (b) a second output audio media data subset generated based on the preferred rate of audio utterance, the second rate of audio utterance, and the second input media data subset.

8. A non-transitory computer readable storage medium comprising instructions, which when executed by one or more processors cause performance of steps of:

dividing input media data having an input normal playback speed into a plurality of input media data subsets each having the same input normal playback speed;

determining a first rate of audio utterance for a first input media data subset in the plurality of media data subsets;

determining a second different rate of audio utterance for a second input media data subset in the plurality of media data subsets;

receiving a preferred rate of audio utterance;

generating, from the plurality of input media data subsets each having the same input normal playback speed, audio output media data comprising a plurality of output media data subsets having at least two different output normal playback speeds but the same preferred rate of audio utterance; the plurality of output media data subsets comprising (a) a first output audio media data subset generated based on the preferred rate of audio utterance, the first rate of audio utterance, and the first input media data subset, and (b) a second output audio media data subset generated based on the preferred rate of audio utterance, the second rate of audio utterance, and the second input media data subset.

9. The medium as recited in claim 8 , wherein the at least two output normal playback speeds vary relative to different rates of audio utterance of the plurality of input media data subsets played at the input normal playback speed.

10. The medium as recited in claim 8 , wherein the steps further comprise storing the first rate of audio utterance and the second rate of audio utterance in a data store.

11. The medium as recited in claim 8 , wherein the steps further comprise embedding one or more tags in a media stream, wherein the one or more tags are generated based on the first rate of audio utterance and the second rate of audio utterance.

12. The medium as recited in claim 8 , wherein at least one of the first rate of audio utterance or the second rate of audio utterance is one of a rate of audio utterance for sentences, a rate of audio utterance for words, or a rate of audio utterance for syllables.

13. The medium as recited in claim 8 , wherein at least one of the first rate of audio utterance or the second rate of audio utterance is derived from one or more tags embedded in a media stream.

14. A non-transitory computer readable storage medium comprising instructions, which when executed by one or more processors cause performance of steps of:

receiving input media data for playing at an input normal playback speed, the input media data comprising a plurality of input media data subsets each having the same input normal playback speed;

receiving a first rate of audio utterance for a first input media data subset in the plurality of media data subsets;

receiving a second different rate of audio utterance for a second input media data subset in the plurality of media data subsets;

receiving a preferred rate of audio utterance;

based at least in part on the preferred rate of audio utterance, the first rate of audio utterance and the second rate of audio utterance, generating output media data comprising a plurality of output media data subsets having at least two different output normal playback speeds but the same preferred rate of audio utterance; the plurality of output media data subsets comprising (a) a first output audio media data subset generated based on the preferred rate of audio utterance, the first rate of audio utterance, and the first input media data subset, and (b) a second output audio media data subset generated based on the preferred rate of audio utterance, the second rate of audio utterance, and the second input media data subset.

15. An apparatus comprising:

a subsystem, implemented at least partially in hardware, that divides input media data having an input normal playback speed into a plurality of input media data subsets each having the same input normal playback speed;

a subsystem, implemented at least partially in hardware, that determines a first rate of audio utterance for a first input media data subset in the plurality of media data subsets;

a subsystem, implemented at least partially in hardware, that determines a second different rate of audio utterance for a second input media data subset in the plurality of media data subsets;

a subsystem, implemented at least partially in hardware, that receives a preferred rate of audio utterance;

a subsystem, implemented at least partially in hardware, that generates, from the plurality of input media data subsets each having the same input normal playback speed, audio output media data comprising a plurality of output media data subsets having at least two different output normal playback speeds but the same preferred rate of audio utterance; the plurality of output media data subsets comprising (a) a first output audio media data subset generated based on the preferred rate of audio utterance, the first rate of audio utterance, and the first input media data subset, and (b) a second output audio media data subset generated based on the preferred rate of audio utterance, the second rate of audio utterance, and the second input media data subset.

16. The apparatus as recited in claim 15 , wherein the at least two output normal playback speeds vary relative to different rates of audio utterance of the plurality of input media data subsets played at the input normal playback speed.

17. The apparatus as recited in claim 15 , further comprising a subsystem, implemented at least partially in hardware, that stores the first rate of audio utterance and the second rate of audio utterance in a data store.

18. The apparatus as recited in claim 15 , further comprising a subsystem, implemented at least partially in hardware, that embeds one or more tags in a media stream, wherein the one or more tags are generated based on the first rate of audio utterance and the second rate of audio utterance.

19. The apparatus as recited in claim 15 , wherein at least one of the first rate of audio utterance or the second rate of audio utterance is one of a rate of audio utterance for sentences, a rate of audio utterance for words, or a rate of audio utterance for syllables.

20. The apparatus as recited in claim 15 , wherein at least one of the first rate of audio utterance or the second rate of audio utterance is derived from one or more tags embedded in a media stream.

21. An apparatus comprising:

a subsystem, implemented at least partially in hardware, that receives input media data for playing at an input normal playback speed, the input media data comprising a plurality of input media data subsets each having the same input normal playback speed;

a subsystem, implemented at least partially in hardware, that receives a first rate of audio utterance for a first input media data subset in the plurality of media data subsets;

a subsystem, implemented at least partially in hardware, that receives a second different rate of audio utterance for a second input media data subset in the plurality of media data subsets;

a subsystem, implemented at least partially in hardware, that receives a preferred rate of audio utterance;

a subsystem, implemented at least partially in hardware, that, based at least in part on the preferred rate of audio utterance, the first rate of audio utterance and the second rate of audio utterance, generates output media data comprising a plurality of output media data subsets having at least two different output normal playback speeds but the same preferred rate of audio utterance; the plurality of output media data subsets comprising (a) a first output audio media data subset generated based on the preferred rate of audio utterance, the first rate of audio utterance, and the first input media data subset, and (b) a second output audio media data subset generated based on the preferred rate of audio utterance, the second rate of audio utterance, and the second input media data subset.

Assignments (10)
CHANGE OF NAME Recorded Sep 27, 2024
From: TIVO SOLUTIONS INC.
To: ADEIA MEDIA SOLUTIONS INC.
Reel/Frame 069067/0489 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: HPS INVESTMENT PARTNERS, LLC
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053458/0749 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053481/0790 →
SECURITY INTEREST Recorded Jun 1, 2020
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS INC.; VEVEO, INC.; INVENSAS CORPORATION; INVENSAS BONDING TECHNOLOGIES, INC.; TESSERA, INC.; TESSERA ADVANCED TECHNOLOGIES, INC.; DTS, INC.; PHORUS, INC.; IBIQUITY DIGITAL CORPORATION
To: BANK OF AMERICA, N.A.
Reel/Frame 053468/0001 →
RELEASE OF SECURITY INTEREST IN PATENT RIGHTS Recorded Nov 25, 2019
From: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
To: TIVO SOLUTIONS INC.
Reel/Frame 051109/0969 →
PATENT SECURITY AGREEMENT Recorded Nov 25, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 051110/0006 →
SECURITY INTEREST Recorded Nov 22, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: HPS INVESTMENT PARTNERS, LLC, AS COLLATERAL AGENT
Reel/Frame 051143/0468 →
CHANGE OF NAME Recorded Jan 25, 2017
From: TIVO INC.
To: TIVO SOLUTIONS INC.
Reel/Frame 041493/0822 →
SECURITY INTEREST Recorded Dec 7, 2016
From: TIVO SOLUTIONS INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 041076/0051 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2014
From: WATTS, ROBERT
To: TIVO INC.
Reel/Frame 032397/0850 →