IP Library Granted Patent US 10,354,676
Granted Patent B2
US 10,354,676 · App. 15/913,197 · Granted Jul 16, 2019

Automatic rate control for improved audio time scaling

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,354,676
App. No.
15/913,197
Granted
Jul 16, 2019
Kind
B2
Abstract

Input media data with an input playing speed is received and divided into input media data subsets. A first rate of audio utterance is determined for a first input media data subset in the media data subsets. A second different rate of audio utterance is determined for a second input media data subset in the media data subsets. Audio output media data is generated with an output playing speed at which audio utterance in the audio output media data is played at a preferred rate of audio utterance. The audio output media data comprises (a) a first output audio media data subset generated based on the preferred rate, the first rate, and the first input media data subset and (b) a second output audio media data subset generated based on the preferred rate, the second rate, and the second input media data subset.

Claims (53)

1. A method comprising:

receiving, from a remote server, input media data comprising audio content;

receiving, from the remote server, a data file comprising, for each of a first media data subset and a second media data subset, an output playing speed and time information, wherein the time information links playing times of the input media data with playing times of the first media data subset and the second media data subset, wherein a first output playing speed is based on a first rate of audio utterance for the first media data subset, and wherein a second output playing speed is based on a second rate of audio utterance for the second media data subset; and

generating output media data based on the input media data and the data file, comprising:

identifying a plurality of input media data subsets from the input media data;

determining, based on the time information, a first input media data subset of the plurality of input media data subsets that corresponds with the first media data subset and a second input media data subset of the plurality of input media data subsets that corresponds with the second media data subset; and

applying the first output playing speed of the first media data subset to the first input media data subset and the second output playing speed of the second media data subset to the second input media data subset.

2. The method of claim 1 , wherein the media stream further comprises input video data, the method further comprising:

determining a normal playback speed of the input media data;

identifying a first portion of the input video data that corresponds to the first media data subset;

identifying a second portion of the input video data that corresponds to the second media data subset;

modifying the first portion of the input video data based on a first difference between the first output playing speed and the normal playback speed; and

modifying the second portion of the input video data based on a second difference between the second output playing speed and the normal playback speed.

3. The method of claim 2 , wherein modifying the first portion of the input video data comprises modifying a first video frame rate of the first portion of the input video data based on the first difference and wherein modifying the second portion of the input video data comprises modifying a second video frame rate of the second portion of the input video data based on the second difference.

4. The method of claim 3 , wherein the modification comprises sub-sampling frames from the input video data when the output playing speed is faster than the normal playback speed.

5. The method of claim 3 , wherein the modification comprises super-sampling frames from the input video data when the output playing speed is slower than the normal playback speed.

6. The method of claim 2 , further comprising altering audio transcription text associated with the first media data subset based on the first difference and altering audio transcription text associated with the second media data subset based on the second difference.

7. The method of claim 1 , wherein applying the first output playing speed of the first media data subset to the first input media data subset further comprises applying a pitch correction to the output media data to approximate one or more pitches present in the first input media data subset.

8. The method of claim 1 , wherein at least one of the first rate of audio utterance or the second rate of audio utterance is one of a rate of audio utterance for sentences, a rate of audio utterance for words, or a rate of audio utterance for syllables.

9. A non-transitory computer readable storage medium comprising instructions, which when executed by one or more processors cause performance of steps of:

receiving, from a remote server, input media data comprising audio content;

receiving, from the remote server, a data file comprising, for each of a first media data subset and a second media data subset, an output playing speed and time information, wherein the time information links playing times of the input media data with playing times of the first media data subset and the second media data subset, wherein a first output playing speed is based on a first rate of audio utterance for the first media data subset, and wherein a second output playing speed is based on a second rate of audio utterance for the second media data subset; and

generating output media data based on the input media data and the data file, comprising:

identifying a plurality of input media data subsets from the input media data;

determining, based on the time information, a first input media data subset of the plurality of input media data subsets that corresponds with the first media data subset and a second input media data subset of the plurality of input media data subsets that corresponds with the second media data subset; and

applying the first output playing speed of the first media data subset to the first input media data subset and the second output playing speed of the second media data subset to the second input media data subset.

10. The non-transitory computer readable storage medium of claim 9 further comprising instructions, which when executed by one or more processors cause performance of steps of:

determining a normal playback speed of the input media data;

identifying a first portion of the input video data that corresponds to the first media data subset;

identifying a second portion of the input video data that corresponds to the second media data subset;

modifying the first portion of the input video data based on a first difference between the first output playing speed and the normal playback speed; and

modifying the second portion of the input video data based on a second difference between the second output playing speed and the normal playback speed.

11. The non-transitory computer readable storage medium of claim 10 wherein modifying the first portion of the input video data comprises modifying a first video frame rate of the first portion of the input video data based on the first difference and wherein modifying the second portion of the input video data comprises modifying a second video frame rate of the second portion of the input video data based on the second difference.

12. The non-transitory computer readable storage medium of claim 11 , wherein the modification comprises sub-sampling frames from the input video data when the output playing speed is faster than the normal playback speed.

13. The non-transitory computer readable storage medium of claim 11 , wherein the modification comprises super-sampling frames from the input video data when the output playing speed is slower than the normal playback speed.

14. The non-transitory computer readable storage medium of claim 10 further comprising instructions, which when executed by one or more processors cause performance of step of altering audio transcription text associated with the first media data subset based on the first difference and altering audio transcription text associated with the second media data subset based on the second difference.

15. The non-transitory computer readable storage medium of claim 9 , wherein applying the first output playing speed of the first media data subset to the first input media data subset further comprises applying a pitch correction to the output media data to approximate one or more pitches present in the first input media data subset.

16. The non-transitory computer readable storage medium of claim 9 , wherein at least one of the first rate of audio utterance or the second rate of audio utterance is one of a rate of audio utterance for sentences, a rate of audio utterance for words, or a rate of audio utterance for syllables.

17. An apparatus comprising:

a subsystem, implemented at least partially in hardware, that receives, from a remote server, input media data comprising audio content;

a subsystem, implemented at least partially in hardware, that receives, from the remote server, a data file comprising, for each of a first media data subset and a second media data subset, an output playing speed and time information, wherein the time information links playing times of the input media data with playing times of the first media data subset and the second media data subset, wherein a first output playing speed is based on a first rate of audio utterance for the first media data subset, and wherein a second output playing speed is based on a second rate of audio utterance for the second media data subset; and

a subsystem, implemented at least partially in hardware, that generates output media data based on the input media data and the data file, wherein the generation of output media data comprises:

identifying a plurality of input media data subsets from the input media data;

determining, based on the time information, a first input media data subset of the plurality of input media data subsets that corresponds with the first media data subset and a second input media data subset of the plurality of input media data subsets that corresponds with the second media data subset; and

applying the first output playing speed of the first media data subset to the first input media data subset and the second output playing speed of the second media data subset to the second input media data subset.

18. The apparatus of claim 17 , further comprising:

a subsystem, implemented at least partially in hardware, that determines a normal playback speed of the input media data;

a subsystem, implemented at least partially in hardware, that identifies a first portion of the input video data that corresponds to the first media data subset;

a subsystem, implemented at least partially in hardware, that identifies a second portion of the input video data that corresponds to the second media data subset;

a subsystem, implemented at least partially in hardware, that modifies the first portion of the input video data based on a first difference between the first output playing speed and the normal playback speed; and

a subsystem, implemented at least partially in hardware, that modifies the second portion of the input video data based on a second difference between the second output playing speed and the normal playback speed.

19. The apparatus of claim 18 , wherein the subsystem, implemented at least partially in hardware, that modifies the first portion of the input video data is configured to modify a first video frame rate of the first portion of the input video data based on the first difference and wherein modifying the second portion of the input video data comprises modifying a second video frame rate of the second portion of the input video data based on the second difference.

20. The apparatus of claim 18 , further comprising a subsystem, implemented at least partially in hardware, that alters audio transcription text associated with the first media data subset based on the first difference and altering audio transcription text associated with the second media data subset based on the second difference.

Assignments (8)
CHANGE OF NAME Recorded Sep 27, 2024
From: TIVO SOLUTIONS INC.
To: ADEIA MEDIA SOLUTIONS INC.
Reel/Frame 069067/0489 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: HPS INVESTMENT PARTNERS, LLC
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053458/0749 →
RELEASE OF SECURITY INTEREST Recorded Jun 5, 2020
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
Reel/Frame 053481/0790 →
SECURITY INTEREST Recorded Jun 1, 2020
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS INC.; VEVEO, INC.; INVENSAS CORPORATION; INVENSAS BONDING TECHNOLOGIES, INC.; TESSERA, INC.; TESSERA ADVANCED TECHNOLOGIES, INC.; DTS, INC.; PHORUS, INC.; IBIQUITY DIGITAL CORPORATION
To: BANK OF AMERICA, N.A.
Reel/Frame 053468/0001 →
PATENT SECURITY AGREEMENT Recorded Nov 25, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 051110/0006 →
SECURITY INTEREST Recorded Nov 22, 2019
From: ROVI SOLUTIONS CORPORATION; ROVI TECHNOLOGIES CORPORATION; ROVI GUIDES, INC.; TIVO SOLUTIONS, INC.; VEVEO, INC.
To: HPS INVESTMENT PARTNERS, LLC, AS COLLATERAL AGENT
Reel/Frame 051143/0468 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2018
From: WATTS, ROBERT
To: TIVO INC.
Reel/Frame 045188/0302 →
CHANGE OF NAME Recorded Mar 13, 2018
From: TIVO INC.
To: TIVO SOLUTIONS INC.
Reel/Frame 045188/0333 →