IP Library › Granted Patent US 11,509,878
Granted Patent B2
US 11,509,878 · App. 16/567,644 · Granted Nov 22, 2022

Methods and apparatus for using track derivations for network based media processing

Inventors: Xin Wang (San Jose, CA); Lulin Chen (San Jose, CA)
Assignee: MEDIATEK Singapore Pte. Ltd.
H04N13/161H04N13/194H04N13/282H04N13/349
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,509,878
App. No.
16/567,644
Granted
Nov 22, 2022
Kind
B2
Abstract

The techniques described herein relate to methods, apparatus, and computer readable media configured to perform media processing. A media processing entity includes at least one processor in communication with a memory, wherein the memory stores computer-readable instructions that, when executed by the at least one processor, cause the at least one processor to perform receiving, from a remote computing device, multi-view multimedia data comprising a hierarchical track structure comprising at least a first track comprising first media data at a first level of the hierarchical track structure, and metadata associated with a second track at a second level in the hierarchical track structure that is different than the first level of the first track. The instructions further cause the processor to perform processing the first media data of the first track based on the metadata associated with the second track to generate second media data for the second track.

Claims (85)

1. A media processing method implemented by a media processing entity comprising at least one processor in communication with a memory, wherein the memory stores computer-readable instructions that, when executed by the at least one processor, cause the at least one processor to perform:

receiving, from a remote computing device, encoded multi-view multimedia data comprising a hierarchical track structure comprising at least:

a first track comprising first video media data at a first level of the hierarchical track structure, wherein the first track comprises a first set of samples that comprise first video media data; and

a second track at a second level in the hierarchical track structure that is different than the first level of the first track, wherein:

the second track comprises a second set of samples, wherein at least one of the samples of the second set of samples comprises a transform property specifying a derivation operation to perform on the first video media of the first track;

the second track is separate from the first track, such that the first track is not interleaved with the second track during encoding so that the first track can be received separately from the second track; and

the second track does not comprise second video media data in the received encoded multi-view multimedia data;

processing the first video media data of the first track, comprising performing the derivation operation specified by the transform property on the first video media data to generate the second video media data of a derived track; and

transmitting, over a network, the derived track comprising the generated second video media data to a second computing device, wherein the second computing device comprises a second media processing entity, a second remote computing device, or both, which are different than the remote computing device and the media processing entity.

2. The method of claim 1 , wherein receiving the multi-view media data from the remote computing device comprises receiving the multi-view media data from a second remote media processing entity.

3. The method of claim 1 , further comprising transmitting, to the second media processing entity, metadata associated with a third track at a third level in the hierarchical track structure that is different than the first level of the first track and the second level of the second track.

4. The method of claim 1 , wherein:

the second level in the hierarchical track structure is above the first level of the first track so that the second video media is different than the first video media data of the first track; and

processing the first video media data of the first track comprises decoding the first video media data of the first track prior to performing the derivation operation to generate the second video media data for the derived track.

5. The method of claim 4 , wherein:

the transform property specifies one or more of:

a stitching operation to stitch images of the first video media data of the first track and map the stitched images onto a projection surface to generate the second video media data;

a reverse projection operation to project images of the first video media data onto a three-dimensional sphere to generate the second video media data;

a reverse packing operation to perform one or more of transforming, resizing, and relocating one or more regions of the first video media data to generate the second video media data;

a reverse sub-picture operation to compose the second video media data from a plurality of tracks, the plurality of tracks comprising the first track and one or more additional tracks;

a selection of one operation to construct sample images from the first video media data to generate the second video media data;

a transcoding operation to transcode the first video media data from a first bitrate to a second bitrate to generate the second video media data;

a scaling operation to scale the first video media data from a first scale to a second scale to generate the second video media data; and

a resizing operation to resize the first video media data from a first width and a first height to a second width and a second height to generate the second video media data.

6. The method of claim 1 , wherein:

the second level in the hierarchical track structure is below the first level of the first track; and

processing the first video media data of the first track comprises encoding the first video media data of the first track to generate the second video media data for the derived track.

7. The method of claim 6 , wherein:

the transform property specifies one or more of:

a projection operation to project images of the first video media data onto a two-dimensional plane to generate the second video media data;

a packing operation to perform one or more of transforming, resizing, and relocating one or more regions of the first video media data to generate the second video media data;

a sub-picture operation to compose a plurality of different media data for a plurality of tracks, the plurality of tracks comprising the second track and one or more additional tracks;

a viewport operation to construct viewport sample images from spherical sample images of the first video media data to generate the second video media data;

a transcoding operation to transcode the first video media data from a first bitrate to a second bitrate to generate the second video media data;

a scaling operation to scale the first video media data from a first scale to a second scale to generate the second video media data; and

a resizing operation to resize the first video media data from a first width and a first height to a second width and a second height to generate the second video media data.

8. The method of claim 1 , wherein the transform property specifies a plurality of output tracks, and specifies how to generate each of the plurality of output tracks.

9. The method of claim 1 , wherein the second track comprises a data structure specifying the transform property to perform on the first video media data to generate the second video media data, the data structure comprising a number of inputs, a number of outputs, and the transform property.

10. The method of claim 9 , wherein the second track comprises the data structure.

11. An apparatus configured to process video data, the apparatus comprising a processor in communication with memory, the processor being configured to execute instructions stored in the memory that cause the processor to:

receive, from a remote computing device, encoded multi-view multimedia data comprising a hierarchical track structure comprising at least:

a first track at a first level of the hierarchical track structure, wherein the first track comprises a first set of samples that comprise first video media data; and

a second track at a second level in the hierarchical track structure that is different than the first level of the first track, wherein:

the second track comprises a second set of samples, wherein at least one of the samples of the second set of samples comprises a transform property specifying a derivation operation to perform on the first video media of the first track;

the second track is separate from the first track, such that the first track is not interleaved with the second track during encoding so that the first track can be received separately from the second track; and

the second track does not comprise second video media data in the received encoded multi-view multimedia data;

process the first video media data of the first track, comprising performing the derivation operation specified by the transform property on the first video media data to generate second video media data of a derived track; and

transmit, over a network, the derived track comprising the generated second video media data to a second computing device, wherein the second computing device comprises a second media processing entity, a second remote computing device, or both, which are different than the remote computing device and the apparatus.

12. The apparatus of claim 11 , wherein receiving the multi-view media data from the remote computing device comprises receiving the multi-view media data from a second remote media processing entity.

13. The apparatus of claim 11 , wherein the instructions further cause the processor to transmit metadata associated with a third track at a third level in the hierarchical track structure that is different than the first level of the first track and the second level of the second track, to the second computing device.

14. The apparatus of claim 11 , wherein:

the second level in the hierarchical track structure is above the first level of the first track so that the second video media is different than the first video media data of the first track; and

processing the first video media data of the first track comprises decoding the first video media data of the first track prior to performing the derivation operation to generate the second video media data for the derived track.

15. The apparatus of claim 14 , wherein:

the transform property specifies one or more of:

a stitching operation to stitch images of the first video media data of the first track and map the stitched images onto a projection surface to generate the second video media data;

a reverse projection operation to project images of the first video media data onto a three-dimensional sphere to generate the second video media data;

a reverse packing operation to perform one or more of transforming, resizing, and relocating one or more regions of the first video media data to generate the second video media data;

a reverse sub-picture operation to compose the second video media data from a plurality of tracks, the plurality of tracks comprising the first track and one or more additional tracks;

a selection of one operation to construct sample images from the first video media data to generate the second video media data;

a transcoding operation to transcode the first video media data from a first bitrate to a second bitrate to generate the second video media data;

a scaling operation to scale the first video media data from a first scale to a second scale to generate the second video media data; and

a resizing operation to resize the first video media data from a first width and a first height to a second width and a second height to generate the second video media data.

16. The apparatus of claim 11 , wherein:

the second level in the hierarchical track structure is below the first level of the first track; and

processing the first video media data of the first track comprises encoding the first video media data of the first track to generate the second video media data for the derived track.

17. The apparatus of claim 16 , wherein:

the transform property specifies one or more of:

a projection operation to project images of the first video media data onto a two-dimensional plane to generate the second video media data;

a packing operation to perform one or more of transforming, resizing, and relocating one or more regions of the first video media data to generate the second video media data;

a sub-picture operation to compose a plurality of different media data for a plurality of tracks, the plurality of tracks comprising the second track and one or more additional tracks;

a viewport operation to construct viewport sample images from spherical sample images of the first video media data to generate the second video media data;

a transcoding operation to transcode the first video media data from a first bitrate to a second bitrate to generate the second video media data;

a scaling operation to scale the first video media data from a first scale to a second scale to generate the second video media data; and

a resizing operation to resize the first video media data from a first width and a first height to a second width and a second height to generate the second video media data.

18. The apparatus of claim 11 , wherein the second track comprises a data structure specifying the transform property to perform on the first video media data to generate the second video media data, the data structure comprising a number of inputs, a number of outputs, and the transform property.

19. At least one computer readable storage medium storing processor-executable instructions that, when executed by at least one processor of a media processing entity, cause the at least one processor to perform:

receiving, from a remote computing device, encoded multi-view multimedia data comprising a hierarchical track structure comprising at least:

a first track comprising first video media data at a first level of the hierarchical track structure, wherein the first track comprises a first set of samples that comprise first video media data; and

a second track at a second level in the hierarchical track structure that is different than the first level of the first track, wherein:

the second track comprises a second set of samples, wherein at least one of the samples of the second set of samples comprises a transform property specifying a derivation operation to perform on the first video media of the first track;

the second track is separate from the first track, such that the first track is not interleaved with the second track during encoding so that the first track can be received separately from the second track; and

the second track does not comprise second video media data in the received encoded multi-view multimedia data;

processing the first video media data of the first track, comprising performing the derivation operation specified by the transform property on the first video media data to generate second video media data of a derived track; and

transmitting, over a network, the derived track comprising the generated second video media data to a second computing device, wherein the second computing device comprises a second media processing entity, a second remote computing device, or both, which are different than the remote computing device and the media processing entity.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 29, 2020
From: WANG, XIN; CHEN, LULIN
To: MEDIATEK SINGAPORE PTE. LTD.
Reel/Frame 053921/0592 →
Continuity (3)
Provisional Application 62741648 · Oct 5, 2018
Provisional Application 62731131 · Sep 14, 2018
Related Publication 20200092530A1 · Mar 19, 2020