IP Library Granted Patent US 12,406,463
Granted Patent B2
US 12,406,463 · App. 18/009,120 · Granted Sep 2, 2025

Representing volumetric video in saliency video streams

Inventors: Ajit Ninan (San Jose, CA); Shwetha Ram (San Jose, CA); Gregory John Ward (San Francisco, CA); Domagoj Baricevic (San Francisco, CA); Vijay Kamarshi (San Francisco, CA)
Assignee: Dolby Laboratories Licensing Corporation
G06V10/462G06T7/20G06T7/40G06T7/55G06V10/25H04N13/383H04N13/388
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,406,463
App. No.
18/009,120
Granted
Sep 2, 2025
Kind
B2
Abstract

Saliency regions are identified in a global scene depicted by volumetric video. Saliency video streams that track the saliency regions are generated. Each saliency video stream tracks a respective saliency region. A saliency stream based representation of the volumetric video is generated to include the saliency video streams. The saliency stream based representation of the volumetric video is transmitted to a video streaming client.

Claims (16)

1. A method for streaming volumetric video, comprising:

identifying a set of one or more saliency regions in a global scene depicted by the volumetric video;

generating a set of one or more saliency video streams that track the set of one or more saliency regions identified in the global scene, wherein each saliency video stream in the set of one or more saliency video streams tracks a respective saliency region in the set of one or more saliency regions;

generating a saliency stream based representation of the volumetric video, wherein the saliency stream based representation includes the set of one or more saliency video streams;

transmitting the saliency stream based representation of the volumetric video to a video streaming client,

wherein the set of one or more saliency video streams includes a first saliency video stream assigned with a first saliency rank and a second saliency video stream assigned with a second saliency rank lower than the first saliency rank, wherein the second saliency video stream is removed from the set of one or more saliency video streams to be transmitted to the video streaming client at a later time, in response to determining that an available data rate has been reduced.

2. The method of claim 1 , wherein the first saliency rank and/or the second saliency rank is ranked by one or more of: input from a content creator or a director, viewer statistical information gathering and analyses, and visual saliency algorithms based at least in part on computer vision techniques.

3. The method of claim 1 , wherein the saliency stream based representation includes at least one base video stream depicting the global scene including image areas other than the set of one or more saliency regions, wherein the at least one base video stream enables the video streaming client to render imagery outside the set of one or more saliency regions depicted by the set of one or more saliency video streams.

4. The method of claim 1 , wherein the saliency stream based representation includes a disocclusion data associated with a saliency video stream in the set of one or more saliency video streams, wherein the disocclusion data includes texture and depth information for image details occluded in a reference view depicted by the saliency video stream, wherein the image details occluded in the reference view become disoccluded in one or more other views adjacent to the reference view.

5. The method of claim 1 , wherein the set of one or more saliency video streams includes the first saliency video stream assigned with a first saliency rank and the second saliency video stream assigned with a second saliency rank lower than the first saliency rank, wherein the second saliency video stream is removed from the set of one or more saliency video streams to be transmitted to the video streaming client at a later time, in response to determining, based on real time viewpoint data, that a virtual view represented by a viewer's viewpoint is directed to a different spatial region away from a saliency region tracked by the second saliency video stream.

6. The method of claim 1 , wherein image metadata is transmitted with the set of one or more saliency video streams to enable the video streaming client to render images derived from the set of saliency video streams and any accompanying disocclusion.

7. The method of claim 1 , wherein viewpoint data of a viewer collected in real time while the viewer is viewing imagery generated from the volumetric video is received from the video streaming client, wherein the viewpoint data is used to select the set of one or more saliency streams in one or more reference views closest to a virtual view represented by a viewpoint, of the viewer, as indicated in the viewpoint data.

8. The method of claim 1 , wherein an input version of the volumetric video is received and used to derive the saliency stream based representation of the volumetric video, wherein the input version of the volumetric video includes a group of pictures (GOP) comprising a plurality of full resolution images depicting the global scene, wherein at least one saliency video stream in the set of one or more saliency video streams is initialized at a starting time point of the GOP.

9. An apparatus comprising a non-transitory computer readable storage medium, storing software instructions, which when executed by one or more processors cause performance of the method recited in claim 1 .

10. A non-transitory computer readable storage medium, storing software instructions, which when executed by one or more processors cause performance of the method recited in claim 1 .

11. A computing device comprising one or more processors and one or more storage media, storing a set of instructions, which when executed by one or more processors cause performance of the method recited in claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 28, 2023
From: NINAN, AJIT; RAM, SHWETHA; WARD, GREGORY JOHN; BARICEVIC, DOMAGOJ; KAMARSHI, VIJAY
To: DOLBY LABORATORIES LICENSING CORPORATION
Reel/Frame 064096/0093 →
Priority Claims (1)
EP 20180178 · Jun 16, 2020 · regional
Continuity (2)
Provisional Application 63039589 · Jun 16, 2020
Related Publication 20230215129A1 · Jul 6, 2023
References Cited (60)
US 7020668B2 · Matsuda · 2006 [cited by examiner]
US 8666146B1 · Smolic · 2014 [cited by examiner]
US 9041709B2 · Bruls · 2015 [cited by applicant]
US 9147221B2 · Grasset · 2015 [cited by examiner]
US 9338431B2 · Lee · 2016 [cited by applicant]
US 9467750B2 · Banica · 2016 [cited by applicant]
US 10051286B2 · Grangetto · 2018 [cited by applicant]
US 10074012B2 · Zhou · 2018 [cited by applicant]
US 10237548B2 · Korneliussen · 2019 [cited by applicant]
US 10362265B2 · Pio · 2019 [cited by examiner]
US 10419737B2 · Pang · 2019 [cited by applicant]
US 20100215251A1 · Klein Gunnewiek · 2010 [cited by applicant]
US 20130127844A1 · Koeppel · 2013 [cited by applicant]
US 20130326583A1 · Freihold · 2013 [cited by applicant]
US 20140176553A1 · Pettersson · 2014 [cited by applicant]
US 20140359152A1 · Heng · 2014 [cited by examiner]
US 20150350594A1 · Mate · 2015 [cited by applicant]
US 20170118540A1 · Thomas · 2017 [cited by examiner]
US 20170318262A1 · Safaei · 2017 [cited by applicant]
US 20170332064A1 · Martineau · 2017 [cited by applicant]
US 20180007352A1 · Chang · 2018 [cited by applicant]
US 20180097867A1 · Pang · 2018 [cited by applicant]
US 20180109817A1 · Wang · 2018 [cited by applicant]
US 20180146198A1 · Atluru · 2018 [cited by examiner]
US 20180199042A1 · Wang · 2018 [cited by applicant]
US 20180359489A1 · Lakshman · 2018 [cited by applicant]
US 20190213784A1 · Schmalstieg · 2019 [cited by applicant]
US 20190373278A1 · Castaneda · 2019 [cited by examiner]
US 20200021791A1 · Hur · 2020 [cited by applicant]
US 20200045290A1 · Ruhm · 2020 [cited by applicant]
US 20200279384A1 · Jia · 2020 [cited by applicant]
US 20200286293A1 · Jia · 2020 [cited by applicant]
US 20200288114A1 · Lakshman · 2020 [cited by examiner]
US 20230224447A1 · Ward · 2023 [cited by examiner]
WO 2009001255A1 · 2008 [cited by applicant]
WO 2013168091A1 · 2013 [cited by applicant]
WO 2017080420A1 · 2017 [cited by applicant]
WO 2019055389A1 · 2019 [cited by applicant]
WO 2019209838A1 · 2019 [cited by applicant]
WO 2019211519A1 · 2019 [cited by applicant]
WO 2019243663A1 · 2019 [cited by applicant]
WO 2020008106A1 · 2020 [cited by applicant]
WO WO2021257639A1 · 2021 [cited by applicant]
ITU-T H.264, “Advanced Video coding for Generic Audiovisual Services” Series H: Audiovisual and Multimedia systems, Infrastructure of Audiovisual services—Coding of Moving Video, Jan. 2012, 680 pages. [cited by applicant]
Korea Aerospace University et al, “KAU Response to Immersive Video CE3: Atlas Padding,” ISO/IEC JTC1/SC29/ WG11 MPEG2020/m52189 (Jan. 2020.), 4 pages. [cited by applicant]
Kwang-Soon Lee, Jeong-Il Seo, “Trends and Prospects of MPEG Immersive Video Standard Technology,” IITP, Weekly Technology Trends, Issue 1969, Oct. 21, 2020, 44 pages. [cited by applicant]
M. Wien, J. M. Boyce, T. Stockhammer, and W.-H. Peng, “Standardization Status of Immersive Video Coding,” IEEE JourEmerg. Select. Topics Circuits Syst., vol. 9, No. 1, Mar. 2019, pp. 5-17, 13 pages. [cited by applicant]
Chen et al., “Test Model 11 of 3D-HEVC and MV-HEVC,” ISO/IEC JTC1/SC29/WG11 N 15141, Feb. 2015, Geneva, Switzerland, <https://mpeg.chiariglione.org/standards/mpeg-h/hevc-reference-software/n15141-test-model-11-3d-hevc-a… [cited by applicant]
U.S. Appl. No. 62/423,287, Prediction and Verifying Regions of Interest Selections, filed Nov. 17, 2016. [cited by applicant]
U.S. Appl. No. 62/518,187, “Coding Multiview Video,” filed Jun. 12, 2017. [cited by applicant]
U.S. Appl. No. 62/811,956, Hole Filling for Depth Image Based Rendering, filed Apr. 1, 2019. [cited by applicant]
U.S. Appl. No. 62/813,527, “Multi-Resolution Multi-View Video Rendering,” filed Mar. 4, 2019. [cited by applicant]
U.S. Appl. No. 63/039,595, “Supporting Multi-View Video Operations with Disocclusion Atlas,” filed Jun. 16, 2020. [cited by applicant]
Shum et al., “A Review of Image-Based Rendering Techniques,” Microsoft Research Document, Feb. 2016, accessed May 15, 2023, <https://www.microsoft.com/en-us/research/wpcontent/uploads/2016/02/review_image_rendering.pdf)… [cited by applicant]
A Review of Image-based Rendering Techniques (https://www.microsoft.com/en-us/research/wpcontent/uploads/2016/02/review_image_rendering.pdf), pp. 1-12, 12 pages. [cited by applicant]
Garcia-Dorado et al., “Automatic urban modeling using volumetric reconstruction with surface graph cuts,” Computers & Graphics. Nov. 1, 2013;37(7): pp. 896-910. (Year: 2013), 16 pages. [cited by applicant]
“HEVC-3D standard,” 2013-2021 Fraunhofer Heinrich Hertz Institute, pp. 1-2, 2 pages. [cited by applicant]
“VPCC Codec Description,” MPEG-I VPCC Standard, Jun. 17, 2020, ISO/IEC JTC 1/SC 29/WG 7, pp. 1-77, 77 pages. [cited by applicant]
Mueller et al., “Shading atlas streaming,” ACM Transactions on Graphics (TOG), 37(6), (Year: 2018), pp. 1-16, 16 pages. [cited by applicant]
N15141, Test Model 11 of 3D-HEVC and MV-HEVC (https://mpeg.chiariglione.org/standards/mpeg-h/hevc-reference-software/n15141-test-model-11-3d-hevc-and-mv-hevc), Feb. 2015, pp. 1-58, 58 pages. [cited by applicant]