IP Library Granted Patent US 12,499,613
Granted Patent B2
US 12,499,613 · App. 18/275,193 · Granted Dec 16, 2025

Athlete's perspective sports game broadcast system based on VR technology

Inventors: Xiao Zhang (Beijing, CN); Zheng Zhang (San Carlos, CA); Tao Xu (Cambridge, MA); Jie Tang (Sammamish, WA); Yaobo Liang (Beijing, CN); Yanan Wei (Redmond, WA)
Assignee: Microsoft Technology Licensing, LLC
G06T15/205G06T7/194G06T2207/20081G06T2207/20084G06T2207/30168G06T2207/30221G06T2210/08H04N21/2187H04N21/2343H04N21/816
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,613
App. No.
18/275,193
Granted
Dec 16, 2025
Kind
B2
Abstract

The present disclosure proposes a method and apparatus for image transmission for three-dimensional (3D) scene reconstruction. A target image may be obtained. A geometry feature and an appearance feature may be disentangled from the target image. An intermediate image may be reconstructed based on the geometry feature and a reference appearance feature. A difference between the intermediate image and the target image may be determined. The geometry feature may be transmitted to a receiving device for 3D scene reconstruction in response to determining that the difference is lower than a predetermined threshold.

Claims (72)

1 . A method for image transmission for three-dimensional (3D) scene reconstruction, comprising:

obtaining a target image;

disentangling a geometry feature and an appearance feature from the target image;

reconstructing an intermediate image based on the geometry feature and a reference appearance feature;

determining a difference between the intermediate image and the target image; and

in response to determining that the difference is not lower than a predetermined threshold:

updating the reference appearance feature; and

transmitting the geometry feature and the updated reference appearance feature to a receiving device for 3D scene reconstruction.

2 . The method of claim 1 , wherein the target image and the reference appearance feature are associated with a same target object.

3 . The method of claim 1 , wherein the obtaining a target image comprises:

receiving an original image taken by a camera; and

extracting the target image from the original image.

4 . The method of claim 3 , wherein the extracting the target image from the original image comprises:

segmenting a foreground image from the original image; and

extracting the target image from the foreground image.

5 . The method of claim 4 , further comprising:

transmitting a predetermined background image to the receiving device.

6 . The method of claim 1 , wherein the disentangling a geometry feature and an appearance feature from the target image is performed through a disentangling model which is based on a neural network.

7 . The method of claim 6 , wherein the disentangling model is trained through:

disentangling a training geometry feature and a training appearance feature from a training image through the disentangling model;

reconstructing a training intermediate image based on the training geometry feature and the training appearance feature;

determining a difference between the training intermediate image and the training image; and

optimizing the disentangling model through minimizing the difference between the training intermediate image and the training image.

8 . The method of claim 1 , wherein the determining a difference between the intermediate image and the target image comprises:

calculating a sub-difference between each pixel in a set of pixels of the intermediate image and a corresponding pixel in a set of pixels of the target image, to obtain a set of sub-differences; and

calculating the difference based on the set of sub-differences.

9 . The method of claim 1 , wherein the updating the reference appearance feature comprises:

updating the reference appearance feature to the appearance feature.

10 . The method of claim 1 , wherein the target image is from a set of target images, the set of target images are associated with a same target object and are from a set of original images taken at a same time, and the updating the reference appearance feature comprises:

generating a comprehensive appearance feature based on a set of appearance features disentangled from the set of target images; and

updating the reference appearance feature to the comprehensive appearance feature.

11 . The method of claim 1 , wherein the reconstructing an intermediate image is performed through differentiable rendering.

12 . The method of claim 1 , further comprising:

transmitting the reference appearance feature to the receiving device before obtaining the target image.

13 . The method of claim 1 , further comprising:

obtaining a second target image;

disentangling a second geometry feature and a second appearance feature from the second target image;

reconstructing a second intermediate image based on the second geometry feature and the second appearance feature;

determining a second difference between the second intermediate image and the second target image; and

in response to determining that the second difference is lower than the predetermined threshold, transmitting the second geometry feature to the receiving device for 3D scene reconstruction.

14 . An apparatus for image transmission for three-dimensional (3D) scene reconstruction, comprising:

at least one processor; and

a memory storing computer-executable instructions that, when executed, cause the at least one processor to:

obtain a target image,

disentangle a geometry feature and an appearance feature from the target image,

reconstruct an intermediate image based on the geometry feature and a reference appearance feature,

determine a difference between the intermediate image and the target image, and

in response to determining that the difference is not lower than a predetermined threshold:

updating the reference appearance feature; and

transmitting the geometry feature and the updated reference appearance feature to a receiving device for 3D scene reconstruction.

15 . The apparatus of claim 14 , wherein the obtaining a target image comprises:

receiving an original image taken by a camera; and

extracting the target image from the original image.

16 . The apparatus of claim 15 , wherein the extracting the target image from the original image comprises:

segmenting a foreground image from the original image; and

extracting the target image from the foreground image.

17 . The apparatus of claim 16 , further comprising:

transmitting a predetermined background image to the receiving device.

18 . A non-transitory machine-readable medium including instructions that, when executed by a machine, cause the machine to perform operations for image transmission for three-dimensional (3D) scene reconstruction, the operations comprising:

obtaining a target image;

disentangling a geometry feature and an appearance feature from the target image;

reconstructing an intermediate image based on the geometry feature and a reference appearance feature;

determining a difference between the intermediate image and the target image; and

in response to determining that the difference is not lower than a predetermined threshold:

updating the reference appearance feature; and

transmitting the geometry feature and the updated reference appearance feature to a receiving device for 3D scene reconstruction.

19 . The non-transitory machine-readable medium of claim 18 , wherein the disentangling a geometry feature and an appearance feature from the target image is performed through a disentangling model which is based on a neural network.

20 . The non-transitory machine-readable medium of claim 19 , wherein the operations further comprise training the disentangling model by:

disentangling a training geometry feature and a training appearance feature from a training image through the disentangling model;

reconstructing a training intermediate image based on the training geometry feature and the training appearance feature;

determining a difference between the training intermediate image and the training image; and

optimizing the disentangling model through minimizing the difference between the training intermediate image and the training image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 16, 2023
From: ZHANG, XIAO; ZHANG, ZHENG; XU, TAO; TANG, JIE; LIANG, YAOBO; WEI, YANAN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 064604/0374 →
Priority Claims (1)
CN 202110218773.7 · Feb 26, 2021 · national
Continuity (1)
Related Publication 20240153199A1 · May 9, 2024
References Cited (15)
US 6307567B1 · Cohen-Or · 2001 [cited by examiner]
US 6510241B1 · Vaillant · 2003 [cited by applicant]
US 20050001841A1 · Francois · 2005 [cited by examiner]
US 20080106544A1 · Lee · 2008 [cited by examiner]
US 20120224629A1 · Bhagavathy et al. · 2012 [cited by applicant]
US 20160197999A1 · Chun · 2016 [cited by examiner]
US 20210279952A1 · Chen · 2021 [cited by examiner]
US 20220198738A1 · Xu · 2022 [cited by examiner]
US 20230379426A1 · Stone · 2023 [cited by examiner]
Notice of Allowance Received for Chinese Application No. 202110218773.7, mailed on Dec. 17, 2024, 8 pages. (English Translation Provided). [cited by applicant]
Notice of First Examination Received for Chinese Application No. 202110218773.7, mailed on Mar. 15, 2024, 8 pages (English Translation Provided). [cited by applicant]
Lorenz, et al., “Unsupervised Part-Based Disentangling of Object Shape and Appearance”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 15, 2019, pp. 10947-10956. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US22/013958”, Mailed Date: Jul. 15, 2022, 10 Pages. [cited by applicant]
Siddeq, “DCT and DST Based Image Compression for 3D Reconstruction”, In Journal of 3D Research, vol. 8, Issue 1, Feb. 4, 2017, 19 Pages. [cited by applicant]
Yao, et al., “3D-Aware Scene Manipulation via Inverse Graphics”, In Repository of arXiv:1808.09351v1, Aug. 28, 2018, 10 Pages. [cited by applicant]