IP Library Granted Patent US 10,991,150
Granted Patent B2
US 10,991,150 · App. 16/181,406 · Granted Apr 27, 2021

View generation from a single image using fully convolutional neural networks

Inventors: Mohamed Elgharib (Doha, QA); Wojciech Matusik (Cambridge, MA); Sung-Ho Bae (Cambridge, MA); Mohamed Hefeeda (Doha, QA)
Assignees: MASSACHUSETTS INSTITUTE OF TECHNOLOGY; QATAR FOUNDATION FOR EDUCATION, SCIENCE AND COMMUNITY DEVELOPMENT
G06T15/205G06N3/08H04N13/15H04N13/161
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,991,150
App. No.
16/181,406
Granted
Apr 27, 2021
Kind
B2
Abstract

A method of rendering a stereoscopic 3D image from a single image, including receiving a collection of pairs of 3D images including an input image and an output image that is a 3D pair of the input image, training a neural network composed of convolutional layers without any fully connected layers, with pairs of 3D images from the collection, to receive an input image and generate an output image that is a 3D pair of the input image, wherein the neural network is provided as an application on a computing device, receiving an input image and generating an output image that is a 3D pair of the input image by the neural network.

Claims (31)

1. A method of rendering a stereoscopic 3D image from a single image, comprising:

receiving a collection of pairs of stereoscopic 3D images comprising an input image and an output image that is a stereoscopic 3D pair of the input image;

training a neural network composed of convolutional layers without any fully connected layers, with pairs of stereoscopic 3D images from the collection, to receive an input image and generate an output image that is a stereoscopic 3D pair of the input image; wherein the neural network is provided as an application on a computing device;

upon receiving an input image, generating an output image that is a stereoscopic 3D pair of the input image by the neural network,

wherein a first half of the convolution lavers downscale their inputs with a preconfigured factor by using a stride value and

wherein a second half of the convolution layers upscale their inputs with the preconfigured factor by using an inverse value of the stride value to reduce distortion in the output image.

2. The method according to claim 1 , wherein the neural network includes an encoding network, a decoding network and a rendering network.

3. The method according to claim 2 , wherein the rendering network includes a softmax layer that normalizes the output of the decoding network.

4. The method according to claim 1 , wherein the neural network includes a conversion block that converts the received input image from RGB to Y, Cb and Cr; wherein Y serves as a luminance channel of the input and Cb and Cr serve as a chrominance channel of the input.

5. The method according to claim 4 , wherein the neural network includes an inversion block that inverts the generated output image from a luminance channel of the output and a chrominance channel of the output back to an RGB image.

6. The method according to claim 4 , wherein the neural network includes a luminance network that processes a luminance channel of the received input image and a chrominance network that processes a chrominance channel of the received input image.

7. The method according to claim 6 , wherein the luminance network and the chrominance network each comprise an encoding network and a decoding network.

8. The method according to claim 6 , wherein the chrominance channel of the received input image is downscaled by a factor of two before processing with the chrominance network.

9. The method according to claim 6 , wherein the luminance network is trained with luminance channels of the images of the collection and the chrominance network is trained with chrominance channels of the images of the collection.

10. The method according to claim 1 , wherein each convolution layer except the last layer is followed by a rectified linear unit (ReLu) layer.

11. The method according to claim 1 , wherein the convolution layers of the neural network are configured to accept images of varying resolutions.

12. A system for rendering a stereoscopic 3D image from a single image, comprising:

a collection of pairs of stereoscopic 3D images comprising an input image and an output image that is a stereoscopic 3D pair of the input image;

a computing device;

a neural network application composed of convolutional layers without any fully connected layers that is configured to be executed on the computing device and programed to:

train the neural network, with pairs of stereoscopic 3D images from the collection, to receive an input image and generate an output image that is a stereoscopic 3D pair of the input image;

upon receiving an input image, generate an output image that is a stereoscopic 3D pair of the input image,

wherein a first half of the convolution lavers downscale their inputs with a preconfigured factor by using a stride value and

wherein a second half of the convolution lavers upscale their inputs with the preconfigured factor by using an inverse value of the stride value, to reduce distortion in the output image.

13. The system according to claim 12 , wherein the neural network includes an encoding network, a decoding network and a rendering network.

14. The system according to claim 13 , wherein the rendering network includes a softmax layer that normalizes the output of the decoding network.

15. The system according to claim 12 , wherein the neural network includes a conversion block that converts the received input image from RGB to Y, Cb and Cr; wherein Y serves as a luminance channel of the input and Cb and Cr serve as a chrominance channel of the input.

16. The system according to claim 15 , wherein the neural network includes an inversion block that inverts the generated output image from a luminance channel of the output and a chrominance channel of the output back to an RGB image.

17. The system according to claim 15 , wherein the neural network includes a luminance network that processes a luminance channel of the received input image and a chrominance network that processes a chrominance channel of the received input image.

18. The system according to claim 17 , wherein the luminance network and the chrominance network each comprise an encoding network and a decoding network.

19. The system according to claim 17 , wherein the chrominance channel of the received input image is downscaled by a factor of two before processing with the chrominance network.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 17, 2025
From: QATAR FOUNDATION FOR EDUCATION, SCIENCE & COMMUNITY DEVELOPMENT
To: HAMAD BIN KHALIFA UNIVERSITY
Reel/Frame 069936/0656 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 6, 2018
From: ELGHARIB, MOHAMED; MATUSIK, WOJCIECH; BAE, SUNG-HO; HEFEEDA, MOHAMED
To: MASSACHUSETTS INSTITUTE OF TECHNOLOGY; QATAR FOUNDATION FOR EDUCATION, SCIENCE AND COMMUNITY DEVELOPMENT
Reel/Frame 047416/0782 →
Continuity (2)
Provisional Application 62668845 · May 9, 2018
Related Publication 20190347847A1 · Nov 14, 2019
Cited By (1)
US 12,423,848