IP Library › Granted Patent US 12,412,315
Granted Patent B2
US 12,412,315 · App. 17/816,787 · Granted Sep 9, 2025

System and method of dual-pixel image synthesis and image background manipulation

Inventors: Abdullah Abuolaim (North York, CA); Mahmoud Afifi (North York, CA); Michael Brown (Toronto, CA)
G06T9/004G06T7/536G06T7/571G06T7/194
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,315
App. No.
17/816,787
Granted
Sep 9, 2025
Kind
B2
Abstract

A system and method of determining synthetic dual-pixel data, performing deblurring, predicting dual pixel views, and view synthesis. The method including: receiving an input image; determining synthetic dual-pixel data using a trained artificial neural network with the input image as input to the trained artificial neural network, the trained artificial neural network includes a latent space encoder, a left dual-pixel view decoder, and a right dual-pixel view decoder; and outputting the synthetic dual-pixel data. In some cases, determination of the synthetic dual-pixel data can include performing reflection removal, defocus deblurring, or view synthesis.

Claims (24)

1. A method of determining synthetic dual-pixel data, the method comprising:

receiving an input image;

determining synthetic dual-pixel data using a trained artificial neural network with the input image as input to the trained artificial neural network, the trained artificial neural network comprises a latent space encoder, a left dual-pixel view decoder, and a right dual-pixel view decoder, the artificial neural network trained with a loss function comprising a dual-pixel-loss, a view difference loss, and a mean-square-error loss between ground truth and estimated dual-pixel views; and

outputting the synthetic dual-pixel data.

2. The method of claim 1 , wherein the artificial neural network is trained by inputting a training dataset of images and optimizing for a loss function that imposes a constraint on dual-pixel view reconstruction and a view difference loss function.

3. The method of claim 2 , wherein the training dataset of images comprises a plurality of scenes, each scene comprising both dual pixel images capturing the scene.

4. The method of claim 1 , wherein the left dual-pixel view decoder and the right dual-pixel view decoder comprise an early-stage weight sharing at the end of the latent space encoder.

5. The method of claim 1 , further comprising performing deblurring of the input image and outputting a deblurred image, wherein the trained artificial neural network further comprises a deblurring decoder, and wherein the deblurred image comprises the output of the deblurring decoder.

6. The method of claim 1 , further comprising predicting dual pixel views of the input image by outputting the output of the left dual-pixel view decoder and the right dual-pixel view decoder.

7. The method of claim 6 , further comprising performing reflection removal, defocus deblurring, or both, using the predicted dual pixel views.

8. The method of claim 1 , further comprising performing view synthesis using the input image, wherein determining the synthetic dual-pixel data comprises passing each of a plurality of rotated views of the input image as input to the trained artificial neural network, and wherein the view synthesis comprises a combination of the output of the left dual-pixel view decoder and the output of the right dual-pixel view decoder for each of the rotated view of the input image.

9. The method of claim 8 , further comprising synthesizing image motion by rotating point spread functions through a plurality of different angles during the view synthesis.

10. A system for determining synthetic dual-pixel data, the system comprising a processing unit and non-transitory computer readable storage media, the non-transitory computer readable storage media comprising instructions for the processing unit to execute:

an input module to receive an input image;

a neural network module to determine synthetic dual-pixel data using a trained artificial neural network with the input image as input to the trained artificial neural network, the trained artificial neural network comprises a latent space encoder, a left dual-pixel view decoder, and a right dual-pixel view decoder, the artificial neural network trained with a loss function comprising a dual-pixel-loss, a view difference loss, and a mean-square-error loss between ground truth and estimated dual-pixel views; and

an output module to output the synthetic dual-pixel data.

11. The system of claim 10 , wherein the artificial neural network is trained by inputting a training dataset of images and optimizing for a function that imposes a constraint on dual-pixel view reconstruction and a view difference loss function.

12. The system of claim 11 , wherein the training dataset of images comprises a plurality of scenes, each scene comprising both dual pixel images capturing the scene.

13. The system of claim 10 , wherein the left dual-pixel view decoder and the right dual-pixel view decoder comprise an early-stage weight sharing at the end of the latent space encoder.

14. The system of claim 10 , wherein the neural network module further performs deblurring of the input image and the output module outputs the deblurred image, wherein the trained artificial neural network further comprises a deblurring decoder, and wherein the deblurred image comprises the output of the deblurring decoder.

15. The system of claim 10 , wherein the neural network module further predicts dual pixel views of the input image by determining the output of the left dual-pixel view decoder and the right dual-pixel view decoder.

16. The system of claim 15 , wherein reflection removal, defocus deblurring, or both, are performed using the predicted dual pixel views.

17. The system of claim 10 , further comprising a synthesis module to perform view synthesis using the input image, wherein determining the synthetic dual-pixel data comprises passing each of a plurality of rotated views of the input image as input to the trained artificial neural network, and wherein the view synthesis comprises a combination of the output of the left dual-pixel view decoder and the output of the right dual-pixel view decoder for each of the rotated view of the input image.

18. The system of claim 17 , wherein the synthesis module further synthesizes image motion by rotating point spread functions through a plurality of different angles during the view synthesis.

Continuity (2)
Provisional Application 63228729 · Aug 3, 2021
Related Publication 20230056657A1 · Feb 23, 2023
References Cited (8)
US 9350977B1 · Prasad · 2016 [cited by examiner]
US 11107228B1 · Shrivastava · 2021 [cited by examiner]
US 20210312242A1 · Jafarkhani · 2021 [cited by examiner]
US 20220375042A1 · Garg · 2022 [cited by examiner]
CN 112792821A · 2021 [cited by examiner]
Pan et al Dual Pixel Exploration: Simultaneous Depth Estimation and Image Restoration, arXiv:201.00301v1 (Year: 2020). [cited by examiner]
Punnappurath et al Reflection Removal Using a Dual-Pixel Sensor, CVPR, pp. 1556-1565 (Year: 2019). [cited by examiner]
Abuolaim et al Defocus Deblurring Using Dual-Pixel Data, arXiv:2005.00305v3 Jun. 16, 2020. [cited by examiner]