IP Library › Granted Patent US 12,610,138
Granted Patent B2
US 12,610,138 · App. 18/347,604 · Granted Apr 21, 2026

Extended depth of field using deep learning

Inventors: Dan C. Lelescu (Grandvaux, CH); Rohit Rajiv Ranade (Pleasanton, CA); Noah Bedard (Los Gatos, CA); Brian McCall (San Jose, CA); Kathrin Berkner Cieslicki (Los Altos, CA); Michael W. Tao (San Jose, CA); Robert K. Molholm (Scotts Valley, CA); Toke Jansen (Holte, DK); Vladimir Krneta (San Jose, CA)
Assignee: Apple Inc.
H04N23/67G06T7/30H04N23/958G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,610,138
App. No.
18/347,604
Granted
Apr 21, 2026
Kind
B2
Abstract

A method for image enhancement includes capturing multiple input images of a scene, including at least a first input image having a first field of view (FOV) captured with a first focal depth and a second input image having a second FOV captured with a second focal depth. The input images in the sequence are preprocessed so as to align the images. The aligned images are processed in a neural network, which generates an output image having an extended depth of field encompassing at least the first and second focal depths.

Claims (27)

1 . A method for image enhancement, comprising:

capturing multiple input images of a scene, including at least a first input image having a first field of view (FOV) captured with a first focal depth and a second input image having a second FOV captured with a second focal depth, which is greater than the first focal depth, wherein the second FOV is wider than the first FOV;

preprocessing the multiple input images so as to align the multiple input images; and

processing the aligned images in a neural network, which generates an output image having an extended depth of field encompassing at least the first and second focal depths.

2 . The method according to claim 1 , wherein capturing the multiple input images comprises capturing at least the first and second input images sequentially using a handheld imaging device.

3 . The method according to claim 1 , wherein capturing the multiple input images comprises capturing at least the first input image using a first camera and at least the second input image using a second camera, different from the first camera.

4 . The method according to claim 1 , wherein the second FOV is at least 15% wider than the first FOV.

5 . The method according to claim 3 , wherein capturing at least the second input image comprises capturing multiple second input images having different, respective focal depths using the second camera.

6 . The method according to claim 1 , wherein the second FOV is shifted transversely relative to the first FOV.

7 . The method according to claim 6 , wherein the second FOV is shifted transversely by at least 1° relative to the first FOV.

8 . The method according to claim 1 , wherein the second focal depth is at least 30% greater than the first focal depth.

9 . The method according to claim 1 , wherein the first input image has a first depth of field, and the second input image has a second depth of field, different from the first depth of field.

10 . The method according to claim 1 , wherein the sequence of images comprises no more than three input images, which are processed to generate the output image.

11 . The method according to claim 10 , wherein the second focal depth is greater than the first focal depth, and wherein the sequence of images comprises a third input image having a third focal depth greater than the second focal depth, wherein the extended depth of field encompasses at least the first and third focal depths.

12 . The method according to claim 1 , wherein preprocessing the input images comprises aligning at least the first and second fields of view of the first and second input images.

13 . The method according to claim 1 , wherein preprocessing the input images comprises warping one or more of the input images so as to register geometrical features among the input images.

14 . The method according to claim 1 , wherein preprocessing the images comprises correcting photometric variations among the input images.

15 . The method according to claim 1 , and comprising training the neural network using a ground-truth image of a test scene having an extended depth of field and a set of training images of the test scene having different, respective focal settings.

16 . The method according to claim 1 , wherein processing the aligned images comprises generating the output image such that an object in the scene that is out of focus in each of the input images is sharply focused in the output image.

17 . The method according to claim 1 , wherein the neural network comprises one or more encoder layers, which encode features of the images, and decoding layers, which process the encoded features to generate the output image, and wherein processing the aligned images comprises implementing the encoder layers in a first processor to produce encoded features of the input images, and conveying the encoded features to a second processor, which implements the decoding layers to decode the encoded features and generate the output image with the extended depth of field.

18 . The method according to claim 1 , wherein preprocessing the input images further comprises estimating a displacement vector and a prediction error, which predict the second input image in terms of the first input image, and wherein processing the aligned images comprises reconstructing the second input image based on the first input image and the estimated displacement vector and prediction error, and processing the first image together with the reconstructed second image to generate the output image with the extended depth of field.

19 . Imaging apparatus, comprising:

an imaging device configured to capture multiple input images of a scene, including at least a first input image having a first field of view (FOV) captured with a first focal depth and a second input image having a second FOV captured with a second focal depth, which is greater than the first focal depth, wherein the second FOV is wider than the first FOV; and

a processor, which is configured to preprocess the multiple input images so as to align the multiple input images, and to process the aligned images in a neural network, which generates an output image having an extended depth of field encompassing at least the first and second focal depths.

20 . The apparatus according to claim 19 , wherein the imaging device comprises a handheld camera, which is configured to capture at least the first and second input images sequentially.

21 . The apparatus according to claim 19 , wherein the imaging device comprises a first camera, which is configured to capture at least the first input image, and a second camera, which is configured to capture at least the second input image.

22 . A computer software product, comprising a non-transitory computer-readable medium in which program instructions are stored, which instructions, when read by a processor, cause the processor to receive multiple input images of a scene, including at least a first input image having a first field of view (FOV) captured with a first focal depth and a second input image having a second FOV captured with a second focal depth, which is greater than the first focal depth, wherein the second FOV is wider than the first FOV, to preprocess the multiple input images so as to align the multiple input images, and to process the aligned images in a neural network, which generates an output image having an extended depth of field encompassing at least the first and second focal settings.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 6, 2023
From: LELESCU, DAN C.; RANADE, ROHIT RAJIV; BEDARD, NOAH; MCCALL, BRIAN; BERKNER CIESLICKI, KATHRIN; TAO, MICHAEL W.; MOLHOLM, ROBERT K.; JANSEN, TOKE; KRNETA, VLADIMIR
To: APPLE INC.
Reel/Frame 064162/0602 →
Continuity (2)
Provisional Application 63408869 · Sep 22, 2022
Related Publication 20240107162A1 · Mar 28, 2024
References Cited (15)
US 11863881B2 · Feng · 2024 [cited by examiner]
US 12198300B2 · Yang · 2025 [cited by examiner]
US 20110182528A1 · Scherteler et al. · 2011 [cited by applicant]
US 20190333199A1 · Ozcan · 2019 [cited by examiner]
US 20210149170A1 · Leshem · 2021 [cited by examiner]
Bhat et al., “Deep Burst Super-Resolution,” CVPR 2021 paper, Open Access Version, Computer Vision Foundation, pp. 9209-9218, year 2021. [cited by applicant]
Godard et al., “Deep Burst Denoising,” arXiv:1712.05790v1, pp. 1-10, Dec. 15, 2017. [cited by applicant]
Haris et al., “Deep Back-Projection Networks for Single Image Super-Resolution,” arXiv:1904.05677v2, pp. 1-14, Jun. 13, 2020. [cited by applicant]
Lim et al., “Enhanced Deep Residual Networks for Single Image Super-Resolution,” arXiv:1707.02921v1, pp. 1-9, Jul. 10, 2017. [cited by applicant]
Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation,” arXiv:1505.04597v1, pp. 1-8, May 18, 2015. [cited by applicant]
Wronski et al., “Handheld Multi-Frame Super-Resolution,” arXiv:1905.03277v2, pp. 1-24, Feb. 16, 2021. [cited by applicant]
Nguyen et al., U.S. Appl. No. 18/184,677, filed Mar. 16, 2023. [cited by applicant]
Pourreza-Shahri , “Automatic Exposure Selection for High Dynamic Range Photography,” IEEE International Conference on Consumer Electronics (ICCE), 2015, pp. 1-2. [cited by applicant]
Jia et al “Dynamic Filter Networks,”, 30th Conference on Neural Information Processing Systems (NIPS 2016), Barcelona, Spain), pp. 1-9. [cited by applicant]
Final Office Action, U.S. Appl. No. 18/184,677, dated Nov. 26, 2025. [cited by applicant]