IP Library › Granted Patent US 12,217,169
Granted Patent B2
US 12,217,169 · App. 17/198,670 · Granted Feb 4, 2025

Local neural implicit functions with modulated periodic activations

Inventors: Ishit bhadresh Mehta (San Diego, CA); Michaël Gharbi (San Francisco, CA); Connelly Barnes (Seattle, WA); Elya Shechtman (Seattle, WA)
Assignee: ADOBE INC.
G06N3/08G06N3/0455G06N3/048G06T3/4007G06T5/77G06T2207/10016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,169
App. No.
17/198,670
Granted
Feb 4, 2025
Kind
B2
Abstract

Systems and methods for signal processing are described. Embodiments receive a digital signal comprising original signal values corresponding to a discrete set of original sample locations, generate modulation parameters based on the digital signal using a modulator network, wherein each of a plurality of modulator layers of the modulator network outputs a set of the modulation parameters, and generate a predicted signal value of the digital signal at an additional location using a synthesizer network, wherein each of a plurality of synthesizer layers of the synthesizer network operates based on the set of the modulation parameters from a corresponding modulator layer of the modulator network.

Claims (44)

1. A method for signal processing, comprising:

receiving a digital signal comprising signal values corresponding to discrete sample locations;

generating modulation parameters based on the digital signal using a modulator network, wherein each of a plurality of modulator layers of the modulator network outputs a set of the modulation parameters; and

generating a predicted signal value of the digital signal at an additional location using a synthesizer network, computing, at a first synthesizer layer, a product of a set of modulation parameters output by a modulator layer and a continuous function of a first set of features to produce a second set of features, wherein the second set of features are input to a second synthesizer layer, wherein each of a plurality of synthesizer layers of the synthesizer network operates based on the set of the modulation parameters from a corresponding modulator layer of the modulator network.

2. The method of claim 1 , further comprising:

encoding at least a portion of a digital image to produce a latent vector, wherein the modulation parameters are generated based on the latent vector.

3. The method of claim 2 , wherein:

the discrete set original sample locations correspond to pixel location within the digital image, and the predicted signal value comprises color values at the additional location.

4. The method of claim 2 , further comprising:

inpainting at least one pixel of the digital image based on the predicted signal value, wherein the additional location comprises a location of the inpainted at least one pixel.

5. The method of claim 2 , further comprising:

generating a refined digital image based on the predicted signal value, wherein the refined digital image comprises a higher resolution than the digital image.

6. The method of claim 1 , further comprising:

interpolating an intermediate frame of a digital video based on the predicted signal value to produce an updated digital video with a higher frame rate than the digital video, wherein the digital signal comprises the digital video.

7. The method of claim 1 , wherein:

the digital signal comprises a digital audio signal.

8. The method of claim 1 , wherein:

the digital signal comprises a three dimensional (3D) image, and the original signal values represent distances from a surface of the 3D image.

9. The method of claim 1 , wherein:

the continuous function comprises a sine function, and the product comprises a Hadamard product.

10. The method of claim 1 , wherein:

the modulator network and the synthesizer network are trained based on training signals other than the digital signal.

11. The method of claim 1 , wherein:

the discrete set of original sample locations does not include the additional location.

12. An apparatus for signal processing, comprising:

at least one processor;

at least one memory;

a modulator network comprising a plurality of modulator layers, wherein each of the plurality of modulator layers of the modulator network is configured to output a different set of modulation parameters based on a same digital signal;

an encoder configured to produce a latent vector representing the digital signal, wherein the modulator network takes the latent vector as input; and

a synthesizer network comprising a plurality of synthesizer layers, wherein the synthesizer network represents a continuous function of a signal parameter of the digital signal, and wherein each of the synthesizer layers is configured to receive the set of modulation parameters from a corresponding modulator layer of the modulator network.

13. The apparatus of claim 12 , wherein:

the modulator layers comprise multi-layer perceptron (MLP) layers with rectified linear unit (ReLU) activation.

14. The method of claim 12 , wherein:

the synthesizer layers comprise MLP layers with sine function activation.

15. A method for training a neural network, comprising:

receiving a digital signal comprising signal values corresponding to discrete sample locations;

generating modulation parameters based on the digital signal using a modulator network, wherein each of a plurality of modulator layers of the modulator network outputs a set of the modulation parameters, encoding the digital signal to produce a latent vector, wherein the modulator network takes the latent vector as input to produce the modulation parameters;

generating a predicted signal value of the digital signal for at least one of the original sample locations using a synthesizer network, wherein each of a plurality of synthesizer layers of the synthesizer network receives the set of the modulation parameters from a corresponding modulator layer of the modulator network;

computing a loss function based on the predicted signal value and a value of the original signal values corresponding to the at least one of the original sample locations; and

training the modulator network and the synthesizer network based on the loss function.

16. The method of claim 15 , wherein:

the predicted signal values correspond to predicted color values of a training image at pixel locations corresponding to the original sample locations.

17. The method of claim 15 , wherein:

the training is based on an auto-encoder training process.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 11, 2021
From: MEHTA, ISHIT BHADRESH; GHARBI, MICHAËL; BARNES, CONNELLY; SHECHTMAN, ELYA
To: ADOBE INC.
Reel/Frame 055563/0787 →
Continuity (1)
Related Publication 20220292341A1 · Sep 15, 2022
References Cited (43)
US 20230419082A1 · Yuan · 2023 [cited by examiner]
Niklaus, Simon, Long Mai, and Feng Liu. “Video frame interpolation via adaptive convolution.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2017. (Year: 2017). [cited by examiner]
Li, Yangyan, et al. “Fpnn: Field probing neural networks for 3d data.” Advances in neural information processing systems 29 (2016). (Year: 2016). [cited by examiner]
Si, Jiong, Sarah L. Harris, and Evangelos Yfantis. “A dynamic ReLU on neural network.” 2018 IEEE 13th Dallas Circuits and Systems Conference (DCAS). IEEE, 2018. (Year: 2018). [cited by examiner]
Chen, Je-Chian, and Yu-Min Wang. “Comparing activation functions in modeling shoreline variation using multilayer perceptron neural network.” Water 12.5 (2020): 1281. (Year: 2020). [cited by examiner]
1Atzmon et al., “SAL: Sign Agnostic Learning of Shapes from Raw Data”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2565-2574, 2020. [cited by applicant]
2Bi, et al, “Deep Reflectance vols. Relightable Reconstructions from Multi-View Photometric Images”, arXiv preprint: arXiv:2007.09892v1 [cs.CV] Jul. 20, 2020, 21 pages. [cited by applicant]
3Chabra, et al., “Deep Local Shapes: Learning Local SDF Priors for Detailed 3D Reconstruction”, arXiv preprint: arXiv:2003.10983v3 [cs.CV] Aug. 21, 2020, 26 pages. [cited by applicant]
4Chen et al., “Learning Implicit Fields for Generative Shape Modeling”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5939-5948, 2019. [cited by applicant]
5Davies, et al., “Overfit Neural Networks as a Compact Shape Representation”, arXiv preprint: arXiv:2009.09808v2 [cs.GR] Oct. 12, 2020, 9 pages. [cited by applicant]
6Glorot, et al., “Understanding the difficulty of training deep feedforward neural networks”, In Proceedings of the thirteenth international conference on artificial intelligence and statistics, pp. 249-256, 2010. [cited by applicant]
7Ha, et al., “HyperNetworks”, arXiv preprint: arXiv:1609.09106v4 [cs:LG] Dec. 1, 2016, 29 pages. [cited by applicant]
8Jiang, et al., “Local Implicit Grid Representations for 3D Scenes”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6001-6010, 2020. [cited by applicant]
9Klocek, et al., “Hypernetwork functional image representation”, In International Conference on Artificial Neural Networks, arXiv:1902.10404v3 [cs.LG] Jun. 3, 2019, pp. 496-510, Springer. [cited by applicant]
10Lapedes et al., “Nonlinear Signal Processing Using Neural Networks: Prediction and System Modelling”, Technical report, 1987, 52 pages. [cited by applicant]
11Liu, et al., “Neural Sparse Voxel Fields”, Advances in Neural Information Processing Systems, 33, 2020, arXiv:2007.11571v2 [cs.CV] Jan. 6, 2021, 22 pages. [cited by applicant]
12Liu, et al., “Deep Learning Face Attributes in the Wild”, In Proceedings of International Conference on Computer Vision (ICCV), Dec. 2015, pp. 3730-3738. [cited by applicant]
13Lombardi, et al., “Neural Volumes: Learning Dynamic Renderable vols. from Images”, arXiv preprint: arXiv:1906.07751v1 [cs.GR] Jun. 18, 2019, 14 pages. [cited by applicant]
14Lorensen et al., “Marching Cubes: a High Resolution 3D Surface Construction Algorithm”, ACM siggraph computer graphics, 21(4):163-169, 1987. [cited by applicant]
15Mahajan, et al., “A Theory of Locally Low Dimensional Light Transport”, In ACM SIGGRAPH 2007 papers, pp. 62-1-62-10, 2007. [cited by applicant]
16Mildenhall, et al., “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis”, arXiv preprint: arXiv:2003.08934v2 [cs.CV] Aug. 3, 2020, 25 pages. [cited by applicant]
17Bemana, et al., “X-Fields: Implicit Neural View-, Light- and Time-Image Interpolation”, ACM Transactions on Graphics (Proc. SIGGRAPH Asia 2020), 39(6), 2020, pp. 257:1-257:15. [cited by applicant]
18Mordvintsev, et al., “Differentiable Image Parameterizations”, Distill, 2018, 27 pages. [cited by applicant]
19Niemeyer, et al., “Differentiable Volumetric Rendering: Learning Implicit 3D Representations without 3D Supervision”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3504-3515… [cited by applicant]
20OLah, et al., “Feature Visualization”, Distill, 2017, 19 pages. [cited by applicant]
21Parascandolo, et al., “Taming the Waves: Sine as Activation Function in Deep Neural Networks”, 2016, 12 pages. [cited by applicant]
22Park, et al., “DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 165-174, 2019. [cited by applicant]
23Radford, et al., “Unsupervised Representation Learning With Deep Convolutional Generative Adversarial Networks”, arXiv preprint: arXiv:1511.06434v2 [cs.LG] Jan. 7, 2016, 16 pages. [cited by applicant]
24Schwarz, et al., “GRAF: Generative Radiance Fields for 3D-Aware Image Synthesis”, arXiv preprint: arXiv:2007.02442v3 [cs.CV] Dec. 15, 2020, 13 pages. [cited by applicant]
25Shi, et al., “A Benchmark Dataset and Evaluation for Non-Lambertian and Uncalibrated Photometric Stereo”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3707-3716, 2016. [cited by applicant]
26Sitzmann, et al., “MetaSDF: Meta-learning Signed Distance Functions”, Advances in Neural Information Processing Systems, 33, arXiv:2006.09662v1 [cs.CV] Jun. 17, 2020, 17 pages. [cited by applicant]
27Sitzmann, et al., “Implicit Neural Representations with Periodic Activation Functions”, arXiv preprint: arXiv:2006.09661v1 [cs.CV] Jun. 17, 2020, 35 pages. [cited by applicant]
28Sitzmann, et al., “DeepVoxels: Learning Persistent 3D Feature Embeddings”, In Proc. Computer Vision and Pattern Recognition (CVPR), IEEE, 2019, pp. 2437-2446. [cited by applicant]
29Sitzmann, et al., “Scene Representation Networks: Continuous 3D-Structure-Aware Neural Scene Representations”, In Advances in Neural Information Processing Systems, arXiv:1906.01618v2 [cs.CV] Jan. 28, 2020, pp. 1121-1… [cited by applicant]
30Sopena, et al., “Neural networks with periodic and monotonic activation functions: a comparative study in classification problems”, 1999, 6 pages. [cited by applicant]
31Stanley, “Compositional pattern producing networks: A novel abstraction of development”, 8(2):131-162, 2007. [cited by applicant]
32Tancik, et al., “Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains”, arXiv preprint: arXiv:2006.10739v1 [cs.CV] Jun. 18, 2020, 24 pages. [cited by applicant]
33Thies, et al., “Deferred Neural Rendering: Image Synthesis using Neural Textures”, ACM Transactions on Graphics (TOG), 38(4):1-12, arXiv:1904.12356v1 [cs.CV] Apr. 28, 2019. [cited by applicant]
34Vaswani, et al., “Attention Is All You Need”, In Advances in neural information processing systems, pp. 5998-6008, arXiv:1706.03762v5 [cs.CL] Dec. 6, 2017. [cited by applicant]
35Wood, et al., “Surface light fields for 3D photography”, In Proceedings of the 27th annual conference on Computer graphics and interactive techniques, pp. 287-296, 2000, 136 pages. [cited by applicant]
36Xue, et al., “Video Enhancement with Task-Oriented Flow”, International Journal of Computer Vision (IJCV), 127 (8):1106-1125, arXiv:1711.09078v3 [cs.CV] Nov. 10, 2019. [cited by applicant]
37Zhou, et al., “Thingi10K: A Dataset of 10,000 3D-Printing Models”, arXiv preprint: arXiv:1605.04797v2 [cs.GR] Jul. 2, 2016, 8 pages. [cited by applicant]
38Mescheder, et al., “Occupancy Networks: Learning 3D Reconstruction in Function Space”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2019, pp. 4460-4470. [cited by applicant]