IP Library › Granted Patent US 12,469,256
Granted Patent B2
US 12,469,256 · App. 17/870,462 · Granted Nov 11, 2025

Performance of complex optimization tasks with improved efficiency via neural meta-optimization of experts

Inventors: Avneesh Sud (Belmont, CA); Andrea Tagliasacchi (Toronto, CA); Ben Usman (Boston, MA)
Assignee: GOOGLE LLC
G06V10/7715G06N5/022G06T7/70G06T2207/20084G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,469,256
App. No.
17/870,462
Granted
Nov 11, 2025
Kind
B2
Abstract

Example systems perform complex optimization tasks with improved efficiency via neural meta-optimization of experts. In particular, provided is a machine learning framework in which a meta-optimization neural network can learn to fuse a collection of experts to provide a predicted solution. Specifically, the meta-optimization neural network can learn to predict the output of a complex optimization process which optimizes over outputs from the collection of experts to produce an optimized output. In such fashion, the meta-optimization neural network can, after training, be used in place of the complex optimization process to produce a synthesized solution from the experts, leading to orders of magnitude faster and computationally more efficient prediction or problem solution.

Claims (46)

1 . A computer-implemented method for performing complex optimization tasks with improved efficiency or accuracy, the method comprising:

obtaining, by a computing system comprising one or more computing devices, a set of input data, wherein the input data comprises a plurality of images that depict a scene;

processing, by the computing system, the input data with one or more existing expert models to generate one or more expert outputs, wherein the one or more expert outputs comprise a plurality of features detected in the plurality of images;

processing, by the computing system, the one or more expert outputs with a meta-optimization neural network to generate a predicted output;

performing, by the computing system, an optimization technique on the one or more expert outputs to generate an optimized output, wherein performing, by the computing system, the optimization technique on the one or more expert outputs to generate the optimized output comprises performing, by the computing system, a bundle adjustment technique on the plurality of features to generate the optimized output, and wherein the optimized output comprises a geometry of the plurality of images relative to the scene; and

modifying, by the computing system, one or more learnable parameters of the meta-optimization neural network based at least in part on a loss function that compares the predicted output with the optimized output.

2 . The computer-implemented method of claim 1 , wherein the optimization technique comprises an iterative minimization technique.

3 . The computer-implemented method of claim 1 , wherein the one or more existing expert models comprise one or more machine-learned expert models which have previously been trained to generate the one or more expert outputs.

4 . The computer-implemented method of claim 1 , wherein the loss function comprises one or more loss terms that encode one or more priors of the optimized output.

5 . The computer-implemented method of claim 1 , wherein:

the input data comprises a sequence of inputs over time; and

the one or more expert outputs comprise a sequence of expert outputs respectively generated over time by the one or more existing experts from the sequence of inputs over time.

6 . The computer-implemented method of claim 1 , wherein the one or more expert outputs have a same data structure as the optimized output and the predicted output.

7 . The computer-implemented method of claim 1 , wherein the one or more expert outputs comprise one or more hyperpriors.

8 . The computer-implemented method of claim 1 , wherein:

the plurality of images depict an object;

the one or more expert outputs comprise an initial predicted pose for the object;

the predicted output comprises a final predicted pose for the object; and

the optimized output further comprises a refined pose for the object.

9 . The computer-implemented method of claim 8 , wherein:

the object comprises a human body; and

the initial predicted pose, the final predicted pose, and the refined pose for the human body are parameterized using joint locations.

10 . The computer-implemented method of claim 1 , wherein the plurality of images comprise a plurality of monocular images.

11 . A computing system, comprising:

one or more processors; and

one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising:

obtaining a set of input data, wherein the input data comprises a plurality of images that depict a scene;

processing the input data with one or more existing expert models to generate one or more expert outputs, wherein the one or more expert outputs comprise a plurality of features detected in the plurality of images; and

processing the one or more expert outputs with a meta-optimization neural network to generate a predicted output;

wherein the meta-optimization neural network has been trained to generate the predicted output by performance of a supervised learning approach relative to optimized outputs generated by performance of an optimization technique on initial inputs generated by the one or more existing expert models, wherein performance of the optimization technique comprises performance of a bundle adjustment technique on initial features generated by the one or more existing expert models to generate the optimized output, and wherein the optimized output comprises a geometry of the plurality of images relative to the scene.

12 . The computer system of claim 11 , wherein the optimization technique comprises an iterative minimization technique.

13 . The computer system of claim 11 , wherein the one or more existing expert models comprise one or more machine-learned expert models which have previously been trained to generate the one or more expert outputs.

14 . The computer system of claim 11 , wherein:

the input data comprises a sequence of inputs over time; and

the one or more expert outputs comprise a sequence of expert outputs respectively generated over time by the one or more existing experts from the sequence of inputs over time.

15 . One or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more processors, cause a computing system to perform operations, the operations comprising:

obtaining, by the computing system, a plurality of images that depict a scene;

processing, by the computing system, the plurality of images with one or more existing expert models to generate a plurality of features detected in the plurality of images;

processing, by the computing system, the plurality of images with a meta-optimization neural network to generate a predicted output, wherein the predicted output comprises a predicted geometry of the plurality of images relative to the scene;

performing, by the computing system, a bundle adjustment technique on the plurality of features to generate the optimized output, wherein the optimized output comprises an optimized geometry of the plurality of images relative to the scene; and

modifying, by the computing system, one or more learnable parameters of the meta-optimization neural network based at least in part on a loss function that compares the predicted geometry of the plurality of images with the optimized geometry of the plurality of images relative to the scene.

16 . The one or more non-transitory computer-readable media of claim 15 , wherein:

the plurality of images depict an object;

the one or more expert outputs comprise an initial predicted pose for the object;

the predicted output comprises a final predicted pose for the object; and

the optimized output further comprises a refined pose for the object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 10, 2022
From: SUD, AVNEESH; USMAN, BEN; TAGLIASACCHI, ANDREA
To: GOOGLE LLC
Reel/Frame 060773/0719 →
Continuity (2)
Provisional Application 63224079 · Jul 21, 2021
Related Publication 20230040793A1 · Feb 9, 2023
References Cited (57)
US 10713794B1 · He · 2020 [cited by examiner]
US 11238650B2 · Li · 2022 [cited by examiner]
US 20190138786A1 · Trenholm · 2019 [cited by examiner]
US 20200357111A1 · Wang · 2020 [cited by examiner]
US 20210192136A1 · Sar Shalom · 2021 [cited by examiner]
US 20220232162A1 · Gupta · 2022 [cited by examiner]
Agarwal et al, “Building Rome in a Day”, Communications of the ACM, vol. 54, No. 10, Oct. 2011, 8 pages. [cited by applicant]
Belagiannis et al, “3D Pictorial Structures for Multiple Human Pose Estimation”, Conference on Computer Vision and Pattern Recognition, Jun. 23-28, 2014, Columbus, Ohio, United States, 8 pages. [cited by applicant]
Branch et al, “A Subspace, Interior, and Conjugate Gradient Method for Large-Scale Bound-Constrained Minimization Problems”, SIAM Journal on Scientific Computing, vol. 21, Issue 12, 1999 29 pages. [cited by applicant]
Bridgeman et al, “Multi-Person 3D Pose Estimation and Tracking in Sports”, Conference on Computer Vision and Pattern Recognition, Jun. 16-20, 2019, Long Beach, California, United States, 10 pages. [cited by applicant]
Byrd et al, “Approximate Solution of the Trust Region Problem by Minimization over Two-Dimensional Subspace”. Mathematical Programming, Oct. 1988, 31 pages. [cited by applicant]
Chen et al, “Cross-View Tracking for Multi-Human 3D Pose Estimation at Over 100 FPS”, Conference on Computer Vision and Pattern Recognition, Jun. 14-19, 2020, Virtual, pp. 3279-3288. [cited by applicant]
Chen et al, “Unsupervised 3D Pose Estimation with Geometric Self-Supervision”, arXiv:1904.04812v1, Apr. 9, 2019, 11 pages. [cited by applicant]
Deng et al, “Vector Neurons: A General Framework for so (3)-Equivariant Networks”, arXiv:2104.12229v1, Apr. 25, 2021, 12 pages. [cited by applicant]
Drover et al, “Can 3D Pose be Learned from 2D Projections Alone?”, European Conference on Computer Vision. Sep. 8-14, 2018, Munich, Germany, 17 pages. [cited by applicant]
Feng et al, “Hgaze Typing: Head-Gesture Assisted Gaze Typing”, Symposium on Eye Tracking Research and Applications, May 25-27, 2021, Virtual, 11 pages. [cited by applicant]
Frisch et al, “Gaussian Mixture Estimation from Weighted Samples”, arXiv:2106.05109v1, Jun. 9, 2021 7 pages. [cited by applicant]
Gleicher, “Animation from Observation: Motion Capture and Motion Editing”, Computer Graphics, vol. 33, No. 4, Jul. 2014, 5 pages. [cited by applicant]
Gu et al, “Home-Based Physical Therapy with an Interactive Computer Vision System”, International Conference on Computer Vision, 2019, Oct. 27-Nov. 2, 2019, Seoul, Korea, 10 pages. [cited by applicant]
Guler et al, “Densepose: Dense Human Pose Estimation in the Wild”, arXiv:1802.00434v1, Feb. 1, 2018, 12 pages. [cited by applicant]
Guler et al, “Holopose: Holistic 3D Human Reconstruction In-The-Wild”, Conference on Computer Vision and Pattern Recognition, Jun. 16-20, 2019, Long Beach, California, United States, pp. 10884-10894. [cited by applicant]
He et al, “Epipolar Transformers”, Conference on Computer Vision and Pattern Recognition, Jun. 14-19, 2020, Virtual, pp. 7779-7788. [cited by applicant]
Ionescu et al, “Human3 6m: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments”, Transactions on Pattern Analysis and Machine Intelligence, Sep. 2014, 15 pages. [cited by applicant]
Iqbal et al, “Weakly Supervised 3D Human Pose Learning via Multi-View Images in the Wild”, arXiv:2003.07581v1, Mar. 17. 2020, 11 pages. [cited by applicant]
Iskakov et al, “Learnable Triangulation of Human Pose”, arXiv:1905.05754v1, May 14, 2019, 9 pages. [cited by applicant]
Jakab et al, “Self-Supervised Learning of Interpretable Keypoints from Unlabelled Videos”. arXiv:1907.02055v2. Dec. 23, 2020, 21 pages. [cited by applicant]
Jedlovec, “Introducing Statcast 2020: Hawk-Eye and Google Cloud”, MLB Technology Blog, Jul. 20, 2020, https://technology.mlblogs.com/introducing-statcast-2020-hawk-eye-and-google-cloud-a5f5c20321b8, retrieved on Aug. 24… [cited by applicant]
Joo et al, “Exemplar Fine-Tuning for 3D Human Model Fitting Towards in-the-Wild 3D Human Pose Estimation”, arXiv:2004.03686v3, Oct. 22, 2021, 21 pages. [cited by applicant]
Joo et al, Panoptic Studio: A Massively Multiview System for Social Interaction Capture, arXiv:1612.03153v1, Dec. 9, 2016, 14 pages. [cited by applicant]
Joseph-Rivlin et al, “Momen(e)t: Flavor the Moments in Learning to Classify Shapes”, arXiv:1812.07431v2, Oct. 3, 2019, 10 pages. [cited by applicant]
Kanazawa et al, “End-to-End Recovery of Human Shape and Pose”, arXiv:1712.06584v2, Jun. 23, 2018, 10 pages. [cited by applicant]
Karashchuk et al, “Anipose: a Toolkit for Robust Markerless 3D Pose Estimation”, Cell Reports, vol. 36, Issue 13, Sep. 28, 2021, 53 pages. [cited by applicant]
Kingma et al, “Adam: A Method for Stochastic Optimization”, arXiv:1412.6980v9, Jan. 30, 2017, 15 pages. [cited by applicant]
Kocabas et al, “Self-Supervised Learning of 3D Human Pose using Multi-View Geometry”, arXiv:1903.02330v2, Apr. 9, 2019, 10 pages. [cited by applicant]
Kocabas et al, “Vibe: Video Inference for Human Body Pose and Shape Estimation”, arXiv:1912.05656v3, Apr. 29, 2020, 12 pages. [cited by applicant]
Kolotouros et al, “Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the Loop”, arXiv:1909.12828v1, Sep. 27, 2019, 10 pages. [cited by applicant]
Kundu et al, “Self-Supervised 3D Human Pose Estimation via Part Guided Novel Image Synthesis”, Conference on Computer Vision and Pattern Recognition, Jun. 14-19, 2020, Virtual, pp. 6152-6162. [cited by applicant]
Ma et al, “Deep Feedback Inverse Problem Solver”, arXiv:2101.07719v1, Jan. 19, 2021, 18 pages. [cited by applicant]
Martinez et al, “A Simple Yet Effective Baseline for 3D Human Pose Estimation”, arXiv:1705.03098v2, Aug. 4, 2017, 10 pages. [cited by applicant]
Mitra et al, “Multiview-Consistent Semi-Supervised Learning for 3D Human Pose Estimation”, Conference on Computer Vision and Pattern Recognition, Jun. 14-19, 2020, Virtual, pp. 6907-6916. [cited by applicant]
Newell et al, “Stacked Hourglass Networks for Human Pose Estimation”, arXiv:1603.06937v2, Jul. 26, 2016, 17 pages. [cited by applicant]
Papandreou et al, “Person-Lab: Person Pose Estimate and Instance Segmentation with a Bottom-Up, Part-Based, Geometric Embedding Model”, arXiv:1803.08225v1, Mar. 22, 2018, 21 pages. [cited by applicant]
Rhodin et al, “Learning Monocular 3D Human Pose Estimation from Multi-View Images”, arXiv:1803.04775v2, Mar. 24, 2018, 10 pages. [cited by applicant]
Rhodin et al, “Unsupervised Geometry-Aware Representation for 3D Human Pose Estimation”, arXiv:1804.01110v1, Apr. 3, 2018, 17 pages. [cited by applicant]
Rosales et al, “Estimating 3D Body Pose using Uncalibrated Cameras”, Boston University Technical Report, No. 2001-008, Jun. 2001, 8 pages. [cited by applicant]
Saraee et al, “Exercisecheck: Data Analytics for a Remote Monitoring and Evaluation Platform for Home-Based Physical Therapy”, Conference on Pervasive Technologies Related to Assistive Environments, Jun. 29-Jul. 1, 2019… [cited by applicant]
Schonemann, “A Generalized Solution of the Orthogonal Procrustes Problem”, Psychometrika, vol. 31, No. 1, Mar. 1966, 10 pages. [cited by applicant]
Takahashi et al, “Human Pose as Calibration Pattern: 3D Human Pose Estimation with Multiple Unsynchronized and Uncalibrated Cameras”, Conference on Computer Vision and Pattern Recognition, Jun. 18-22, 2018, Salt Lake Ci… [cited by applicant]
Tu et al, “Voxelpose: Towards Multi-Camera 3D Human Pose Estimation in Wild Environment”, arXiv:2004.06239v4, Aug. 24, 2020, 17 pages. [cited by applicant]
Wandt et al, “CanonPose: Self-Supervised Monocular 3d Human Pose Estimation in the Wild”, Conference on Computer Vision and Pattern Recognition, Jun. 19-25, 2021, Nashville, Tennessee, United States, pp. 13294-13304. [cited by applicant]
Wandt et al, “Repent: Weakly Supervised Training of an Adversarial Reprojection Network for 3D Human Pose Estimation”, arXiv:1902.09868v2, Mar. 12, 2019, 10 pages. [cited by applicant]
Xie et al, “MetaFuse: A Pre-Trained Fusion Model for Human Pose Estimation”, Conference on Computer Vision and Pattern Recognition, Jun. 14-19, 2020, Virtual, pp. 13686-13695. [cited by applicant]
Zanfir et al, “Neural Descent for Visual 3D Human Pose and Shape”, arXiv:2008.06910v2, Jun. 14, 2021, 10 pages. [cited by applicant]
Zanfir et al, “Weakly Supervised 3D Human Pose and Shape Reconstruction with Normalizing Flows”, European Conference on Computer Vision, Aug. 23-28, 2020, Virtual, 17 pages. [cited by applicant]
Zhou et al, “Fast Global Registration”, European Conference on Computer Vision, Oct. 8-16, 2016, Amsterdam, The Netherlands, 16 pages. [cited by applicant]
Zhou et al, “On the Continuity of Rotation Representations in Neural Networks”, arXiv:1812.07035v4, Jun. 8, 2020, 13 pages. [cited by applicant]
Zhou et al, “Towards 3D Human Pose Estimation in the Wild: A Weakly-Supervised Approach”, arXiv:1704.02447v2, Jul. 30, 2017, 10 pages. [cited by applicant]