IP Library Granted Patent US 12,499,678
Granted Patent B2
US 12,499,678 · App. 18/230,414 · Granted Dec 16, 2025

Unsupervised pre-training of geometric vision models

Inventors: Romain Brégier (Grenoble, FR); Yohann Cabon (Montbonnot-Saint-Martin, FR); Thomas Lucas (Grenoble, FR); Jérôme Revaud (Meylan, FR); Philippe Weinzaepfel (Montonnot-Saint-Martin, FR); Boris Chidlovskii (Meylan, FR); Vincent Leroy (Laval, FR); Leonid Antsfeld (Saint Ismier, FR); Gabriela Csurka Khedari (Crolles, FR)
Assignee: NAVER CORPORATION
G06V10/82G06N3/0455G06N3/088G06T7/70G06V10/44G06V10/7753G06V10/776G06V10/80G06V10/96G06V20/64G06V40/10G06T2207/20081G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,678
App. No.
18/230,414
Granted
Dec 16, 2025
Kind
B2
Abstract

A method includes: performing unsupervised pre-training of a model, the model including and a decoder including: obtaining a first image and a second image under different conditions or from different viewpoints; encoding, by the encoder, the first image into a representation of the first image and the second image into a representation of the second image; transforming the representation of the first image into a transformed representation; decoding, by the decoder, the transformed representation into a reconstructed image, where the transforming of the representation of the first image and the decoding of the transformed representation is based on the representation of the first image and the representation of the second image; and adjusting one or more parameters of at least one of the encoder and the decoder based on minimizing a loss; and fine-tuning the model, initialized with a set of task specific encoder parameters, for a geometric vision task.

Claims (121)

1 . A computer-implemented machine learning method of training a task specific machine learning model for a downstream geometric vision task, the method comprising:

performing unsupervised pre-training of a machine learning model, the machine learning model comprising an encoder having a set of encoder parameters and a decoder having a set of decoder parameters,

wherein the performing of the unsupervised pre-training of the machine learning model includes:

obtaining a pair of unannotated images including a first image and a second image,

wherein the first and second images depict a same scene and are taken under different conditions or from different viewpoints;

encoding, by the encoder, the first image into a representation of the first image and the second image into a representation of the second image;

transforming the representation of the first image into a transformed representation;

decoding, by the decoder, the transformed representation into a reconstructed image,

wherein the transforming of the representation of the first image and the decoding of the transformed representation is based on the representation of the first image and the representation of the second image; and

adjusting one or more parameters of at least one of the encoder and the decoder based on minimizing a loss;

constructing the task specific machine learning model for the downstream geometric vision task based on the pre-trained machine learning model,

the task specific machine learning model comprising a task specific encoder having a set of task specific encoder parameters;

initializing the set of task specific encoder parameters with the set of encoder parameters of the pre-trained machine learning model; and

fine-tuning the task specific machine learning model, initialized with the set of task specific encoder parameters, for the downstream geometric vision task.

2 . The method of claim 1 , wherein the unsupervised pre-training of the machine learning model is a cross-view alignment pre-training, and wherein the transforming of the representation of the first image includes applying a transformation to the representation of the first image to generate the transformed representation, the transformation being determined based on the representation of the first image and the representation of the second image such that the transformed representation approximates the representation of the second image.

3 . The method of claim 1 , wherein the unsupervised pre-training of the machine learning model is a cross-view alignment pre-training, and wherein the loss is based on a metric quantifying a difference between the reconstructed image and the second image.

4 . The method of claim 1 , wherein the unsupervised pre-training of the machine learning model is a cross-view alignment pre-training, and wherein the loss is based on a metric quantifying a difference between the transformed representation and the representation of the second image.

5 . The method of claim 2 , wherein the representation of the first image is a first set of n vectors {x 1,i } i=1 . . . n , each x 1,i ∈ K , wherein the representation of the second image is a second set of n vectors {x 2,i } i=1 . . . n , each x 2,i ∈ K , wherein the applying of the transformation includes decomposing each vector of the first and second sets of vectors in a D-dimensional equivariant part and a (K−D)-dimensional invariant part and applying a (D×D)-dimensional transformation matrix Ω to the equivariant part of each vector of the first set of vectors, wherein 0<D≤K.

6 . The method of claim 5 , wherein the transformation is a D-dimensional rotation and Ω is a D-dimensional rotation matrix, and wherein Ω is set based on aligning the equivariant parts of the vectors of the first set of vectors with the equivariant parts of the respective vectors of the second set of vectors.

7 . The method of claim 5 , further comprising determining Ω based on the equation:

Ω

=

arg

min

Ω

^

SO

(

D

)

i

=

1

n

Ω

ˆ

x

1

,

i

e

q

u

i

v

-

x

2

,

i

e

q

u

i

v

2

where x 1,i equiv denotes the equivariant part of vector x 1,i , x 2,i equiv denotes the equivariant part of vector x 2,i , and SO(D) denotes the D-dimensional rotation group.

8 . The method of claim 1 , wherein the unsupervised pre-training is a cross-view completion pre-training, and wherein the performing of the cross-view completion pre-training of the machine learning model further comprises:

splitting the first image into a first set of non-overlapping patches and splitting the second image into a second set of non-overlapping patches; and

masking ones of the patches of the first set of patches,

wherein the encoding of the first image into the representation of the first image includes, encoding, by the encoder, each unmasked patch of the first set of patches into a corresponding representation of the respective unmasked patch, thereby generating a first set of patch representations,

wherein the encoding the second image into the representation of the second image includes, encoding, by the encoder, each patch of the second set of patches into a corresponding representation of the respective patch, thereby generating a second set of patch representations,

wherein the decoding of the transformed representation includes, generating, by the decoder, for each masked patch of the first set of patches, a predicted reconstruction for the respective masked patch based on the transformed representation and the second set of patch representations, and

wherein the loss function is based on a metric quantifying the difference between each masked patch and its respective predicted reconstruction.

9 . The method of claim 8 , wherein the transforming of the representation of the first image into the transformed representation further includes, for each masked patch of the first set of patches, padding the first set of patch representations with a respective learned representation of the masked patch.

10 . The method of claim 9 , wherein each learned representation includes a set of representation parameters.

11 . The method of claim 9 , wherein the generating of the predicted reconstruction of a masked patch of the first set of patches includes decoding, by the decoder, the learned representation of the masked patch into the predicted reconstruction of the masked patch, where the decoder receives the first and second sets of patch representations as input data and decodes the learned representation of the masked patch based on the input data, and wherein the method further includes adjusting the learned representations of the masked patches by adjusting the respective set of representation parameters.

12 . The method of claim 11 wherein the adjusting the respective set of representation parameters includes adjusting the set of representation parameters based on minimizing the loss.

13 . The method of claim 1 , wherein fine-tuning the task specific machine learning model includes performing supervised fine-tuning of the task specific machine learning model.

14 . The method of claim 13 wherein the supervised fine-tuning includes:

i) applying the task specific machine learning model to one or more annotated images to create output data according to the downstream geometric vision task,

wherein the annotated images are annotated with ground truth data for the downstream geometric vision task; and

ii) adjusting the task specific machine learning model by adjusting the set of task specific encoder parameters of the task specific machine learning model based on minimizing a task specific loss, the task specific loss being based on a metric quantifying a difference between the created output data and the ground truth data.

15 . The method of claim 14 further comprising repeating i) and ii) for a predetermined number of times or until a minimum value of the task specific loss is reached.

16 . The method of claim 14 further comprising applying the fine-tuned task specific machine learning model to one or more new images to extract prediction data from the one or more new images for the downstream geometric vision task.

17 . The method of claim 16 , wherein the downstream geometric vision task is relative pose estimation.

18 . The method of claim 17 wherein the applying of the fine-tuned task specific machine learning model includes applying the fine-tuned task specific machine learning model to a new pair of images to extract a relative rotation and a relative translation between the images of the new pair of images as the prediction data,

wherein the images of the pair of new images depict two views of a same scene.

19 . The method of claim 16 , wherein the downstream geometric vision task is depth estimation.

20 . The method of claim 19 wherein the applying of the fine-tuned task specific machine learning model includes applying the fine-tuned task specific machine learning model to one or more new images to extract one or more depth maps from the one or more new images as the prediction data.

21 . The method of claim 16 , wherein the downstream geometric vision task is optical flow estimation.

22 . The method of claim 21 wherein the applying of the fine-tuned task specific machine learning model includes applying the fine-tuned task specific machine learning model to a new image pair including a new first image and a new second image to generate a plurality of pixel pairs as the prediction data,

wherein each pixel pair includes a pixel in the new first image and a corresponding pixel in the new second image,

where the new first and second images depict a same scene under different conditions or from different viewpoints.

23 . A computer-implemented machine learning method for generating prediction data according to a downstream geometric vision task, the method comprising:

training a first task specific machine learning model for the downstream geometric vision task using cross-view alignment pre-training of a first machine learning model;

training a second task specific machine learning model for the downstream geometric vision task using cross-view completion pre-training of a second pretext machine learning model;

generating first prediction data according to the downstream geometric vision task by applying the trained first task specific machine learning model to at least one image;

generating second prediction data according to the downstream geometric vision task by applying the trained second task specific machine learning model to the at least one image;

determining a first confidence value for the first prediction data and a second confidence value for the second prediction data; and

generating resulting prediction data according to the geometric vision task by fusing together the first and second prediction data based on the first and second confidence values.

24 . A computing system comprising:

one or more processors; and

memory including code that, when executed by the one or more processors, perform to:

perform unsupervised pre-training of a machine learning model, the machine learning model comprising an encoder having a set of encoder parameters and a decoder having a set of decoder parameters,

wherein the performance of the unsupervised pre-training of the machine learning model includes:

obtaining a pair of unannotated images including a first image and a second image,

wherein the first and second images depict a same scene and are taken under different conditions or from different viewpoints;

encoding, by the encoder, the first image into a representation of the first image and the second image into a representation of the second image;

transforming the representation of the first image into a transformed representation;

decoding, by the decoder, the transformed representation into a reconstructed image,

wherein the transforming of the representation of the first image and the decoding of the transformed representation is based on the representation of the first image and the representation of the second image; and

adjusting one or more parameters of at least one of the encoder and the decoder based on minimizing a loss;

construct the task specific machine learning model for the downstream geometric vision task based on the pre-trained machine learning model,

the task specific machine learning model comprising a task specific encoder having a set of task specific encoder parameters;

initialize the set of task specific encoder parameters with the set of encoder parameters of the pre-trained machine learning model; and

fine-tune the task specific machine learning model, initialized with the set of task specific encoder parameters, for the downstream geometric vision task.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2024
From: NAVER LABS CORPORATION
To: NAVER CORPORATION
Reel/Frame 068820/0495 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 31, 2023
From: BRÉGIER, ROMAIN; CABON, YOHANN; LUCAS, THOMAS; REVAUD, JÉRÔME; WEINZAEPFEL, PHILIPPE; CHIDLOVSKII, BORIS; LEROY, VINCENT; ANTSFELD, LEONID; CSURKA KHEDARI, GABRIELA
To: NAVER CORPORATION; NAVER LABS CORPORATION
Reel/Frame 064765/0514 →
Priority Claims (1)
EP 22306534 · Oct 11, 2022 · regional
Continuity (1)
Related Publication 20240135695A1 · Apr 25, 2024
References Cited (102)
US 20190228267A1 · Singh · 2019 [cited by examiner]
US 20190228313A1 · Lee · 2019 [cited by examiner]
US 20210060433A1 · Projansky · 2021 [cited by examiner]
US 20240037940A1 · Wu · 2024 [cited by examiner]
US 20240078726A1 · Weber · 2024 [cited by examiner]
CA 3090504A1 · 2021 [cited by examiner]
CN 112990218A · 2021 [cited by examiner]
CN 116012903A · 2023 [cited by examiner]
GB 2585197A · 2021 [cited by examiner]
KR 20210053202A · 2021 [cited by examiner]
U.S. Appl. No. 18/239,739, filed Aug. 29, 2023, Philippe Weinzaepfel. [cited by applicant]
Alaaeldin El-Nouby, Gautier Izacard, Hugo Touvron, Ivan Laptev, Herve Jegou, and Edouard Grave. Are Large-scale Datasets Necessary for Self-Supervised Pre-training? arXiv preprint arXiv:2112.10740, 2021. [cited by applicant]
Alexander Kapitanov, Andrey Makhlyarchuk, and Karina Kvanchiani. Hagrid—hand gesture recognition image, arXiv:2206.08219v1; Jun. 16, 2022. [cited by applicant]
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language. arXiv preprint arXiv:2202.03555, 2022. [cited by applicant]
Angjoo Kanazawa, Michael J Black, David W Jacobs, and Jitendra Malik. End-to-end recovery of human shape and pose. In CVPR, 2018. [cited by applicant]
Arora, Vaibhav and Antsfeld, Leonid and Chidlovskii, Boris and Csurka, Gabriela and Revaud Jerome. CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion. In NeurIPS, 2022. [cited by applicant]
Sara Atito, Muhammad Awais, and Josef Kittler. SiT: Self-supervised vision Transformer. arXiv preprint arXiv:2104.03602, 2021. [cited by applicant]
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6M: Large scale datasets and predictive methods for 3D human sensing in natural environments. IEEE Trans. PAMI, 2014. [cited by applicant]
Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer. Masked Feature Prediction for Self-Supervised Visual Pre-Training. In CVPR, 2022. [cited by applicant]
Christian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan Russell, Max Argus, and Thomas Brox. Freihand: A dataset for markerless capture of hand pose and shape from single rgb images. In ICCV, 2019. [cited by applicant]
Christian Zimmermann, Max Argus, and Thomas Brox. Contrastive representation learning for hand shape estimation. In Pattern Recognition, 2021. [cited by applicant]
Christoph Feichtenhofer, Haoqi Fan, Yanghao Li, and Kaiming He. Masked autoencoders as spatiotemporal learners. NeurIPS, 2022. [cited by applicant]
Dai et al., “Scannet: Richly-annotated 3d reconstructions of indoor scenes”, CVPR, 2017. [cited by applicant]
Dosovitskiy et al., An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, ICLR, 2021. [cited by applicant]
Dushyant Mehta, Helge Rhodin, Dan Casas, Pascal Fua, Oleksandr Sotnychenko, Weipeng Xu, and Christian Theobalt. Monocular 3d human pose estimation in the wild using improved CNN supervision. In 3DV, 2017. [cited by applicant]
European Search Report for European Application No. 22306534.3 dated May 12, 2023. [cited by applicant]
Fangzhou Hong, Liang Pan, Zhongang Cai, and Ziwei Liu. Versatile multi-modal pre-training for human-centric perception. arXiv preprint arXiv:2203.13815, 2022. [cited by applicant]
Fisher Yu, Ari Seff, Yinda Zhang, Shuran Song, Thomas Funkhouser, and Jianxiong Xiao. Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop. arXiv preprint arXiv:1506.03365, 2015. [cited by applicant]
Gregory Rogez, James S Supancic, and Deva Ramanan. Understanding everyday hands in action from rgb-d images. In ICCV, 2015. [cited by applicant]
Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge Belongie. The inaturalist species classification and detection dataset. In CVPR, 2018. [cited by applicant]
Gul Varol, Javier Romero, Xavier Martin, Naureen Mahmood, Michael J. Black, Ivan Laptev, and Cordelia Schmid. Learning from synthetic humans. In CVPR, 2017. [cited by applicant]
Gyeongsik Moon, Hongsuk Choi, and Kyoung Mu Lee. Accurate 3d hand pose estimation for whole-body 3d human mesh estimation. In CVPR, 2022. [cited by applicant]
Gyeongsik Moon, Shoou-I Yu, He Wen, Takaaki Shiratori, and Kyoung Mu Lee. Interhand2.6m: A dataset and baseline for 3d interacting hand pose estimation from a single rgb image. In ECCV, 2020. [cited by applicant]
Hanbyul Joo, Natalia Neverova, and Andrea Vedaldi. Exemplar fine-tuning for 3d human model fitting towards in-the wild 3d human pose estimation. In 3DV, 2021. [cited by applicant]
Helge Rhodin et al: “Unsupervised Geometry-Aware Representation for 3D Human Pose Estimation” Cornell University, ARXIV XP080867288, Apr. 3, 2018. [cited by applicant]
Jae Shin Yoon, Zhixuan Yu, Jaesik Park, and Hyun Soo Park. HUMBI: A large multiview dataset of human body expressions. In CVPR, 2021. [cited by applicant]
Hongsuk Choi, Gyeongsik Moon, and Kyoung Mu Lee. Pose2mesh: Graph convolutional network for 3D human pose and mesh recovery from a 2d human pose. In ECCV, 2020. [cited by applicant]
Zhixuan Yu, Jae Shin Yoon, In Kyu Lee, Prashanth Venkatesh, Jaesik Park, Jihun Yu, and Hyun Soo Park. HUMBI: A large multiview dataset of human body expressions. In CVPR, 2020. [cited by applicant]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pretraining of deep bidirectional transformers for language understanding. In NAACLHLT, 2019. [cited by applicant]
Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, Olivier Hénaff, Matthew M. Botvinick, Andrew Zisserman, Or… [cited by applicant]
Javier Romero, Dimitrios Tzionas, and Michael J Black. Embodied hands: Modeling and capturing hands and bodies together. arXiv preprint arXiv:2201.02610, 2022. [cited by applicant]
Jean-Bastien Grill, Florian Strub, Florent Altche, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Remi … [cited by applicant]
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR. IEEE, 2009. [cited by applicant]
Joao Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman. A short note on the kinetics-700 human action dataset. arXiv preprint arXiv:1907.06987, 2019. [cited by applicant]
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross Girshick. Masked Autoencoders are Scalable Vision Learners. In CVPR, 2021. [cited by applicant]
Ke Gong, Xiaodan Liang, Yicheng Li, Yimin Chen, Ming Yang, and Liang Lin. Instance-level human parsing via part grouping network. In ECCV, 2018. [cited by applicant]
Kevin Lin, Lijuan Wang, and Zicheng Liu. End-to-end human pose and mesh reconstruction with transformers. In CVPR, 2021. [cited by applicant]
Liang Zheng, Zhi Bie, Yifan Sun, Jingdong Wang, Chi Su, Shengjin Wang, and Qi Tian. Mars: A video benchmark for large-scale person re-identification. In ECCV, 2016. [cited by applicant]
Liu et al., “On the variance of the adaptive learning rate and beyond” ICLR, 2020. [cited by applicant]
Long Zhao, Yuxiao Wang, Jiaping Zhao, Liangzhe Yuan, Jennifer J Sun, Florian Schroff, Hartwig Adam, Xi Peng, Dimitris Metaxas, and Ting Liu. Learning view-disentangled human pose representation by contrastive cross-view… [cited by applicant]
Loshchilov and Hutter, “Decoupled Weight Decay Regularization”, ICLR, 2019. [cited by applicant]
Luca Schmidtke, Athanasios Vlontzos, Simon Ellershaw, Anna Lukens, Tomoki Arichi, and Bernhard Kainz. Unsupervised human pose estimation through transforming shape templates. In CVPR, 2021. [cited by applicant]
Mahmoud Assran, Mathilde Caron, Ishan Misra, Piotr Bojanowski, Florian Bordes, Pascal Vincent, Armand Joulin, Michael Rabbat, and Nicolas Ballas. Masked siamese networks for label-efficient learning. In ECCV, 2022. [cited by applicant]
Mahmoud Assran, Randall Balestriero, Quentin Duval, Florian Bordes, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, and Nicolas Ballas. The hidden uniform cluster prior in self-supervised learning. In ICL… [cited by applicant]
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra. Habitat: A Platform for Embodied Al Resea… [cited by applicant]
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative Pretraining From Pixels. In ICML, 2020. [cited by applicant]
Mathilde Caron, Hugo Touvron, Ishan Misra, Herve Jegou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging Properties in Self-Supervised Vision Transformers. In ICCV, 2021. [cited by applicant]
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. In NeurIPS, 2020. [cited by applicant]
Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. Deep clustering for unsupervised learning of visual features. In ECCV, 2018. [cited by applicant]
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. Smpl: A skinned multiperson linear model. ACM TOG, 2015. [cited by applicant]
Natalia Neverova, David Novotny, Marc Szafraniec, Vasil Khalidov, Patrick Labatut, and Andrea Vedaldi. Continuous surface embeddings. In NeurIPS, 2020. [cited by applicant]
Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black. Amass: Archive of motion capture as surface shapes. In ICCV. 2019. [cited by applicant]
Nikos Kolotouros, Georgios Pavlakos, Michael J Black, and Kostas Daniilidis. Learning to reconstruct 3D human pose and shape via model-fitting in the loop. In ICCV, 2019. [cited by applicant]
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognitio… [cited by applicant]
Paszke et al., “PyTorch: An Imperative Style, High-Performance Deep Learning Library”, NeurIPS, 2019. [cited by applicant]
Pathak et al., “Context Encoders: Feature Learning by Inpainting”, CVPR, 2016. [cited by applicant]
Philippe Weinzaepfel et al: “CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion” Cornell University, Oct. 19, 2022. [cited by applicant]
Ramakrishnan et al., “Habitat-matterport 3d dataset (HM3d): 1000 large-scale 3d environments for embodied AI”, NeurIPS datasets and benchmarks, 2021. [cited by applicant]
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced d… [cited by applicant]
Rasmus Rothe, Radu Timofte, and Luc Van Gool. Dex: Deep expectation of apparent age from a single image. In ICCVW, 2015. [cited by applicant]
Rene Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vision Transformers for Dense Prediction. In ICCV, 2021. [cited by applicant]
Riza Alp Guler, Natalia Neverova, and lasonas Kokkinos. Densepose: Dense human pose estimation in the wild. In CVPR, 2018. [cited by applicant]
Roberto Martin-Martin, Mihir Patel, Hamid Rezatofighi, Abhijeet Shenoi, JunYoung Gwak, Eric Frankel, Amir Sadeghian, and Silvio Savarese. Jrdb: A dataset and benchmark of egocentric robot visual perception of humans in … [cited by applicant]
Romain Bregier: “Deep Regression on Manifolds: A 3D Rotation Case Study” Cornell University, Oct. 12, 2021. [cited by applicant]
Roman Bachmann, David Mizrahi, Andrei Atanov, and Amir Zamir. MultiMAE: Multi-modal Multi-task Masked Autoencoders. In ECCV, 2022. [cited by applicant]
Schönemann, A generalized solution of the orthogonal Procrustes problem, Psychometrika, 1966. [cited by applicant]
Senthil Purushwalkam and Abhinav Gupta. Demystifying contrastive self-supervised learning: Invariances, augmentations and dataset biases. In NeurIPS, 2020. [cited by applicant]
Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. Honnotate: A method for 3d annotation of hand and object poses. In CVPR, 2020. [cited by applicant]
Shuhei Tsuchida, Satoru Fukayama, Masahiro Hamasaki, and Masataka Goto. Aist dance video database: Multi-genre, multi-dancer, and multi-camera database for dance information processing. In ISMIR, 2019. [cited by applicant]
Straub et al., “The Replica dataset: A digital replica of indoor spaces”, arXiv:1906.05797, 2019. [cited by applicant]
Szot et al., “Habitat 2.0: Training home assistants to rearrange their habitat”, NeurIPS, 2021. [cited by applicant]
Timo von Marcard, Roberto Henschel, Michael Black, Bodo Rosenhahn, and Gerard Pons-Moll. Recovering accurate 3d human pose in the wild using imus and a moving camera. In ECCV, 2018. [cited by applicant]
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In ICML, 2020. [cited by applicant]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural … [cited by applicant]
Tomas Jakab, Ankush Gupta, Hakan Bilen, and Andrea Vedaldi. Self-supervised learning of interpretable keypoints from unlabelled videos. In CVPR, 2020. [cited by applicant]
Tomas Jakab, Ankush Gupta, Hakan Bilen, and Andrea Vedaldi. Unsupervised learning of object landmarks through conditional image generation. NeurIPS, 2018. [cited by applicant]
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. [cited by applicant]
Umar Iqbal, Anton Milan, and Juergen Gall. Posetrack: Joint multi-person pose estimation and tracking. In CVPR. Apr. 7, 2017. [cited by applicant]
Umar Iqbal, Kevin Xie, Yunrong Guo, Jan Kautz, and Pavlo Molchanov. Kama: 3d keypoint aware body mesh articulation. In 2021 International Conference on 3D Vision (3DV), pp. 689-699. IEEE, 2021. [cited by applicant]
Umeyama, Least-squares estimation of transformation parameters between two point patterns, TPAMI, 1991. [cited by applicant]
Vaswani et al., Attention is all you need, NeurIPS, 2017. [cited by applicant]
Wandt Bastian et al: “CanonPose: Self-Supervised Monocular 3D Human Pose Estimation in the Wild” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 20, 2021. [cited by applicant]
Weiyu Zhang, Menglong Zhu, and Konstantinos G. Derpanis. From actemes to action: A strongly-supervised representation for detailed action understanding. In ICCV, 2013. [cited by applicant]
Chenxin Tao et al: “Siamese Image Modeling for Self-Supervised Vision Representation Learning” Cornell University, Jul. 5, 2022. [cited by applicant]
Ziwei Chen, Xiaofeng Wang, Wankou Yang, Qiang Li. Liftedcl: Lifting contrastive learning for human-centric perception. In ICLR, 2023. [cited by applicant]
Yao Feng, Vasileios Choutas, Timo Bolkart, Dimitrios Tzionas, and Michael J Black. Collaborative regression of expressive bodies using moderation. In 3DV, 2021. [cited by applicant]
Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao. Vitpose: Simple vision transformer baselines for human pose estimation. In NeurIPS, 2022. [cited by applicant]
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. NeurIPS, 2022. [cited by applicant]
Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai, and Han Hu. SimMIM: A Simple Framework for Masked Image Modeling. In CVPR, 2022. [cited by applicant]
Zhenguang Liu, Runyang Feng, Haoming Chen, ShuangWu, Yixing Gao, Yunjun Gao, and XiangWang. Temporal feature alignment and mutual information maximization for video based human pose estimation. In CVPR, 2022. [cited by applicant]
Zhirong Wu, Yuanjun Xiong, Stella X. Yu, and Dahua Lin. Unsupervised feature learning via non-parametric instance level discrimination. In CVPR, 2018. [cited by applicant]
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Large-scale celebfaces attributes (celeba) dataset. Retrieved Feb. 28, 2024. [cited by applicant]