IP Library › Granted Patent US 12,638,859
Granted Patent B2
US 12,638,859 · App. 19/362,617 · Granted May 26, 2026

Bipedal action model for humanoid robot

Inventors: Corey Lynch (San Jose, CA); Toki Migimatsu (San Jose, CA); Michael Ahn (San Jose, CA)
Assignee: FIGURE AI
G05D1/495B62D57/032G06F40/40G06V10/766G06V10/774G05D2101/15G05D2109/12G05D2111/52G05D2111/58
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,638,859
App. No.
19/362,617
Filed
Oct 20, 2025
Granted
May 26, 2026
Kind
B2
Art Unit
3657
USPC
700/245
Abstract

The present disclosure provides a humanoid robot system comprising a mechanical structure including a torso, two arms, and two legs providing at least 30 degrees of freedom, actuators coupled to the degrees of freedom, a sensor suite comprising at least one camera and proprioceptive sensors including joint encoders and an inertial measurement unit, a computing system comprising at least one processor and memory storing instructions which, when executed, implement a hierarchical bipedal action model including a Beta model configured to receive multimodal input data and generate a token sequence indicative of task intent and environmental state, and an Alpha model configured to condition on the token sequence and current robot pose data to output continuous action chunks comprising sequences of future target joint states over a finite horizon, and a low-level controller configured to convert the continuous action chunks into actuator control signals for execution.

Claims (82)

1 . A method for controlling a humanoid robot using a hierarchical bipedal action model (BAM), comprising:

obtaining a hierarchical BAM that is generated by:

collecting training data comprising: (i) obtaining video data from a 3 rd party database, and (ii) obtaining real-world robot demonstrations;

preprocessing the training data to form distinct segments that have natural language descriptions;

training, using the preprocessed training data, both a Beta model and an Alpha model; and

deploying the trained hierarchical BAM on a humanoid robot; and

controlling the humanoid robot to perform an autonomous task using the trained hierarchical BAM.

2 . The method of claim 1 , wherein deploying the hierarchical BAM comprises loading the Alpha model on a GPU that is positioned within the humanoid robot.

3 . The method of claim 1 , wherein at least one of the Alpha model and the Beta model is a diffusion model.

4 . The method of claim 1 , wherein the Beta model operates at a frequency that is less than 50 Hz.

5 . The method of claim 4 , wherein the Beta model is trained using a cross-entropy loss function and has more than 500 million parameters.

6 . The method of claim 1 , wherein the Alpha model operates at a frequency that is greater than 50 Hz.

7 . The method of claim 6 , wherein the Alpha model is trained using a regression-based loss function and has fewer than 500 million parameters.

8 . The method of claim 1 , wherein the Alpha and Beta models are trained end-to-end, and wherein said training includes allowing error gradients from the output of the Alpha model to be backpropagated through the Beta model.

9 . The method of claim 1 , further comprising the step of executing a safety verification that rejects or truncates outputs from the Alpha model that violate a predetermined constraint.

10 . The method of claim 1 , wherein the step of collecting training data includes capturing first-person video data synchronized with head and hand positions using a virtual reality or augmented reality headset.

11 . The method of claim 1 , wherein the natural language descriptions are generated using an AI model.

12 . The method of claim 1 , wherein obtaining a hierarchical BAM includes the step of obtaining at least one pre-trained model having a set of parameters; and

wherein training the hierarchical BAM includes using supervised learning to modify the set of parameters based in part on the preprocessed training data.

13 . The method of claim 2 , wherein deploying the trained hierarchical BAM includes loading the Beta model on a GPU that is not positioned within the humanoid robot.

14 . A method for controlling a robot using a hierarchical action model, comprising:

obtaining a hierarchical action model by:

collecting training data comprising obtaining: (i) video data from a 3 rd party database, (ii) first-person video data, and (iii) data from a robot demonstration;

training, using the training data, a hierarchical action model that includes an Alpha model and a Beta model, and wherein the Alpha model has a first number of parameters and the Beta model has a second number of parameters that is larger than the first number of parameters; and

deploying the trained hierarchical action model on a robot; and

controlling the robot to perform an autonomous task using the trained hierarchical action model.

15 . The method of claim 14 , wherein controlling the robot includes the step of using a retrieval-augmented generation technique to obtain additional real-time knowledge from external sources.

16 . The method of claim 14 , wherein at least one of the Alpha and Beta models is a diffusion model.

17 . The method of claim 16 , wherein the Alpha model is trained using reinforcement learning.

18 . The method of claim 17 , wherein the Beta model is trained using unsupervised learning.

19 . The method of claim 14 , wherein controlling the robot includes generating continuous actions comprising a first set of X, Y, Z floating-point coordinates and a first set of X, Y, Z floating-point rotations.

20 . The method of claim 14 , wherein deploying the hierarchical action model comprises loading the Alpha model on a GPU that is positioned within the robot.

21 . The method of claim 20 , wherein deploying the trained hierarchical action model includes loading the Beta model on a GPU that is not positioned within the robot.

22 . The method of claim 14 , wherein the first-person video data is captured using a virtual reality or augmented reality headset.

23 . A method for controlling a robot using an action model, comprising:

obtaining an action model by:

collecting training data comprising obtaining: (i) video data from a 3 rd party database, (ii) first-person video data, and (iii) data from a robot demonstration;

training an action model using the training data, and wherein the action model includes more than 1 billion parameters; and

deploying the trained action model on a robot; and

using the trained action model to generate continuous meter-actions.

24 . The method of claim 23 , wherein the first-person video data is captured using a virtual reality or augmented reality headset.

25 . The method of claim 23 , wherein the continuous actions include a first set of X, Y, Z floating-point coordinates and a first set of X, Y, Z floating-point rotations.

26 . The method of claim 23 , wherein the action model is trained using reinforcement learning and the deployment of the action model includes loading the model on a GPU located within the robot.

27 . The method of claim 23 , further comprising using a retrieval-augmented generation technique to obtain additional real-time knowledge from sources that are external to the robot.

28 . The method of claim 23 , further comprising the step of executing a safety verification that rejects or truncates outputs from the action model that violate a predetermined constraint.

29 . The method of claim 1 , wherein controlling the humanoid robot includes the step of using a retrieval-augmented generation technique to obtain additional real-time knowledge from external sources.

30 . The method of claim 1 , wherein the Beta model is trained using unsupervised learning and the Alpha model is trained using reinforcement learning.

31 . The method of claim 1 , wherein the hierarchical BAM further comprises a Gamma model that includes more than 1 trillion parameters, and wherein the Gamma model is not deployed on the humanoid robot.

32 . The method of claim 1 , wherein controlling the humanoid robot includes using the trained hierarchical BAM to generate continuous actions.

33 . The method of claim 1 , wherein controlling the humanoid robot includes using the trained hierarchical BAM to generate X, Y, Z floating-point coordinates and X, Y, Z floating-point rotations.

34 . The method of claim 1 , wherein the humanoid robot includes at least thirty electric actuators, and wherein at least ten of the thirty electric actuators include a strain wave gearbox.

35 . The method of claim 1 , wherein a majority of the humanoid robot is covered with a deformable textile that includes polymers.

36 . The method of claim 35 , wherein the humanoid robot includes an illumination assembly configured to communicate the humanoid robot's status.

37 . The method of claim 14 , wherein the Beta model operates at a frequency that is less than 50 Hz and has more than 500 million parameters.

38 . The method of claim 37 , wherein the Alpha model operates at a frequency that is greater than 50 Hz and has fewer than 500 million parameters.

39 . The method of claim 38 , wherein the Alpha and Beta models are trained end-to-end, and wherein said training includes allowing error gradients from the output of the Alpha model to be backpropagated through the Beta model.

40 . The method of claim 39 , wherein the Beta model is trained using a cross-entropy loss function and the Alpha model is trained using a regression-based loss function.

41 . The method of claim 14 , wherein the robot includes at least thirty electric actuators, and wherein at least one of the thirty electric actuators includes a strain wave gearbox.

42 . The method of claim 41 , wherein a majority of the robot is covered with a deformable textile that includes polymers.

43 . The method of claim 14 , wherein the hierarchical action model further comprises a Gamma model that includes more than 1 trillion parameters, and wherein the Gamma model is locally hosted relative to the robot.

44 . The method of claim 23 , wherein the action model includes an Alpha model that operates at a frequency that is greater than 50 Hz and has fewer than 500 million parameters.

45 . The method of claim 44 , wherein the action model includes a Beta model that operates at a frequency that is less than 50 Hz and has more than 500 million parameters.

46 . The method of claim 45 , wherein the Beta model is trained using a cross-entropy loss function and the Alpha model is trained using a regression-based loss function.

47 . The method of claim 23 , wherein the action model includes an Alpha model and a Beta model, and said Alpha and Beta models are trained end-to-end, and wherein said training includes allowing error gradients from the output of the Alpha model to be backpropagated through the Beta model.

48 . The method of claim 23 , wherein the robot includes at least thirty electric actuators, and wherein at least one of the thirty electric actuators includes a strain wave gearbox.

49 . The method of claim 48 , wherein a majority of the robot is covered with a deformable textile that includes polymers.

50 . The method of claim 23 , wherein the action model further comprises a Gamma model that includes more than 1 trillion parameters, and wherein the Gamma model is locally hosted relative to the robot.

51 . A method for controlling a robot using a model, comprising:

obtaining a model by:

collecting training data comprising obtaining: (i) video data from a 3 rd party database, and (ii) data from a robot demonstration;

training a model using the training data, and wherein the trained model includes more than 1 billion parameters; and

deploying the trained model on a robot; and

using the trained model to generate X, Y, Z floating-point coordinates and X, Y, Z floating-point rotations.

52 . The method of claim 51 , further comprising using a retrieval-augmented generation technique to obtain additional real-time knowledge from sources that are external to the robot.

53 . The method of claim 51 , wherein the model includes: (i) a Beta model that operates at a frequency that is less than 50 Hz, and (ii) an Alpha model that operates at a frequency that is more than 50 Hz.

54 . The method of claim 53 , wherein the Alpha and Beta models are trained end-to-end, and wherein said training includes allowing error gradients from the output of the Alpha model to be backpropagated through the Beta model.

55 . The method of claim 51 , wherein the model includes: (i) an Alpha model that has fewer than 500 million parameters, and (ii) a Beta model that has more than 500 million parameters.

56 . The method of claim 55 , wherein the Alpha and Beta models are trained end-to-end, and wherein said training includes allowing error gradients from the output of the Alpha model to be backpropagated through the Beta model.

57 . The method of claim 56 , wherein the model includes: (i) an Alpha model that is trained using a regression-based loss function, and (ii) a Beta model that is trained using a cross-entropy loss function.

58 . The method of claim 51 , wherein the robot includes at least thirty electric actuators, and wherein at least one of the thirty electric actuators includes a strain wave gearbox.

59 . The method of claim 58 , wherein a majority of the robot is covered with a deformable textile that includes polymers.

60 . The method of claim 51 , wherein the model further comprises a Gamma model that includes more than 1 trillion parameters, and wherein the Gamma model is locally hosted relative to the robot.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2026
From: LYNCH, COREY; MIGIMATSU, TOKI; AHN, MICHAEL
To: FIGURE AI INC.
Reel/Frame 073745/0023 →
Continuity (10)
Continuation In Part 19325486 · Sep 10, 2025
Provisional Application 63883647 · Sep 17, 2025
Provisional Application 63819533 · Jun 6, 2025
Provisional Application 63776429 · Mar 24, 2025
Provisional Application 63760617 · Feb 19, 2025
Provisional Application 63750617 · Jan 28, 2025
Provisional Application 63725279 · Nov 26, 2024
Provisional Application 63722057 · Nov 18, 2024
Provisional Application 63715270 · Nov 1, 2024
Related Publication 20260126805A1 · May 7, 2026
References Cited (327)
US 5673367A · Buckley · 1997 [cited by applicant]
US 7024276B2 · Ito · 2006 [cited by applicant]
US 7072741B2 · Nagashima · 2006 [cited by applicant]
US 7308336B2 · Takenaka · 2007 [cited by applicant]
US 7319918B2 · Takenaka · 2008 [cited by applicant]
US 7386364B2 · Mikami · 2008 [cited by applicant]
US 7664569B2 · Shimizu · 2010 [cited by applicant]
US D641808S · Matsuda · 2011 [cited by applicant]
US 8224652B2 · Wang · 2012 [cited by applicant]
US D677743S · Koshiishi · 2013 [cited by applicant]
US D687908S · Hoang · 2013 [cited by applicant]
US D689566S · Wong · 2013 [cited by applicant]
US 8942849B2 · Maisonnier · 2015 [cited by applicant]
US 9205560B1 · Edsinger · 2015 [cited by applicant]
US 9302393B1 · Rosen · 2016 [cited by examiner]
US 9494415B2 · Sweetser · 2016 [cited by applicant]
US 9569976B2 · Krauss · 2017 [cited by applicant]
US D794692S · Haranaka · 2017 [cited by applicant]
US D795320S · Liu · 2017 [cited by applicant]
US D795321S · Liu · 2017 [cited by applicant]
US 9789607B1 · Whitman · 2017 [cited by applicant]
US 9842585B2 · Huang · 2017 [cited by applicant]
US 9868210B1 · Whitman · 2018 [cited by applicant]
US D835214S · Xiong · 2018 [cited by applicant]
US D838759S · Kowalski · 2019 [cited by applicant]
US D841708S · Koshiishi · 2019 [cited by applicant]
US 10203209B2 · Roumeliotis · 2019 [cited by applicant]
US 10349245B2 · Tokuchi · 2019 [cited by applicant]
US D866684S · Früh · 2019 [cited by applicant]
US D868866S · Gable · 2019 [cited by applicant]
US D872152S · Xiong · 2020 [cited by applicant]
US D873320S · Clerc · 2020 [cited by applicant]
US 10545497B1 · Cui · 2020 [cited by applicant]
US 10571896B2 · Benaim · 2020 [cited by applicant]
US D885451S · Chen · 2020 [cited by applicant]
US D888120S · Hurst · 2020 [cited by applicant]
US D892886S · Klassen · 2020 [cited by applicant]
US D892887S · Klassen · 2020 [cited by applicant]
US D893573S · Yan · 2020 [cited by applicant]
US D898789S · Nazarikhorram · 2020 [cited by applicant]
US D911459S · Xiong · 2021 [cited by applicant]
US 10946528B2 · Gupta · 2021 [cited by applicant]
US 10960539B1 · Kalakrishnan · 2021 [cited by applicant]
US D932531S · Xu · 2021 [cited by applicant]
US 11188821B1 · Kalakrishnan · 2021 [cited by applicant]
US 11247738B2 · Lavalley · 2022 [cited by applicant]
US 11416003B2 · Whitman · 2022 [cited by applicant]
US 11498223B2 · Williams · 2022 [cited by applicant]
US 11546504B2 · Kim · 2023 [cited by applicant]
US 11600010B2 · Doutre · 2023 [cited by applicant]
US D985643S · Li · 2023 [cited by applicant]
US 11645444B2 · Scheutz · 2023 [cited by applicant]
US 11736677B2 · Grunnet-Jepsen · 2023 [cited by applicant]
US 11833680B2 · Deits · 2023 [cited by applicant]
US 11851120B2 · Fay · 2023 [cited by applicant]
US 11999423B2 · Whitman · 2024 [cited by applicant]
US 12054208B2 · Swilling · 2024 [cited by applicant]
US 12070863B2 · Whitman · 2024 [cited by applicant]
US 12077229B2 · Whitman · 2024 [cited by applicant]
US 12097626B2 · Ikeda · 2024 [cited by applicant]
US D1051193S · Mahoor · 2024 [cited by applicant]
US 12134181B2 · Klingensmith · 2024 [cited by applicant]
US 12172537B2 · Gonano · 2024 [cited by applicant]
US 12205214B2 · Starke · 2025 [cited by applicant]
US 12214497B2 · Whitman · 2025 [cited by applicant]
US 12235652B2 · Whitman · 2025 [cited by applicant]
US 12240117B2 · Chebotar · 2025 [cited by applicant]
US D1069875S · Belon · 2025 [cited by applicant]
US 12290940B1 · Abate · 2025 [cited by applicant]
US D1082881S · Wang · 2025 [cited by applicant]
US 12365094B2 · Mccall · 2025 [cited by applicant]
US 12403611B2 · Mccall · 2025 [cited by applicant]
US 12466075B2 · Stathis · 2025 [cited by examiner]
US 12482243B1 · Cass · 2025 [cited by examiner]
US 12515706B2 · Li · 2026 [cited by examiner]
US 20020103576A1 · Takamura · 2002 [cited by examiner]
US 20030060930A1 · Fujita · 2003 [cited by examiner]
US 20040015265A1 · Asano · 2004 [cited by examiner]
US 20040019485A1 · Kobayashi · 2004 [cited by examiner]
US 20040039483A1 · Kemp · 2004 [cited by examiner]
US 20040230340A1 · Fukuchi · 2004 [cited by examiner]
US 20060217838A1 · Sugino · 2006 [cited by applicant]
US 20090059033A1 · Shimada · 2009 [cited by applicant]
US 20100280662A1 · Abdallah · 2010 [cited by applicant]
US 20110058800A1 · Lee · 2011 [cited by applicant]
US 20110071671A1 · Ihrke · 2011 [cited by applicant]
US 20120072215A1 · Yu · 2012 [cited by applicant]
US 20120078419A1 · Kim · 2012 [cited by applicant]
US 20120155775A1 · Ahn · 2012 [cited by examiner]
US 20120215539A1 · Juneja · 2012 [cited by applicant]
US 20120310412A1 · Seo · 2012 [cited by applicant]
US 20130175816A1 · Kawasaki · 2013 [cited by applicant]
US 20130345863A1 · Linder · 2013 [cited by applicant]
US 20140039675A1 · Ead · 2014 [cited by applicant]
US 20150192399A1 · Raab · 2015 [cited by applicant]
US 20150290795A1 · Oleynik · 2015 [cited by applicant]
US 20160008988A1 · Kennedy · 2016 [cited by applicant]
US 20160064263A1 · Hosek · 2016 [cited by applicant]
US 20170028551A1 · Hemken · 2017 [cited by examiner]
US 20170028563A1 · Hemken · 2017 [cited by applicant]
US 20170080582A1 · Mugnier · 2017 [cited by applicant]
US 20170125008A1 · Maisonnier · 2017 [cited by applicant]
US 20170326736A1 · Nagatsuka · 2017 [cited by applicant]
US 20180136912A1 · Venkataramani · 2018 [cited by applicant]
US 20180182260A1 · Ciniello · 2018 [cited by applicant]
US 20180232201A1 · Holtmann · 2018 [cited by applicant]
US 20180357552A1 · Campos · 2018 [cited by applicant]
US 20190005374A1 · Shankar · 2019 [cited by examiner]
US 20190079924A1 · Sugiura · 2019 [cited by applicant]
US 20190100263A1 · Amino · 2019 [cited by applicant]
US 20190163985A1 · Wang · 2019 [cited by examiner]
US 20190329413A1 · Johnson · 2019 [cited by applicant]
US 20190371307A1 · Zhao · 2019 [cited by applicant]
US 20200009739A1 · Moon · 2020 [cited by applicant]
US 20200086479A1 · Messier · 2020 [cited by applicant]
US 20200117214A1 · Jonak · 2020 [cited by examiner]
US 20200117336A1 · Mani · 2020 [cited by examiner]
US 20200368616A1 · Delamont · 2020 [cited by examiner]
US 20210331313A1 · Klingensmith · 2021 [cited by examiner]
US 20210390420A1 · Barnett · 2021 [cited by examiner]
US 20220226996A1 · Ishizuka · 2022 [cited by applicant]
US 20220294062A1 · Kamon · 2022 [cited by applicant]
US 20220314448A1 · Wales · 2022 [cited by examiner]
US 20220388174A1 · Stathis · 2022 [cited by applicant]
US 20220390952A1 · Yu · 2022 [cited by applicant]
US 20220410380A1 · Lu · 2022 [cited by applicant]
US 20230048725A1 · Barbour · 2023 [cited by applicant]
US 20230126906A1 · Kranski · 2023 [cited by examiner]
US 20230129990A1 · Liu · 2023 [cited by examiner]
US 20230143315A1 · Whitman · 2023 [cited by applicant]
US 20230168921A1 · Kim · 2023 [cited by examiner]
US 20230173683A1 · Gomez · 2023 [cited by applicant]
US 20230182296A1 · Sermanet · 2023 [cited by applicant]
US 20230222454A1 · Cella · 2023 [cited by examiner]
US 20230306302A1 · Mori · 2023 [cited by examiner]
US 20230347514A1 · Xiao · 2023 [cited by applicant]
US 20230390948A1 · Hsu · 2023 [cited by applicant]
US 20240091964A1 · Smith · 2024 [cited by applicant]
US 20240095077A1 · Singh · 2024 [cited by examiner]
US 20240181637A1 · Gillett · 2024 [cited by applicant]
US 20240193848A1 · Wang · 2024 [cited by examiner]
US 20240199068A1 · Pavone · 2024 [cited by examiner]
US 20240217104A1 · Neville · 2024 [cited by applicant]
US 20240228191A1 · Kumar · 2024 [cited by applicant]
US 20240289973A1 · Liu · 2024 [cited by examiner]
US 20240416963A1 · Li · 2024 [cited by examiner]
US 20250050507A1 · Camasmie · 2025 [cited by applicant]
US 20250147517A1 · Swilling · 2025 [cited by applicant]
US 20250196326A1 · Katz · 2025 [cited by applicant]
US 20250196327A1 · Geating · 2025 [cited by applicant]
US 20250218160A1 · Wang · 2025 [cited by examiner]
CN 102357889 · 2012 [cited by applicant]
CN 209615545 · 2019 [cited by applicant]
CN 210998685 · 2020 [cited by applicant]
CN 115649316 · 2023 [cited by applicant]
CN 218802294 · 2023 [cited by applicant]
CN 117301022 · 2023 [cited by applicant]
CN 117462367 · 2024 [cited by applicant]
KR 3020240036125 · 2024 [cited by applicant]
SU 1734994 · 1992 [cited by applicant]
WO 2023107501 · 2023 [cited by applicant]
WO 2023110778 · 2023 [cited by applicant]
WO 2023246994 · 2023 [cited by applicant]
WO 2023246995 · 2023 [cited by applicant]
WO 2024058844 · 2024 [cited by applicant]
WO 2024085904 · 2024 [cited by applicant]
WO 2024112350 · 2024 [cited by applicant]
WO 2024112351 · 2024 [cited by applicant]
WO 2024123766 · 2024 [cited by applicant]
WO 2024163992 · 2024 [cited by applicant]
WO D243074010 · 2024 [cited by applicant]
WO 2025019583 · 2025 [cited by applicant]
WO 2025042802 · 2025 [cited by applicant]
WO 2025072321 · 2025 [cited by applicant]
Jeung et al., Realization of human neck motion with novel robotic mechanism, 2016, IEEE, p. 482-486 (Year: 2016). [cited by applicant]
Barker et al., Natural head movement for HRI with a muscular-skeletal head and neck robot, 2017, IEEE, p. 587-592 (Year: 2017). [cited by applicant]
Gao et al., Development of a low motion-noise humanoid neck: Statics analysis and experimental validation, 2010, IEEE, p. 1203-1208 (Year: 2010). [cited by applicant]
International Search Report for PCT/US2025/012544. [cited by applicant]
International Search Report for PCT/US2025/010425. [cited by applicant]
International Search Report for PCT/US2025/011450. [cited by applicant]
Keselman et al., “Intel RealSense stereoscopic depth cameras,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pp. 1-10, 2017. [cited by applicant]
Brown et al., “Language Models are Few-Shot Learners,” arXiv:2005.14165v4 (Jul. 22, 2020). [cited by applicant]
Bjorck, Johan, et al. “Gr00t n1: An open foundation model for generalist humanoid robots.” arXiv preprint arXiv:2503.14734 (Mar. 18, 2025). [cited by applicant]
Chang et al., “A Survey on Evaluation of Large Language Models,” ACM Trans. Intell. Syst. Technol., vol. 15, No. 3, Article 39. (Mar. 2024). [cited by applicant]
Chen et al., “InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks,” https://github.com/OpenGVLab/InternVL (2024). [cited by applicant]
Gia et al., “Densely Connected Feature Pyramid Network for Image Segmentation,” 2020 8th International Conference on Digital Home (ICDH). [cited by applicant]
Jin et al., “Unified Language-Vision Pretraining in LLM With Dynamic Discrete Visual Tokenization,” arXiv:2309.04669v3 [cs.CV] Mar. 22, 2024. [cited by applicant]
Kim et al., “Parallel Feature Pyramid Network for Object Detection,” ECCV 2018. [cited by applicant]
Kim et al., “Giving Robots a Hand: Learning Generalizable Manipulation with Eye-in-Hand Human Video Demonstrations,” arXiv:2307.05959v1 (Jul. 12, 2023). [cited by applicant]
Kirillov et al., “Panoptic Feature Pyramid Networks,” IEEE Xplore (2018). [cited by applicant]
Lee et al., “Learning Robot Activities from First-Person Human Videos Using Convolutional Future Regression,” IEEE Xplore (2017). [cited by applicant]
Li et al., “Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-Training Paradigm,” arXiv:2110.05208v2 (Mar. 14, 2022). [cited by applicant]
Li et al., BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation, Proceedings of the 39 th International Conference on Machine Learning, Baltimore, Maryland, USA, PMLR … [cited by applicant]
Lin et al., “Feature Pyramid Networks for Object Detection,” IEEE Xplore (2016). [cited by applicant]
Lin et al., “Feature Pyramid Networks for Object Detection,” IEEE Xplore (2024). [cited by applicant]
Liu et al., “A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends,” Journal of Latex Class Files, vol. 14, No. (Aug. 8, 2021). [cited by applicant]
Liu et al., “Imitation from Observation: Learning to Imitate Behaviors from Raw Video via Context Translation,” arXiv:1707.03374v2 (Jun. 18, 2018). [cited by applicant]
Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” arXiv:1907.11692v1 (Jul. 26, 2019). [cited by applicant]
Liu et al., “Improved Baselines with Visual Instruction Tuning,” IEEE Xplore (2024). [cited by applicant]
Liu et al., “Visual Instruction Tuning,” 37th Conference on Neural Information Processing Systems (NeurIPS 2023). [cited by applicant]
Lin et al., “VILA: On Pre-training for Visual Language Models,” https://github.com/Efficient-Large-Model/VILA (2024). [cited by applicant]
Mandi et al., “Towards More Generalizable One-shot Visual Imitation Learning,” bencharXiv: 2110.13423v2 [cs.RO] Feb. 8, 2022. [cited by applicant]
Maniparambil et al., “Do Vision and Language Encoders Represent the World Similarly?” IEEE Xplore (2024). [cited by applicant]
Nvidia “Object Detection Synthetic DataGeneration,” Nvidia Corporation, (Nov. 9, 2024). [cited by applicant]
Peng et al., “DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills,” ACM Trans. Graph., vol. 37, No. 4, Article 143. Publication date: Aug. 2018. [cited by applicant]
Radford et al., “Improving Language Understanding by Generative Pre-Training,” 2018. [cited by applicant]
Radford et al., “Language Models are Unsupervised Multitask Learners,” 2019. [cited by applicant]
Radford et al., “Learning Transferable Visual Models From Natural Language Supervision,” Proceedings of the 38 th International Conference on Machine Learning, PMLR 139, 2021. [cited by applicant]
Raffel et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” Journal of Machine Learning Research 21 (2020) 1-67. [cited by applicant]
Ramachandruni et al., “Attentive Task-Net: Self Supervised Task-Attention Network for Imitation Learning using Video Demonstration,” 2020 IEEE International Conference on Robotics and Automation (ICRA) May 31—Paris, Fra… [cited by applicant]
Rombach et al., “High-Resolution Image Synthesis with Latent Diffusion Models,” IEEE Xplore (2022). [cited by applicant]
Sanh et al., “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter,” arXiv:1910.01108v4 (Mar. 1, 2020). [cited by applicant]
Schulman et al., “Proximal Policy Optimization Algorithms,” arXiv:1707.06347v2 (Aug. 28, 2017). [cited by applicant]
Sharma et al., “Third-Person Visual Imitation Learning via Decoupled Hierarchical Controller,” 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada. [cited by applicant]
Sieb et al., “Graph-Structured Visual Imitation,” 3rd Conference on Robot Learning (CoRL 2019), Osaka, Japan. (May 12, 2020). [cited by applicant]
Smith et al., “AVID: Learning Multi-Stage Tasks via Pixel-Level Translation of Human Videos,”arXiv:1912.04443v3 (Jun. 21, 2020). [cited by applicant]
Stadie et al., “Third-Person Imitation Learning,” arXiv:1703.01703v2 (Sep. 22, 2019). [cited by applicant]
Touvron et al., “Llama 2: Open Foundation and Fine-Tuned Chat Models,” arXiv:2307.09288v2 (Jul. 19, 2023). [cited by applicant]
Vaswani et al., “Attention Is All You Need,” 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA. (Jun. 12, 2017). [cited by applicant]
Wang et al., “Structbert: Incorporating Language Structures Into Pre-Training for Deep Language Understanding,,” arXiv:1908.04577v3 (Sep. 27, 2019). [cited by applicant]
Xiong et al., “Learning by Watching: Physical Imitation of Manipulation Skills from Human Videos,” arXiv:2101.07241v2 (Nov. 14, 2021). [cited by applicant]
Yao et al., “Filip: Fine-Grained Interactive Language-Image Pre-Training,” arXiv:2111.07783v1 (Nov. 9, 2021). [cited by applicant]
Yin et al., “A Survey on Multimodal Large Language Models,” arXiv:2306.13549v2 (Apr. 1, 2024). [cited by applicant]
Sun et al., “Learning by Watching via Keypoint Extraction and Imitation Learning,” Machines 2022, 10, 1049. https://doi.org/10.3390/machines10111049 (Nov. 9, 2022). [cited by applicant]
Hebi Robotics, “T-Series Actuator,” Jan. 29, 2024. [cited by applicant]
Available online at https://youtu.be/YsdnsNjvwKo?si=bu2dXk8mQaL86C2M, at least as early as Aug. 23, 2023. [cited by applicant]
Available online at https://www.youtube.com/watch?v=q8ldbodRG14, at least as early as May 22, 2019. [cited by applicant]
Available online at https://youtu.be/GtPs_ygfaEA?si=7lv6MEFvFoaacKfa, at least as early as Aug. 15, 2023. [cited by applicant]
Available online at https://www.youtube.com/watch?v=G6JE7mNYz2A, at least as early as Oct. 17, 2024. [cited by applicant]
Available online at https://www.youtube.com/watch?v=FuNFr7V7KFQ, at least as early as Aug. 19, 2024. [cited by applicant]
Available online at https://www.youtube.com/watch?v=GzX1qOIO1bE, at least as early as May 13, 2024. [cited by applicant]
Available online at https://youtu.be/_MBd_XfXy9M?si=PbEHUJpRUFqaxS3J, at least as early as Jun. 26, 2023. [cited by applicant]
Available online at https://youtu.be/SHPxcRBIXNO?si=VbJqbK7jzUqtZGmn, at least as early as Sep. 26, 2023. [cited by applicant]
Available online at https://youtu.be/BvFxD-8AhJA?si=Vx1F4a76tbQDUX48, at least as early as Nov. 16, 2023. [cited by applicant]
Available online at https://www.youtube.com/watch?v=jWTWWuzB6Cg, at least as early as Aug. 27, 2024. [cited by applicant]
Available online at https://www.youtube.com/watch?v=B-ebMigAHzQ, at least as early as Sep. 30, 2024. [cited by applicant]
Available online at https://youtu.be/XiQkeWOFwmk?si=1qOPC8gXgmmGvXRT, at least as early as May 16, 2023. [cited by applicant]
Available online at https://youtu.be/cpraXaw7dyc?si=JvPaT6eMA18psrmU, at least as early as Dec. 13, 2023. [cited by applicant]
Available online at https://www.youtube.com/watch?v=DrNcXgoFv20, at least as early as Oct. 18, 2024. [cited by applicant]
Available online at https://youtu.be/BNSZ8Fwcd20?si=_YnVgjYblVuhASk1, at least as early as Oct. 27, 2023. [cited by applicant]
Available online at https://youtu.be/SS3Ga2HQQ0s?si=Dwr3sJuCsOeUoSLj, at least as early as Nov. 20, 2023. [cited by applicant]
Available online at https://www.youtube.com/watch?v=iWC8rSjDywU, at least as early as Oct. 18, 2024. [cited by applicant]
Available online at https://youtu.be/sihlDeJ4Hmk?si=fJsKpvRFPzFejmS6, at least as early as Dec. 27, 2023. [cited by applicant]
Available online at https://www.youtube.com/watch?v=zkBnFPBV3f0, at least as early as Jul. 11, 2013. [cited by applicant]
Available online at https://www.youtube.com/watch?v=oXBYZxa25vc&t=1s, at least as early as Apr. 3, 2013. [cited by applicant]
Available online at https://www.youtube.com/watch?v=LBeml9AmTT4, at least as early as Apr. 7, 2015. [cited by applicant]
Available online at https://www.youtube.com/watch?v=IE-YBaYjbqY, at least as early as Dec. 10, 2013. [cited by applicant]
Available online at https://www.youtube.com/watch?v=y-j4dixQQml&t=222s, at least as early as May 22, 2012. [cited by applicant]
Available online at https://www.youtube.com/watch?v=Bmglbk_Op64&t=1s, at least as early as Nov. 10, 2011. [cited by applicant]
Available online at https://www.youtube.com/watch?v=20GHG-R9eFI, at least as early as Mar. 6, 2023. [cited by applicant]
Available online at https://www.youtube.com/watch?v=bUrLuUxv9gE, at least as early as Aug. 30, 2024. [cited by applicant]
Available online at https://www.youtube.com/watch?v=- 9EM5_VFIt8, at least as early as Apr. 16, 2024. [cited by applicant]
Available online at https://www.youtube.com/watch?v=29ECwExc -_M&t=2s, at least as early as Apr. 17, 2024. [cited by applicant]
Available online at https://www.youtube.com/watch?v=67CUudkjEG4, at least as early as Oct. 26, 2009. [cited by applicant]
Available online at https://www.youtube.com/watch?v=yBmatGQ0giY&t=1s, at least as early as Aug. 11, 2022. [cited by applicant]
Available online at https://www.youtube.com/watch?v=bdVrWxjK2vo, at least as early as Sep. 17, 2024. [cited by applicant]
Available online at https://www.youtube.com/watch?v=qw2y0kceAv0, at least as early as Oct. 15, 2024. [cited by applicant]
Available online at https://www.youtube.com/watch?v=B_I2k7MZEKg, at least as early as Jun. 30, 2024. [cited by applicant]
Available online at https://www.youtube.com/watch?v=CbA9wA9etGA, at least as early as Sep. 19, 2024. [cited by applicant]
Available online at https://www.youtube.com/watch?v=zLhA-RWBBYU, at least as early as Jul. 5, 2024. [cited by applicant]
Available online at https://www.youtube.com/watch?v=_mQJw8VhZ7w&t=111s, as least as early as Oct. 5, 2022. [cited by applicant]
Available online at https://www.youtube.com/watch?v=1fC7b2LjVW4, at least as early as Jul. 12, 2016. [cited by applicant]
Available online at https://www.youtube.com/watch?v=UPOLcE1vwA0, at least as early as Apr. 28, 2016. [cited by applicant]
Available online at https://www.youtube.com/watch?v=UBbk18oZbTc, at least as early as Oct. 14, 2024. [cited by applicant]
Available online at https://www.youtube.com/watch?v=UHe1zSQwep0, at least as early as Oct. 14, 2024. [cited by applicant]
Available online at https://www.youtube.com/watch?v=MCbGeC-kuBM, at least as early as Aug. 5, 2024. [cited by applicant]
Available online at https://www.youtube.com/watch?v=ujdK3yd2gHY, at least as early as Jul. 2, 2024. [cited by applicant]
Available online at https://www.youtube.com/watch?v=-HizP4UQvug, at least as early as Apr. 25, 2024. [cited by applicant]
Available online at https://www.youtube.com/watch?v=ioOkbUQqmZ0, at least as early as Nov. 9, 2022. [cited by applicant]
Available online at https://www.youtube.com/watch?v=zmqWU2dQKZ8, at least as early as Oct. 24, 2024. [cited by applicant]
Available online at https://www.youtube.com/watch?v=q8ldbodRG14, at least as early as Feb. 26, 2024. [cited by applicant]
Available online at https://www.youtube.com/watch?v=CUhuhleQNos, at least as early as May 22, 2019. [cited by applicant]
Available online at https://www.youtube.com/watch?v=dY57qnD_O7U, at least as early as Jul. 27, 2021. [cited by applicant]
Pateromichelakis et al., Head-eyes system and gaze analysis of the humanoid robot Romeo, 2014, IEEE, p. 1374-1379 (Year: 2014). [cited by applicant]
Ye, Seonghyeon, et al. “Latent action pretraining from videos.” arXiv preprint arXiv:2410.11758 (2024). [cited by applicant]
Zhang et al., “An Object Attribute Guided Framework for Robot Learning Manipulations from Human Demonstration Videos,” 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) Macau, China, (Nov. … [cited by applicant]
Zhang et al., “MM-LLMs: Recent Advances in MultiModal Large Language Models,” arXiv:2401.13601v5 (May 28, 2024). [cited by applicant]
Zhang et al., “Llama-Adapter: Efficient Fine-Tuning of Large Language Models With Zero-Initialized Attention,” arXiv:2303.16199v3 (Sep. 18, 2024). [cited by applicant]
Zhou et al., “Watch, Try, Learn: Meta-Learning From Demonstrations and Rewards,” arXiv:1906.03352v4 (Jan. 30, 2020). [cited by applicant]
International Search Report for PCT/US2025/016930. [cited by applicant]
Nakada et al., Deep Learning of Neuromuscular and Visuomotor Control of a Biomimetic Simulated Humanoid, 2020, IEEE, p. 3952-3959 (Year: 2020). [cited by applicant]
Lim et al., Basic emotional walking using a biped humanoid robot, 1999, IEEE, p. 954-959 (Year: 1999). [cited by applicant]
Albers et al., Upper Body of a new Humanoid Roboy—the Design of Armar III, 2006, IEEE, p. 308-309 (Year: 2006). [cited by applicant]
International Search Report for PCT/US2025/023064. [cited by applicant]
Englsberger et al., “Overview of the Torque-Controlled Humanoid Robot TORO,” 2014 14th IEEE-RAS International Conference on Humanoid Robots (Humanoids), Nov. 18-20, 2014. Madrid, Spain. [cited by applicant]
Mikayla Tetteh-Martey, (date posted Nov. 3, 2024), Blurring Lines: Resurgence of ‘I, Robot’, Cornellsun.com, URL: (https:// www.cornellsun.com/article/2024/11/blurring-lines-resurgence-of-i-robot), (Year: 2024). [cited by applicant]
Mike Oitzman, (date posted Aug. 6, 2024), Figure 02 humanoid robot is ready to get to work, therobotreport.com, URL: (https://www.therobotreport.com/figure-02-humanoid-robot-is-ready-to-get-to-work/), (Year: 2024). [cited by applicant]
International Search Report for PCT/US2025/019793. [cited by applicant]
International Search Report for PCT/US2025/025005. [cited by applicant]
International Search Report for PCT/US2025/024817. [cited by applicant]
Cheng et al., “Human Posture Estimation Using Voxel Data for “Smart” Airbag Systems: Issues and Framework,” IEEE, p. 84-89 (2004). [cited by applicant]
Droeschel et al., “Learning to Interpret Pointing Gestures with a Time-of-Flight Camera,” IEEE, p. 481-488 (2025). [cited by applicant]
Frohlich et al., “Design and Impementation of a Spherical Joint for Mobile Manipulators,” IEEE, p. 251-258 (2025). [cited by applicant]
Netzev et al., “Many Faced Robot—Design and Manufacturing of a parametric, Modular and Open Source Robot Head,” IEEE, p. 342-348 (2019). [cited by applicant]
Netzev et al., Design and implementation of a spherical joint for mobile manipulators, 2019, IEEE, p. 342-348 (Year: 2019). [cited by applicant]
Haddadin et al., The “DLR crash report”: Towards a standard crash-testing protocol for robot safety—Part II: Discussions, 2009, IEEE, p. 280-287 (Year: 2009). [cited by applicant]
Yaghoubi et al., Region-Based CNNs for Pedestrian Gender Recognition in Visual Surveillance Environments, 2019, IEEE, p. 1-5 (Year: 2019). [cited by applicant]
Pateromichelakis et al., Head-eyes system and gaze analysis of the humanoid robot Romeo, 2014, IEEE, p. 1374-131379 (Year: 2014). [cited by applicant]
Mokhtari et al., Taban:A Retro-Projected Social Robotic—Head for Human-Robot Interaction, 2019, IEEE, p. 46-51 (Year: 2019). [cited by applicant]
International Search Report for PCT/US25/23325. [cited by applicant]
Kim, Moo Jin, et al. “Openvla: An open-source vision-language-action model.” arXiv preprint arXiv:2406.09246 (Jun. 13, 2024). [cited by applicant]
Advancing Physical AI with NVIDIA Cosmos World Foundation Model Platform, avaiable at https://developer.nvidia.com/blog/advancing-physical-ai-with-nvidia-cosmos-world-foundation-model-platform/ (Jan. 9, 2025). [cited by applicant]
Building a Synthetic Motion Generation Pipeline for Humanoid Robot Learning, avaiable at https://developer.nvidia.com/blog/building-a-synthetic-motion-generation-pipeline-for-humanoid-robot-learning/ (Mar. 18, 2025). [cited by applicant]
Brohan, Anthony, et al. “Rt-1: Robotics transformer for real-world control at scale.” arXiv preprint arXiv:2212.06817 (Dec. 13, 2022). [cited by applicant]
Zitkovich, Brianna, et al. “Rt-2: Vision-language-action models transfer web knowledge to robotic control.” Conference on Robot Learning. PMLR, (Jul. 28, 2023). [cited by applicant]
Lynch, Corey, et al. “Interactive language: Talking to robots in real time.” IEEE Robotics and Automation Letters (Oct. 12, 2023). [cited by applicant]
Team, Octo Model, et al. “Octo: An open-source generalist robot policy.” arXiv preprint arXiv:2405.12213 (May 20, 2024). [cited by applicant]
Lynch, Corey, Kamelia Aryafar, and Josh Attenberg. “Images don't lie: Transferring deep visual semantic features to large-scale multimodal learning to rank.” Proceedings of the 22nd ACM SIGKDD international conference o… [cited by applicant]
Sermanet, Pierre, et al. “Time-contrastive networks: Self-supervised learning from video.” 2018 IEEE international conference on robotics and automation (ICRA). IEEE, (Mar. 20, 2018). [cited by applicant]
Dwibedi, Debidatta, et al. “Learning actionable representations from visual observations.” 2018 IEEE/RSJ international conference on intelligent robots and systems (IROS). IEEE, Feb. 2, 2018). [cited by applicant]
Lynch, Corey, et al. “Learning latent plans from play.” Conference on robot learning. Pmlr, (Dec. 20, 2020). [cited by applicant]
Pirk, Sören, et al. “Online object representations with contrastive learning.” arXiv preprint arXiv:1906.04312 (Jun. 10, 2019). [cited by applicant]
Gupta, Abhishek, et al. “Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning.” arXiv preprint arXiv:1910.11956 (Oct. 25, 2019). [cited by applicant]
Lynch, Corey, and Pierre Sermanet. “Language conditioned imitation learning over unstructured data.” arXiv preprint arXiv:2005.07648 (Jul. 7, 2020). [cited by applicant]
Jang, Eric, et al. “Bc-z: Zero-shot task generalization with robotic imitation learning.” Conference on Robot Learning. PMLR, (Feb. 4, 2022). [cited by applicant]
Gupta, Abhishek, et al. “Bootstrapped autonomous practicing via multi-task reinforcement learning.” arXiv preprint arXiv:2203.15755 (Mar. 29, 2022). [cited by applicant]
Heravi, Negin, et al. “Visuomotor control in multi-object scenes using object-aware representations.” arXiv preprint arXiv:2205.06333 (May 12, 2022). [cited by applicant]
Ding, Tianli, et al. “Goalseye: Learning high speed precision table tennis on a physical robot.” arXiv preprint arXiv:2210.03662 (Oct. 13, 2022). [cited by applicant]
Driess, Danny, et al. “Palm-e: An embodied multimodal language model.” (Mar. 6, 2023). [cited by applicant]
Wiedebach, Georg, et al. “Walking on partial footholds including line contacts with the humanoid robot atlas.” 2016 IEEE-RAS 16th International Conference on Humanoid Robots (Humanoids). IEEE, 2016. [cited by applicant]
Griffin, Robert J., et al. “Walking stabilization using step timing and location adjustment on the humanoid robot, atlas.” 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2017. [cited by applicant]
Griffin, Robert J., et al. “Straight-leg walking through underconstrained whole-body control.” 2018 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2018. [cited by applicant]
Griffin, Robert J., et al. “Capture point trajectories for reduced knee bend using step time optimization.” 2017 IEEE-RAS 17th International Conference on Humanoid Robotics (Humanoids). IEEE, 2017. [cited by applicant]
Griffin, Robert J., et al. “Footstep planning for autonomous walking over rough terrain.” 2019 IEEE-RAS 19th international conference on humanoid robots (humanoids). IEEE, 2019. [cited by applicant]
Garcia, Gabriel, Robert Griffin, and Jerry Pratt. “0-step capturability, motion decomposition and global feedback control of the 3D variable height-inverted pendulum.” arXiv preprint arXiv:1912.06078 (2019). [cited by applicant]
Dafarra, Stefano, et al. “Non-linear trajectory optimization for large step-ups: Application to the humanoid robot atlas.” 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2020. [cited by applicant]
Calvert, Duncan, et al. “A fast, autonomous, bipedal walking behavior over rapid regions.” 2022 IEEE-RAS 21st International Conference on Humanoid Robots (Humanoids). IEEE, 2022. [cited by applicant]