IP Library › Granted Patent US 12,367,662
Granted Patent B2
US 12,367,662 · App. 18/159,189 · Granted Jul 22, 2025

Systems and methods for video models with procedure understanding

Inventors: Roberto Martin-Martin (Austin, TX); Silvio Savarese (San Francisco, CA); Honglu Zhou (San Francisco, CA); Juan Carlos Niebles Duque (Mountain View, CA)
Assignee: Salesforce, Inc.
G06V10/774G06F40/30G06F40/40G06V10/7753G06V10/776G06V10/82G06V20/41G06V20/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,662
App. No.
18/159,189
Granted
Jul 22, 2025
Kind
B2
Abstract

Embodiments described herein provide systems and methods for training video models to perform a task from an input instructional video. A procedure knowledge graph (PKG) may be generated with nodes representing procedure steps, and edges representing relationships between the steps. The PKG may be generated based on text and/or video training data which includes procedures (e.g., instructional videos). Using the PKG, a video model may be trained using the PKG to provide supervisory training signals for a number of tasks. Once the model is trained, it may be fine-tuned for a specific task which benefits from the model being trained in a way that makes the model embed procedural information when encoding videos.

Claims (60)

1. A method of training a neural network based video model to perform a task from an input instructional video, the method comprising:

receiving, via a data interface, at least one instructional video relating to completing a task having a series of steps and at least one text-based procedure description describing the series of steps;

generating, based on the instructional video and the text-based procedure description, a procedure knowledge graph having a plurality of nodes representing the series of steps and a plurality of edges representing sequential relationships between the series of steps;

generating, by the neural network based video model, a first training output according to a first training objective from an input of a video segment from the at least one instructional video;

generating a plurality of pseudo-labels corresponding to the first training objective based on the procedure knowledge graph;

computing a first loss corresponding to the first training objective based on the first training output and the plurality of pseudo-labels; and

updating parameters of the neural network based video model based on the computed first loss via backpropagation.

2. The method of claim 1 , wherein generating, by the neural network based video model, the first training output comprises:

generating, by a pre-trained video foundation encoder and a light-weight encoder, a latent representation of the input of the video segment;

generating, by a decoder corresponding to the first training objective, the first training output; and

updating parameters in the light-weight encoder based on the first loss while keeping the pre-trained video foundation encoder frozen.

3. The method of claim 1 , wherein the generating the plurality of pseudo-labels includes performing one or more queries on the procedure knowledge graph according to the first training objective.

4. The method of claim 3 , wherein the first training output includes a plurality of outputs, each of the plurality of outputs corresponding to a respective pseudo-label of the plurality of pseudo-labels.

5. The method of claim 4 , wherein the plurality of outputs are associated with different queries performed on the procedure knowledge graph.

6. The method of claim 1 , wherein generating, by the neural network based video model, the first training output includes generating a latent representation of the input of the video segment, further comprising:

training a task-specific neural network model which takes the latent representation as an input, and outputs a prediction associated with the input of the video segment.

7. The method of claim 1 , further comprising:

generating a second training output according to a second training objective from the input of the video segment;

generating a second plurality of pseudo-labels corresponding to the second training objective based on the procedure knowledge graph;

computing a second loss corresponding to the second training objective based on the second training output and the second plurality of pseudo-labels; and

updating parameters of the neural network based video model based on the computed second loss via backpropagation.

8. The method of claim 7 , wherein the first training objective or the second training objective is one of video-node matching, video-task matching, task context learning or node relation learning.

9. A system for training a neural network based video model to perform a task from an input instructional video, the system comprising:

a memory that stores the video model and a plurality of processor executable instructions;

a communication interface that receives at least one instructional video relating to completing a task having a series of steps and at least one text-based procedure description describing the series of steps; and

one or more hardware processors that read and execute the plurality of processor-executable instructions from the memory to perform operations comprising:

generating, based on the instructional video and the text-based procedure description, a procedure knowledge graph having a plurality of nodes representing the series of steps and a plurality of edges representing sequential relationships between the series of steps;

generating, by the neural network based video model, a first training output according to a first training objective from an input of a video segment from the at least one instructional video;

generating a plurality of pseudo-labels corresponding to the first training objective based on the procedure knowledge graph;

computing a first loss corresponding to the first training objective based on the first training output and the plurality of pseudo-labels; and

updating parameters of the neural network based video model based on the computed first loss via backpropagation.

10. The system of claim 9 , wherein generating, by the neural network based video model, the first training output comprises:

generating, by a pre-trained video foundation encoder and a light-weight encoder, a latent representation of the input of the video segment;

generating, by a decoder corresponding to the first training objective, the first training output; and

updating parameters in the light-weight encoder based on the first loss while keeping the pre-trained video foundation encoder frozen.

11. The system of claim 9 , wherein the generating the plurality of pseudo-labels includes performing one or more queries on the procedure knowledge graph according to the first training objective.

12. The system of claim 11 , wherein the first training output includes a plurality of outputs, each of the plurality of outputs corresponding to a respective pseudo-label of the plurality of pseudo-labels.

13. The system of claim 12 , wherein the plurality of outputs are associated with different queries performed on the procedure knowledge graph.

14. The system of claim 9 , wherein generating, by the neural network based video model, the first training output includes generating a latent representation of the input of the video segment, further comprising:

training a task-specific neural network model which takes the latent representation as an input, and outputs a prediction associated with the input of the video segment.

15. The system of claim 9 , the operations further comprising:

generating a second training output according to a second training objective from the input of the video segment;

generating a second plurality of pseudo-labels corresponding to the second training objective based on the procedure knowledge graph;

computing a second loss corresponding to the second training objective based on the second training output and the second plurality of pseudo-labels; and

updating parameters of the neural network based video model based on the computed second loss via backpropagation.

16. The system of claim 15 , wherein the first training objective or the second training objective is one of video-node matching, video-task matching, task context learning or node relation learning.

17. A non-transitory machine-readable medium comprising a plurality of machine-executable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform operations comprising:

receiving, via a data interface, at least one instructional video relating to completing a task having a series of steps and at least one text-based procedure description describing the series of steps;

generating, based on the instructional video and the text-based procedure description, a procedure knowledge graph having a plurality of nodes representing the series of steps and a plurality of edges representing sequential relationships between the series of steps;

generating, by a neural network based video model, a first training output according to a first training objective from an input of a video segment from the at least one instructional video;

generating a plurality of pseudo-labels corresponding to the first training objective based on the procedure knowledge graph;

computing a first loss corresponding to the first training objective based on the first training output and the plurality of pseudo-labels; and

updating parameters of the neural network based video model based on the computed first loss via backpropagation.

18. The non-transitory machine-readable medium of claim 17 , wherein the generating the plurality of pseudo-labels includes performing one or more queries on the procedure knowledge graph according to the first training objective.

19. The non-transitory machine-readable medium of claim 18 , wherein the first training output includes a plurality of outputs, each of the plurality of outputs corresponding to a respective pseudo-label of the plurality of pseudo-labels.

20. The non-transitory machine-readable medium of claim 17 , further comprising:

generating a second training output according to a second training objective from the input of the video segment;

generating a second plurality of pseudo-labels corresponding to the second training objective based on the procedure knowledge graph;

computing a second loss corresponding to the second training objective based on the second training output and the second plurality of pseudo-labels; and

updating parameters of the neural network based video model based on the computed second loss via backpropagation.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE NAME OF THE ASSIGNEE FROM SALESFORCE.COM, INC. TO SALESFORCE, INC. AS INDICATED ON THE EXECUTED ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 062582 FRAME: 0214. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 14, 2023
From: MARTIN-MARTIN, ROBERTO; SAVARESE, SILVIO; ZHOU, HONGLU; NIEBLES DUQUE, JUAN CARLOS
To: SALESFORCE, INC.
Reel/Frame 063525/0885 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 3, 2023
From: MARTIN-MARTIN, ROBERTO; SAVARESE, SILVIO; ZHOU, HONGLU; NIEBLES DUQUE, JUAN CARLOS
To: SALESFORCE.COM, INC.
Reel/Frame 062582/0214 →
Continuity (2)
Provisional Application 63383399 · Nov 11, 2022
Related Publication 20240161464A1 · May 16, 2024
References Cited (35)
US 11580302B2 · Taylor · 2023 [cited by examiner]
US 11625620B2 · Singaraju · 2023 [cited by examiner]
US 11687795B2 · Ferreira Moreno · 2023 [cited by examiner]
US 11755924B2 · Madden · 2023 [cited by examiner]
US 11847164B2 · Wang · 2023 [cited by examiner]
US 12046023B2 · Vongkulbhisal · 2024 [cited by examiner]
US 20150066477A1 · Hu · 2015 [cited by examiner]
US 20200057946A1 · Singaraju · 2020 [cited by examiner]
US 20200265324A1 · Ferreira Moreno · 2020 [cited by examiner]
US 20210034985A1 · Vongkulbhisal · 2021 [cited by examiner]
US 20210216717A1 · Wang · 2021 [cited by examiner]
US 20220036001A1 · Taylor · 2022 [cited by examiner]
US 20220100800A1 · Georgopoulos · 2022 [cited by examiner]
US 20220261599A1 · Kastaniotis · 2022 [cited by examiner]
US 20230169361A1 · Mitra · 2023 [cited by examiner]
US 20230206069A1 · Zhang · 2023 [cited by examiner]
US 20230401424A1 · Jiang · 2023 [cited by examiner]
US 20240135110A1 · Mujica-Parodi, III · 2024 [cited by examiner]
US 20240289647A1 · Zhao · 2024 [cited by examiner]
CN 107015963A · 2017 [cited by examiner]
CN 110070023B · 2020 [cited by examiner]
CN 108615047B · 2022 [cited by examiner]
CN-107015963-A (machine translation) (Year: 2017). [cited by examiner]
CN-108615047-B (machine translation) (Year: 2022). [cited by examiner]
CN-110070023-B (machine translation) (Year: 2020). [cited by examiner]
Gao et al., “A Survey of Graph Neural Networks for Recommender Systems: Challenges, Methods, and Directions.” arXiv preprint arXiv:2109.12843 (2022). (Year: 2022). [cited by examiner]
Yu et al., “Knowledge Embedding Based Graph Convolutional Network.” arXiv preprint arXiv:2006.07331 (2021). (Year: 2021). [cited by examiner]
Mehta et al., “Open-Domain Trending Hashtag Recommendation for Videos,” 2021 IEEE International Symposium on Multimedia (ISM), Naple, Italy, 2021, pp. 174-181, doi: 10.1109/ISM52913.2021.00035. (Year: 2021). [cited by examiner]
Xudong Lin, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach, Shih-Fu Chang, and Lorenzo Torresani. Learning to recognize procedural activities with distant supervision. In Proceedings of the IEEE/CVF Conference on Compu… [cited by applicant]
Muheng Li, Lei Chen, Yueqi Duan, Zhilan Hu, Jianjiang Feng, Jie Zhou, and Jiwen Lu. Bridge-prompt: Towards ordinal action understanding in instructional videos. In Proceedings of the IEEE/CVF Conference on Computer Visi… [cited by applicant]
Reza Ghoddoosian, Saif Sayed, and Vassilis Athitsos. Hierarchical modeling for task recognition and action segmentation in weakly-labeled instructional videos. In Proceedings of the IEEE/CVF Winter Conference on Applica… [cited by applicant]
Ludan Ruan and Qin Jin. Survey: Transformer based videolanguage pre-training. AI Open, 2022. [cited by applicant]
Madeline C Schiappa, Yogesh S Rawat, and Mubarak Shah. Self-supervised learning for videos: A survey. arXiv preprint arXiv:2207.00419, 2022. [cited by applicant]
Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong. Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in Neural Information… [cited by applicant]
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. End-to-end learning of visual representations from uncurated instructional videos. In Proceedings of the IEEE/CVF Confe… [cited by applicant]
Cited By (1)
US 12,718,564