IP Library › Granted Patent US 12,288,397
Granted Patent B2
US 12,288,397 · App. 18/464,493 · Granted Apr 29, 2025

Automated digital document generation from digital videos

Inventors: Niyati Himanshu Chhaya (Hyderabad, IN); Tripti Shukla (Lucknow, IN); Jeevana Kruthi Karnuthala (Kurnool, IN); Bhanu Prakash Reddy Guda (Pittsburgh, PA); Ayudh Saxena (Pune, IN); Abhinav Bohra (Ahmedabad, IN); Abhilasha Sancheti (Bhilwara, IN); Aanisha Bhattacharyya (Hooghly, IN)
Assignee: Adobe Inc.
G06V20/47G06F16/73G06F40/166G06N20/00G06V10/86G06V20/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,288,397
App. No.
18/464,493
Granted
Apr 29, 2025
Kind
B2
Abstract

Techniques are described that support automated generation of a digital document from digital videos using machine learning. The digital document includes textual components that describe a sequence of entity and action descriptions from the digital video. These techniques are usable to generate a single digital document based on a plurality of digital videos as well as incorporate user-specified constraints in the generation of the digital document.

Claims (37)

1. A method of generating a digital document, the method comprising:

receiving, by a computing device, at least one constraint to be used in controlling the generation of the digital document;

locating, by the computing device, action clips that depict actions from at least one digital video;

determining, by the computing device, frames from the action clips of the at least one digital video;

forming, by the computing device, textual components that include entity and respective action descriptions based on the frames using a model trained using machine-learning;

generating, by the computing device, a digital document based on the textual components and using the at least one constraint; and

rendering the digital document by a display device.

2. The method as described in claim 1 , wherein the at least one constraint is received from a user interface having user controls that support user inputs.

3. The method as described in claim 1 , wherein the at least one constraint is selected from length, format, semantics, or layout of the digital document.

4. The method as described in claim 1 , wherein the at least one constraint includes length, and the length is determined as a number of steps in the digital document.

5. The method as described in claim 1 , wherein the at least one constraint includes layout of the digital document, and the layout of the digital document is selected from templates.

6. The method as described in claim 1 , further comprising generating, by the computing device, an action graph based on the action clips.

7. The method as described in claim 6 , further comprising selecting, by the computing device, a path based on the action graph.

8. The method as described in claim 1 , wherein the forming by the machine-learning model includes processing the frames along with respective portions of transcripts generated from the at least one digital video.

9. The method as described in claim 1 , wherein the generating the digital document includes selecting a digital image from the action clips and including the digital image as part of the digital document.

10. A digital document generation system comprising:

a processing system;

a non-transitory computer readable media communicatively coupled to the processing system;

a user interface supporting user inputs to specify at least one constraint that is used in controlling the generation of a digital document;

an action detection module implemented by the processing system to locate action clips that depict actions from at least one digital video;

a frame location module implemented by the processing system to locate frames from the at least one digital video; and

a decoding module implemented by the processing system to form textual components to be used in generating the digital document, the textual components including a sequence of entity and respective action descriptions based on the located frames using a model trained using machine-learning.

11. The system as described in claim 10 , wherein the at least one constraint is selected from length, format, semantics, or layout of the digital document.

12. The system as described in claim 10 , wherein the at least one constraint includes length, and the length is determined as a number of steps in the digital document.

13. The system as described in claim 10 , further comprising an action graph generation module implemented by the processing system to generate an action graph based on the action clips.

14. The system as described in claim 13 , further comprising a path selection module implemented by the processing system to select a path based on the action graph, wherein the frames are located based on a mapping nodes of the path to the action clips.

15. The system as described in claim 10 , further comprising a search module implemented by the processing system to generate a search result that references the at least one digital video and wherein the action detection module is configured to locate the action clips from the at least one digital video.

16. A system comprising:

means for inputting at least one constraint to be used in controlling generation of a digital document;

means for locating action clips from at least one digital video;

means for locating frames from the action clips of the at least one digital video;

means for forming textual components that include a sequence of entity and respective action descriptions based on the located frames using a model trained using machine-learning; and

means for generating the digital document based on the textual components, the means for generating the digital document using the at least one constraint in generation of the digital document.

17. The system as described in claim 16 , wherein the at least one constraint is selected from length, format, semantics, or layout of the digital document.

18. The system as described in claim 16 , wherein the at least one constraint includes length, and the length is determined as a number of steps in the digital document.

19. The system as described in claim 16 , further comprising a means for generating an action graph based on the action clips.

20. The system as described in claim 19 , further comprising a means for selecting a path based on the action graph.

Continuity (2)
Continuation 17691526 · Mar 10, 2022
Related Publication 20230419666A1 · Dec 28, 2023
References Cited (37)
US 6636238B1 · Amir et al. · 2003 [cited by examiner]
US 7640494B1 · Chen et al. · 2009 [cited by applicant]
US 11783584B2 · Chhaya · 2023 [cited by applicant]
US 20100050083A1 · Axen et al. · 2010 [cited by applicant]
US 20170062013A1 · Carter et al. · 2017 [cited by applicant]
US 20170262159A1 · Denoue et al. · 2017 [cited by examiner]
US 20230290146A1 · Chhaya et al. · 2023 [cited by examiner]
“Article Generator”, Free Article Generator Online [retrieved Feb. 16, 2022]. Retrieved from the Internet <https://articlegenerator.org/>., 2 Pages. [cited by applicant]
“Articleforge 3.0”, Glimpse.ai [retrieved Feb. 16, 2022]. Retrieved from the Internet <https://www.articleforge.com/>., 2015, 6 Pages. [cited by applicant]
“Document Management & Workflow Software”, R2 Docuo [retrieved Feb. 16, 2022]. Retrieved from the Internet <https://www.r2docuo.com/en/>., 2012, 7 Pages. [cited by applicant]
“Docupilot”, Flackon Inc. [retrieved Feb. 16, 2022]. Retrieved from the Internet <https://docupilot.app/>., 2018, 5 Pages. [cited by applicant]
“Nintex Process Management and Workflow Automation Software”, Nintex UK Ltd [retrieved Feb. 16, 2022]. Retrieved from the Internet <https://www.nintex.com/>., 2005, 4 Pages. [cited by applicant]
“RecipeGPT”, SMU Living Analytics Research Centre [retrieved Feb. 16, 2022]. Retrieved from the Internet <https://recipegpt.org/>., 2020, 1 Page. [cited by applicant]
“Speech to Text”, Microsoft [retrieved Feb. 17, 2022]. Retrieved from the Internet <https://azure.microsoft.com/en-in/services/cognitive-services/speech-to-text/#faq>., 2021, 11 Pages. [cited by applicant]
U.S. Appl. No. 17/691,526, “Non-Final Office Action”, U.S. Appl. No. 17/691,526, Apr. 13, 2023, 6 pages. [cited by applicant]
U.S. Appl. No. 17/691,526, “Notice of Allowance”, U.S. Appl. No. 17/691,526, Aug. 16, 2023, 7 pages. [cited by applicant]
Bień, Michał, et al., “RecipeNLG: A Cooking Recipes Dataset for Semi-Structured Text Generation”, INLG [retrieved Feb. 16, 2022]. Retrieved from the Internet <https://aclanthology.org/2020.inlg-1.4.pdf>., Dec. 2020, 7 P… [cited by applicant]
Canny, John , “A Computational Approach to Edge Detection”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. PAMI-8, No. 6 [retrieved Feb. 17, 2022]. Retrieved from the Internet <https://doi.org/10.1… [cited by applicant]
Damen, Dima , et al., “The EPIC-KITCHENS Dataset: Collection, Challenges and Baselines”, Cornell University arXiv, arXiv.org [retrieved Feb. 17, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/2005.00343.pdf>.… [cited by applicant]
Devlin, Jacob , et al., “BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding”, Cornell University, arXiv Preprint, arXiv.org [retrieved on Jun. 14, 2023]. Retrieved from the Internet <https:… [cited by applicant]
Gildea, Daniel , et al., “Automatic Labeling of Semantic Roles”, Computational Linguistics vol. 28, No. 3 [retrieved Feb. 17, 2022]. Retrieved from the Internet <http://I3s.de/˜zerr/clusterlabelling/s1.pdf>., Sep. 2002,… [cited by applicant]
Hartigan, J. A. , et al., “Algorithm AS 136: A K-Means Clustering Algorithm”, Journal of the Royal Statistical Society vol. 28, No. 1 [retrieved Feb. 17, 2022]. Retrieved from the Internet <https://www.stat.cmu.edu/˜rnu… [cited by applicant]
Haussmann, Steven , et al., “FoodKG: A Semantics-Driven Knowledge Graph for Food Recommendation”, International Semantic Web Conference [retrieved Feb. 17, 2022]. Retrieved from the Internet <http://www.cs.rpi.edu/˜zaki… [cited by applicant]
He, Kaiming , et al., “Deep Residual Learning for Image Recognition”, Proceedings of the IEEE conference on computer vision and pattern recognition, 2016 [retrieved Feb. 18, 2022], Retrieved from the Internet: <https://… [cited by applicant]
Lee, Helena H, et al., “RecipeGPT: Generative Pre-training Based Cooking Recipe Generation and Evaluation System”, Companion Proceedings of the Web Conference 2020 [retrieved Feb. 16, 2022]. Retrieved from the Internet … [cited by applicant]
Marín, Javier , et al., “Recipe1M+: A Dataset for Learning Cross-Modal Embeddings for Cooking Recipes and Food Images”, IEEE Transactions on Pattern Analysis and Machine Intelligence vol. 43, No. 1 [retrieved Feb. 16, 2… [cited by applicant]
Nishimura, Taichi , et al., “Procedural Text Generation from a Photo Sequence”, Proceedings of the 12th International Conference on Natural Language Generation [retrieved Feb. 17, 2022]. Retrieved from the Internet <htt… [cited by applicant]
Radford, Alec , et al., “Learning Transferable Visual Models From Natural Language Supervision”, Cornell University arXiv, arXiv.org [retrieved Jun. 29, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/2103.000… [cited by applicant]
Salvador, Amaia , et al., “Inverse Cooking: Recipe Generation From Food Images”, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) [retrieved Feb. 17, 2022]. Retrieved from the Internet <https://arxi… [cited by applicant]
Sener, Fadime , et al., “Zero-Shot Anticipation for Instructional Activities”, Cornell University arXiv, arXiv.org [retrieved Feb. 17, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1812.02501.pdf>., Oct. 20,… [cited by applicant]
Shi, Peng , et al., “Simple BERT Models for Relation Extraction and Semantic Role Labeling”, Cornell University arXiv, arXiv.org [retrieved Feb. 16, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1904.05255.p… [cited by applicant]
Ushiku, Atsushi , et al., “Procedural Text Generation from an Execution Video”, International Joint Conference on Natural Language Processing [retrieved Feb. 17, 2022]. Retrieved from the Internet <http://www.Ista.media… [cited by applicant]
Vaswani, Ashish , et al., “Attention is all you Need”, Cornell University arXiv Preprint, arXiv.org [retrieved Aug. 9, 2023]. Retrieved from the Internet <https://arxiv.org/pdf/1706.03762.pdf>., Dec. 6, 2017, 15 pages. [cited by applicant]
Xu, Frank F, et al., “A Benchmark for Structured Procedural Knowledge Extraction from Cooking Videos”, Cornell University arXiv, arXiv.org [retrieved Feb. 17, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/20… [cited by applicant]
Zhou, Luowei , et al., “End-to-End Dense Video Captioning with Masked Transformer”, Cornell University arXiv, arXiv.org [retrieved Feb. 17, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1804.00819.pdf>., Apr… [cited by applicant]
Zhou, Luowei , et al., “Towards Automatic Learning of Procedures from Web Instructional Videos”, Cornell University arXiv, arXiv.org [retrieved Feb. 17, 2022]. Retrieved from the Internet <https://arxiv.org/pdf/1703.097… [cited by applicant]
Zhou, Luowei , et al., “YouCook2 Dataset”, University of Michigan [retrieved Feb. 17, 2022]. Retrieved from the Internet <http://youcook2.eecs.umich.edu/>., 2017, 4 Pages. [cited by applicant]