IP Library › Granted Patent US 11,783,584
Granted Patent B2
US 11,783,584 · App. 17/691,526 · Granted Oct 10, 2023

Automated digital document generation from digital videos

Inventors: Niyati Himanshu Chhaya (Hyderabad, IN); Tripti Shukla (Lucknow, IN); Jeevana Kruthi Karnuthala (Kurnool, IN); Bhanu Prakash Reddy Guda (Pittsburgh, PA); Ayudh Saxena (Pune, IN); Abhinav Bohra (Ahmedabad, IN); Abhilasha Sancheti (Bhilwara, IN); Aanisha Bhattacharyya (Hooghly, IN)
Assignee: Adobe Inc.
G06V20/47G06F16/73G06F40/166G06N20/00G06V10/86G06V20/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,783,584
App. No.
17/691,526
Granted
Oct 10, 2023
Kind
B2
Abstract

Techniques are described that support automated generation of a digital document from digital videos using machine learning. The digital document includes textual components that describe a sequence of entity and action descriptions from the digital video. These techniques are usable to generate a single digital document based on a plurality of digital videos as well as incorporate user-specified constraints in the generation of the digital document.

Claims (42)

1. A method implemented by a computing device, the method comprising:

locating, by the computing device, action clips including frames that depict actions from a plurality of digital videos;

generating, by the computing device, an action graph based on the action clips;

selecting, by the computing device, a path based on the action graph;

finding, by the computing device, frames from the plurality of digital videos by mapping nodes of the path to the action clips;

forming, by the computing device, textual components that include a sequence of entity and respective action descriptions based on the frames using a model trained using machine-learning;

generating, by the computing device, a digital document based on the textual components; and

rendering the digital document by a display device.

2. The method as described in claim 1 , wherein the forming employs at least one user-specified constraint.

3. The method as described in claim 2 , wherein the at least one user-specified constraint includes length of the digital document, a number of steps, semantics, or layout.

4. The method as described in claim 1 , wherein the locating includes:

extracting key clips from the plurality of digital videos; and

detecting the action clips as including frames that depict actions from the key clips.

5. The method as described in claim 4 , wherein the locating the action clips is performed using a binary classifier that computes a combined representation based on the key clips and respective portions of transcripts generated from the plurality of digital videos.

6. The method as described in claim 1 , wherein the action graph includes nodes representing actions and edges having weights based on probabilities of transition between the nodes, respectively.

7. The method as described in claim 6 , wherein the selecting the path includes traversing the nodes of the action graph based on respective said probabilities.

8. The method as described in claim 1 , wherein the forming by the machine-learning model includes processing the frames along with respective portions of transcripts generated from the plurality of digital videos.

9. The method as described in claim 1 , wherein the generating the digital document includes selecting a digital image from the action clips based on contribution of the digital image to the path and including the digital image as part of the digital document.

10. The method as described in claim 1 , wherein the forming includes selecting a digital image from at least one said frame included within a defined portion of a respective said digital video.

11. A system comprising:

a processing system;

a non-transitory computer readable media communicatively coupled to the processing system;

an action detection module implemented by the processing system to locate action clips including frames that depict actions from a digital video;

an action graph generation module implemented by the processing system to generate an action graph based on the action clips;

a path selection module implemented by the processing system to select a path based on the action graph;

a frame location module implemented by the processing system to locate frames from the digital video based on a mapping nodes of the path to the action clips; and

a decoding module implemented by the processing system to form textual components that include a sequence of entity and respective action descriptions based on the located frames using a model trained using machine-learning.

12. The system as described in claim 11 , further comprising a search module implemented by the processing system to generate a search result that references a plurality of said digital videos and wherein the action detection module is configured to locate the action clips from the plurality of said digital videos.

13. The system as described in claim 11 , further comprising a transcription module implemented by the processing system to extract a transcript from the digital video.

14. The system as described in claim 11 , further comprising a key clip extraction module implemented by the processing system to extract key clips from the digital video and an action clip detection module implemented by the processing system to detect the action clips as including frames that depict actions from the key clips.

15. The system as described in claim 14 , wherein the action clip detection module implements a binary classifier that computes a combined representation based on the key clips and respective portions of a transcript generated from the digital video.

16. The system as described in claim 11 , wherein the action graph includes nodes representing actions and edges having weights based on probabilities of transitions between the nodes, respectively.

17. The system as described in claim 16 , wherein the path selection module is configured to select the path by traversing the nodes of the action graph based on respective said probabilities.

18. The system as described in claim 11 , wherein the decoding module is configured to process the frames along with respective portions of transcripts generated from the digital video.

19. A system comprising:

means for locating action clips from a plurality of digital videos;

means for generating an action graph based on the action clips;

means for selecting a path based on the action graph;

means for locating frames from the plurality of digital videos based on the path;

means for forming textual components that include a sequence of entity and respective action descriptions based on the located frames using a model trained using machine-learning; and

means for generating a digital document based on the textual components.

20. The system as described in claim 19 , wherein the model is implemented using an encoder-decoder network as part of machine learning.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 17, 2022
From: CHHAYA, NIYATI HIMANSHU; SHUKLA, TRIPTI; KARNUTHALA, JEEVANA KRUTHI; GUDA, BHANU PRAKASH REDDY; SAXENA, AYUDH; BOHRA, ABHINAV; SANCHETI, ABHILASHA; BHATTACHARYYA, AANISHA
To: ADOBE INC.
Reel/Frame 059291/0024 →
Continuity (1)
Related Publication 20230290146A1 · Sep 14, 2023
Cited By (1)
US 12,288,397