IP Library › Granted Patent US 12,238,451
Granted Patent B2
US 12,238,451 · App. 18/055,301 · Granted Feb 25, 2025

Predicting video edits from text-based conversations using neural networks

Inventors: Uttaran Bhattacharya (Sunnyvale, CA); Gang Wu (San Jose, CA); Viswanathan Swaminathan (Saratoga, CA); Stefano Petrangeli (Mountain View, CA)
Assignee: Adobe Inc.
H04N7/002G06T11/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,238,451
App. No.
18/055,301
Granted
Feb 25, 2025
Kind
B2
Abstract

Embodiments are disclosed for predicting, using neural networks, editing operations for application to a video sequence based on processing conversational messages by a video editing system. In particular, in one or more embodiments, the disclosed systems and methods comprise receiving an input including a video sequence and text sentences, the text sentences describing a modification to the video sequence, mapping, by a first neural network content of the text sentences describing the modification to the video sequence to a candidate editing operation, processing, by a second neural network, the video sequence to predict parameter values for the candidate editing operation, and generating a modified video sequence by applying the candidate editing operation with the predicted parameter values to the video sequence.

Claims (77)

1. A computer-implemented method comprising:

receiving an input including a video sequence and text sentences, the text sentences describing a modification to the video sequence;

mapping, by a first neural network, content of the text sentences describing the modification to the video sequence to a candidate video editing operation;

processing, by a second neural network, the video sequence to predict parameter values for the candidate video editing operation; and

generating a modified video sequence by applying the candidate video editing operation with the predicted parameter values to the video sequence.

2. The computer-implemented method of claim 1 , wherein mapping the content of the text sentences describing the modification to the video sequence to the candidate video editing operation comprises:

mapping the content of the text sentences to a reference sentence; and

identifying a video editing operation associated with the reference sentence as the candidate video editing operation.

3. The computer-implemented method of claim 2 , wherein mapping the content of the text sentences to the reference sentence comprises:

generating, by a sentence transformer, sentence features for the text sentences;

calculating cosine similarity values between the sentence features for the text sentences and reference sentence features for reference sentences, wherein each reference sentence of the reference sentences is associated with a video editing operation; and

identifying the reference sentence having a highest calculated cosine similarity with the sentence features for the text sentences.

4. The computer-implemented method of claim 1 , wherein processing the video sequence to predict the parameter values for the candidate video editing operation comprises:

for each frame of the video sequence:

generating an RGB feature vector and an optical flow feature vector,

concatenating the RGB feature vector and the optical flow feature vector to create a concatenated feature vector, and

passing the concatenated feature vector through an editing parameters prediction network to predict the parameter values for the candidate video editing operation.

5. The computer-implemented method of claim 4 , wherein the predicted parameter values for the candidate video editing operation include mean parameter values and standard deviation parameter values.

6. The computer-implemented method of claim 1 , further comprising:

receiving a second input including second text sentences, the second text sentences describing a second modification to trim the video sequence by a first amount of time;

detecting shot boundaries within the video sequence;

determining that an end time of a shot boundary is within the first amount of time; and

trimming a second amount of time from the video sequence starting at the end time of the shot boundary, wherein the second amount of time is different from the first amount of time.

7. The computer-implemented method of claim 1 , wherein generating the modified video sequence by applying the candidate video editing operation with the predicted parameter values for the candidate video editing operation to the video sequence comprises:

adjusting a brightness parameter in response to mapping the text sentences to a brightness editing operation.

8. The computer-implemented method of claim 1 , further comprising:

receiving a second input including a second video sequence;

processing, by the second neural network, the second video sequence to predict parameter values for one or more video editing operations; and

generating a modified second video sequence by applying the one or more video editing operations with the predicted parameter values to the second video sequence.

9. A non-transitory computer-readable storage medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:

receiving an input including a video sequence and text sentences, the text sentences describing a modification to the video sequence;

mapping, by a first neural network, content of the text sentences describing the modification to the video sequence to a candidate video editing operation;

processing, by a second neural network, the video sequence to predict parameter values for the candidate video editing operation; and

generating a modified video sequence by applying the candidate video editing operation with the predicted parameter values to the video sequence.

10. The non-transitory computer-readable storage medium of claim 9 , wherein to map the content of the text sentences describing the modification to the video sequence to the candidate video editing operation the instructions further cause the processing device to perform operations comprising:

mapping the content of the text sentences to a reference sentence; and

identifying a video editing operation associated with the reference sentence as the candidate video editing operation.

11. The non-transitory computer-readable storage medium of claim 10 , wherein to map the content of the text sentences to the reference sentence the instructions further cause the processing device to perform operations comprising:

generating, by a sentence transformer, sentence features for the text sentences;

calculating cosine similarity values between the sentence features for the text sentences and reference sentence features for reference sentences, wherein each reference sentence of the reference sentences is associated with a video editing operation; and

identifying the reference sentence having a highest calculated cosine similarity with the sentence features for the text sentences.

12. The non-transitory computer-readable storage medium of claim 9 , wherein to process the video sequence to predict the parameter values for the candidate video editing operation the instructions further cause the processing device to perform operations comprising:

for each frame of the video sequence:

generating an RGB feature vector and an optical flow feature vector,

concatenating the RGB feature vector and the optical flow feature vector to create a concatenated feature vector, and

passing the concatenated feature vector through an editing parameters prediction network to predict the parameter values for the candidate video editing operation.

13. The non-transitory computer-readable storage medium of claim 12 , wherein the predicted parameter values for the candidate video editing operation include mean parameter values and standard deviation parameter values.

14. The non-transitory computer-readable storage medium of claim 9 , wherein the instructions further cause the processing device to perform operations comprising:

receiving a second input including second text sentences, the second text sentences describing a second modification to trim the video sequence by a first amount of time;

detecting shot boundaries within the video sequence;

determining that an end time of a shot boundary is within the first amount of time; and

trimming a second amount of time from the video sequence starting at the end time of the shot boundary, wherein the second amount of time is different from the first amount of time.

15. The non-transitory computer-readable storage medium of claim 9 , wherein to generate the modified video sequence by applying the candidate video editing operation with the predicted parameter values for the candidate video editing operation to the video sequence the instructions further cause the processing device to perform operations comprising:

adjusting a brightness parameter in response to mapping the text sentences to a brightness editing operation.

16. The non-transitory computer-readable storage medium of claim 9 , wherein the instructions further cause the processing device to perform operations comprising:

receiving a second input including a second video sequence;

processing, by the second neural network, the second video sequence to predict parameter values for one or more video editing operations; and

generating a modified second video sequence by applying the one or more video editing operations with the predicted parameter values to the second video sequence.

17. A system comprising:

a memory component; and

a processing device coupled to the memory component, the processing device to perform operations comprising:

receiving an input including a video sequence and text sentences, the text sentences describing a modification to the video sequence;

mapping, by a first neural network, content of the text sentences describing the modification to the video sequence to a candidate video editing operation;

processing, by a second neural network, the video sequence to predict parameter values for the candidate video editing operation; and

generating a modified video sequence by applying the candidate video editing operation with the predicted parameter values to the video sequence.

18. The system of claim 17 , wherein to map the content of the text sentences describing the modification to the video sequence to the candidate video editing operation the processing device further performs operations comprising:

mapping the content of the text sentences to a reference sentence; and

identifying a video editing operation associated with the reference sentence as the candidate video editing operation.

19. The system of claim 18 , wherein to map the content of the text sentences to the reference sentence the processing device further performs operations comprising:

generating, by a sentence transformer, sentence features for the text sentences;

calculating cosine similarity values between the sentence features for the text sentences and reference sentence features for reference sentences, wherein each reference sentence of the reference sentences is associated with a video editing operation; and

identifying the reference sentence having a highest calculated cosine similarity with the sentence features for the text sentences.

20. The system of claim 17 , wherein to process the video sequence to predict the parameter values for the candidate video editing operation the processing device further performs operations comprising:

for each frame of the video sequence:

generating an RGB feature vector and an optical flow feature vector,

concatenating the RGB feature vector and the optical flow feature vector to create a concatenated feature vector, and

passing the concatenated feature vector through an editing parameters prediction network to predict the parameter values for the candidate video editing operation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 17, 2022
From: BHATTACHARYA, UTTARAN; WU, GANG; SWAMINATHAN, VISWANATHAN; PETRANGELI, STEFANO
To: ADOBE INC.
Reel/Frame 061817/0091 →
Continuity (1)
Related Publication 20240163393A1 · May 16, 2024
References Cited (30)
US 11138693B2 · Mejjati et al. · 2021 [cited by applicant]
US 11170389B2 · Wang et al. · 2021 [cited by applicant]
US 11574477B2 · Wu et al. · 2023 [cited by applicant]
US 20220130077A1 · Rajarathnam · 2022 [cited by examiner]
US 20230196817A1 · Oh et al. · 2023 [cited by applicant]
US 20230260502A1 · Gabrys · 2023 [cited by examiner]
US 20230376690A1 · Bellegarda · 2023 [cited by examiner]
Bau, D. et al., “Semantic Photo Manipulation with a Generative Image Prior,” ACM Trans. Graph. 38, 4, Article 59, Aug. 2019, pp. 59:1-59:11. [cited by applicant]
Bhattacharya, U. et al., “HighlightMe: Detecting Highlights From Human-Centric Videos,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2021, pp. 8157-8167. [cited by applicant]
Bhattacharya, U. et al., “Show Me What I Like: Detecting User-Specific Video Highlights Using Content-Based Multi-Head Attention,” MM '22: Proceedings of the 30th ACM International Conference on Multimedia, Oct. 2022, p… [cited by applicant]
Carreira, J. et al., “Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset”, arXiv:1705.07750v1 [cs.CV], May 22, 2017, 10 pages. [cited by applicant]
Castellano, B., “scenedetect 0.6.3: pip install scenedetect,” https://pypi.org/project/scenedetect/, retrieved on Mar. 14, 2024, 5 pages. [cited by applicant]
Fu, T.-J., et al., “M3L: Language-based Video Editing via Multi-Modal Multi-Level Transformers,” arXiv:2104.01122 [cs.CV], Apr. 2021, pp. 1-12. [cited by applicant]
Hasler, D. et al., “Measuring Colorfulness in Natural Images,” Proceedings of SPIE—The International Society for Optical Engineering 5007, Jun. 2003, pp. 87-95. [cited by applicant]
Heo, M. et al., “VITA: Video Instance Segmentation via Object Token Association,” 36th Conference on Neural Information Processing Systems (NeurIPS 2022)., arXiv:2206.04403v2 [cs.CV], Oct. 2022, pp. 1-14. [cited by applicant]
Hwang, S. et al., “Video Instance Segmentation using Inter-Frame Communication Transformers,” 35th Conference on Neural Information Processing Systems (NeurIPS 2021), Dec. 2021, pp. 1-12. [cited by applicant]
Leake, M. et al., “Computational Video Editing for Dialogue-Driven Scenes,” ACM Transactions on Graphics, vol. 36, No. 4, Article 130, Jul. 2017, pp. 130:1-130: 14. [cited by applicant]
Liu, J. et al., “Violin: A Large-Scale Dataset for Video-and-Language Inference,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, p. 10897-10907. [cited by applicant]
Liu, X. et al., “Learning to Predict Layout-to-image Conditional Convolutions for Semantic Image Synthesis,” 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Oct. 2019, pp. 1-11. [cited by applicant]
Liu, X. et al., “Open-Edit: Open-Domain Image Manipulation With Open-Vocabulary Instructions<” Computer Vision-ECCV 2020: 16th European Conference, Proceedings, Part XI 16, Aug. 23-28, 2020, pp. 1-17. [cited by applicant]
Mejjati, Y.A. et al., “Look here! A parametric learning based approach to redirect visual attention,” Computer Vision—ECCV 2020, 16th European Conference, Proceedings, Part XXIII, Aug. 23-28, 2020, pp. 1-16. [cited by applicant]
NASA Ames Research Center, “Luminance Contrast,” https://colorusage.arc.nasa.gov/luminance_cont.php, retrieved on Mar. 14, 2024, 5 pages. [cited by applicant]
Reimers, N., “SentenceTransformers Documentation,” www.sbert.net, retrieved on Mar. 14, 2024, pp. 1-8. [cited by applicant]
Smith, J., “Calculating Color Temperature and Illuminance using the Taos TCS3414CS Digital Color Sensor,” Intelligent Opto Sensor Designer's Notebook, No. 25, Feb. 2009, pp. 1-7. [cited by applicant]
Tong, H. et al., “Blur Detection for Digital Images Using Wavelet Transform,” 2004 IEEE International Conference on Multimedia and Expo (ICME) (IEEE Cat. No.04TH8763), 2004, pp. 17-20. [cited by applicant]
W3C, “Techniques For Accessibility Evaluation And Repair Tools: W3C Working Draft,” https://www.w3.org/TR/2000/WD-AERT-20000426, Apr. 26, 2000, 58 pages. [cited by applicant]
Wang, Z. et al., “Image Quality Assessment: From Error Visibility to Structural Similarity,” IEEE Transactions on Image Processing, vol. 13, No. 4, Apr. 2004, pp. 600-612. [cited by applicant]
Xu, R. et al., “Closing-the-Loop: A Data-Driven Framework for Effective Video Summarization,” 2020 IEEE International Symposium on Multimedia (ISM), Dec. 2020, pp. 1-8. [cited by applicant]
Xu, W. et al., “Symmetric Regularization based BERT for Pair-wise Semantic Reasoning,” arXiv:1909.03405v3 [cs.CL], Jun. 17, 2021, pp. 1-8. [cited by applicant]
Zhang, Z. et al., “Semantics-Aware BERT for Language Understanding,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, No. 05, Apr. 2020, pp. 9628-9635. [cited by applicant]