IP Library Granted Patent US 12,438,995
Granted Patent B1
US 12,438,995 · App. 18/959,339 · Granted Oct 7, 2025

Integration of video language models with AI for filmmaking

Inventor: Benjamin Geza Affleck-Boldt (West Hollywood, CA)
Assignee: FIN BONE, LLC
H04N5/2224G06F16/7837G06T7/521
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,438,995
App. No.
18/959,339
Granted
Oct 7, 2025
Kind
B1
Abstract

A method integrates video LLMs with AI algorithms for filmmaking by processing filmmaking metadata and Lidar data to simulate professional techniques. A system includes processors and memory to process filmmaking metadata, integrate Lidar data, and enhance video LLMs with advanced filmmaking capabilities. A computer-readable medium contains instructions for adapting video LLMs to generate content simulating professional filmmaking techniques using metadata and Lidar data.

Claims (37)

1. A computer-implemented method for integrating one or more existing video large language models (LLMs) with custom AI algorithms for filmmaking, the method comprising:

interfacing with an existing video LLM;

receiving detailed metadata related to professional filmmaking techniques, including camera settings, shot composition, and lighting setups;

processing the received metadata to adapt the existing video LLM to generate video content that simulates professional filmmaking techniques;

receiving Lidar data captured from a lidar sensor, the Lidar data including spatial coordinates, distance measurements and relative positional information of objects within a scene; and

integrating the Lidar data with the processed metadata by combining the spatial coordinates and distance measurements from the Lidar data with the filmmaking metadata to enhance the generated video content by providing a three-dimensional spatial understanding of the scene indicating depths and positional relationships among objects; and

applying transfer learning techniques to the existing video LLM based on the processed metadata and Lidar data to refine its video content generation capabilities.

2. The method of claim 1 , further comprising fine-tuning the existing video LLM with a dataset enriched with detailed filmmaking metadata and Lidar data to match criteria associated with professional filmmaking criteria.

3. The method of claim 1 , further comprising utilizing specialized interface protocols to enable efficient knowledge transfer between the custom AI algorithms and the existing video LLM.

4. The method of claim 1 , further comprising simulating professional filmmaking techniques within the generated video content based on the processed metadata and integrated Lidar data.

5. The method of claim 1 , further comprising dynamically adjusting the generated video content based on scene changes documented in the metadata to maintain narrative coherence and adhere to professional filmmaking standards.

6. The method of claim 1 , further comprising systematically altering key filmmaking variables in the metadata to simulate an impact of each element on a video output, thereby teaching the existing video LLM to apply these variables effectively.

7. A computing system for enhancing existing video large language models (LLMs) with advanced filmmaking capabilities, the system comprising:

one or more processors; and

one or more memories having stored thereon computer-executable instructions that, when executed, cause the system to:

receive and process metadata related to professional filmmaking techniques including camera settings, shot composition and lighting setups;

receive and integrate Lidar data including three-dimensional positional and distance information of objects within a scene;

interface with an existing video LLM by transmitting the processed metadata and integrated Lidar data as input to the existing video LLM via one or more interface protocols to enable the video LLM to generate video content that includes three-dimensional spatial information and object-position relationships derived from the Lidar data and filmmaking metadata; and

apply transfer learning techniques to the existing video LLM based on the processed metadata and Lidar data to refine its video content generation capabilities.

8. The system of claim 7 , further comprising instructions to apply transfer learning techniques to the existing video LLM based on the processed metadata and integrated Lidar data.

9. The system of claim 7 , further comprising instructions to fine-tune the existing video LLM with a dataset enriched with detailed filmmaking metadata and Lidar data.

10. The system of claim 7 , further comprising instructions to utilize specialized interface protocols to enable efficient knowledge transfer between one or more custom AI algorithms and the existing video LLM.

11. The system of claim 7 , further comprising instructions to simulate professional filmmaking techniques within generated video content based on the processed metadata and integrated Lidar data.

12. The system of claim 7 , further comprising instructions to adjusting generated video content based on scene changes documented in the metadata to maintain narrative coherence and adhere to filmmaking criteria.

13. The system of claim 7 , further comprising instructions to systematically altering key filmmaking variables in the metadata to simulate an impact of each element on a video output, thereby teaching the existing video LLM to apply these variables effectively.

14. A computer-readable medium having stored thereon instructions that when executed by a processor cause a system to perform:

interfacing with an existing video large language model (LLM) by transmitting filmmaking data processed by AI algorithms to the existing video LLM via interface protocols;

receiving detailed metadata related to professional filmmaking techniques including camera settings, shot composition and lighting setups;

processing the received metadata to adapt the existing video LLM to generate video content that simulates professional filmmaking techniques;

receiving Lidar data from a Lidar sensor capturing spatial coordinates and depth information for objects in a scene;

integrating the Lidar data with the processed metadata by combining information from the Lidar data and metadata to provide a three-dimensional representation of object positions and spatial relationships within the generated video content to enhance the generated video content; and

applying transfer learning techniques to the existing video LLM based on the processed metadata and Lidar data to refine its video content generation capabilities.

15. The computer-readable medium of claim 14 , further comprising instructions that cause the system to fine-tune the existing video LLM with a dataset enriched with detailed filmmaking metadata and Lidar data.

16. The computer-readable medium of claim 14 , further comprising instructions that cause the system to simulate professional filmmaking techniques within the generated video content based on the processed metadata and integrated Lidar data.

17. The computer-readable medium of claim 14 , further comprising instructions that cause the system to dynamically adjust the generated video content based on scene changes documented in the metadata to maintain narrative coherence and adhere to professional filmmaking standards.

18. The computer-readable medium of claim 14 , further comprising instructions that cause the system to systematically alter key filmmaking variables in the metadata to simulate an impact of each element on a video output, thereby teaching the existing video LLM to apply these variables effectively.

19. The computer-readable medium of claim 14 , further comprising instructions that cause the system to utilize specialized interface protocols to enable efficient knowledge transfer to the existing video LLM.

Assignments (3)
CHANGE OF NAME Recorded Dec 23, 2025
From: FIN BONE, LLC
To: INTERPOSITIVE, LLC
Reel/Frame 074049/0975 →
CHANGE OF NAME Recorded Nov 24, 2025
From: FIN BONE, LLC
To: INTERPOSITIVE, LLC
Reel/Frame 073322/0142 →
NUNC PRO TUNC ASSIGNMENT Recorded Jul 2, 2025
From: AFFLECK-BOLDT, BENJAMIN GEZA
To: FIN BONE, LLC
Reel/Frame 071595/0698 →
Continuity (1)
Provisional Application 63657756 · Jun 7, 2024
References Cited (44)
US 10645356B1 · Suhy et al. · 2020 [cited by applicant]
US 11288864B2 · George et al. · 2022 [cited by applicant]
US 11398255B1 · Mann et al. · 2022 [cited by applicant]
US 11450053B1 · Georgis et al. · 2022 [cited by applicant]
US 11461963B2 · Manivasagam et al. · 2022 [cited by applicant]
US 11570378B2 · Newman · 2023 [cited by applicant]
US 11830159B1 · Mann et al. · 2023 [cited by applicant]
US 11928799B2 · Chopra et al. · 2024 [cited by applicant]
US 12051205B1 · Deutsch et al. · 2024 [cited by applicant]
US 12133030B2 · Zink et al. · 2024 [cited by applicant]
US 12236517B2 · Bradley et al. · 2025 [cited by applicant]
US 20150286644A1 · Turner et al. · 2015 [cited by applicant]
US 20160071544A1 · Waterston et al. · 2016 [cited by applicant]
US 20160205379A1 · Kurihara · 2016 [cited by examiner]
US 20180124382A1 · Smith · 2018 [cited by examiner]
US 20180136332A1 · Barfield, Jr. et al. · 2018 [cited by applicant]
US 20200082431A1 · Rajasekharan et al. · 2020 [cited by applicant]
US 20200111447A1 · Yaacob et al. · 2020 [cited by applicant]
US 20200364877A1 · Bradski et al. · 2020 [cited by applicant]
US 20220197306A1 · Cella et al. · 2022 [cited by applicant]
US 20230124190A1 · Chui et al. · 2023 [cited by applicant]
US 20230281913A1 · Rematas et al. · 2023 [cited by applicant]
US 20230342481A1 · Nikoghossian et al. · 2023 [cited by applicant]
US 20240005960A1 · Murarka et al. · 2024 [cited by applicant]
US 20240105231A1 · Ratias · 2024 [cited by applicant]
US 20240134926A1 · Tunnicliffe et al. · 2024 [cited by applicant]
US 20240177412A1 · Ranganath et al. · 2024 [cited by applicant]
US 20240290119A1 · Fashandi et al. · 2024 [cited by applicant]
US 20240320918A1 · Amador et al. · 2024 [cited by applicant]
US 20240346731A1 · Graham et al. · 2024 [cited by applicant]
US 20240362897A1 · Klinghoffer · 2024 [cited by examiner]
US 20240394511A1 · Thevenin · 2024 [cited by examiner]
US 20240412542A1 · O'Neill · 2024 [cited by applicant]
US 20240419923A1 · Chollampatt Muhammed Ashraf · 2024 [cited by examiner]
US 20250014606A1 · Wong · 2025 [cited by examiner]
US 20250032945A1 · Mechlowicz · 2025 [cited by applicant]
US 20250148671A1 · Sainz-Nieto · 2025 [cited by applicant]
US 20250190761A1 · Filip et al. · 2025 [cited by applicant]
US 20250204987A1 · Mohareri et al. · 2025 [cited by applicant]
GB 2623644A · 2024 [cited by applicant]
Hong, Wenyi, et al. “Cogvideo: Large-scale pretraining for text-to-video generation via transformers.” arXiv preprint arXiv: 2205.15868 (2022). (Year: 2022). [cited by examiner]
Lin, Han, et al. “Videodirectorgpt: Consistent multi-scene video generation via IIm-guided planning.” arXiv preprint arXiv:2309.15091 (2023). (Year: 2023). [cited by examiner]
Jiang et al., “Example-driven Virtual Cinematography by Learning Camera Behaviors”, ACM Trans. Graph. 39:4 (2020). [cited by applicant]
Roush et al., LLM as an Art Director (LaDi): Using LLM's to improve Text-to-Media Generators, 2023 (Year: 2023). [cited by applicant]