IP Library Granted Patent US 12,725,604
Granted Patent B2
US 12,725,604 · App. 18/429,124 · Granted Sep 1, 2026

Video scene describer

Inventors: Oron Nir (Herzliya, IL); Shemer Shmuel Steinlauf (Tel-Aviv, IL); Eliyahu Strugo (Tel Aviv, IL)
Assignee: Microsoft Technology Licensing, LLC
G10L13/02G06V20/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,604
App. No.
18/429,124
Granted
Sep 1, 2026
Kind
B2
Abstract

Examples of the present disclosure describe a video scene describer. The video scene describer receives video content data as input and provides audio description (AD) data as output. The video scene describer utilizes one or more components using artificial intelligence (AI) and/or algorithms to analyze and describe the video content data. For example, the video scene describer may include a video indexer component to identify and describe particular aspects of the video content data and generates video insights data based on the analysis. The video indexer provides video insights data to a large language model (LLM) component. The video scene describer may additionally include a visual-language model system, which includes a visual encoder, a relation aggregator, a transformer encoder, and/or transformer. The LLM component synthesizes the video insights data and video embedding data, along with any prompt (e.g., a request or question) or dialogue context, to provide the AD data.

Claims (50)

1 . A system comprising:

a processing system; and

memory comprising executable instructions that when executed, perform operations, comprising:

receiving, by a visual-language model system, video content data;

generating, by the visual-language model system, video embedding data based at least in part on the video content data;

receiving video insights data comprising speech-to-text (STT) data, optical character recognition (OCR) data, and facial recognition data;

providing the video embedding data and the video insights data to a large language model (LLM) component;

receiving, from the LLM component, audio description (AD) data based at least in part on the video embedding data and the video insights data, wherein the AD data is received from the LLM component in an audio format and describes at least one insight of the video insights data; and

providing the AD data to a device.

2 . The system of claim 1 , wherein the LLM component generates the AD data based on an auto-recursive algorithm, the operations further comprising:

generating at least one previous AD data, wherein at least one of the video embedding data, the video insights data, or the AD data correspond to a given shot, and wherein the at least one previous AD data corresponds to at least one previous shot before the given shot.

3 . The system of claim 2 , wherein generating the AD data using the auto-recursive algorithm comprises generating the AD data based on at least one of the video embedding data, the video insights data, or the at least one previous AD data.

4 . The system of claim 1 , wherein the video embedding data comprises at least one of audio embeddings or RGB embeddings.

5 . The system of claim 1 , wherein the video insights data further comprises at least one of transcripts, objects, clothing, age, gender, emotion, or landmarks.

6 . The system of claim 1 , wherein the visual-language model system comprises at least one of a visual encoder, a transformer encoder, a relation aggregator, or a transformer.

7 . The system of claim 6 , wherein the transformer encoder uses cross attention to combine visual and language data of video content data into a unified video embedding.

8 . The system of claim 1 , wherein the video embedding data comprises at least one vector of numbers.

9 . The system of claim 1 , wherein generating the AD data comprises concatenating the video embedding data with the video insights data using a plurality of delimiters.

10 . The system of claim 1 , wherein generating the AD data comprises performing cross-attention on the video embedding data and the video insights data.

11 . The system of claim 1 , wherein the AD data comprises a textual description of at least one of:

audio elements of the video content data;

visual elements of the video content data;

explicit elements of the video content data; or

implicit elements of the video content data.

12 . A system comprising:

a processing system; and

memory comprising executable instructions that when executed, perform operations, comprising:

receiving, by a visual-language model system, video content data;

providing, to a large language model (LLM) component, video embedding data of the video content data;

receiving, from a video indexer component, video insights data of the video content data;

providing, to the LLM component, input data comprising the video embedding data, the video insights data, and first audio description (AD) data of the video content data;

receiving, from the LLM component, second AD data of the video content data based at least in part on the input data, wherein the second AD data is received from the LLM component in an audio format and describes at least one insight of the video insights data; and

providing the second AD data to a device.

13 . The system of claim 12 , wherein the at least one of the video embedding data, the video insights data, or the second AD data correspond to a given shot, and wherein the first AD data corresponds to at least one previous shot before the given shot.

14 . The system of claim 12 , wherein the video embedding data, the video insights data, and the second AD data correspond to a given frame, and wherein the first AD data corresponds to previous frames before the given frame.

15 . The system of claim 12 , the operations further comprising:

providing the second AD data to a narrator tool.

16 . A system comprising:

a processing system; and

memory comprising executable instructions that when executed, perform operations, comprising:

receiving, by a visual-language model system, video content data;

generating, by the visual-language model system, video embedding data based on the video content data;

receiving, from a video indexer component, video insights data based on the video content data;

providing, to a large language model (LLM) component, the video embedding data and the video insights data;

creating concatenated data by concatenating, by the LLM component, the video embedding data with the video insights data using a plurality of delimiters; and

providing, by the LLM component, audio description (AD) data for presentation based on the concatenated data, wherein the AD data is received from the LLM component in an audio format and describes at least one insight of the video insights data.

17 . The system of claim 16 , wherein the plurality of delimiters indicate the separation of the video insights data from the video embedding data.

18 . The system of claim 16 , wherein the concatenating further comprises concatenating at least one previous AD data with the video embedding data and the video insights data.

19 . The system of claim 16 , wherein the video insights data comprises at least one of optical character recognition (OCR) data, speech-to-text (STT) data, audio effects data, emotion data, keywords, object tracking data, or topics inference data.

20 . The system of claim 16 , wherein the video embedding data comprises at least one of audio embeddings or RGB embeddings.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2024
From: NIR, ORON; STEINLAUF, SHEMER SHMUEL; STRUGO, ELIYAHU
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 066321/0632 →
Continuity (1)
Related Publication 20250246176A1 · Jul 31, 2025
References Cited (28)
US 6877134B1 · Fuller · 2005 [cited by examiner]
US 10560734B2 · Jassin et al. · 2020 [cited by applicant]
US 10762375B2 · Ronen et al. · 2020 [cited by applicant]
US 10936630B2 · Ronen et al. · 2021 [cited by applicant]
US 11363084B1 · Mishra · 2022 [cited by examiner]
US 12518740B1 · Carre · 2026 [cited by examiner]
US 20130067333A1 · Brenneman · 2013 [cited by examiner]
US 20180132011A1 · Shichman · 2018 [cited by examiner]
US 20180356893A1 · Soni · 2018 [cited by examiner]
US 20210174146A1 · Nir et al. · 2021 [cited by applicant]
US 20210303924A1 · Sakthivel · 2021 [cited by examiner]
US 20210383171A1 · Lee · 2021 [cited by examiner]
US 20220366131A1 · Ekron · 2022 [cited by examiner]
US 20230306056A1 · Lee · 2023 [cited by examiner]
US 20240045902A1 · Park · 2024 [cited by examiner]
US 20240127804A1 · Shirodkar · 2024 [cited by examiner]
US 20240212249A1 · Ume · 2024 [cited by examiner]
US 20250119625A1 · Bhattacharyya · 2025 [cited by examiner]
US 20250156567A1 · Belgi · 2025 [cited by examiner]
Extended European Search Report Received in European Patent Application No. 25150369.4, mailed on May 27, 2025, 10 pages. [cited by applicant]
Han, et al., “AutoAD II: The Sequel—Who, When, and What in Movie Audio Description”, Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 1, 2023, pp. 13599-13609. [cited by applicant]
Zhang, et al., “MM-Narrator: Narrating Long-form Videos with Multimodal In-Context Learning”, In Repository of arXiv:2311.17435v1, Nov. 29, 2023, 17 Pages. [cited by applicant]
Li, et al., “Blip-2: Bootstrapping Language-image Pre-training with Frozen Image Encoders and Large Language Models”, lin Repository of arXiv:2301.12597v3, Jun. 15, 2023, 13 Pages. [cited by applicant]
Li, et al., “Uniformerv2: Spatiotemporal Learning by Arming Image Vits with Video Uniformer”, In Repository of arXiv:2211.09552v1, Nov. 17, 2023, pp. 1-24. [cited by applicant]
Li, et al., “Videochat: Chat-Centric Video Understanding”, In Repository of arXiv:2305.06355v1, May 10, 2023, pp. 1-17. [cited by applicant]
Nir, et al., “CAST: Character labeling in Animation using Self-supervision by Tracking”, In Proceedings of Computer Graphics Forum, vol. 41, Issue 2, May 2022, pp. 135-145. [cited by applicant]
Touvron, et al., “Llama 2: Open Foundation and Fine-Tuned Chat Models”, In Repository of arXiv:2307.09288v2, Jul. 19, 2023, pp. 1-77. [cited by applicant]
Wang, et al., “InternVideo: General Video Foundation Models via Generative and Discriminative Learning”, In Repository of arXiv:2212.03191v2, Dec. 7, 2022, pp. 1-18. [cited by applicant]