IP Library › Granted Patent US 12,519,906
Granted Patent B2
US 12,519,906 · App. 18/226,740 · Granted Jan 6, 2026

Mapping video conferencing content to video frames

Inventors: Ajay Jain (San Jose, CA); Sanjeev Tagra (Bothell, WA); Sachin Soni (New Delhi, IN); David Kong (Princeton, NJ)
Assignee: Adobe Inc.
H04N7/155G06V20/49
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,519,906
App. No.
18/226,740
Granted
Jan 6, 2026
Kind
B2
Abstract

Methods and systems for mapping video conferencing content to video frames are provided. In an example method, a processing device receives video conference information and a digital video, the digital video including a plurality of frames. The processing device segments the video conference information into one or more video-conference time segments and the digital video into one or more digital-video time segments. The processing device associates each video-conference time segment with a digital-video time segment. The processing device maps first content information of a first video-conference time segment of the one or more video-conference time segments onto a first digital-video time segment associated with the first video-conference time segment based a first identifier of the first content information. The processing device causes the first content information to be displayed during a displaying of the digital video.

Claims (60)

1 . A method comprising:

receiving video conference information and a digital video, the digital video including a plurality of frames, wherein the video conference information includes a video stream and an audio stream;

segmenting the video conference information into one or more video-conference time segments;

segmenting the digital video into one or more digital-video time segments;

associating each video-conference time segment with a digital-video time segment;

mapping first content information of a first video-conference time segment of the one or more video-conference time segments onto a first digital-video time segment associated with the first video-conference time segment based on a first identifier of the first content information; and

causing the first content information to be displayed during a displaying of the digital video.

2 . The method of claim 1 , wherein:

each video-conference time segment includes content information including the video conference information associated with the video-conference time segment; and

each digital-video time segment includes one or more frames of the plurality of frames of the digital video.

3 . The method of claim 1 , wherein causing the digital video to be displayed comprises:

determining, by a video production server, the first content information using the first identifier of the first content information; and

displaying the digital video, by the video production server.

4 . The method of claim 3 , wherein the first content information displayed during the displaying of the digital video includes at least textual chat messages or transcribed spoken words.

5 . The method of claim 3 , wherein the first content information displayed during the displaying of the digital video is displayed during one or more frames included in the first digital-video time segment.

6 . The method of claim 1 , wherein the video conference information includes at least one of chat message information, reaction information, or annotation information.

7 . The method of claim 1 , further comprising:

identifying, from the audio stream, one or more spoken words, wherein the first content information includes the one or more spoken words; and

identifying, from the video stream, a first frame of the plurality of frames of the digital video.

8 . The method of claim 1 , wherein:

the video conference information is received from a video conferencing platform; and

the digital video is received from a video production server.

9 . The method of claim 1 , further comprising:

identifying the first content information from the first video-conference time segment of the one or more video-conference time segments, comprising:

extracting undesignated content information from the first video-conference time segment;

classifying the undesignated content information using at least one of a textual classifier, a video classifier, or an audio classifier;

designating the undesignated content information as the first content information; and

embedding the first content information in a data structure.

10 . The method of claim 9 , wherein:

at least one of the textual classifier, the video classifier, or the audio classifier is based on a machine learning model trained to classify the undesignated content information;

the machine learning model generates a probability that the undesignated content information is relevant; and

designating the undesignated content information as the first content information s responsive to the probability exceeding a predefined threshold probability.

11 . The method of claim 1 , further comprising:

identifying second content information from the first video-conference time segment;

mapping the second content information of the first video-conference time segment of the one or more video-conference time segments onto the first digital-video time segment associated with the first video-conference time segment based on a second identifier of the second content information; and

grouping the first content information and the second content information on the first digital-video time segment associated with the first video-conference time segment.

12 . A system, comprising:

a memory component; and

one or more processing devices coupled to the memory component configured to perform operations comprising:

receiving, from a segmentation module, time-segmented video conference information, comprising one or more video-conference time segments, wherein the time-segmented video conference information is based on at least a video stream and an audio stream;

receiving one or more frames from a digital video;

determining, for each video-conference time segment, the one or more frames from the digital video that are included in the video-conference time segment;

determining, from content information included in each video-conference time segment, at least one corresponding frame from the one or more frames included in the video-conference time segment; and

storing the correspondence between the content information and the at least one corresponding frame in a memory device, including an identifier of the content information.

13 . The system of claim 12 , wherein the content information included in each video-conference time segment includes at least one of chat message information, reaction information, or annotation information.

14 . The system of claim 12 , wherein the time-segmented video conference information is based on the video stream and at least a portion of the digital video is included in the video stream.

15 . The system of claim 12 , wherein the one or more frames from the digital video are received from a video production server.

16 . A non-transitory computer-readable storage medium storing processor-executable instructions configured to cause one or more processors to perform operations including:

receiving time-segmented video conference information based on a video stream and an audio stream, comprising one or more video-conference time segments;

receiving one or more frames from a digital video;

determining, for each video-conference time segment, the one or more frames from the digital video that are included in the video-conference time segment;

determining, from first content information included in a first video-conference time segment, at least one corresponding frame from the one or more frames included in the first video-conference time segment; and

storing the correspondence between the first content information and the at least one corresponding frame, including a first identifier of the first content information.

17 . The non-transitory computer-readable storage medium of claim 16 , wherein the first content information included in each video-conference time segment includes at least one of chat message information, reaction information, or annotation information.

18 . The non-transitory computer-readable storage medium of claim 16 , wherein the time-segmented video conference information is based on the video stream and at least a portion of the digital video is included in the video stream.

19 . The non-transitory computer-readable storage medium of claim 16 , wherein the one or more frames from the digital video are received from a video production server.

20 . The non-transitory computer-readable storage medium of claim 16 , further comprising:

identifying second content information from the first video-conference time segment;

determining, from the second content information included in the first video-conference time segment, the at least one corresponding frame from the one or more frames included in the first video-conference time segment; and

grouping the first content information and the second content information on the at least one corresponding frame, including the first identifier of the first content information and a second identifier of the second content information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 26, 2023
From: JAIN, AJAY; TAGRA, SANJEEV; SONI, SACHIN; KONG, DAVID
To: ADOBE INC.
Reel/Frame 064395/0524 →
Continuity (1)
Related Publication 20250039335A1 · Jan 30, 2025
References Cited (3)
US 20100306796A1 · McDonald · 2010 [cited by examiner]
US 20230300430A1 · Mishra · 2023 [cited by examiner]
Adobe Blog, “Adobe and Microsoft Announce New, Deeper Integrations to Turbocharge the Modern Workplace” Available online athttps://blog.adobe.com/en/publish/2022/05/24/adobe-microsoft-announce-new-deeper-integrations-to… [cited by applicant]