IP Library Granted Patent US 12,430,932
Granted Patent B2
US 12,430,932 · App. 17/832,636 · Granted Sep 30, 2025

Video frame type classification for a communication session

Inventors: Renjie Tao (Sunnyvale, CA); Ling Tsou (Lawndale, CA)
Assignee: Zoom Communications, Inc.
G06V20/63G06V10/82G06V20/41G06V20/46G06V30/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,932
App. No.
17/832,636
Granted
Sep 30, 2025
Kind
B2
Abstract

Methods and systems provide for providing video frame type classification in a communication session. In one embodiment, the system receives video content of a communication session with a number of participants; extracts frames from the video content; classifies the frames of the video content based on image analysis; and transmits, to one or more client devices, the classification of the frames of the video content.

Claims (47)

1. A method, comprising:

receiving video content of a communication session comprising a plurality of participants;

extracting high-resolution versions and low-resolution versions of frames from the video content;

classifying the low-resolution versions of the frames of the video content into black frames, slide frames, and demo frames based on image analysis;

identifying text in the low-resolution versions of the frames;

extracting textual content from the high-resolution versions of the frames; and

transmitting, to one or more client devices, the extracted textual content.

2. The method of claim 1 , wherein the black frames are devoid of content, the slide frames include presentation slides, a face frame that does not include a screen share, and the demo frames include a demonstration.

3. The method of claim 2 , wherein textual content of the face frame is used for one or more of: a sentiment analysis presentation, and an engagement analysis presentation.

4. The method of claim 2 , wherein textual content of the demo frames are used for presentation of an analysis of a product demonstration duration within the communication session.

5. The method of claim 2 , wherein textual content of the demo frame is used for training data for one or more natural language parsing (NLP) tasks.

6. The method of claim 1 , wherein classifying the frames of the video content is performed using a convolutional neural network (CNN).

7. The method of claim 1 , further comprising:

post-processing the frames based on a classification of the frames.

8. The method of claim 7 , wherein the post-processing comprises:

determining a time between two neighboring frames that does not meet a length threshold; and

removing noise between the two neighboring frames.

9. The method of claim 1 , further comprising:

determining one or more differences in a classification in neighboring frames.

10. The method of claim 9 , further comprising:

segmenting the communication session into topic segments based on the determined differences in the classification in neighboring frames.

11. The method of claim 9 , further comprising:

presenting, to the one or more client devices, a visual indication of the differences in the classification in neighboring frames throughout the video content of the communication session.

12. A communication system comprising one or more processors configured to perform operations comprising:

receiving video content of a communication session comprising a plurality of participants;

extracting high-resolution versions and low-resolution versions of frames from the video content;

classifying the low-resolution versions of the frames of the video content into black frames that are devoid of content, slide frames, and demo frames based on image analysis;

identifying text in the low-resolution versions of the frames;

extracting textual content from the high-resolution versions of the frames; and

transmitting, to one or more client devices, the extracted textual content.

13. The communication system of claim 12 , wherein the slide frames include presentation slides and the demo frames include a demonstration.

14. The communication system of claim 12 , wherein the frames include a face frame that does not include a screen share, wherein textual content of the face frame is used for one or more of: a sentiment analysis presentation, and an engagement analysis presentation.

15. The communication system of claim 12 , wherein textual content of the demo frames are used for presentation of an analysis of a product demonstration duration within the communication session.

16. The communication system of claim 12 , wherein textual content of the demo frames are used for training data for one or more natural language parsing (NLP) tasks.

17. The communication system of claim 12 , wherein classifying the frames of the video content is performed using a convolutional neural network (CNN).

18. The communication system of claim 12 , further comprising:

post-processing the frames based on a classification of the frames.

19. The communication system of claim 18 , wherein the post-processing comprises:

determining a time between two neighboring frames that does not meet a length threshold; and

removing noise between the two neighboring frames.

20. A non-transitory computer-readable medium containing instructions that when executed by a processor, cause the processor to perform operations comprising:

receiving video content of a communication session comprising a plurality of participants;

extracting high-resolution versions and low-resolution versions of frames from the video content;

classifying the low-resolution versions of the frames of the video content into black frames that are devoid of content, slide frames, and demo frames based on image analysis;

identifying text in the low-resolution versions of the frames;

extracting textual content from the high-resolution versions of the frames; and

transmitting, to one or more client devices, the extracted textual content.

Assignments (2)
CHANGE OF NAME Recorded Jan 7, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 069839/0593 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2022
From: TAO, RENJIE; TSOU, LING
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 060811/0784 →
Continuity (1)
Related Publication 20230394851A1 · Dec 7, 2023
References Cited (26)
US 5852435A · Vigneaux · 1998 [cited by examiner]
US 11188746B1 · Patel et al. · 2021 [cited by applicant]
US 11580737B1 · Miller-Smith et al. · 2023 [cited by applicant]
US 11849196B2 · Parmar et al. · 2023 [cited by applicant]
US 11915429B2 · Vartakavi et al. · 2024 [cited by applicant]
US 20020145622A1 · Kauffman · 2002 [cited by examiner]
US 20040216173A1 · Horoszowski · 2004 [cited by examiner]
US 20090116695A1 · Anchyshkin et al. · 2009 [cited by applicant]
US 20100164836A1 · Liberatore · 2010 [cited by applicant]
US 20100165081A1 · Jung et al. · 2010 [cited by applicant]
US 20110081075A1 · Adcock · 2011 [cited by examiner]
US 20140099034A1 · Rafati et al. · 2014 [cited by applicant]
US 20160182757A1 · Yoo · 2016 [cited by applicant]
US 20160247024A1 · Loui et al. · 2016 [cited by applicant]
US 20190384965A1 · Rodriguez et al. · 2019 [cited by applicant]
US 20210076105A1 · Parmar · 2021 [cited by examiner]
US 20210210097A1 · Diamant · 2021 [cited by examiner]
US 20220051011A1 · Patel et al. · 2022 [cited by applicant]
US 20230394854A1 · Polavaram et al. · 2023 [cited by applicant]
CN 114494951A · 2022 [cited by examiner]
WO 2015073501A2 · 2015 [cited by applicant]
WO 2021051024A1 · 2021 [cited by applicant]
WO 2022031283A1 · 2022 [cited by applicant]
Zhu, Xingquan, et al. “ClassMiner: Mining Medical Video Content Structure and Events Towards Efficient Access and Scalable Skimming.” DMKD. 2002. (Year: 2002). [cited by examiner]
International Search Report and Written Opinion mailed on Sep. 12, 2023 in corresponding PCT Application No. PCT/US2023/024304. [cited by applicant]
Honglin Li et al: “Hierarchical Segmentation of Presentation Videos through Visual and Text Analysis”, Signal Processing and Information Technology, 2006 IEEE International Symposium On, IEEE, PI, Aug. 1, 2006 (Aug. 1, … [cited by applicant]