IP Library Patent Application 18937782
Patent Application
App. No. 18/937,782

Video-Based And Transcript-Based Segmentation Of Communication Session Content

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/937,782
Abstract

Methods and systems provide for video-based and transcript-based segmentation of communication session content. The method may include obtaining a transcript associated with video content of a communication session and performing video-based segmentation on the video content to determine a category label from a list of category labels for each video frame of the video content. The video content may include topic segments of consecutive frames associated with a same category label. The method may further include performing transcript-based segmentation to divide one of the topic segments into a first topic segment associated with a first title and a second topic segment associated with a second title.

Claims (47)

1 . A method, comprising:

obtaining a transcript associated with video content of a communication session;

performing video-based segmentation on the video content to determine a category label from a list of category labels for each video frame of the video content, the video content comprising topic segments of consecutive frames associated with a same category label; and

performing transcript-based segmentation to divide one of the topic segments into a first topic segment associated with a first title and a second topic segment associated with a second title.

2 . The method of claim 1 , further comprising:

receiving, from a client device, a request to search for specified text within the video content;

determining, based on the request, one or more topic segments related to the request; and

presenting, to the client device, content from at least one of the one or more topic segments related to the request.

3 . The method of claim 2 , further comprising:

presenting, to the client device, a portion of the transcript associated with the at least one of the one or more topic segments related to the request.

4 . The method of claim 1 , further comprising:

presenting, to a client device associated with a user, the video content and a timeline associated with the video content, wherein the timeline is visually segmented into the topic segments.

5 . The method of claim 1 , further comprising:

determining, during the video-based segmentation, a title for at least one of the topic segments.

6 . The method of claim 5 , wherein the title is extracted using optical character recognition.

7 . The method of claim 1 , wherein the performing transcript-based segmentation comprises determining the one of the topic segments is longer than a preset threshold.

8 . The method of claim 1 , wherein the performing transcript-based segmentation comprises determining the first title by extracting text from the transcript.

9 . The method of claim 1 , wherein the performing transcript-based segmentation comprises determining the first title from a list of title candidates.

10 . The method of claim 1 , further comprising:

merging two of the topic segments into a third topic segment, wherein the third topic segment comprises frames associated with a first category label and frames associated with a second category label.

11 . An apparatus, comprising:

one or more processors configured to execute instructions to:

obtain a transcript associated with video content of a communication session;

perform video-based segmentation on the video content to determine a category label from a list of category labels for each video frame of the video content, the video content comprising topic segments of consecutive frames associated with a same category label; and

perform transcript-based segmentation to divide one of the topic segments into a first topic segment associated with a first title and a second topic segment associated with a second title.

12 . The apparatus of claim 11 , wherein the one or more processors are further configured to execute instructions to:

receive, from a client device, a request to search for specified text within the video content;

determine, based on the request, one or more topic segments related to the request; and

present, to the client device, content from at least one of the one or more topic segments related to the request.

13 . The apparatus of claim 12 , wherein the one or more processors are further configured to execute instructions to:

present, to the client device, a portion of the transcript associated with the at least one of the one or more topic segments related to the request.

14 . The apparatus of claim 11 , wherein the one or more processors are further configured to execute instructions to:

determine, during the video-based segmentation, a title for at least one of the topic segments.

15 . The apparatus of claim 11 , wherein the instructions to perform transcript-based segmentation comprise instructions to determine the first title by extracting text from the transcript.

16 . The apparatus of claim 11 , wherein the instructions to perform transcript-based segmentation comprise instructions to determine the first title from a list of title candidates.

17 . The apparatus of claim 11 , wherein the one or more processors are further configured to execute instructions to:

merge two of the topic segments into a third topic segment, wherein the third topic segment comprises frames associated with a first category label and frames associated with a second category label.

18 . One or more non-transitory computer readable medium storing instructions operable to cause one or more processors to perform operations comprising:

obtaining a transcript associated with video content of a communication session;

performing video-based segmentation on the video content to determine a category label from a list of category labels for each video frame of the video content, the video content comprising topic segments of consecutive frames associated with a same category label; and

performing transcript-based segmentation to divide one of the topic segments into a first topic segment associated with a first title and a second topic segment associated with a second title.

19 . The one or more non-transitory computer readable medium of claim 18 , further comprising:

receiving, from a client device, a request to search for specified text within the video content;

determining, based on the request, one or more topic segments related to the request; and

presenting, to the client device, content from at least one of the one or more topic segments related to the request.

20 . The one or more non-transitory computer readable medium of claim 19 , further comprising:

presenting, to the client device, a portion of the transcript associated with the at least one of the one or more topic segments related to the request.

Assignments (2)
CHANGE OF NAME Recorded Jan 7, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 069839/0593 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 5, 2024
From: MILLER-SMITH, ANDREW; TAO, RENJIE; TSOU, LING
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 069148/0262 →