IP Library Patent Application 17832642
Patent Application
App. No. 17/832,642

VIDEO-BASED SEARCH RESULTS WITHIN A COMMUNICATION SESSION

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/832,642
Abstract

Methods and systems provide for video-based search results within a communication session. In one embodiment, the system receives video content of a communication session with a number of participants; extracts, via optical character recognition (“OCR”), textual content from the frames of the video content, each piece of textual content including a timestamp representing a temporal location of the frame within the video content; receives, from a client device associated with a user, a request to search for specified text within the video content; in response to receiving the request, determines one or more matching pieces of textual content which match to the specified text; and presents, to the client device, the matching pieces of textual content.

Claims (43)

1 . A method, comprising:

receiving video content of a communication session between a plurality of participants;

extracting, via optical character recognition (OCR), a plurality of textual content from the frames of the video content, each piece of textual content comprising a timestamp representing a temporal location of the frame within the video content;

receiving, from a client device associated with a user, a request to search for specified text within the video content;

in response to receiving the request, determining one or more matching pieces of textual content which match to the specified text; and

presenting, to the client device, the matching pieces of textual content.

2 . The method of claim 1 , wherein at least a subset of the plurality of textual content comprises one or more titles detected within the frames of the video content.

3 . The method of claim 1 , wherein the specified text within the request comprises at least one of: one or more words, one or more phrases, one or more numbers, and one or more symbols.

4 . The method of claim 1 , wherein determining one or more matching pieces of text comprises determining one or more exact matches with the specified text.

5 . The method of claim 1 , wherein determining one or more matching pieces of text comprises determining one or more exact matches with a spell-corrected version of the specified text.

6 . The method of claim 1 , wherein determining one or more matching pieces of text comprises determining one or more non-exact matches with the specified text.

7 . The method of claim 6 , wherein the non-exact match is based on entity extraction techniques.

8 . The method of claim 6 , wherein the non-exact match is based on relationship embedding techniques.

9 . The method of claim 6 , wherein the non-exact match is based on matching synonyms.

10 . The method of claim 1 , further comprising:

ranking the matching pieces of textual content based on a relevance score; and

wherein the matching pieces of textual content are presented to the client device in order of ranking.

11 . The method of claim 10 , wherein the relevance score is based on one or more of: the specified text, user preferences, user behavior, user search history, and popularity of the matching piece of textual content.

12 . The method of claim 1 , wherein the matching pieces of textual content are presented to the client device in chronological order based on the associated timestamps.

13 . A communication system comprising one or more processors configured to perform the operations of:

receiving video content of a communication session between a plurality of participants;

extracting, via optical character recognition (OCR), a plurality of textual content from the frames of the video content, each piece of textual content comprising a timestamp representing a temporal location within the video content;

receiving, from a client device associated with a user, a request to search for specified text within the video content;

in response to receiving the request, determining one or more matching pieces of textual content which match to the specified text; and

presenting, to the client device, the matching pieces of textual content.

14 . The communication system of claim 13 , wherein presenting the matching pieces of textual content comprises:

presenting the frame associated with each matching piece of textual content, the matching piece of textual content being visually highlighted within the presented frame.

15 . The communication system of claim 13 , wherein presenting the matching pieces of textual content comprises:

presenting the full textual content from the frame associated with each matching piece of textual content, the matching piece of textual content being visually highlighted within the presented full textual content.

16 . The communication system of claim 13 , wherein presenting the matching pieces of textual content comprises:

presenting a subset of the textual content from the frame associated with each matching piece of textual content, the matching piece of textual content being visually highlighted within the presented subset of the textual content.

17 . The communication system of claim 16 , wherein the one or more processors are further configured to perform the operation of:

identifying, from the frame associated with each matching piece of textual content, a contextual portion of the textual content representing a context for the matching piece of textual content within a prespecified threshold distance from the matching piece of textual content,

wherein the presented subset of the textual content is the contextual portion of the textual content.

18 . The communication system of claim 16 , wherein the presented subset is determined based on the available space within a window for presenting the subset.

19 . The communication system of claim 13 , wherein presenting the matching pieces of textual content comprises:

presenting one or more frames associated with the matching pieces of textual content and one or more pieces of textual content associated with the frames, the matching pieces of textual content being visually highlighted within the pieces of textual content associated with the frames.

20 . A non-transitory computer-readable medium containing instructions comprising:

instructions for receiving video content of a communication session between a plurality of participants;

instructions for extracting, via optical character recognition (OCR), a plurality of textual content from the frames of the video content, each piece of textual content comprising a timestamp representing a temporal location within the video content;

instructions for receiving, from a client device associated with a user, a request to search for specified text within the video content;

in response to receiving the request, instructions for determining one or more matching pieces of textual content which match to the specified text; and

instructions for presenting, to the client device, the matching pieces of textual content.

Assignments (2)
CHANGE OF NAME Recorded Jan 7, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 069839/0593 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2022
From: TAO, RENJIE; TSOU, LING
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 060811/0809 →