IP Library › Granted Patent US 11,782,979
Granted Patent B2
US 11,782,979 · App. 17/114,922 · Granted Oct 10, 2023

Method and apparatus for video searches and index construction

Inventors: Yiliang Lyu (Hangzhou, CN); Mingqian Tang (Hangzhou, CN); Zhen Han (Hangzhou, CN); Yulin Pan (Hangzhou, CN)
Assignee: ALIBABA GROUP HOLDING LIMITED
G06F16/73G06F16/71G06F40/30G06V20/41G06V20/49G11B27/10G06V2201/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,782,979
App. No.
17/114,922
Granted
Oct 10, 2023
Kind
B2
Abstract

Embodiments of the disclosure provide methods and apparatuses for video searches and methods and apparatuses for index construction. In one embodiment, the method comprises: upon receiving a search request input by a user to search for a target video, processing, based on a pre-configured algorithm, multimodal search data for the target video included in the search request; providing a processing result of the multimodal search data with regard to a corresponding pre-constructed index to search to obtain the target video.

Claims (69)

1. A method comprising:

receiving a search request input by a user to search for a target video, the search request including multimodal search data for the target video;

obtaining a processing result, utilizing a pre-configured algorithm, based on the multimodal search data, the processing result comprising one or more semantic labels for the multimodal search data generated by the pre-configured algorithm; and

using the processing result to search a corresponding pre-constructed index, the pre-constructed index based on the processing result to obtain the target video.

2. The method of claim 1 , the multimodal search data comprising text data and the obtaining a processing result of the multimodal search data comprising: processing the text data based on a pre-configured text algorithm to obtain a semantic text label of the text data.

3. The method of claim 1 , the multimodal search data comprising image data and the obtaining a processing result of the multimodal search data comprising:

processing the image data based on a pre-configured image algorithm to obtain a semantic image label of the image data; and

processing the image data based on a pre-configured vectorization model to obtain a vectorized description of the image data.

4. The method of claim 1 , the multimodal search data comprising video data, the method further comprising:

processing the video data into video metadata and video stream data; and

segmenting the video stream data into a sequence of video frames based on a pre-configured segmenting manner.

5. The method of claim 4 , the obtaining a processing result of the multimodal search data based on the multimodal search data comprising:

processing the video metadata based on a text algorithm to obtain a semantic text label of the video metadata;

processing video frames in the sequence of video frames based on a pre-configured video algorithm to obtain semantic video labels of the video frames; and

processing the video frames based on a vectorization model to obtain vectorized descriptions of the video frames.

6. The method of claim 2 , the providing the processing result of the multimodal search data with regard to a corresponding pre-constructed index to search to obtain the target video comprising providing a semantic text label of text data with regard to a corresponding pre-constructed inverted index to search to obtain the target video.

7. The method of claim 3 , the providing the processing result of the multimodal search data with regard to a corresponding pre-constructed index to search to obtain the target video comprising:

providing the semantic image label with regard to a corresponding pre-constructed inverted index to search to obtain a first initial video;

providing the vectorized description of the image data with regard to a corresponding pre-constructed vector index to search to obtain a second initial video; and

obtaining the target video based on the first initial video and the second initial video.

8. The method of claim 5 , the providing the processing result of the multimodal search data with regard to a corresponding pre-constructed index to search to obtain the target video comprising:

providing the semantic text label of the video metadata with regard to a corresponding pre-constructed inverted index to search to obtain a third initial video;

providing the vectorized descriptions of the video frames with regard to a corresponding pre-constructed vector index to search to obtain a fourth initial video; and

obtaining the target video based on the third initial video and the fourth initial video.

9. The method of claim 7 , the providing the semantic image label with regard to obtain a first initial video comprising:

combining the semantic image label and the semantic text label of the text data to generate a combined label; and

providing the combined label with regard to the corresponding pre-constructed inverted index to search to obtain the first initial video.

10. The method of claim 8 , the providing the semantic text label of the video metadata with regard to a corresponding pre-constructed inverted index to search to obtain a third initial video comprising:

combining the semantic text label of the video metadata, the semantic text label of the text data, and the semantic video labels of the video frames to generate a combined label; and

providing the combined label with regard to the corresponding pre-constructed inverted index to search to obtain the third initial video.

11. The method of claim 1 , the index being constructed by:

obtaining video data;

processing the video data into video metadata and video stream data;

processing the video metadata based on a pre-configured text algorithm to obtain a text processing result of the video metadata;

processing the video stream data based on a pre-configured video algorithm and vectorization model, respectively, to obtain a video processing result and a vectorization processing result of the video stream data;

constructing an inverted index based on the text processing result and the video processing result; and

constructing a vector index based on the vectorization processing result.

12. An apparatus comprising:

a processor; and

a storage medium for tangibly storing thereon program logic for execution by the processor, the stored program logic comprising:

logic, executed by the processor, for receiving a search request input by a user to search for a target video, the search request including multimodal search data for the target video,

logic, executed by the processor, for obtaining a processing result, utilizing a pre-configured algorithm, based on the multimodal search data, the processing result comprising one or more semantic labels for the multimodal search data generated by the pre-configured algorithm; and

logic, executed by the processor, for using the processing result to search a corresponding pre-constructed index, the pre-constructed index based on the processing result to obtain the target video.

13. The apparatus of claim 12 , the multimodal search data comprising text data; and the logic for obtaining a processing result of the multimodal search data comprising:

logic, executed by the processor, for processing the text data based on a pre-configured text algorithm to obtain a semantic text label of the text data.

14. The apparatus of claim 12 , the multimodal search data image data;

and the logic for obtaining a processing result of the multimodal search data comprising:

logic, executed by the processor, for processing the image data based on a pre-configured image algorithm to obtain a semantic image label of the image data, and

logic, executed by the processor, for processing the image data based on a pre-configured vectorization model to obtain a vectorized description of the image data.

15. The apparatus of claim 12 , the multimodal search data video data; and

the logic for obtaining a processing result of the multimodal search data comprising:

logic, executed by the processor, for processing the video data into video metadata and video stream data, and

logic, executed by the processor, for segmenting the video stream data into a sequence of video frames based on a pre-configured segmenting manner.

16. The apparatus of claim 15 , the logic for obtaining a processing result of the multimodal search data comprising:

logic, executed by the processor, for processing the video metadata based on a text algorithm to obtain a semantic text label of the video metadata,

logic, executed by the processor, for processing video frames in the sequence of video frames based on a pre-configured video algorithm to obtain semantic video labels of the video frames, and

logic, executed by the processor, for processing the video frames based on a vectorization model to obtain vectorized descriptions of the video frames.

17. The apparatus of claim 12 , the logic for providing the processing result of the multimodal search data with regard to a corresponding pre-constructed index to search to obtain the target video comprising: providing a semantic text label of text data with regard to a corresponding pre-constructed inverted index to search to obtain the target video.

18. A non-transitory computer-readable storage medium for tangibly storing computer program instructions capable of being executed by a computer processor, the computer program instructions defining the steps of:

receiving a search request input by a user to search for a target video, the search request including multimodal search data for the target video;

obtaining a processing result, utilizing a pre-configured algorithm, based on the multimodal search data the processing result comprising one or more semantic labels for the multimodal search data generated by the pre-configured algorithm; and

using the processing result to search a corresponding pre-constructed index, the pre-constructed index based on the processing result to obtain the target video.

19. The computer-readable storage medium of claim 17 , the multimodal search data comprising text data and the obtaining a processing result of the multimodal search data comprising: processing the text data based on a pre- configured text algorithm to obtain a semantic text label of the text data.

20. The computer-readable storage medium of claim 17 , the multimodal search data comprising image data and the obtaining a processing result of the multimodal search data comprising:

processing the image data based on a pre-configured image algorithm to obtain a semantic image label of the image data; and

processing the image data based on a pre-configured vectorization model to obtain a vectorized description of the image data.

21. The computer-readable storage medium of claim 17 , the multimodal search data comprising video data, the method further comprising:

processing the video data into video metadata and video stream data; and

segmenting the video stream data into a sequence of video frames based on a pre-configured segmenting manner.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 18, 2020
From: LYU, YILIANG; TANG, MINGQIAN; HAN, ZHEN; PAN, YULIN
To: ALIBABA GROUP HOLDING LIMITED
Reel/Frame 054688/0705 →
Priority Claims (1)
CN 201911398726.4 · Dec 30, 2019 · national
Continuity (1)
Related Publication 20210200802A1 · Jul 1, 2021
Cited By (1)
US 12,536,217