IP Library › Granted Patent US 11,748,987
Granted Patent B2
US 11,748,987 · App. 17/234,783 · Granted Sep 5, 2023

Method and system for performing content-aware deduplication of video files

Inventor: John V Missale (Secaucus, NJ)
Assignee: Larsen & Toubro Infotech Ltd
G06V20/46G06F16/71G06F16/7837G06F16/7847G06F18/22G06N7/01G06V10/74G06V20/41G06V30/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,748,987
App. No.
17/234,783
Granted
Sep 5, 2023
Kind
B2
Abstract

The invention relates to a method and system for performing content-aware deduplication of video files and content storage cost optimization. The method includes pre-processing video files into a plurality of groups of video files based on type of genre and run-time of a video. The genre of a plurality of video files is automatically detected using a sliding-window similarity index, which is utilized to improve accuracy of genre detection. After the pre-processing step, each group of the plurality of groups of video files are simultaneously fed into a plurality of machine learning (ML) instances and models which measure a degree of similarity corresponding to each group of video files by detecting one or more conditions that exists in the video files. The one or more conditions are detected by performing deep inspection of content in the video files using hash-based active recognition of objects.

Claims (49)

1. A method for performing content-aware deduplication of video files, the method comprising:

pre-processing video files into a plurality of groups of video files based on type of genre and run-time of a video, wherein the pre-processing comprises automatically detecting genre of a plurality of video files using a sliding-window similarity index, wherein the sliding-window similarity index is utilized to improve accuracy of genre detection;

feeding each group of the plurality of groups of video files simultaneously into a plurality of machine learning (ML) instances and models; and

measuring, by the plurality of ML instances and models, a degree of similarity corresponding to each group of video files by detecting at least one condition that exists in the video files, wherein the detecting comprises performing deep inspection of content in the video files using hash-based active recognition of objects.

2. The method as claimed in claim 1 , wherein a video file comprises a movie and the genre is at least one of Drama, Horror and Western.

3. The method as claimed in claim 1 , wherein the genre is automatically detected using Multi-label Logistic Regression.

4. The method as claimed in claim 1 , wherein the measuring further comprises dynamically fine-tuning a threshold of the plurality of ML instances and models for detecting the at least one condition for each type of genre.

5. The method as claimed in claim 1 , wherein a condition is at least one of additional audio content, different languages, different textual captions, recording with different encoding equipment, different frame rates and resolutions, different scene environmental locations, different scene order, different intent, blurred background, deleted frames, inserted frames, background hidden by the addition of objects, scenes with different spectral composition, different amounts of participation of a celebrity or object, and different background audio.

6. The method as claimed in claim 1 , wherein a degree of similarity is measured based on time code start and end points using metadata to detect the at least one condition and visually verifying the detected at least one condition.

7. The method as claimed in claim 5 , wherein a degree of similarity is measured based on additional audio data using audio fingerprinting, decoding and similarity.

8. The method as claimed in claim 5 , wherein a degree of similarity is measured based on different languages using Optical Character Recognition (OCR).

9. The method as claimed in claim 5 , wherein a degree of similarity is measured based on different textual captions by: detecting the text using OCR, vectorizing the detected text, and comparing the vectorized text using cosine similarity.

10. The method as claimed in claim 5 , wherein a degree of similarity is measured on video files that are recorded with different encoding equipment using metadata.

11. The method as claimed in claim 5 , wherein a degree of similarity is measured based on different frame rates and resolutions using metadata.

12. The method as claimed in claim 1 , wherein the deep inspection of content comprises detecting scenes with celebrities, objects, captions, language, and perceptual differences in the video files.

13. The method as claimed in claim 1 , wherein the deep inspection of content comprises automatically detecting and removing artifacts in a video file, wherein artifacts comprise at least one of black frames, color bars, countdown slates and any abnormalities that may cause visual degradation in video quality.

14. The method as claimed in claim 1 , wherein the deep inspection of content comprises extracting metadata from a video file and writing the metadata back to a Media Asset Management (MAM) system to improve the descriptive taxonomy and search capability of the MAM system.

15. The method as claimed in claim 1 , wherein the deep inspection of content is used for efficient and automatic content identification and verification across a content supply chain to greatly improve the identification and performance of video content of a video file.

16. The method as claimed in claim 1 , wherein the deep inspection of content comprises verifying if any inserted content in a video file has up to date usage rights or whether additional rights need to be obtained by a content provider for distribution.

17. The method as claimed in claim 1 , wherein the deep inspection of content comprises detecting and classifying disaster conditions in live video in the video files to trigger specific first responders' attention.

18. The method as claimed in claim 1 , wherein the deep inspection of content comprises detecting semantic conditions in the video files, wherein a semantic condition comprises at least one of emotion and behavior.

19. The method as claimed in claim 1 further comprises computing hashes for detecting the degree of similarity based on Hamming Distance using md5 File Hash, wherein the hashes are recorded on the Blockchain to prevent black box attacks using Generative Adversarial Networks (GANs).

20. The method as claimed in claim 1 , wherein the content-aware deduplication of video files achieves content storage cost optimization.

21. The method as claimed in claim 20 , wherein the content storage cost optimization comprises organizing content maintenance for unorganized content by separating said content based on at least one category of the video files and detecting original video files from a given set of video files, wherein a category is at least one of movies, episodes/serials, trailers, user generated content, video blogs/video logs (vlogs), wildlife films, and advertisements (ads).

22. A system for performing content-aware deduplication of video files, the system comprising:

a memory;

a processor communicatively coupled to the memory, the processor configured to:

pre-process video files into a plurality of groups of video files based on type of genre and run-time of a video, wherein the processor is configured to automatically detect genre of a plurality of video files using a sliding-window similarity index, wherein the sliding-window similarity index is utilized to improve accuracy of genre detection;

feed each group of the plurality of groups of video files simultaneously into a plurality of machine learning (ML) instances and models; and

measure, by the plurality of ML instances and models, a degree of similarity corresponding to each group of video files by detecting at least one condition that exists in the video files, wherein the processor is configured to perform deep inspection of content in the video files using hash-based active recognition of objects.

23. The system as claimed in claim 22 , wherein the genre is automatically detected using Multi-label Logistic Regression.

24. The system as claimed in claim 22 , wherein the processor is further configured to dynamically fine-tune a threshold of the plurality of ML instances and models for detecting the at least one condition for each type of genre.

25. The system as claimed in claim 22 , wherein a condition is at least one of additional audio content, different languages, different textual captions, recording with different encoding equipment, different frame rates and resolutions, different scene environmental locations, different scene order, different intent, blurred background, deleted frames, inserted frames, background hidden by the addition of objects, scenes with different spectral composition, different amounts of participation of a celebrity or object, and different background audio.

26. The system as claimed in claim 22 , wherein a degree of similarity is measured based on time code start and end points using metadata to detect the at least one condition and visually verifying the detected at least one condition.

27. The system as claimed in claim 25 , wherein a degree of similarity is measured based on additional audio data using audio fingerprinting, decoding and similarity.

28. The system as claimed in claim 25 , wherein a degree of similarity is measured based on different languages using Optical Character Recognition (OCR).

29. The system as claimed in claim 25 , wherein a degree of similarity is measured based on different textual captions by: detecting the text using OCR, vectorizing the detected text, and comparing the vectorized text using cosine similarity.

30. The system as claimed in claim 25 , wherein a degree of similarity is measured on video files that are recorded with different encoding equipment using metadata.

31. The system as claimed in claim 25 , wherein a degree of similarity is measured based on different frame rates and resolutions using metadata.

32. The system as claimed in claim 22 , wherein the processor is configured to detect scenes with celebrities, objects, captions, language, and perceptual differences in the video files.

33. The system as claimed in claim 22 , wherein the processor is configured to automatically detect and remove artifacts in a video file, wherein artifacts comprise at least one of black frames, color bars, countdown slates and any abnormalities that may cause visual degradation in video quality.

34. The system as claimed in claim 22 , wherein the processor is configured to extract metadata from a video file and write the metadata back to a Media Asset Management (MAM) system to improve the descriptive taxonomy and search capability of the MAM system.

35. The system as claimed in claim 22 , wherein the deep inspection of content is used for efficient and automatic content identification and verification across a content supply chain to greatly improve the identification and performance of video content of a video file.

36. The system as claimed in claim 22 , wherein the processor is configured to verify if any inserted content in a video file has up to date usage rights or whether additional rights need to be obtained by a content provider for distribution.

37. The system as claimed in claim 22 , wherein the processor is configured to detect and classify disaster conditions in live video in the plurality of video files to trigger specific first responders' attention.

38. The system as claimed in claim 22 , wherein the processor is configured to detect semantic conditions in the plurality of video files, wherein a semantic condition comprises at least one of emotion and behavior.

39. The system as claimed in claim 22 , wherein the processor is further configured to compute hashes for detecting the degree of similarity based on Hamming Distance using md5 File Hash, wherein the hashes are recorded on the Blockchain to prevent black box attacks using Generative Adversarial Networks (GANs).

40. The system as claimed in claim 22 , wherein the content-aware deduplication of video files achieves content storage cost optimization.

41. The system as claimed in claim 40 , wherein the processor is configured to organize content maintenance for unorganized content by separating said content based on at least one category of the video files and detect original video files from a given set of video files, wherein a category is at least one of movies, episodes/serials, trailers, user generated content, video blogs/video logs (vlogs), wildlife films, and advertisements (ads).

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 19, 2021
From: MISSALE, JOHN V
To: LARSEN & TOUBRO INFOTECH LTD
Reel/Frame 055964/0512 →
Continuity (1)
Related Publication 20220335245A1 · Oct 20, 2022