IP Library Granted Patent US 11,615,622
Granted Patent B2
US 11,615,622 · App. 16/929,250 · Granted Mar 28, 2023

Systems, methods, and devices for determining an introduction portion in a video program

Inventors: Ehsan Younessian (Washington, DC); Faisal Ishtiaq (Plainfield, IL); Anthony Braskich (Kildeer, IL); Benjamin Spier (Silver Spring, MD)
Assignee: Comcast Cable Communications, LLC
G06V20/48G06N20/00G06V20/41H04N21/4394H04N21/442H04N21/44008H04N21/4665
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,615,622
App. No.
16/929,250
Granted
Mar 28, 2023
Kind
B2
Abstract

Systems, methods, and devices relating to determining an introduction portion in a video program are described herein. A method may determine first and second hard-matching pairs of video segments in first and second video content such that video fingerprints of the first hard-matching pair match and video fingerprints of the second hard-matching pair also match. The method may classify a third pair of video segments in the first and second video content, sequentially between the first and second hard-matching pairs, as a soft-matching pair of video segments of an introduction portion. The method may use the classification of the third pair of video segments as a soft-matching pair to determine a model configured to determine that a pair of video segments in two video content items are a soft-matching pair of video segments of an introduction portion.

Claims (76)

1. A method comprising:

receiving first video content comprising video segments and second video content comprising video segments, wherein the second video content is associated with the first video content;

determining a first hard-matching pair of video segments of the first video content and the second video content, wherein video fingerprints of the first hard-matching pair of video segments match;

determining a second hard-matching pair of video segments in the first video content and the second video content, wherein video fingerprints of the second hard-matching pair of video segments match;

classifying a third pair of video segments in the first video content and the second video content as a soft-matching pair of video segments of an introduction portion of at least one of the first video content or the second video content, wherein the third pair of video segments is sequentially between the first hard-matching pair of video segments and the second hard-matching pair of video segments, wherein video fingerprints of the third pair of video segments do not match; and

determining, based on the classifying the third pair of video segments as a soft-matching pair of video segments of an introduction portion of at least one of the first video content or the second video content, a model configured to determine that a pair of video segments in two video content items are a soft-matching pair of video segments of an introduction portion of at least one of the two video content items.

2. The method of claim 1 , wherein the first video content comprises at least a portion of a first episode of a video program series and the second video content comprises at least a portion of a second episode of the video program series.

3. The method of claim 1 , wherein the first video content comprises target video content in which the introduction portion is not known and the second video content comprises reference video content in which the introduction portion is known.

4. The method of claim 1 , further comprising:

determining the model via machine learning, wherein a training data input for the machine learning comprises the video fingerprints of the third pair of video segments, and a training data output for the machine learning comprises the classification of the third pair of video segments as a soft-matching pair of video segments of an introduction portion of at least one of the first video content or the second video content.

5. The method of claim 4 , wherein the model comprises a regressor model and a training data input for determining the regressor model comprises a difference between the video fingerprints of the third pair of video segments.

6. The method of claim 1 , wherein:

a difference between lengths of the first hard-matching pair of video segments satisfies a length threshold, and

a difference between lengths of the second hard-matching pair of video segments satisfies the length threshold.

7. The method of claim 6 , wherein a difference between lengths of the third pair of video segments does not satisfy the length threshold.

8. The method of claim 1 , wherein the video segments of the first video content comprise respective shots in the first video content and the video segments of the second video content comprise respective shots in the second video content.

9. The method of claim 1 , wherein a video fingerprint of a video segment comprises an alphanumeric value, and a matching pair of video fingerprints each comprise the same alphanumeric value.

10. A method comprising:

determining one or more soft-matching pairs of video segments among a plurality of video content items, wherein each of the one or more soft-matching pairs of video segments comprises a first video segment of one of the plurality of video content items and a second video segment of a different one of the plurality of video content items, wherein a characteristic of the first and second video segments of each soft-matching pair does not match, and wherein each of the one or more soft matching pairs of video segments is located within the corresponding video content items between two hard- matching pairs of video segments of the video content items; and

determining, based on the determining the one or more soft-matching pairs of video segments, a model configured to determine that a pair of video segments comprises common video content.

11. The method of claim 10 , wherein the plurality of video contents items comprises different episodes of one or more video programs.

12. The method of claim 11 , wherein the first video segment and the second video segment of each soft-matching pair are associated with two episodes of a same video program.

13. The method of claim 10 , wherein the first characteristic comprises audio elements, an audio fingerprint, closed captioning data, subtitle data, on-screen text, or a detected visual feature.

14. The method of claim 10 , wherein common video content comprises at least one of an introduction portion, a closing portion, or an advertisement.

15. A device comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors cause the device to:

receive first video content comprising video segments and second video content comprising video segments, wherein the second video content is associated with the first video content;

determine a first hard-matching pair of video segments of the first video content and the second video content, wherein video fingerprints of the first hard-matching pair of video segments match;

determine a second hard-matching pair of video segments in the first video content and the second video content, wherein video fingerprints of the second hard-matching pair of video segments match;

classify a third pair of video segments in the first video content and the second video content as a soft-matching pair of video segments of an introduction portion of at least one of the first video content or the second video content, wherein the third pair of video segments is sequentially between the first hard-matching pair of video segments and the second hard-matching pair of video segments, wherein video fingerprints of the third pair of video segments do not match; and

determine, based on the classifying the third pair of video segments as a soft-matching pair of video segments of an introduction portion of at least one of the first video content or the second video content, a model configured to determine that a pair of video segments in two video content items are a soft-matching pair of video segments of an introduction portion of at least one of the two video content items.

16. The device of claim 15 , wherein the first video content comprises at least a portion of a first episode of a video program series and the second video content comprises at least a portion of a second episode of the video program series.

17. The device of claim 15 , wherein the first video content comprises target video content in which the introduction portion is not known and the second video content comprises reference video content in which the introduction portion is known.

18. The device of claim 15 , wherein the instructions, when executed by the one or more processors, further cause the device to:

determine the model via machine learning, wherein a training data input for the machine learning comprises the video fingerprints of the third pair of video segments, and a training data output for the machine learning comprises the classification of the third pair of video segments as a soft-matching pair of video segments of an introduction portion of at least one of the first video content or the second video content.

19. The device of claim 18 , wherein the model comprises a regressor model and a training data input for determining the regressor model comprises a difference between the video fingerprints of the third pair of video segments.

20. The device of claim 15 , wherein:

a difference between lengths of the first hard-matching pair of video segments satisfies a length threshold, and

a difference between lengths of the second hard-matching pair of video segments satisfies the length threshold.

21. The device of claim 20 , wherein a difference between lengths of the third pair of video segments does not satisfy the length threshold.

22. The device of claim 15 , wherein the video segments of the first video content comprise respective shots in the first video content and the video segments of the second video content comprise respective shots in the second video content.

23. The device of claim 15 , wherein a video fingerprint of a video segment comprises an alphanumeric value, and a matching pair of video fingerprints each comprise the same alphanumeric value.

24. A device comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors cause the device to:

determine one or more soft-matching pairs of video segments among a plurality of video content items, wherein each of the one or more soft-matching pairs of video segments comprises a first video segment of one of the plurality of video content items and a second video segment of a different one of the plurality of video content items, wherein a characteristic of the first and second video segments of each soft-matching pair does not match, and wherein each of the one or more soft matching pairs of video segments is located within the corresponding video content items between two hard- matching pairs of video segments of the video content items; and

determine, based on the determining the one or more soft-matching pairs of video segments, a model configured to determine that a pair of video segments comprises common video content.

25. The device of claim 24 , wherein the plurality of video contents items comprises different episodes of one or more video programs.

26. The device of claim 25 , wherein the first video segment and the second video segment of each soft-matching pair are associated with two episodes of a same video program.

27. The device of claim 24 , wherein the first characteristic comprises audio elements, an audio fingerprint, closed captioning data, subtitle data, on-screen text, or a detected visual feature.

28. The device of claim 24 , wherein common video content comprises at least one of an introduction portion, a closing portion, or an advertisement.

29. A non-transitory computer-readable medium storing instructions that, when executed, cause:

receiving first video content comprising video segments and second video content comprising video segments, wherein the second video content is associated with the first video content;

determining a first hard-matching pair of video segments of the first video content and the second video content, wherein video fingerprints of the first hard-matching pair of video segments match;

determining a second hard-matching pair of video segments in the first video content and the second video content, wherein video fingerprints of the second hard- matching pair of video segments match;

classifying a third pair of video segments in the first video content and the second video content as a soft-matching pair of video segments of an introduction portion of at least one of the first video content or the second video content, wherein the third pair of video segments is sequentially between the first hard-matching pair of video segments and the second hard-matching pair of video segments, wherein video fingerprints of the third pair of video segments do not match; and

determining, based on the classifying the third pair of video segments as a soft- matching pair of video segments of an introduction portion of at least one of the first video content or the second video content, a model configured to determine that a pair of video segments in two video content items are a soft-matching pair of video segments of an introduction portion of at least one of the two video content items.

30. The non-transitory computer readable medium of claim 29 , wherein the first video content comprises at least a portion of a first episode of a video program series and the second video content comprises at least a portion of a second episode of the video program series.

31. The non-transitory computer readable medium of claim 29 , wherein the first video content comprises target video content in which the introduction portion is not known and the second video content comprises reference video content in which the introduction portion is known.

32. The non-transitory computer readable medium of claim 29 , wherein the instructions, when executed, further cause:

determining the model via machine learning, wherein a training data input for the machine learning comprises the video fingerprints of the third pair of video segments, and a training data output for the machine learning comprises the classification of the third pair of video segments as a soft-matching pair of video segments of an introduction portion of at least one of the first video content or the second video content.

33. The non-transitory computer readable medium of claim 32 , wherein the model comprises a regressor model and a training data input for determining the regressor model comprises a difference between the video fingerprints of the third pair of video segments.

34. The non-transitory computer readable medium of claim 29 , wherein:

a difference between lengths of the first hard-matching pair of video segments satisfies a length threshold, and

a difference between lengths of the second hard-matching pair of video segments satisfies the length threshold.

35. The non-transitory computer readable medium of claim 34 , wherein a difference between lengths of the third pair of video segments does not satisfy the length threshold.

36. The non-transitory computer readable medium of claim 29 , wherein the video segments of the first video content comprise respective shots in the first video content and the video segments of the second video content comprise respective shots in the second video content.

37. The non-transitory computer readable medium of claim 29 , wherein a video fingerprint of a video segment comprises an alphanumeric value, and a matching pair of video fingerprints each comprise the same alphanumeric value.

38. A non-transitory computer-readable medium storing instructions that, when executed, cause:

determining one or more soft-matching pairs of video segments among a plurality of video content items, wherein each of the one or more soft-matching pairs of video segments comprises a first video segment of one of the plurality of video content items and a second video segment of a different one of the plurality of video content items, wherein a characteristic of the first and second video segments of each soft-matching pair does not match, and wherein each of the one or more soft matching pairs of video segments is located within the corresponding video content items between two hard-matching pairs of video segments of the video content items; and

determining, based on the determining the one or more soft-matching pairs of video segments, a model configured to determine that a pair of video segments comprises common video content.

39. The non-transitory computer readable medium of claim 38 , wherein the plurality of video contents items comprises different episodes of one or more video programs.

40. The non-transitory computer readable medium of claim 39 , wherein the first video segment and the second video segment of each soft-matching pair are associated with two episodes of a same video program.

41. The non-transitory computer readable medium of claim 38 , wherein the first characteristic comprises audio elements, an audio fingerprint, closed captioning data, subtitle data, on-screen text, or a detected visual feature.

42. The non-transitory computer readable medium of claim 38 , wherein common video content comprises at least one of an introduction portion, a closing portion, or an advertisement.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2020
From: YOUNESSIAN, EHSAN; ISHTIAQ, FAISAL; BRASKICH, ANTHONY; SPIER, BARUCH
To: COMCAST CABLE COMMUNICATIONS, LLC
Reel/Frame 053222/0477 →
Continuity (1)
Related Publication 20220019809A1 · Jan 20, 2022