IP Library Granted Patent US 12,260,639
Granted Patent B2
US 12,260,639 · App. 17/454,666 · Granted Mar 25, 2025

Annotating a video with a personalized recap video based on relevancy and watch history

Inventors: Keith Gregory Frost (Delaware, OH); Vincent Tkac (Delaware, OH); Andrew C Myers (Columbus, OH); Joshua M Rice (Marysville, OH)
Assignee: International Business Machines Corporation
G06V20/20G06F18/22G06V20/46G06V20/48G10L15/26H04N21/44204
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,639
App. No.
17/454,666
Granted
Mar 25, 2025
Kind
B2
Abstract

A method for automatically annotating an intended video with at least one personalized recap video based on previously viewed videos is provided. The method may include automatically tracking user viewership of the previously viewed videos, and in response to detecting an intention to view the intended video: automatically identifying and extracting a subset of video footage from one or more of the previously viewed videos based on the tracked user viewership and based on a determined relevancy of the subset of video footage to content in the intended video; generating the at least one personalized recap video by compiling and sorting the extracted subset of video footage from the one or more previously viewed videos into a compilation video; and annotating the intended video with the at least one personalized video by presenting the at least one personalized recap video on the intended video.

Claims (56)

1. A computer-implemented method for automatically annotating an intended video with at least one personalized recap video based on previously viewed videos, comprising:

automatically tracking user viewership of the previously viewed videos, and in response to detecting an intention to view the intended video, automatically generating the at least one personalized recap video based on a configurable amount of time elapsed since the user viewership of a previously viewed video immediately preceding the intended video, wherein automatically generating the at least one personalized recap video further comprises:

automatically identifying and extracting a subset of video footage from one or more of the previously viewed videos based on the tracked user viewership and based on a determined relevancy of the subset of video footage to unwatched content in the intended video, wherein automatically identifying and extracting the subset of video footage further comprises automatically identifying a topic and an entity appearing at various points in the unwatched content, and extracting the subset of video footage matching and comprising the identified topic and the entity from one or more previously viewed videos;

generating the at least one personalized recap video by compiling and sorting the extracted subset of video footage from the one or more previously viewed videos into a compilation video; and

annotating the intended video with the at least one personalized video by presenting the at least one personalized recap video on the intended video.

2. The computer-implemented method of claim 1 , wherein the previously viewed videos and the intended video are part of an episodic series of videos, wherein the previously viewed videos precede the intended video, and wherein the intended video comprises an unviewed video.

3. The computer-implemented method of claim 2 , further comprising:

in response to receiving the episodic series of videos:

transcribing audio from each video associated with the episodic series of videos using a speech-to-text algorithm to individually produce an audio transcript of each video with timestamp data; and

for each video, identifying entities and objects appearing at various points in a video using an image recognition algorithm and correlating the identified entities and objects with the transcribed audio and timestamp data for the video, wherein the identified entities identify characters in the episodic series of videos.

4. The method of claim 3 , further comprising:

for each video associated with the episodic series of videos, using machine learning to perform topic modeling on each video based on the audio transcript and the identified entities and objects to identify topics and changes in topics for each video, wherein the topics represents scenes in each video, and wherein the changes in topics represent changes in scenes for each video.

5. The computer-implemented method of claim 4 , wherein automatically identifying and extracting the subset of video footage from one or more of the previously viewed videos based on the tracked user viewership and based on the determined relevancy of the subset of video footage to the unwatched content in the intended video further comprises:

determining a match between the topics and the identified entities between the one or more previously viewed videos and the intended video, wherein the determined match is based on a threshold confidence score.

6. The computer-implemented method of claim 1 , wherein annotating the intended video with the at least one personalized video further comprises:

presenting the at least one personalized recap video as an introduction on the intended video.

7. The computer-implemented method of claim 1 , further comprising:

generating a plurality of personalized recap videos for the intended video; and

intersplicing each personalized recap video associated with the plurality of personalized recap videos at different times on the intended video based on the determined relevancy of each personalized recap video to a scene in the intended video.

8. A computer system for automatically annotating an intended video with at least one personalized recap video based on previously viewed videos, comprising:

one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage devices, and program instructions stored on at least one of the one or more storage devices for execution by at least one of the one or more processors via at least one of the one or more memories, wherein the computer system is capable of performing a method comprising:

automatically tracking user viewership of the previously viewed videos, and in response to detecting an intention to view the intended video, automatically generating the at least one personalized recap video based on a configurable amount of time elapsed since the user viewership of a previously viewed video immediately preceding the intended video, wherein automatically generating the at least one personalized recap video further comprises:

automatically identifying and extracting a subset of video footage from one or more of the previously viewed videos based on the tracked user viewership and based on a determined relevancy of the subset of video footage to unwatched content in the intended video, wherein automatically identifying and extracting the subset of video footage further comprises automatically identifying a topic and an entity appearing at various points in the unwatched content, and extracting the subset of video footage matching and comprising the identified topic and the entity from one or more previously viewed videos;

generating the at least one personalized recap video by compiling and sorting the extracted subset of video footage from the one or more previously viewed videos into a compilation video; and

annotating the intended video with the at least one personalized video by presenting the at least one personalized recap video on the intended video.

9. The computer system of claim 8 , wherein the previously viewed videos and the intended video are part of an episodic series of videos, wherein the previously viewed videos precede the intended video, and wherein the intended video comprises an unviewed video.

10. The computer system of claim 9 , further comprising:

in response to receiving the episodic series of videos:

transcribing audio from each video associated with the episodic series of videos using a speech-to-text algorithm to individually produce an audio transcript of each video with timestamp data; and

for each video, identifying entities and objects appearing at various points in a video using an image recognition algorithm and correlating the identified entities and objects with the transcribed audio and timestamp data for the video, wherein the identified entities identify characters in the episodic series of videos.

11. The computer system of claim 10 , further comprising:

for each video associated with the episodic series of videos, using machine learning to perform topic modeling on each video based on the audio transcript and the identified entities and objects to identify topics and changes in topics for each video, wherein the topics represents scenes in each video, and wherein the changes in topics represent changes in scenes for each video.

12. The computer system of claim 11 , wherein automatically identifying and extracting the subset of video footage from one or more of the previously viewed videos based on the tracked user viewership and based on the determined relevancy of the subset of video footage to the unwatched content in the intended video further comprises:

determining a match between the topics and the identified entities between the one or more previously viewed videos and the intended video, wherein the determined match is based on a threshold confidence score.

13. The computer system of claim 8 , wherein annotating the intended video with the at least one personalized video further comprises:

presenting the at least one personalized recap video as an introduction on the intended video.

14. The computer system of claim 8 , further comprising:

generating a plurality of personalized recap videos for the intended video; and

intersplicing each personalized recap video associated with the plurality of personalized recap videos at different times on the intended video based on the determined relevancy of each personalized recap video to a scene in the intended video.

15. A computer program product for automatically annotating an intended video with at least one personalized recap video based on previously viewed videos, comprising:

one or more tangible computer-readable storage devices and program instructions stored on at least one of the one or more tangible computer-readable storage devices, the program instructions executable by a processor, the program instructions comprising:

automatically tracking user viewership of the previously viewed videos, and in response to detecting an intention to view the intended video, automatically generating the at least one personalized recap video based on a configurable amount of time elapsed since the user viewership of a previously viewed video immediately preceding the intended video, wherein automatically generating the at least one personalized recap video further comprises:

automatically identifying and extracting a subset of video footage from one or more of the previously viewed videos based on the tracked user viewership and based on a determined relevancy of the subset of video footage to unwatched content in the intended video, wherein automatically identifying and extracting the subset of video footage further comprises automatically identifying a topic and an entity appearing at various points in the unwatched content, and extracting the subset of video footage matching and comprising the identified topic and the entity from one or more previously viewed videos;

generating the at least one personalized recap video by compiling and sorting the extracted subset of video footage from the one or more previously viewed videos into a compilation video; and

annotating the intended video with the at least one personalized video by presenting the at least one personalized recap video on the intended video.

16. The computer program product of claim 15 , wherein the previously viewed videos and the intended video are part of an episodic series of videos, wherein the previously viewed videos precede the intended video, and wherein the intended video comprises an unviewed video.

17. The computer program product of claim 16 , further comprising:

in response to receiving the episodic series of videos:

transcribing audio from each video associated with the episodic series of videos using a speech-to-text algorithm to individually produce an audio transcript of each video with timestamp data; and

for each video, identifying entities and objects appearing at various points in a video using an image recognition algorithm and correlating the identified entities and objects with the transcribed audio and timestamp data for the video, wherein the identified entities identify characters in the episodic series of videos.

18. The computer program product of claim 17 , further comprising:

for each video associated with the episodic series of videos, using machine learning to perform topic modeling on each video based on the audio transcript and the identified entities and objects to identify topics and changes in topics for each video, wherein the topics represents scenes in each video, and wherein the changes in topics represent changes in scenes for each video.

19. The computer program product of claim 18 , wherein automatically identifying and extracting the subset of video footage from one or more of the previously viewed videos based on the tracked user viewership and based on the determined relevancy of the subset of video footage to the unwatched content in the intended video further comprises:

determining a match between the topics and the identified entities between the one or more previously viewed videos and the intended video, wherein the determined match is based on a threshold confidence score.

20. The computer program product of claim 15 , wherein annotating the intended video with the at least one personalized video further comprises:

presenting the at least one personalized recap video as an introduction on the intended video.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 12, 2021
From: FROST, KEITH GREGORY; TKAC, VINCENT; MYERS, ANDREW C; RICE, JOSHUA M
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 058097/0185 →
Continuity (1)
Related Publication 20230154184A1 · May 18, 2023
References Cited (28)
US 6535639B1 · Uchihachi · 2003 [cited by applicant]
US 6961954B1 · Maybury · 2005 [cited by applicant]
US 10311913B1 · Shekhar · 2019 [cited by applicant]
US 10341728B2 · Girlando · 2019 [cited by applicant]
US 10390082B2 · Song · 2019 [cited by examiner]
US 10555023B1 · McCarthy · 2020 [cited by examiner]
US 10565177B2 · Gausman · 2020 [cited by applicant]
US 10984667B2 · Raynaud · 2021 [cited by applicant]
US 20020051077A1 · Liou · 2002 [cited by examiner]
US 20100104261A1 · Liu · 2010 [cited by examiner]
US 20160092561A1 · Liu · 2016 [cited by examiner]
US 20160173941A1 · Gilson · 2016 [cited by examiner]
US 20170040037A1 · Hunt · 2017 [cited by examiner]
US 20170289617A1 · Song · 2017 [cited by examiner]
US 20180343482A1 · Loheide · 2018 [cited by examiner]
US 20190104342A1 · Catalano · 2019 [cited by examiner]
US 20190306215A1 · Bastide · 2019 [cited by examiner]
US 20200194035A1 · Catalano · 2020 [cited by examiner]
EP 2240251A1 · 2018 [cited by applicant]
WO 2009102991A1 · 2009 [cited by applicant]
Baraldi, et al., “A Deep Siamese Network for Scene Detection in Broadcast Videos,” [Conference Paper], Oct. 2015, 5 pages, DOI:10.1145/2733373.2806316, ResearchGate, Retrieved from the Internet: <URL: https://www.resear… [cited by applicant]
Baraldi, et al., “Recognizing and Presenting the Storytelling Video Structure with Deep Multimodal Networks,” Sep. 2014, IEEE Transactions on Multimedia, vol. 13, No. 9, Sep. 2014, pp. 1-14, arXiv:1610.0137v2 [cs.CV] No… [cited by applicant]
Disclosed Anonymously, “Hybrid push/on-demand content provision,” IP.com Prior Art Database Technical Disclosure, Sep. 17, 2009, 3 Pages. IP.com No. IPCOM000187738D, Retrieved from the Internet: <URL: https://priorart.i… [cited by applicant]
Disclosed Anonymously, “Mechanism for Queuing Live Stream Content in a Media Content Display Queue,” IP.com Prior Art Database Technical Disclosure, Jan. 5, 2018, 9 pages, IP.com No. IPCOM000252367D, Retrieved from the … [cited by applicant]
Disclosed Anonymously, “Method and System for Personalized Video Recap of Serialized Video Content,” IP.com Prior Art Database Technical Disclosure, Aug. 16, 2018, 5 pages, IP.com No. IPCOM000254959D, Retrieved from the… [cited by applicant]
Disclosed Anonymously, “System Method or Apparatus for Exchanging Knowledge, Information, Products or Any Entity(ies) of Value and Real Time Market and/or Individual Sensitive or Responsive System of Education,” IP.com … [cited by applicant]
Popescu-Belis, et al., “Building and Using a Corpus of Shallow Dialogue Annotated Meetings,” Proceedings of the Fourth International Conference on Language Resources and Evaluation (LREC'04), May 2004, 6 pages, European… [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing”, Recommendations of the National Institute of Standards and Technology, NIST Special Publication 800-145, Sep. 2011, 7 pages. [cited by applicant]