IP Library › Granted Patent US 12,432,428
Granted Patent B2
US 12,432,428 · App. 18/644,707 · Granted Sep 30, 2025

Method and system for producing interaction-aware video synopsis

Inventors: Ariel Naim (Modi'in, IL); Igal Dvir (Modi'in, IL); Shmuel Peleg (Modi'in, IL)
Assignee: BRIEFCAM LTD.
H04N21/8549G06T7/194G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,432,428
App. No.
18/644,707
Granted
Sep 30, 2025
Kind
B2
Abstract

A method and system for generating an interaction-aware video synopsis are provided herein. The method may include the following steps: obtaining, a source video containing a plurality of source objects; extracting the source objects from the source video; detecting at least one object-interaction between at least two source objects; generating synopsis objects by sampling respective source objects; and generating a synopsis video having an overall play time shorter than the overall play time of the source video, by determining a play time for each one of the synopsis objects, wherein a relative display time in the synopsis video of the at least two synopsis objects created from the at least two source objects, detected as having an object-interaction therebetween, is the same as the relative display time of the at least two source objects in the source video.

Claims (56)

1. A method of generating an interaction-aware video synopsis, the method comprising:

obtaining, using a computer processor, a source video containing a plurality of source objects;

extracting, using the computer processor, the source objects from the source video;

detecting, using the computer processor, at least one object-interaction between at least two source objects, wherein said object interaction comprises an interaction between the at least two objects in the scene which is visible in the source video;

generating, using the computer processor, synopsis objects by sampling respective source objects; and

generating, using the computer processor, a synopsis video having an overall play time shorter than the overall play time of the source video, by determining a play time for each one of the synopsis objects,

wherein at least two synopsis objects which are played at least partially simultaneously in the synopsis video, are generated from source objects that are captured at different times in the source video, and

maintaining a relative display time in the synopsis video of at least two of the synopsis objects created from the at least two source objects, responsive to the detecting of the at least one object-interaction between the at least two source objects, so the relative display time is the same as the relative display time of the at least two source objects in the source video,

wherein the maintaining of the relative display time in the synopsis video of the at least one object-interaction between the at least two source objects is achieved by either:

a) representing multiple interacting object tubes as if they are a single combined tube, and whenever a new appearance time is determined for the single combined tube, all interacting tubes therein maintain an original relative time; or

b) marking interacting objects and constraining the generating of the video synopsis, by keeping the relative time of iterating objects also in the generated video synopsis.

2. The method according to claim 1 , wherein the extracting of the source objects from the source video is carried out by training a machine learning model to distinguish between foreground objects and a background object.

3. The method according to claim 1 , wherein the detection of the at least one object-interaction between at least two source objects is carried out by training a machine learning model to detect object interaction.

4. The method according to claim 1 , wherein each one of the source objects is represented by a respective tube and wherein the method further comprising merging the respective tubes for the at least two source objects, detected as having an object-interaction therebetween into a single tube.

5. The method according to claim 4 , further comprising using the single tube to guarantee that the relative display time in the synopsis video of the at least two synopsis objects created from the at least two source objects, detected as having an object-interaction therebetween is the same as the relative display time of the at least two source objects in the source video.

6. The method according to claim 1 , wherein each one of the source objects is represented by a respective tube and wherein the method further comprising tagging the respective tubes for the at least two source objects, detected as having an object-interaction therebetween, as associated with a same object interaction.

7. The method according to claim 6 , further comprising using the tagging to guarantee that the relative display time in the synopsis video of the at least two synopsis objects created from the at least two source objects, detected as having an object-interaction therebetween is the same as the relative display time of the at least two source objects in the source video.

8. A non-transitory computer readable medium for generating an interaction-aware video synopsis, the computer readable medium comprising a set of instructions that, when executed, cause at least one computer processor to:

obtain a source video containing a plurality of source objects;

extract the source objects from the source video;

detect at least one object-interaction between at least two source objects wherein said object interaction comprises an interaction between the at least two objects in the scene which is visible in the source video;

generate synopsis objects by sampling respective source objects; and

generate a synopsis video having an overall play time shorter than the overall play time of the source video, by determining a play time for each one of the synopsis objects,

wherein at least two synopsis objects which are played at least partially simultaneously in the synopsis video, are generated from source objects that are captured at different times in the source video, and

maintain a relative display time in the synopsis video of at least two of the synopsis objects created from the at least two source objects, responsive to the detecting of the at least one object-interaction between the at least two source objects, so the relative display time is the same as the relative display time of the at least two source objects in the source video,

wherein the maintaining of the relative display time in the synopsis video of the at least one object-interaction between the at least two source objects is achieved by either:

a) representing multiple interacting object tubes as if they are a single combined tube, and whenever a new appearance time is determined for the single combined tube, all interacting tubes therein maintain an original relative time; or

b) marking interacting objects and constraining the generating of the video synopsis, by keeping the relative time of iterating objects also in the generated video synopsis.

9. The non-transitory computer readable medium according to claim 8 , wherein the extracting of the source objects from the source video is carried out by training a machine learning model to distinguish between foreground objects and a background object.

10. The non-transitory computer readable medium according to claim 8 , wherein the detection of the at least one object-interaction between at least two source objects is carried out by training a machine learning model to detect object interaction.

11. The non-transitory computer readable medium according to claim 8 , wherein each one of the source objects is represented by a respective tube and wherein the computer readable medium further comprising a set of instructions that, when executed, cause the at least one computer processor to merge the respective tubes for the at least two source objects, detected as having an object-interaction therebetween into a single tube.

12. The non-transitory computer readable medium according to claim 11 , and wherein the computer readable medium further comprising a set of instructions that, when executed, cause the at least one computer processor to use the single tube to guarantee that the relative display time in the synopsis video of the at least two synopsis objects created from the at least two source objects, detected as having an object-interaction therebetween is the same as the relative display time of the at least two source objects in the source video.

13. The non-transitory computer readable medium according to claim 8 , wherein each one of the source objects is represented by a respective tube and wherein the computer readable medium further comprising a set of instructions that, when executed, cause the at least one computer processor to tag the respective tubes with a tag for the at least two source objects, detected as having an object-interaction therebetween, as associated with a same object interaction.

14. The non-transitory computer readable medium according to claim 13 , and wherein the computer readable medium further comprising a set of instructions that, when executed, cause the at least one computer processor to use the tag to guarantee that the relative display time in the synopsis video of the at least two synopsis objects created from the at least two source objects, detected as having an object-interaction therebetween is the same as the relative display time of the at least two source objects in the source video.

15. A system for generating an interaction-aware video synopsis, the system comprising:

a server comprising:

a processing device;

a memory device; and

an interface for communicating with a video camera,

wherein the memory device comprising a set of instructions that, when executed, cause at the processing device to:

obtain from the video camera a source video containing a plurality of source objects;

extract the source objects from the source video;

detect at least one object-interaction between at least two source objects, wherein said object interaction comprises an interaction between the at least two objects in the scene which is visible in the source video;

generate synopsis objects by sampling respective source objects; and

generate a synopsis video having an overall play time shorter than the overall play time of the source video, by determining a play time for each one of the synopsis objects,

wherein at least two synopsis objects which are played at least partially simultaneously in the synopsis video, are generated from source objects that are captured at different times in the source video, and

maintain a relative display time in the synopsis video of at least two of the synopsis objects created from the at least two source objects, responsive to the detecting of the at least one object-interaction between the at least two source objects, so the relative display time is the same as the relative display time of the at least two source objects in the source video,

wherein the maintaining of the relative display time in the synopsis video of the at least one object-interaction between the at least two source objects is achieved by either:

a) representing multiple interacting object tubes as if they are a single combined tube, and whenever a new appearance time is determined for the single combined tube, all interacting tubes therein maintain an original relative time; or

b) marking interacting objects and constraining the generating of the video synopsis, by keeping the relative time of iterating objects also in the generated video synopsis.

16. The system according to claim 15 , wherein the extracting of the source objects from the source video is carried out by training a machine learning model to distinguish between foreground objects and a background object.

17. The system according to claim 15 , wherein the detection of the at least one object-interaction between at least two source objects is carried out by training a machine learning model to detect object interaction.

18. The system according to claim 15 , wherein each one of the source objects is represented by a respective tube and wherein the memory device further comprising a set of instructions that, when executed, cause the processing device to merge the respective tubes for the at least two source objects, detected as having an object-interaction therebetween into a single tube.

19. The system according to claim 18 , and wherein the memory device further comprising a set of instructions that, when executed, cause the processing device to use the single tube to guarantee that the relative display time in the synopsis video of the at least two synopsis objects created from the at least two source objects, detected as having an object-interaction therebetween is the same as the relative display time of the at least two source objects in the source video.

20. The system according to claim 15 , wherein each one of the source objects is represented by a respective tube and wherein the memory device further comprising a set of instructions that, when executed, cause the processing device to tag the respective tubes with a tag for the at least two source objects, detected as having an object-interaction therebetween, as associated with a same object interaction.

21. The system according to claim 20 , and wherein the memory device further comprising a set of instructions that, when executed, cause the processing device to use the tag to guarantee that the relative display time in the synopsis video of the at least two synopsis objects created from the at least two source objects, detected as having an object-interaction therebetween is the same as the relative display time of the at least two source objects in the source video.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 24, 2024
From: NAIM, ARIEL; DVIR, IGAL; PELEG, SHMUEL
To: BRIEFCAM LTD.
Reel/Frame 067515/0225 →
Continuity (2)
Provisional Application 63497773 · Apr 24, 2023
Related Publication 20240357218A1 · Oct 24, 2024
References Cited (27)
US 5768447A · Irani · 1998 [cited by examiner]
US 8102406B2 · Peleg · 2012 [cited by examiner]
US 8311277B2 · Peleg · 2012 [cited by examiner]
US 8514248B2 · Peleg · 2013 [cited by examiner]
US 8818038B2 · Peleg · 2014 [cited by examiner]
US 8949235B2 · Peleg · 2015 [cited by examiner]
US 9877086B2 · Richardson · 2018 [cited by examiner]
US 10958854B2 · Elboher · 2021 [cited by examiner]
US 11328160B2 · Guo · 2022 [cited by examiner]
US 11620335B2 · Kim · 2023 [cited by examiner]
US 20050138674A1 · Howard · 2005 [cited by examiner]
US 20050229227A1 · Rogers · 2005 [cited by examiner]
US 20160323658A1 · Richardson · 2016 [cited by examiner]
US 20210286598A1 · Luo · 2021 [cited by examiner]
A. Rav-Acha, Y. Pritch, and S. Peleg. Making a Long Video Short. In CVPR, Jun. 2006. [cited by applicant]
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in Advances in neural information processing systems, 2015, pp. 91-99. [cited by applicant]
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg, “Ssd: Single shot multibox detector,” in European conference on computer vision. Springer, 2016, pp. 21-37. [cited by applicant]
K. He, G. Gkioxari, P. Dollar, and R. Girshick, “Mask r-cnn,” in Computer Vision (ICCV), 2017 IEEE International Conference on. IEEE, 2017, pp. 2980-2988. [cited by applicant]
W. Luo, J. Xing, A. Milan, X. Zhang, W. Liu, X. Zhao, and T.-K. Kim, “Multiple object tracking: A literature review,” arXiv preprint arXiv:1409.7618, 2014. [cited by applicant]
Cao, Z., Simon, T., Wei, S.E., Sheikh, Y.: “Realtime multi-person 2d pose estimation using part affinity fields”. In: CVPR (2017). [cited by applicant]
Chen, Y., Wang, Z., Peng, Y., Zhang, Z., Yu, G., Sun, J.: “Cascaded pyramid network for multi-person pose estimation”. In: CVPR (2018). [cited by applicant]
Jeelani, I., Asadi, K., Ramshankar, H., Han, K., and Albert, A. 2021. “Real-time vision based worker localization & hazard detection for construction.” Automation in Construction. 121(2021): 103448. [cited by applicant]
Li Xuelong et al; Video Synopsis in Complex Situations; IEEE Transactions on Image Processing; IEEE USA, vol. 27, No. 8, Aug. 1, 2018. [cited by applicant]
Yang Yoonsik et al; Scene Adaptive Online Surveillance Video Synopsis via Dynamic Tube Rearrangement Using Octree; IEEE Transactions on Image Processing, IEEE USA, vol. 30, Sep. 29, 2021. [cited by applicant]
Yunzuo Zhang et al, Object interaction-based surveillance video synopsis; Applied Intelligence, vol. 53, No. 4, Feb. 1, 2023. [cited by applicant]
Santos Daniel F. S. et al; Video Segmentation Learning Using Cascade Residual Convolutional Neural Network; 2013 32 [cited by applicant]
Qi Siyuan et al; Learning Human-Object Interactions by Graph Parsing Neural Networks, Oct. 5, 2018, SAT 2015 18 [cited by applicant]