IP Library › Granted Patent US 12,277,768
Granted Patent B2
US 12,277,768 · App. 18/357,627 · Granted Apr 15, 2025

Method and electronic device for generating a segment of a video

Inventors: Jayesh Rajkumar Vachhani (Bengaluru, IN); Sourabh Vasant Gothe (Bengaluru, IN); Barath Raj Kandur Raja (Bengaluru, IN); Pranay Kashyap (Bengaluru, IN); Rakshith Srinivas (Bengaluru, IN); Rishabh Khurana (Bengaluru, IN)
Assignee: Samsung Electronics Co., Ltd.
G06V20/49G06V40/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,277,768
App. No.
18/357,627
Granted
Apr 15, 2025
Kind
B2
Abstract

A method for generating at least one segment of a video by an electronic device is provided. The method includes identifying at least one of a context associated with the video and an interaction of a user in connection with the video, analyzing at least one parameter in at least one frame of the video with reference to at least one of the context and the interaction of the user, wherein the at least one parameter includes at least one of a subject, an environment, an action of the subject, and an object, determining the at least one frame in which a change in the at least one parameter occurs, and generating at least one segment of the video comprising the at least one frame in which the parameter changed as a temporal boundary of the at least one segment.

Claims (42)

1. A method performed by an electronic device for generating at least one segment of a video, the method comprising:

identifying, by the electronic device, at least one of a context associated with the video and an interaction of a user in connection with the video;

analyzing, by the electronic device, at least one parameter in at least one frame of the video with reference to at least one of the context and the interaction of the user, wherein the at least one parameter comprises at least one of a subject, an environment, an action of the subject, and an object;

determining, by the electronic device, the at least one frame in which a change in the at least one parameter occurs;

generating, by the electronic device, at least one segment of the video comprising the at least one frame in which the at least one parameter changed as a temporal boundary of the at least one segment; and

ranking, by the electronic device, the generated at least one segment of the video based on an affinity score and an auxiliary score,

wherein the affinity score is computed between the video and the generated at least one segment of the video based on the generated at least one segment of the video and a computed category probability.

2. The method as claimed in claim 1 , wherein the context comprises at least one of a discussion regarding the video in a social media application, capturing a screenshot of the frame of the video, sharing the video, editing the video, or generating a story based on the video.

3. The method as claimed in claim 1 , wherein the segment comprises at least one of a short video story, a combination of video clips, or a video clip.

4. The method as claimed in claim 1 , wherein a length of the at least one segment is less than a length of the video.

5. The method as claimed in claim 1 , wherein the interaction of the user in connection with the video includes interaction of the user with at least one of the subject, the environment, an action of the subject, or the object.

6. The method as claimed in claim 1 , wherein predicting the temporal boundary comprises establishing, by the electronic device, a relationship between frames of the video by utilizing a Temporal self-similarity Matrix (TSSM).

7. The method as claimed in claim 1 ,

wherein the method further comprises performing at least one action, by the electronic device, using the temporal boundary of the at least one segment, and

wherein the at least one action comprises at least one of a video screenshot, a smart share suggestion for the video, a boundary aware story in a gallery, or a tap and download video cuts in a video editor.

8. The method as claimed in claim 1 , wherein the at least one segment of the video is generated with or without the context.

9. The method as claimed in claim 1 , wherein the generated at least one segment of the video is appended with a predefined number of frames before a first frame to generate the at least one segment of the video.

10. The method as claimed in claim 1 , wherein the at least one segment of the video is generated by:

refining taxonomy free boundaries of the video;

predicting a boundary category;

generating and filtering taxonomy free proposals based on the refined taxonomy free boundaries and the boundary category; and

generating temporal segment from at least one generated taxonomy free proposal.

11. An electronic device, comprising:

a processor;

a memory; and

a video segment generator, coupled with the processor and the memory, configured to:

identify at least one of a context associated with a video and an interaction of a user in connection with the video,

analyze at least one parameter in at least one frame of the video with reference to at least one of the context and the interaction of the user, wherein the at least one parameter comprises at least one of a subject, an environment, an action of the subject, and an object in connection with the context or taxonomy free,

determine the at least one frame in which a change in the at least one parameter occurs,

generate at least one segment of the video comprising the at least one frame in which the at least one parameter changed as a temporal boundary of the at least one segment, and

rank the generated at least one segment of the video based on an affinity score and an auxiliary score,

wherein the affinity score is computed between the video and the generated at least one segment of the video based on the generated at least one segment of the video and a computed category probability.

12. The electronic device as claimed in claim 11 , wherein the context comprises at least one of a discussion regarding the video in a social media application, capturing a screenshot of the frame of the video, sharing the video, editing the video, or generating a story based on the video.

13. The electronic device as claimed in claim 11 , wherein the segment comprises at least one of a short video story, a combination of video clips, or a video clip.

14. The electronic device as claimed in claim 11 , wherein a length of the at least one segment is less than a length of the video.

15. The electronic device as claimed in claim 11 , wherein the interaction of the user in connection with the video includes interaction of the user with at least one of the subject, the environment, an action of the subject, or the object.

16. The electronic device as claimed in claim 11 , wherein the video segment generator predicts the temporal boundary by establishing a relationship between frames of the video by utilizing a Temporal self-similarity Matrix (TSSM).

17. The electronic device as claimed in claim 11 ,

wherein the video segment generator performs at least one action using the temporal boundary of the at least one segment, and

wherein the at least one action comprises at least one of a video screenshot, a smart share suggestion for the video, a boundary aware story in a gallery, or a tap and download video cuts in a video editor.

18. The electronic device as claimed in claim 11 , wherein the video segment generator is further configured to identify the context based on the interaction of the user.

19. The electronic device as claimed in claim 11 , wherein the auxiliary score is calculated based on the context associated with the video and a boundary category.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 24, 2023
From: VACHHANI, JAYESH RAJKUMAR; GOTHE, SOURABH VASANT; KANDUR RAJA, BARATH RAJ; KASHYAP, PRANAY; SRINIVAS, RAKSHITH; KHURANA, RISHABH
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 064360/0409 →
Priority Claims (2)
IN 202241027488 · May 12, 2022 · national
IN 2022 41027488 · Jan 18, 2023 · national
Continuity (2)
Continuation PCTIB2023054910 · May 12, 2023
Related Publication 20230368534A1 · Nov 16, 2023
References Cited (34)
US 8238718B2 · Toyama et al. · 2012 [cited by applicant]
US 9472188B1 · Ouimette · 2016 [cited by examiner]
US 9846815B2 · Karsh et al. · 2017 [cited by applicant]
US 10229326B2 · Grundmann et al. · 2019 [cited by applicant]
US 10452712B2 · Mei et al. · 2019 [cited by applicant]
US 10949674B2 · Hwangbo et al. · 2021 [cited by applicant]
US 20030185450A1 · Garakani · 2003 [cited by examiner]
US 20160014482A1 · Chen et al. · 2016 [cited by applicant]
US 20160334970A1 · Mysore Veera · 2016 [cited by examiner]
US 20170109584A1 · Yao et al. · 2017 [cited by applicant]
US 20170311368A1 · Kandur Raja · 2017 [cited by examiner]
US 20180213288A1 · Patry et al. · 2018 [cited by applicant]
US 20190114487A1 · Vijayanarasimhan et al. · 2019 [cited by applicant]
US 20190206445A1 · Matias et al. · 2019 [cited by applicant]
US 20200074185A1 · Rhodes et al. · 2020 [cited by applicant]
US 20200334290A1 · Dontcheva · 2020 [cited by examiner]
US 20200364464A1 · Vijayanarasimhan · 2020 [cited by examiner]
US 20210021912A1 · Rozhenkov · 2021 [cited by examiner]
US 20210201047A1 · Hwangbo · 2021 [cited by examiner]
US 20220075513A1 · Walker et al. · 2022 [cited by applicant]
US 20220076025A1 · Shin · 2022 [cited by examiner]
CN 113259761A · 2021 [cited by applicant]
CN 114283351A · 2022 [cited by applicant]
EP 3491546A1 · 2019 [cited by applicant]
WO 2018022853A1 · 2018 [cited by applicant]
Lin et al., BSN: Boundary Sensitive Network for Temporal Action Proposal Generation, arXiv:1806.02964v3, Sep. 26, 2018. [cited by applicant]
Lin et al., BMN: Boundary-Matching Network for Temporal Action Proposal Generation, arXiv:1907.09702v1, Jul. 23, 2019. [cited by applicant]
Kang et al., Winning the CVPR'2021 Kinetics—GEBD Challenge: Contrastive Learning Approach, arXiv:2106.11549v1, Jun. 22, 2021. [cited by applicant]
Shou, Mike Zheng et al., Generic event boundary detection: A benchmark for event segmentation, Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021. [cited by applicant]
Wang, Limin et al., Temporal segment networks: Towards good practices for deep action recognition, European conference on computer vision. Springer, Cham, arXiv:1608.00859v1, Aug. 2, 2016. [cited by applicant]
International Search Report dated Aug. 29, 2023, issued in International Application No. PCT/IB2023/054910. [cited by applicant]
Wang Yuxuan et al., GEB+: A benchmark for generic event boundary captioning, grounding and text-based retrieval, Apr. 10, 2022, XP 093241121. [cited by applicant]
Indian Office Action dated Dec. 23, 2024, issued in Indian Application No. 202241027488. [cited by applicant]
European Search Report dated Jan. 29, 2025, issued in European Application No. 23803138.9. [cited by applicant]