IP Library Granted Patent US 12,700,238
Granted Patent B1
US 12,700,238 · App. 18/456,703 · Granted Aug 4, 2026

Ensemble of machine learning models for automatic scene change detection

Inventors: Shixing Chen (Seattle, WA); Muhammad Raffay Hamid (Seattle, WA); Vimal Bhat (Redmond, WA); Shiva Krishnamurthy (Sammamish, WA)
Assignee: Amazon Technologies, Inc.
G06V20/49G06F18/213G06N5/04G06N20/20G10L25/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,700,238
App. No.
18/456,703
Filed
Aug 28, 2023
Granted
Aug 4, 2026
Kind
B1
Art Unit
2645
USPC
382/224
Abstract

Techniques for automatic scene change detection are described. As one example, a computer-implemented method includes receiving a request to train an ensemble of machine learning models on a training dataset of videos having labels that indicate scene changes to detect a scene change in a video, partitioning each video file of the training dataset of videos into a plurality of shots, training the ensemble of machine learning models into a trained ensemble of machine learning models based at least in part on the plurality of shots of the training dataset of videos and the labels that indicate scene changes, receiving an inference request for an input video, partitioning the input video into a plurality of shots, generating, by the trained ensemble of machine learning models, an inference of one or more scene changes in the input video based at least in part on the plurality of shots of the input video, and transmitting the inference to a client application or to a storage location.

Claims (30)

1 . A computer-implemented method comprising:

training a machine learning model, on a training dataset of scene change data, to predict when a shot of a video is an end of a scene in the video;

generating, by the trained machine learning model based on an input of shot representations of a plurality of shots of an input video, a prediction that a shot of the plurality of shots is an end of a scene in the input video;

inserting secondary content into the input video at a shot boundary of the shot indicated by the prediction as the end of the scene to generate an output video; and

transmitting the output video to an application or to a storage location.

2 . The computer-implemented method of claim 1 , wherein the shot representations of the plurality of shots of the input video comprise audio features and video features.

3 . The computer-implemented method of claim 2 , further comprising concatenating the audio features and the video features into a multimodal shot representation, wherein the input of the trained machine learning model comprises the multimodal shot representation.

4 . The computer-implemented method of claim 2 , further comprising extracting voice features from corresponding audio of the input video.

5 . The computer-implemented method of claim 2 , wherein the trained machine learning model comprises a video machine learning model that inputs the audio features of the input video, and an audio machine learning model that inputs the video features of the input video.

6 . The computer-implemented method of claim 1 , wherein the training dataset comprises a timestamp for a training video.

7 . The computer-implemented method of claim 1 , wherein the shot predicted to be the end of the scene is the end of a first scene that is followed by a second different scene in the input video.

8 . The computer-implemented method of claim 1 , wherein the scene in the input video is a set of shots filming action in a same time and location.

9 . The computer-implemented method of claim 1 , further comprising:

receiving a request for the output video from a client device; and

sending the output video to the client device.

10 . A non-transitory computer-readable medium storing code that, when executed by a device, causes the device to perform a method comprising:

training a machine learning model, on a training dataset of scene change data, to predict when a shot of a video is an end of a scene in the video;

generating, by the trained machine learning model based on an input of shot representations of a plurality of shots of an input video, a prediction that a shot of the plurality of shots is an end of a scene in the input video;

inserting secondary content into the input video at a shot boundary of the shot indicated by the prediction as the end of the scene to generate an output video; and

transmitting the output video to an application or to a storage location.

11 . The non-transitory computer-readable medium of claim 10 , wherein the shot representations of the plurality of shots of the input video comprise audio features and video features.

12 . The non-transitory computer-readable medium of claim 11 , wherein the method further comprises concatenating the audio features and the video features into a multimodal shot representation, wherein the input of the trained machine learning model comprises the multimodal shot representation.

13 . The non-transitory computer-readable medium of claim 11 , wherein the method further comprises extracting voice features from corresponding audio of the input video.

14 . The non-transitory computer-readable medium of claim 11 , wherein the trained machine learning model comprises a video machine learning model that inputs the audio features of the input video, and an audio machine learning model that inputs the video features of the input video.

15 . The non-transitory computer-readable medium of claim 10 , wherein the training dataset comprises a timestamp for a training video.

16 . The non-transitory computer-readable medium of claim 10 , wherein the shot predicted to be the end of the scene is the end of a first scene that is followed by a second different scene in the input video.

17 . The non-transitory computer-readable medium of claim 10 , wherein the scene in the input video is a set of shots filming action in a same time and location.

18 . The non-transitory computer-readable medium of claim 10 , wherein the method further comprises:

receiving a request for the output video from a client device; and

sending the output video to the client device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2023
From: CHEN, SHIXING; HAMID, MUHAMMAD RAFFAY; BHAT, VIMAL; KRISHNAMURTHY, SHIVA
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 064721/0432 →
Continuity (1)
Continuation 17107514 · Nov 30, 2020
References Cited (58)
US 10701394B1 · Caballero · 2020 [cited by examiner]
US 10902616B2 · Brown et al. · 2021 [cited by applicant]
US 11526698B2 · Lee et al. · 2022 [cited by applicant]
US 11586902B1 · Sather et al. · 2023 [cited by applicant]
US 11640529B2 · Payne et al. · 2023 [cited by applicant]
US 20160337705A1 · Jones et al. · 2016 [cited by examiner]
US 20170103264A1 · Javan et al. · 2017 [cited by applicant]
US 20170124400A1 · Yehezkel et al. · 2017 [cited by applicant]
US 20180139458A1 · Wang · 2018 [cited by examiner]
US 20190147105A1 · Chu et al. · 2019 [cited by examiner]
US 20190163977A1 · Chen · 2019 [cited by examiner]
US 20190294927A1 · Guttmann · 2019 [cited by applicant]
US 20190311202A1 · Lee et al. · 2019 [cited by examiner]
US 20190354765A1 · Chan et al. · 2019 [cited by applicant]
US 20200090001A1 · Zargahi et al. · 2020 [cited by examiner]
US 20200134316A1 · Krishnamurthy · 2020 [cited by examiner]
US 20200193163A1 · Chang · 2020 [cited by examiner]
US 20200293783A1 · Ramaswamy et al. · 2020 [cited by applicant]
US 20200304755A1 · Narayan · 2020 [cited by examiner]
US 20200394458A1 · Yu et al. · 2020 [cited by applicant]
US 20210004589A1 · Turkelson · 2021 [cited by examiner]
US 20210049468A1 · Karras et al. · 2021 [cited by applicant]
US 20210064965A1 · Pardeshi et al. · 2021 [cited by applicant]
US 20210073944A1 · Liu et al. · 2021 [cited by applicant]
US 20210089779A1 · Chan · 2021 [cited by examiner]
US 20210117728A1 · Lee et al. · 2021 [cited by applicant]
US 20210132688A1 · Kim et al. · 2021 [cited by applicant]
US 20210142066A1 · Jayaram · 2021 [cited by examiner]
US 20210142160A1 · Mohseni et al. · 2021 [cited by applicant]
US 20210192748A1 · Morales Morales · 2021 [cited by examiner]
US 20210192756A1 · Huang · 2021 [cited by examiner]
US 20220021716A1 · Nagendran et al. · 2022 [cited by applicant]
US 20220067386A1 · Rotman · 2022 [cited by examiner]
US 20220101112A1 · Brown · 2022 [cited by examiner]
CA 3130573A1 · 2020 [cited by applicant]
CA 3155314A1 · 2021 [cited by examiner]
CN 111566664A · 2020 [cited by examiner]
WO WO2018083668A1 · 2018 [cited by examiner]
WO 2018154494A1 · 2018 [cited by applicant]
WO 2018217635A1 · 2018 [cited by applicant]
WO 2020012069A1 · 2020 [cited by applicant]
WO 2020068140A1 · 2020 [cited by applicant]
WO 2020068784A1 · 2020 [cited by applicant]
WO 2021041078A1 · 2021 [cited by applicant]
WO 2021096776A1 · 2021 [cited by applicant]
WO 2021105157A1 · 2021 [cited by applicant]
WO WO2022023806A1 · 2022 [cited by examiner]
Carreira et al., “Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Feb. 12, 2018, pp. 1-10. [cited by applicant]
Deng et al., “ImageNet: A Large-Scale Hierarchical Image Database”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2009, 9 pages. [cited by applicant]
Hara et al., “Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Apr. 2, 2018, 10 pages. [cited by applicant]
He et al., “Momentum Contrast for Unsupervised Visual Representation Learning”, Facebook AI Research (FAIR), Mar. 23, 2020, 12 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Dec. 10, 2015, pp. 1-12. [cited by applicant]
Hershey et al., “CNN Architectures for Large-Scale Audio Classification”, In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, Jan. 10, 2017, 5 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 17/236,688, Oct. 5, 2022, 6 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/107,514, Jun. 7, 2023, 25 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/236,688, Apr. 28, 2023, 5 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/236,688, Aug. 3, 2023, 2 pages. [cited by applicant]
Rao et al., “A Local-to-Global Approach to Multi-modal Movie Scene Segmentation”, CVPR 2020, Computer Vision Foundation, 2020, pp. 10146-10155. [cited by applicant]