IP Library › Granted Patent US 12,567,169
Granted Patent B2
US 12,567,169 · App. 18/466,970 · Granted Mar 3, 2026

Artificial intelligence smoothed object detection and tracking in video

Inventors: Neil Leonard Padgett (Toronto, CA); Russ Maschmeyer (Berkeley, CA); Eric Andrew Florenzano (San Francisco, CA); Brennan Letkeman (Calgary, CA); James Lepp (Ottawa, CA); Diego Macario Bello (Montreal, CA)
Assignee: Shopify Inc.
G06T7/70G06T7/20H04N19/159
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,567,169
App. No.
18/466,970
Granted
Mar 3, 2026
Kind
B2
Abstract

Methods and systems for object detection and tracking in video that use at least two different AI-assisted object detection algorithms. A first AI-assisted object detection algorithm selected to be used to detect an object in a video frame and determine a mask defining location of the object on the basis that the video frame is a keyframe. A second AI-assisted object detection algorithm may be used to track location of the mask in temporally subsequent frames until the next keyframe is detected.

Claims (32)

1 . A computer-implemented method, comprising:

identifying an object in a first video frame of a video by applying a first AI-assisted object detection algorithm that outputs a mask defining location of the object in the first video frame, wherein the first AI-assisted object detection algorithm includes a promptable image segmentation model;

tracking the mask and defining its location in at least one temporally subsequent frame of the video using a second AI-assisted object detection algorithm, wherein the second AI-assisted object detection algorithm includes a diffusion-based generative model; and

determining that a further frame of the video is a keyframe and, on that basis, re-applying the first AI-assisted objected detection algorithm to identify the object in the further frame and to output a next mask defining the location of the object in the further frame before repeating the tracking and defining using the second AI-assisted object detection algorithm for one or more frames subsequent to the keyframe.

2 . The method of claim 1 , wherein the first AI-assisted object detection algorithm determines the mask defining location of the object in the first video frame without reference to preceding frames of the video or previously-determined masks for the video, and wherein the second AI-assisted object detection algorithm tracks the mask and defines its location in the at least one temporally subsequent frame of the video based on a location of the mask in a preceding frame and a determination of correspondence between the preceding frame and the at least one temporally subsequent frame.

3 . The method of claim 1 , wherein determining that the further frame is a keyframe includes receiving keyframe identification data from at least one of a video encoder and a video decoder regarding the video.

4 . The method of claim 3 , wherein the keyframe identification data includes data identifying frames within decoded video that are keyframes.

5 . The method of claim 3 , wherein receiving keyframe identification data includes extracting the keyframe identification data from decoded video.

6 . The method of claim 1 , wherein determining that the further frame is a keyframe in based on determining that the further frame was intra-coded by an encoder.

7 . The method of claim 1 , wherein determining that the subsequent frame is a keyframe includes performing scene change analysis on the video and determining that the subsequent frame is a scene change.

8 . The method of claim 1 , wherein determining that the further frame is a keyframe includes tracking the mask and defining its location in the further frame of the video using the second AI-assisted object detection algorithm, applying the first AI-assisted object detection algorithm to identify a new mask defining a location of the object in the further frame, comparing the new mask to the location of the mask determined by the second AI-assisted object detection algorithm, and determining that an error measurement exceeds a threshold value.

9 . The method of claim 8 , wherein the operations of tracking the mask and defining its location in the further frame, applying the first AI-assisted object detection algorithm to identify the new mask, comparing the new mask to the location of the mask determined by the second AI-assisted object detection algorithm, and determining the error measurement are carried out only on frames of the video that were intra-coded.

10 . The method of claim 1 , further comprising displaying the video on a display screen with a visual overlay indicating the object based on the mask and the next mask.

11 . A non-transitory processor-readable medium storing processor-executable instructions that, when executed by one or more processors, are to cause the one or more processors to:

identify an object in a first video frame of a video by applying a first AI-assisted object detection algorithm that outputs a mask defining location of the object in the first video frame, wherein the first AI-assisted object detection algorithm includes a promptable image segmentation model;

track the mask and define its location in at least one temporally subsequent frame of the video using a second AI-assisted object detection algorithm, wherein the second AI-assisted object detection algorithm includes a diffusion-based generative model; and

determine that a further frame of the video is a keyframe and, on that basis, re-apply the first AI-assisted objected detection algorithm to identify the object in the further frame and to output a next mask defining the location of the object in the further frame before repeating the tracking and defining using the second AI-assisted object detection algorithm for one or more frames subsequent to the keyframe.

12 . The non-transitory processor-readable medium of claim 11 , wherein the first AI-assisted object detection algorithm determines the mask defining location of the object in the first video frame without reference to preceding frames of the video or previously-determined masks for the video, and wherein the second AI-assisted object detection algorithm tracks the mask and defines its location in the at least one temporally subsequent frame of the video based on a location of the mask in a preceding frame and a determination of correspondence between the preceding frame and the at least one temporally subsequent frame.

13 . The non-transitory processor-readable medium of claim 11 , wherein the instructions, when executed, are to cause the one or more processors to determine that the further frame is a keyframe by at least receiving keyframe identification data from at least one of a video encoder and a video decoder regarding the video.

14 . The non-transitory processor-readable medium of claim 13 , wherein the keyframe identification data includes data identifying frames within decoded video that are keyframes.

15 . The non-transitory processor-readable medium of claim 13 , wherein receiving keyframe identification data includes extracting the keyframe identification data from decoded video.

16 . The non-transitory processor-readable medium of claim 11 , wherein the instructions, when executed, are to cause the one or more processors to determine that the further frame is a keyframe by determining that the further frame was intra-coded by an encoder.

17 . The non-transitory processor-readable medium of claim 11 , wherein the instructions, when executed, are to cause the one or more processors to determine that the subsequent frame is a keyframe by performing scene change analysis on the video and determining that the subsequent frame is a scene change.

18 . The non-transitory processor-readable medium of claim 11 , wherein the instructions, when executed, are to cause the one or more processors to determine that the further frame is a keyframe by tracking the mask and defining its location in the further frame of the video using the second AI-assisted object detection algorithm, applying the first AI-assisted object detection algorithm to identify a new mask defining a location of the object in the further frame, comparing the new mask to the location of the mask determined by the second AI-assisted object detection algorithm, and determining that an error measurement exceeds a threshold value.

19 . The non-transitory processor-readable medium of claim 18 , wherein the operations of tracking the mask and defining its location in the further frame, applying the first AI-assisted object detection algorithm to identify the new mask, comparing the new mask to the location of the mask determined by the second AI-assisted object detection algorithm, and determining the error measurement are carried out only on frames of the video that were intra-coded.

20 . The non-transitory processor-readable medium of claim 11 , wherein the instructions, when executed, are to cause the one or more processors to display the video on a display screen with a visual overlay indicating the object based on the mask and the next mask.

21 . A computing system, comprising:

one or more processors; and

memory coupled to at least one of the one or more processors, the memory storing computer-executable instructions that, when executed by the one or more processors, are to cause the one or more processors to:

identify an object in a first video frame of a video by applying a first AI-assisted object detection algorithm that outputs a mask defining location of the object in the first video frame, wherein the first AI-assisted object detection algorithm includes a promptable image segmentation model;

track the mask and define its location in at least one temporally subsequent frame of the video using a second AI-assisted object detection algorithm, wherein the second AI-assisted object detection algorithm includes a diffusion-based generative model; and

determine that a further frame of the video is a keyframe and, on that basis, re-apply the first AI-assisted objected detection algorithm to identify the object in the further frame and to output a next mask defining the location of the object in the further frame before repeating the tracking and defining using the second AI-assisted object detection algorithm for one or more frames subsequent to the keyframe.

Assignments (7)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 17, 2024
From: SHOPIFY QUEBEC INC.
To: SHOPIFY INC.
Reel/Frame 066155/0028 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 17, 2024
From: SHOPIFY (USA) INC.
To: SHOPIFY INC.
Reel/Frame 066155/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 12, 2023
From: PADGETT, NEIL LEONARD; LETKEMAN, BRENNAN; LEPP, JAMES
To: SHOPIFY INC.
Reel/Frame 065199/0775 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 12, 2023
From: BELLO, DIEGO MACARIO
To: SHOPIFY QUEBEC INC.
Reel/Frame 065199/0849 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 12, 2023
From: MASCHMEYER, RUSS; FLORENZANO, ERIC ANDREW
To: SHOPIFY (USA) INC.
Reel/Frame 065199/0719 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2023
From: MASCHMEYER, RUSS; FLORENZANO, ERIC ANDREW
To: SHOPIFY (USA) INC.
Reel/Frame 065122/0686 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2023
From: MASCHMEYER, RUSS; FLORENZANO, ERIC ANDREW
To: SHOPIFY (USA) INC.
Reel/Frame 065122/0579 →
Continuity (1)
Related Publication 20250095185A1 · Mar 20, 2025
References Cited (23)
US 20100251287A1 · Deshpande et al. · 2010 [cited by applicant]
US 20120128242A1 · Hampapur · 2012 [cited by examiner]
US 20140343399A1 · Posse · 2014 [cited by examiner]
US 20140359679A1 · Shivadas · 2014 [cited by examiner]
US 20190130191A1 · Zhou et al. · 2019 [cited by applicant]
US 20190130580A1 · Chen et al. · 2019 [cited by applicant]
US 20210158536A1 · Li · 2021 [cited by examiner]
US 20220036084A1 · Migdal · 2022 [cited by applicant]
US 20220222832A1 · Fu · 2022 [cited by examiner]
US 20220262011A1 · Zhang et al. · 2022 [cited by applicant]
US 20220309633A1 · Davies · 2022 [cited by examiner]
US 20230064431A1 · Kansara · 2023 [cited by applicant]
US 20230109379A1 · Kreis · 2023 [cited by examiner]
US 20230162502A1 · Patel · 2023 [cited by examiner]
US 20230230250A1 · Vianello · 2023 [cited by examiner]
US 20230237775A1 · Portail · 2023 [cited by examiner]
Bertasius, Gedas, and Lorenzo Torresani. “Classifying, segmenting, and tracking object instances in video with mask propagation.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020. … [cited by examiner]
Zhu, Xizhou, et al. “Deep feature flow for video recognition.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2017. (Year: 2017). [cited by examiner]
Kirillov, Alexander, et al. “Segment anything.” Proceedings of the IEEE/CVF international conference on computer vision. 2023. (Year: 2023). [cited by examiner]
International Searching Authority; PCT Written Opinion and International Search Report relating to Application No. PCT/CA2024/050247 dated May 8, 2024. [cited by applicant]
Yang, J. et al, “Track Anything: Segment Anything Meets Videos”, arXiv Computer Vision and Pattern Recognition, pp. 1-7, [retrieved on May 1, 2021, retrieved from: https://arxiv.org/pdf/2304.11968] *entire document* dat… [cited by applicant]
“Emergent Correspondence from Image Diffusion”, Tang, et al., https://arxiv.org/pdf/2306.03881.pdf, Jun. 6, 2023. [cited by applicant]
“Segment Anything”, Kirillov, et al., https://arxiv.org/pdf/2304.02643.pdf, Apr. 5, 2023. [cited by applicant]