IP Library Granted Patent US 12,573,201
Granted Patent B2
US 12,573,201 · App. 18/180,039 · Granted Mar 10, 2026

Signature-based object tracking in video feed

Inventor: Sudhanshu Uday Sohoni (Bothell, WA)
Assignee: Microsoft Technology Licensing, LLC
G06V20/46G06V10/225G06V10/267G06V20/52G06V40/161
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,201
App. No.
18/180,039
Granted
Mar 10, 2026
Kind
B2
Abstract

A device may calculate a first signature for a first region of a first video frame of a video feed based on content of the first video frame. The first region is associated with a detected object of interest within the video feed. The device may calculate a second signature for a second region of a second video frame of the video feed based on content of the second video frame. The second video frame is subsequent to the first video frame within the video feed, and the second region is associated with the detected object of interest within the video feed. The device may determine a measure of difference between the first signature and the second signature. Based on the amount of difference between the first signature and the second signature exceeding a threshold, the device may trigger a signature-based object detection on the video feed.

Claims (50)

1 . A method, implemented at a processor system, comprising:

calculating a first signature for a first bounding box within a first video frame of a video feed based on first content of the first video frame, the first bounding box encompassing at least a portion of a detected object of interest within the video feed;

calculating a second signature for a second bounding box within a second video frame of the video feed based on second content of the second video frame, the second video frame being subsequent to the first video frame within the video feed, the second bounding box encompassing at least a portion of the detected object of interest;

determining a measure of difference between the first signature and the second signature; and

based on an amount of difference between the first signature and the second signature exceeding a threshold, triggering an object detection on the video feed.

2 . The method of claim 1 , wherein calculating the first signature comprises:

dividing the first bounding box into a plurality of sub-regions;

calculating a plurality of sub-region signatures corresponding to the plurality of sub-regions, respectively; and

concatenating the plurality of sub-region signatures.

3 . The method of claim 2 , wherein calculating the plurality of sub-region signatures comprises, for each sub-region in the plurality of sub-regions:

identifying a plurality of pixel groups in the sub-region, each pixel group comprising a subject pixel and a plurality of neighbor pixels that bound the subject pixel;

generating a plurality of pixel group signatures, each pixel group signature corresponding to a different pixel group in the plurality of pixel groups; and

creating a sub-region signature based on combining the plurality of pixel group signatures.

4 . The method of claim 3 , wherein each pixel group signature indicates a contrasting pixel based on an edge angle relative to the subject pixel.

5 . The method of claim 1 , wherein each of the first bounding box and the second bounding box includes an expanded region exceeding an initial bounding box of the detected object of interest.

6 . The method of claim 1 , further comprising, based on the object detection on the video feed failing to locate an object of interest, triggering an orientation-based object detection that rotates a video frame.

7 . The method of claim 1 , further comprising, based on the object detection on the video feed failing to locate an object of interest, generating a signature on an entire frame of the video feed.

8 . The method of claim 1 , wherein triggering the object detection on the video feed comprises triggering a face detection using an artificial intelligence model or a machine learning model.

9 . The method of claim 1 , further comprising, subsequent to triggering the object detection on the video feed, triggering an interval-based object detection on the video feed.

10 . A computer system comprising:

a processor system; and

a computer storage medium that stores computer-executable instructions that are executable by the processor system to at least:

calculate a first signature for a first bounding box within a first video frame of a video feed based on first content of the first video frame, the first bounding box encompassing at least a portion of a detected object of interest within the video feed, the first signature characterizing a strength of edges within the first bounding box;

calculate a second signature for a second bounding box within a second video frame of the video feed based on second content of the second video frame, the second video frame being subsequent to the first video frame within the video feed, the second bounding box encompassing at least a portion of the detected object of interest, and the second signature characterizing a strength of edges within the second bounding box;

determine a measure of difference between the first signature and the second signature; and

based on an amount of difference between the first signature and the second signature exceeding a threshold, trigger an object detection on the video feed.

11 . The computer system of claim 10 , wherein calculating the first signature comprises:

dividing the first bounding box into a plurality of sub-regions;

calculating a plurality of sub-region signatures corresponding to the plurality of sub-regions, respectively; and

concatenating the plurality of sub-region signatures.

12 . The computer system of claim 10 , wherein each of the first bounding box and the second bounding box corresponds to an initial bounding box of the detected object of interest.

13 . The computer system of claim 10 , wherein each of the first bounding box and the second bounding box includes an expanded region exceeding an initial bounding box of the detected object of interest.

14 . The computer system of claim 10 , the computer-executable instructions also executable by the processor system to at least, based on the object detection on the video feed failing to locate an object of interest, triggering an orientation-based object detection that rotates a video frame.

15 . The computer system of claim 10 , the computer-executable instructions also executable by the processor system to at least, based on the object detection on the video feed failing to locate an object of interest, generating a signature on an entire frame of the video feed.

16 . The computer system of claim 10 , wherein triggering the object detection on the video feed comprises triggering a face detection using an artificial intelligence model or a machine learning model.

17 . A computer system comprising:

a processor system; and

a computer storage medium that stores computer-executable instructions that are executable by the processor system to at least:

calculate a first signature for a first bounding box within a first video frame of a video feed based on first content of the first video frame, the first bounding box encompassing at least a portion of a detected object of interest within the video feed;

calculate a second signature for a second bounding box within a second video frame of the video feed based on second content of the second video frame, the second video frame being subsequent to the first video frame within the video feed, the second bounding box encompassing at least a portion of the detected object of interest;

determine a measure of difference between the first signature and the second signature; and

based on an amount of difference between the first signature and the second signature exceeding a threshold, trigger an object detection on the video feed.

18 . The computer system of claim 17 , wherein calculating the first signature comprises:

dividing the first bounding box into a plurality of sub-regions;

calculating a plurality of sub-region signatures corresponding to the plurality of sub-regions, respectively; and

concatenating the plurality of sub-region signatures.

19 . The computer system of claim 18 , wherein calculating the plurality of sub-region signatures comprises, for each sub-region in the plurality of sub-regions:

identifying a plurality of pixel groups in the sub-region, each pixel group comprising a subject pixel and a plurality of neighbor pixels that bound the subject pixel;

generating a plurality of pixel group signatures, each pixel group signature corresponding to a different pixel group in the plurality of pixel groups; and

creating a sub-region signature based on combining the plurality of pixel group signatures.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2023
From: SOHONI, SUDHANSHU UDAY
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 062911/0900 →
Continuity (2)
Provisional Application 63428625 · Nov 29, 2022
Related Publication 20240177487A1 · May 30, 2024
References Cited (47)
US 8103062B2 · Abe · 2012 [cited by examiner]
US 9111129B2 · Kim · 2015 [cited by examiner]
US 9858679B2 · Zhang · 2018 [cited by examiner]
US 10318813B1 · Pereira et al. · 2019 [cited by applicant]
US 10873719B2 · Ohmura · 2020 [cited by examiner]
US 11676383B2 · Subramanian · 2023 [cited by examiner]
US 11676384B2 · Subramanian · 2023 [cited by examiner]
US 11804059B2 · Wu · 2023 [cited by examiner]
US 12131518B2 · Danielsson · 2024 [cited by examiner]
US 12315292B2 · Ayyar · 2025 [cited by examiner]
US 20070053660A1 · Abe · 2007 [cited by examiner]
US 20090245573A1 · Saptharishi et al. · 2009 [cited by applicant]
US 20130120635A1 · Kim · 2013 [cited by applicant]
US 20160125232A1 · Zhang · 2016 [cited by examiner]
US 20170359546A1 · Ohmura · 2017 [cited by examiner]
US 20190065861A1 · Savvides et al. · 2019 [cited by applicant]
US 20190073520A1 · Ayyar · 2019 [cited by examiner]
US 20190370980A1 · Hollander et al. · 2019 [cited by applicant]
US 20210201009A1 · Wu · 2021 [cited by examiner]
US 20220083783A1 · Subramanian · 2022 [cited by examiner]
US 20220198778A1 · Danielsson · 2022 [cited by examiner]
US 20220292286A1 · Subramanian · 2022 [cited by examiner]
US 20220292287A1 · Subramanian · 2022 [cited by examiner]
US 20230090941A1 · Li · 2023 [cited by examiner]
US 20230360360A1 · Öhrn · 2023 [cited by examiner]
CN 104008371A · 2014 [cited by examiner]
CN 105631418A · 2016 [cited by examiner]
CN 109951666A · 2019 [cited by applicant]
CN 110969101A · 2020 [cited by examiner]
CN 111614959A · 2020 [cited by examiner]
EP 2076022A2 · 2009 [cited by applicant]
JP 6906973B2 · 2021 [cited by applicant]
WO WO2017027212A1 · 2017 [cited by examiner]
WO WO2022099988A1 · 2022 [cited by examiner]
CN-104008371-A (machine translation) (Year: 2014). [cited by examiner]
CN-110969101-A (machine translation) (Year: 2020). [cited by examiner]
WO-2022099988-A1 (machine translation) (Year: 2022). [cited by examiner]
CN-105631418-A (machine translation) (Year: 2016). [cited by examiner]
CN-111614959-A (machine translation) (Year: 2020). [cited by examiner]
WO-2017027212-A1 (machine translation) (Year: 2017). [cited by examiner]
Lyu et al., “High-speed object tracking with its application in golf playing.” International Journal of Social Robotics 9, No. 3 (2017): 449-461. (Year: 2017). [cited by examiner]
Creusen, et al., “ViCoMo: visual context modeling for scene understanding in video surveillance”, Journal of Electronic Imaging, vol. 22, Issue 4, Sep. 2013, 20 pages. [cited by applicant]
Dalal, et al., “Histograms of oriented gradients for human detection”, IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05), vol. 1, 2005, pp. 886-893. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US23/036728, Feb. 26, 2024, 17 pages. [cited by applicant]
Vrecko, et al., “A computer vision integration model for a multi-modal cognitive system”, RSJ International Conference on Intelligent Robots and Systems—IEEE, 2009, pp. 3140-3147. [cited by applicant]
U.S. Appl. No. 63/428,625, filed Nov. 29, 2022. [cited by applicant]
International Preliminary Report on Patentability received for PCT Application No. PCT/US23/036728, mailed on Jun. 12, 2025, 11 pages. [cited by applicant]