IP Library Granted Patent US 12,437,520
Granted Patent B2
US 12,437,520 · App. 17/681,786 · Granted Oct 7, 2025

Augmentation method of visual training data for behavior detection machine learning

Inventors: Simon Polak (Petach Tikva, IL); Shiri Gordon (Tel Aviv, IL); Menashe Rothschild (Tel-Aviv, IL); Asaf Birenzvieg (Hod-HaSharon, IL)
Assignee: Viisights Solutions Ltd.
G06V10/7747G06T7/20G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,520
App. No.
17/681,786
Granted
Oct 7, 2025
Kind
B2
Abstract

Disclosed herein are methods and systems for training machine learning (ML) models to classify activity of objects, comprising selecting a set of frames depicting one or more objects from one or more video sequences each comprising a plurality of consecutive frames, associating the object(s) with each pixel included in a bounding box of the object(s) identified in each frame of the set, computing a motion mask for each frame of the set indicating whether each pixel associated with the object(s) in the frame is changed or unchanged compared to a corresponding pixel in a preceding frame, augmenting an image of the object(s) in each frame of a subset of frames of the set to depict only the changed pixels associated with the object(s) by cutting out the unchanged pixels, and training one or more ML models, using the set of frames, to classify one or more activities of the object(s).

Claims (41)

1. A computer implemented method of training machine learning (ML) models to classify activity of objects, comprising:

using at least one processor for:

selecting a set of frames depicting at least one object from at least one video sequence comprising a plurality of consecutive frames;

associating the at least one object with each of a plurality of pixels included in a bounding box of the at least one object identified in each of the frames of the set;

computing a motion mask for each frame of the set indicating whether each pixel associated with the at least one object in a respective bounding box in a respective frame is changed or unchanged compared to a corresponding pixel in a preceding frame of the set;

augmenting an image of the at least one object in each frame of a subset of frames selected from the set of frames to depict only the changed pixels within said respective bounding box and associated with the at least one object by cutting out the unchanged pixels associated with the at least one object; and

training at least one ML model, using the set of frames, to classify at least one activity of the at least one object based on identified moving portions of said at least one object according to said depicted changed pixels;

wherein the motion mask computation and pixel augmentation are performed to improve ML model training accuracy by emphasizing temporally changing object features.

2. The computer implemented method of claim 1 , wherein the at least one activity is a member of a group consisting of: an action conducted by the at least one object, an activity the at least one object is engaged in, an interaction of the at least one object with at least one another object, an event in which the at least one object participates, and a behavior of the at least one object.

3. The computer implemented method of claim 1 , wherein the bounding box of the at least one object is identified using at least one visual analysis algorithm.

4. The computer implemented method of claim 1 , wherein the bounding box of the at least one object is defined manually by at least one annotator.

5. The computer implemented method of claim 1 , wherein each of the changed pixels is a member of a group consisting of: a moved pixel, a displaced pixel, and a transformed pixel.

6. The computer implemented method of claim 1 , wherein the changed and unchanged pixels in each frame compared to its preceding frame are identified using at least one motion detection algorithm.

7. The computer implemented method of claim 1 , wherein the subset of frames is selected from the set by computing a random value for each frame of the set using at least one random number module, and including the respective frame in the subset in case the random value computed for the respective frame does not exceed a predefined probability threshold.

8. The computer implemented method of claim 1 , wherein cutting out the unchanged pixels comprises blackening each unchanged pixel associated with the at least one object.

9. The computer implemented method of claim 1 , further comprising augmenting independently a plurality of subsets of images for a plurality of objects, each of the plurality of subsets comprises images of a respective set of frames selected from the plurality of consecutive frames for a respective one of the plurality of objects.

10. A system for training machine learning (ML) models to classify activity of objects, comprising:

a non-transitory computer readable medium storing executable code;

at least one processor adapted to execute said code, the code comprising:

code instructions to select a set of frames depicting at least one object from at least one video sequence comprising a plurality of consecutive frames;

code instructions to associate the at least one object with each of a plurality of pixels included in a bounding box of the at least one object identified in each of the frames of the set;

code instructions to compute a motion mask for each frame of the set indicating whether each pixel associated with the at least one object in a respective bounding box in a respective frame is changed or unchanged compared to a corresponding pixel in a preceding frame of the set;

code instructions to augment an image of the at least one object in each frame of a subset of frames selected from the set of frames to depict only the changed pixels within said respective bounding box and associated with the at least one object by cutting out the unchanged pixels associated with the at least one object; and

code instructions to train at least one ML model, using the set of frames, to classify at least one activity of the at least one object based on identified moving portions of said at least one object according to said depicted changed pixels;

wherein the motion mask computation and pixel augmentation are performed to improve ML model training accuracy by emphasizing temporally changing object features.

11. The system of claim 10 , wherein the at least one activity is a member of a group consisting of: an action conducted by the at least one object, an activity the at least one object is engaged in, an interaction of the at least one object with at least one another object, an event in which the at least one object participates, and a behavior of the at least one object.

12. The system of claim 10 , wherein the bounding box of the at least one object is created using at least one visual analysis algorithm.

13. The system of claim 10 , wherein the bounding box of the at least one object is defined manually by at least one annotator.

14. The system of claim 10 , wherein each of the changed pixels is a member of a group consisting of: a moved pixel, a displaced pixel, and a transformed pixel.

15. The system of claim 10 , wherein the changed and unchanged pixels in each frame compared to its preceding frame are identified using at least one motion detection algorithm.

16. The system of claim 10 , wherein the subset of frames is selected from the set by computing a random value for each frame of the set using at least one random number module, and including the respective frame in the subset in case the random value computed for the respective frame does not exceed a predefined probability threshold.

17. The system of claim 10 , wherein cutting out the unchanged pixels comprises blackening each unchanged pixel associated with the at least one object.

18. The system of claim 10 , wherein the code further comprises code instructions to augment independently a plurality of subsets of frames for a plurality of objects, each of the plurality of subsets comprises images of a respective set of frames selected from the plurality of consecutive frames for a respective one of the plurality of objects.

19. The computer implemented method of claim 1 , wherein:

said at least one object is a person,

said identified moving portions of said at least one object are at least one body part of said person, participating in performing a dynamic activity by said person, and

said augmented image depicts said at least one body part of said person and not depicting other body parts of said person, not participating in performing said dynamic activity.

20. The system of claim 10 , wherein:

said at least one object is a person,

said identified moving portions of said at least one object are at least one body part of said person, participating in performing a dynamic activity by said person, and

said augmented image depicts said at least one body part of said person and not depicting other body parts of said person, not participating in performing said dynamic activity.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 1, 2026
From: VIISIGHTS SOLUTIONS LTD.
To: MILESTONE SYSTEMS A/S
Reel/Frame 073931/0307 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2022
From: POLAK, SIMON; GORDON, SHIRI; ROTHSCHILD, MENASHE; BIRENZVIEG, ASAF
To: VIISIGHTS SOLUTIONS LTD.
Reel/Frame 059180/0537 →
Continuity (1)
Related Publication 20230274536A1 · Aug 31, 2023
References Cited (5)
US 20140247362A1 · Li · 2014 [cited by examiner]
US 20180032845A1 · Polak · 2018 [cited by examiner]
US 20190294881A1 · Polak · 2019 [cited by examiner]
US 20190377974A1 · Mathieu · 2019 [cited by examiner]
US 20200143171A1 · Lee · 2020 [cited by examiner]