IP Library Granted Patent US 12,437,421
Granted Patent B2
US 12,437,421 · App. 18/326,571 · Granted Oct 7, 2025

Image processing device and method of detecting objects crossing a crossline and a direction the objects crosses the crossline

Inventor: Martin Sjöborg (Lund, SE)
Assignee: AXIS AB
G06T7/246G06T7/73G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/20212G06T2207/30196G06T2207/30232
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,421
App. No.
18/326,571
Granted
Oct 7, 2025
Kind
B2
Abstract

A device and method for detecting objects crossing a crossline, and the direction, captured by a video camera are described. A sequence of video image frames of the scene is captured, and a combined image frame is created by extracting two or more lines of pixels of each image frame and arranging them adjacent to each other, wherein the lines of pixels of each video image frame are parallel and correspond to the crossline in the scene. The combined image frame is sent to a machine learning model that detects a combined image frame representing an object crossing a crossline and a direction the object crosses during capturing of the sequence of video image frames. A detection of any object crossing the crossline and a direction of crossing is received from the machine learning model.

Claims (42)

1. A method for detecting objects crossing a crossline in a scene captured by a video camera and a direction the object crosses the crossline, the method comprising:

capturing a sequence of video image frames of the scene;

creating a combined image frame by:

extracting two or more lines of pixels of each image frame of the sequence of video image frames, wherein the two or more lines of pixels of each video image frame are parallel to, and correspond to, the crossline in the scene such that if a representation of an object crosses the two or more lines of pixels in the sequence of video image frames, this corresponds to the object having crossed the crossline in the scene, and

constructing the combined image frame by arranging the extracted two or more lines of pixels of each video image frame of the sequence of video image frames adjacent to each other;

sending the combined image frame to a machine learning model trained to detect a combined image frame created from a sequence of video image frames that represents an object crossing a crossline in a scene and a direction the object crosses the crossline during capturing of the sequence of video image frames; and

receiving from the machine learning model, a detection of any object crossing the crossline in the scene and a direction the object crosses the crossline during capturing of the sequence of video image frames.

2. The method of claim 1 , wherein, in constructing the combined image frame, the extracted two or more lines of pixels are arranged adjacent to each other in the same order as the order in which the sequence of video image frames from which the two or more lines of pixels are extracted.

3. The method of claim 1 , wherein, on condition that an object crosses the crossline, the combined image frame comprises a combined object consisting of representations of portions of the object in the extracted two or more lines of pixels, and wherein, in the act of sending the combined image, the combined image frame is sent to a machine learning model that detects a combined object in a combined image frame, which combined object represents an object crossing a crossline in a scene and a direction the object crosses the crossline during capturing of the sequence of video image frames.

4. The method of claim 1 , wherein the machine learning model is a neural network.

5. The method of claim 1 , wherein the longest distance between two lines of the extracted two or more lines is less than or equal to 20 lines, and preferably less than or equal to 10 lines.

6. The method of claim 1 , wherein the extracted two or more lines are adjacent.

7. The method of claim 1 , wherein four or more lines are extracted.

8. A non-transitory computer-readable storage medium having stored thereon instructions for implementing a method when executed in a device having an image sensor and a processor, the method for detecting objects crossing a crossline in a scene captured by a video camera and a direction the object crosses the crossline, the method comprising:

capturing a sequence of video image frames of the scene;

creating a combined image frame by:

extracting two or more lines of pixels of each image frame of the sequence of video image frames, wherein the two or more lines of pixels of each video image frame are parallel to, and correspond to, the crossline in the scene such that if a representation of an object crosses the two or more lines of pixels in the sequence of video image frames, this corresponds to the object having crossed the crossline in the scene, and

constructing the combined image frame by arranging the extracted two or more lines of pixels of each video image frame of the sequence of video image frames adjacent to each other;

sending the combined image frame to a machine learning model trained to detect a combined image frame created from a sequence of video image frames that represents an object crossing a crossline in a scene and a direction the object crosses the crossline during capturing of the sequence of video image frames; and

receiving from the machine learning model, a detection of any object crossing the crossline in the scene and a direction the object crosses the crossline during capturing of the sequence of video image frames.

9. An image processing device for detecting objects crossing a crossline in a scene captured by a video camera and a direction the object crosses the crossline, the image processing device comprising:

an image sensor for capturing a sequence of video image frames of the scene; and

circuitry arranged to execute:

a creating function arranged to create a combined image frame by:

extracting two or more lines of pixels of each video image frame of the sequence of video image frames, wherein the two or more lines of pixels of each video image frame are parallel to, and correspond to, the crossline in the scene such that if a representation of an object crosses the two or more lines of pixels in the sequence of video image frames, this corresponds to the object having crossed the crossline in the scene, and

constructing the combined image frame by arranging the extracted two or more lines of pixels of the sequence of video image frames adjacent to each other;

a sending function arranged to send the combined image frame to a machine learning model trained to detect a combined image frame created from a sequence of video image frames that represents an object crossing a crossline in a scene and a direction the object crosses the crossline during capturing of the sequence of video image frames; and

a receiving function arranged to receive, from the machine learning model, a detection of any object crossing the crossline in the scene and a direction the object crosses the crossline during capturing of the sequence of video image frames.

10. The image processing device of claim 9 , wherein, in the creating function, the extracted two or more lines of pixels are arranged adjacent to each other in the same order as the order in the sequence of video image frames of the image video frames from which the two or more lines of pixels are extracted.

11. The image processing device of claim 9 , wherein, on condition that an object crosses the crossline, the combined image frame comprises a combined object consisting of representations of portions of an object in the extracted two or more lines, and wherein, in the sending function, the combined image frame is sent to a machine learning model that detects a combined object in a combined image frame, which combined object represents an object crossing a crossline in a scene and a direction the object crosses the crossline during capturing of the sequence of video image frames.

12. The image processing device of claim 9 , wherein the machine learning model is a neural network.

13. A method of training a machine learning model to detect that an object crosses a crossline in a scene captured by a video camera and a direction the object crosses the crossline, the method comprising:

creating a combined image frame for training by:

extracting two or more lines of pixels of each image frame of a training sequence of video image frames, wherein the two or more lines of pixels of each video image frame are parallel to, and correspond to, the crossline in the training scene, such that if a representation of an object crosses the two or more lines of pixels in the training sequence of video image frames, this corresponds to the object having crossed the crossline in the training scene, and wherein a known object crosses the crossline in the training scene in a known direction during capturing of the training sequence of video image frames, and

constructing the combined image frame for training by arranging the extracted two or more lines of pixels of each video image frame of the sequence of video image frames adjacent to each other;

labelling the combined image frame for training with information specifying that the combined image frame for training represents the known object crossing the crossline in the training scene in the known direction during capturing of the training sequence of video image frames; and

having the labelled combined image frame for training as input, training the machine learning model to detect a combined image frame created from a sequence of video image frames that represents an object crossing a crossline in a scene and a direction the object crosses the crossline during capturing of the sequence of video image frames.

14. The method of claim 13 , wherein, in constructing the combined image frame for training, the extracted two or more lines of pixels are arranged adjacent to each other in the same order as the order in which the training sequence of video image frames from which the two or more lines of pixels are extracted.

15. The method of claim 13 , wherein the combined image frame for training comprises a combined object for training consisting of representations of portions of the known object in the extracted two or more lines of pixels,

wherein, in the act of labelling the combined image frame for training, the information specifying that combined image frame for training represents the known object crossing the crossline in the training scene in the known direction during capturing of the training sequence of video image frames specifies that the combined object for training represents the known object crossing the crossline in the training scene in the known direction during capturing of the training sequence of video image frames, and

wherein the act of training the machine learning model comprises:

having the labelled training image as input, training the machine learning model to detect a combined object in a combined image frame created from a sequence of video image frames, which combined object represents an object crossing a crossline in a scene and a direction the object crosses the crossline during capturing of the sequence of video image frames.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 31, 2023
From: SJÖBORG, MARTIN
To: AXIS AB
Reel/Frame 063813/0477 →
Priority Claims (1)
EP 22181445 · Jun 28, 2022 · regional
Continuity (1)
Related Publication 20230419508A1 · Dec 28, 2023
References Cited (20)
US 5103433A · Imhof · 1992 [cited by examiner]
US 6545705B1 · Sigel · 2003 [cited by examiner]
US 11288820B2 · Boult · 2022 [cited by examiner]
US 20140132771A1 · Gustafsson · 2014 [cited by examiner]
US 20160110613A1 · Ghanem et al. · 2016 [cited by applicant]
US 20170169560A1 · Johansson · 2017 [cited by examiner]
US 20180096484A1 · Yasuda · 2018 [cited by examiner]
US 20180260632A1 · Ebihara · 2018 [cited by examiner]
US 20190005336A1 · Schulte · 2019 [cited by applicant]
US 20190378283A1 · Boult · 2019 [cited by applicant]
US 20200175693A1 · Takada · 2020 [cited by examiner]
US 20210027067A1 · Druihle · 2021 [cited by examiner]
US 20210034877A1 · Druihle · 2021 [cited by examiner]
US 20210133483A1 · Prabhu · 2021 [cited by examiner]
US 20210334984A1 · Gauci · 2021 [cited by examiner]
US 20220020159A1 · Falk · 2022 [cited by examiner]
US 20220375025A1 · Ardö · 2022 [cited by examiner]
US 20230046840A1 · Ramanathan · 2023 [cited by examiner]
US 20230419508A1 · Sjöborg · 2023 [cited by examiner]
US 20240203221A1 · Yamada · 2024 [cited by examiner]