IP Library Granted Patent US 12,511,912
Granted Patent B2
US 12,511,912 · App. 18/187,508 · Granted Dec 30, 2025

Object detection and tracking

Inventors: Darya Frolova (Hod Hasharon, IL); Shahar Ben Ezra (Hod Hasharon, IL); Pavel Kisilev (Hod Hasharon, IL); Xiaoli She (Shenzhen, CN); Yu Xie (Shenzhen, CN)
Assignee: Shenzhen Yinwang Intelligent Technologies Co., Ltd.
G06V20/58G06T5/70G06V10/82G06V20/588G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,511,912
App. No.
18/187,508
Granted
Dec 30, 2025
Kind
B2
Abstract

A method and a computing device for object detection and tracking from a video input are described. The method and the computing device may be used to, for example, track objects of interest, such as lane markings, in traffic. A plurality of frames corresponding to a video may be analyzed in a spatiotemporal domain by a neural network. The neural network may be trained using data synthesized in the spatiotemporal domain.

Claims (63)

1 . A method, comprising:

obtaining a plurality of frames, corresponding to a video, comprising features of interest;

forming, based on the plurality of frames comprising features of interest, a spatiotemporal data volume, wherein two dimensions of the spatiotemporal data volume correspond to spatial dimensions of the plurality of frames, and wherein one dimension of the spatiotemporal data volume corresponds to a temporal dimension of the plurality of frames;

slicing the spatiotemporal data volume along a plurality of surfaces, producing a plurality of spatiotemporal images, wherein each spatiotemporal image in the plurality of spatiotemporal images corresponds to the spatiotemporal data volume along a corresponding surface in the plurality of surfaces;

enhancing the features of interest, of the plurality of frames, in the plurality of spatiotemporal images using a neural network, thereby producing enhanced features of interest of a processed plurality of spatiotemporal images of the plurality of frames, wherein the features of interest comprise at least one geometric shape in the plurality of spatiotemporal images, and wherein the enhancing the features of interest in the plurality of spatiotemporal images to produce the enhanced features of interest comprises performing at least one operation taken from the group consisting of:

connecting disconnected parts of the at least one geometrical shape in the plurality of spatiotemporal images;

extracting the at least one geometrical shape in the plurality of spatiotemporal images; and

classifying the at least one geometrical shape in the plurality of spatiotemporal images; and

projecting the enhanced features of interest of the processed plurality of spatiotemporal images onto the plurality of frames.

2 . The method according to claim 1 , wherein the obtaining the plurality of frames that comprises features of interest comprises:

obtaining a plurality of input frames corresponding to the video; and

performing feature extraction on the plurality of input frames, thereby producing the plurality of frames and the features of interest in the plurality of frames.

3 . The method according to claim 1 , wherein the features of interest correspond to objects of interest in vehicle traffic.

4 . The method according to claim 3 , wherein the objects of interest in vehicle traffic comprise at least one object taken from the group consisting of:

a lane marking;

a segment of road; and

an object in traffic to be tracked.

5 . The method according to claim 1 , wherein the neural network comprises a convolutional neural network.

6 . The method according to claim 1 , wherein the neural network is configured to be trained using synthesized spatiotemporal images comprising synthesized features of interest.

7 . A non-transitory computer readable medium comprising program code configured to perform a method comprising:

obtaining a plurality of frames, corresponding to a video, comprising features of interest;

forming, based on the plurality of frames comprising features of interest, a spatiotemporal data volume, wherein two dimensions of the spatiotemporal data volume correspond to spatial dimensions of the plurality of frames, and wherein one dimension of the spatiotemporal data volume corresponds to a temporal dimension of the plurality of frames;

slicing the spatiotemporal data volume along a plurality of surfaces, producing a plurality of spatiotemporal images, wherein each spatiotemporal image in the plurality of spatiotemporal images corresponds to the spatiotemporal data volume along a corresponding surface in the plurality of surfaces;

enhancing the features of interest, of the plurality of frames, in the plurality of spatiotemporal images using a neural network, thereby producing enhanced features of interest of a processed plurality of spatiotemporal images of the plurality of frames, wherein the features of interest comprise at least one geometric shape in the plurality of spatiotemporal images, and wherein the enhancing the features of interest in the plurality of spatiotemporal images to produce the enhanced features of interest comprises performing at least one operation taken from the group consisting of:

connecting disconnected parts of the at least one geometrical shape in the plurality of spatiotemporal images;

extracting the at least one geometrical shape in the plurality of spatiotemporal images; and

classifying the at least one geometrical shape in the plurality of spatiotemporal images; and

projecting the enhanced features of interest of the processed plurality of spatiotemporal images onto the plurality of frames.

8 . The non-transitory computer-readable medium according to claim 7 , wherein the obtaining the plurality of frames that comprises features of interest comprises:

obtaining a plurality of input frames corresponding to the video; and

performing feature extraction on the plurality of input frames, thereby producing the plurality of frames and the features of interest in the plurality of frames.

9 . The non-transitory computer-readable medium according to claim 7 , wherein the features of interest correspond to objects of interest in vehicle traffic.

10 . The non-transitory computer-readable medium according to claim 9 , wherein the objects of interest in vehicle traffic comprise at least one object taken from the group consisting of:

a lane marking;

a segment of road; and

an object in traffic to be tracked.

11 . The non-transitory computer-readable medium according to claim 7 , wherein the neural network comprises a convolutional neural network.

12 . The non-transitory computer-readable medium according to claim 7 , wherein the neural network is configured to be trained using synthesized spatiotemporal images comprising synthesized features of interest.

13 . A computing device comprising:

a processor; and

a non-transitory computer-readable medium including computer-executable instructions that, when executed by the processor, facilitate carrying out a method comprising:

obtaining a plurality of frames, corresponding to a video, comprising features of interest;

forming, based on the plurality of frames comprising features of interest, a spatiotemporal data volume, wherein two dimensions of the spatiotemporal data volume correspond to spatial dimensions of the plurality of frames, and wherein one dimension of the spatiotemporal data volume corresponds to a temporal dimension of the plurality of frames;

slicing the spatiotemporal data volume along a plurality of surfaces, producing a plurality of spatiotemporal images, wherein each spatiotemporal image in the plurality of spatiotemporal images corresponds to the spatiotemporal data volume along a corresponding surface in the plurality of surfaces;

enhancing the features of interest, of the plurality of frames, in the plurality of spatiotemporal images using a neural network, thereby producing enhanced features of interest of a processed plurality of spatiotemporal images of the plurality of frames, wherein the features of interest comprise at least one geometric shape in the plurality of spatiotemporal images, and wherein the enhancing the features of interest in the plurality of spatiotemporal images to produce the enhanced features of interest comprises performing at least one operation taken from the group consisting of:

connecting disconnected parts of the at least one geometrical shape in the plurality of spatiotemporal images;

extracting the at least one geometrical shape in the plurality of spatiotemporal images; and

classifying the at least one geometrical shape in the plurality of spatiotemporal images; and

projecting the enhanced features of interest of the processed plurality of spatiotemporal images onto the plurality of frames.

14 . The computing device according to claim 13 , wherein the obtaining the plurality of frames that comprises features of interest comprises:

obtaining a plurality of input frames corresponding to the video; and

performing feature extraction on the plurality of input frames, thereby producing the plurality of frames and the features of interest in the plurality of frames.

15 . The computing device according to claim 13 , wherein the features of interest correspond to objects of interest in vehicle traffic.

16 . The computing device according to claim 15 , wherein the objects of interest in vehicle traffic comprise at least one object taken from the group consisting of:

a lane marking;

a segment of road; and

an object in traffic to be tracked.

17 . The computing device according to claim 13 , wherein the neural network comprises a convolutional neural network.

18 . The computing device according to claim 13 , wherein the neural network is configured to be trained using synthesized spatiotemporal images comprising synthesized features of interest.

19 . A vehicle comprising the computing device according to claim 13 .

20 . The vehicle according to claim 19 , wherein the obtaining the plurality of frames that comprises features of interest comprises:

obtaining a plurality of input frames corresponding to the video; and

performing feature extraction on the plurality of input frames, thereby producing the plurality of frames and the features of interest in the plurality of frames.

Assignments (3)
CHANGE OF NAME Recorded May 1, 2026
From: SHENZHEN YINWANG INTELLIGENT TECHNOLOGIES CO., LTD.
To: YINWANG INTELLIGENT TECHNOLOGIES CO., LTD.
Reel/Frame 075316/0074 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2026
From: FROLOVA, DARYA; KISILEV, PAVEL; SHE, XIAOLI; XIE, YU
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 075669/0746 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 12, 2024
From: HUAWEI TECHNOLOGIES CO., LTD.
To: SHENZHEN YINWANG INTELLIGENT TECHNOLOGIES CO., LTD.
Reel/Frame 069336/0125 →
Continuity (2)
Continuation PCTCN2020116762 · Sep 22, 2020
Related Publication 20230237811A1 · Jul 27, 2023
References Cited (32)
US 8335350B2 · Wu · 2012 [cited by applicant]
US 8358808B2 · Malinovskiy et al. · 2013 [cited by applicant]
US 8384787B2 · Wu · 2013 [cited by applicant]
US 8538082B2 · Zhao et al. · 2013 [cited by applicant]
US 9672430B2 · Sull et al. · 2017 [cited by applicant]
US 20130051612A1 · Prokhorov · 2013 [cited by examiner]
US 20160254024A1 · Wu · 2016 [cited by examiner]
US 20160379055A1 · Loui et al. · 2016 [cited by applicant]
US 20200394412A1 · Carreira · 2020 [cited by examiner]
US 20210357647A1 · Na · 2021 [cited by examiner]
WO WO2019137912A1 · 2019 [cited by examiner]
Tran et al, Learning Spatiotemporal Features with 3D Convolutional Networks, 2015, IEEE International Conference on Computer Vision, pp. 1-10. (Year: 2015). [cited by examiner]
Le et al, Video Salient Object Detection using Spatiotemporal Deep Features, 2018, arXiv:1708.01447v3, pp. 1-14. (Year: 2018). [cited by examiner]
Rapantzikos et al, Spatiotemporal visual attention architecture for video analysis, 2004, IEEE 6th Workshop on Multimedia Signal Processing, pp. 1-5. (Year: 2005). [cited by examiner]
Wan et al, Action Recognition Based on Two-Stream Convolutional Networks, 2020, IEEE Digital Object Identifier, 8 (2020): 85284-85293. (Year: 2020). [cited by examiner]
Aslan et al., “Deep Convolutional Generative Adversarial Networks Based Flame Detection in Video,” arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, XP081025665, total 5 pages … [cited by applicant]
Borkar et al., “An Efficient Method to Generate Ground Truth for Evaluating Lane Detection Systems,” 2010 IEEE International Conference on Acoustics, Speech and Signal Processing, Dallas, TX, USA, IEEE—Institute of Elec… [cited by applicant]
Borkar et al., “A Novel Lane Detection System With Efficient Ground Truth Generation,” IEEE Transactions on Intelligent Transportation Systems, vol. 13, No. 1, Mar. 2012, XP011427541, pp. 365-374, IEEE—Institute of Elec… [cited by applicant]
Jung et al., “Efficient Lane Detection Based on Spatiotemporal Images,” in IEEE Transactions on Intelligent Transportation Systems, vol. 17, No. 1, pp. 1-7, IEEE—Institute of Electrical and Electronics Engineers, New Yo… [cited by applicant]
Das et al., “Enhanced Algorithm of Automated Ground Truth Generation and Validation for Lane Detection System by M2BMT,” in IEEE Transactions on Intelligent Transportation Systems, vol. 18, No. 4, pp. 1-10, IEEE—Institu… [cited by applicant]
Garnett et al., “3D-LaneNet: End-to-End 3D Multiple Lane Detection,” arXiv: 1811.10203v3, total 14 pages, (Sep. 10, 2019). [cited by applicant]
Behrendt et al., “Deep Learning Lane Marker Segmentation From Automatically Generated Labels,” 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada, pp. 777-782, IEEE—In… [cited by applicant]
Behrendt et al., “Unsupervised Labeled Lane Markers Using Maps,” 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), pp. 832-839, IEEE—Institute of Electrical and Electronics Engineers, New York,… [cited by applicant]
Sarraf et al., “Ground Truth and Performance Evaluation of Lane Border Detection,” International Conference on Computer Vision and Graphics, pp. 1-8 (Sep. 15, 2014). [cited by applicant]
Zou et al., “Robust Lane Detection from Continuous Driving Scenes Using Deep Neural Networks,” IEEE Transactions on Vehicular Technology, pp. 1-15, IEEE—Institute of Electrical and Electronics Engineers, New York, New Y… [cited by applicant]
Tapia-Espinoza et al., “Robust Lane Sensing and Departure Warning under Shadows and Occlusions,” Sensors 2013, vol. 13, doi:10.3390/s130303270, ISSN 1424-8220, pp. 3270-3298 (Mar. 11, 2013). [cited by applicant]
Nowruzi et al., “How much real data do we actually need: Analyzing object detection performance using synthetic and real data,” ICML Workshop on AI for Autonomous Driving, total 11 pages (Jul. 16, 2019). [cited by applicant]
Neven et al., “Towards End-to-End Lane Detection: an Instance Segmentation Approach,” CVPR, total 7 pages (Feb. 15, 2018). [cited by applicant]
Pan et al., “Spatial as Deep: Spatial CNN for Traffic Scene Understanding,” arXiv:1712.06080v1, total 8 pages (Dec. 17, 2017). [cited by applicant]
Hou et al., “Learning Lightweight Lane Detection CNNs by Self Attention Distillation,” ICCV, total 11 pages (Aug. 2, 2019). [cited by applicant]
Ko et al., “Key Points Estimation and Point Instance Segmentation Approach for Lane Detection,” arXiv:2002.06604v4, pp. 1-10 (Sep. 14, 2020). [cited by applicant]
Rav-Acha et al., “Spatio-Temporal Video Warping,” SIGGRAPH '05: ACM SIGGRAPH 2005 Sketches, total 1 page (Jul. 31, 2005). [cited by applicant]