IP Library › Granted Patent US 12,361,703
Granted Patent B2
US 12,361,703 · App. 17/625,390 · Granted Jul 15, 2025

Computer software module arrangement, a circuitry arrangement, an arrangement and a method for improved object detection

Inventors: Ashkan Kalantari (Malmö, SE); Héctor Caltenco (Oxie, SE); Saeed Bastani (Dalby, SE); Yun Li (Lund, SE)
Assignee: Telefonaktiebolaget LM Ericsson (Publ)
G06V10/95G06V10/17G06V10/454G06V10/82G06V20/52G06V10/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,703
App. No.
17/625,390
Granted
Jul 15, 2025
Kind
B2
Abstract

An object detection arrangement having a controller configured to: a) receive a plurality of image data streams; b) perform feature extraction on each of the received a plurality of images providing a plurality of feature data streams; and to c) perform a common feature extraction based on the plurality of feature data streams providing as common feature data stream for object detection.

Claims (28)

1. An object detection arrangement comprising:

a plurality of image capturing devices; and

a controller configured to:

receive a plurality of image data streams, each of the plurality of image data streams originating from one of the plurality of image capturing devices;

perform local feature extraction on each of the received plurality of images to provide a plurality of feature data streams;

perform a common feature extraction at least by applying a plurality of neural network layers exhibiting temporal dynamic behavior to the plurality of feature data streams, the plurality of neural network layers being arranged such that one of the plurality of neural network layers receives input from another of the plurality of neural network layers not immediately preceding the one of the plurality of neural network layers, the common feature extraction being based on the plurality of feature data streams to provide a common feature data stream for object detection; and

at least one of the plurality of image capturing devices is configured to perform the local feature extraction on a corresponding image data stream.

2. The object detection arrangement according to claim 1 , wherein the controller is further configured to perform object detection based on the common feature data stream.

3. The object detection arrangement according to claim 2 , wherein the object detection is based on a Deep Neural Network model.

4. The object detection arrangement according to claim 1 , wherein at least one of the plurality of neural network layers exhibiting temporal dynamic behavior is a recurrent neural network (RNN).

5. The object detection arrangement according to claim 1 , wherein one image data stream comprises images at least partially overlapping images in another image data stream.

6. The object detection arrangement according to claim 1 , wherein the controller is further configured to combine the plurality of feature data streams into a combined feature data stream.

7. The object detection arrangement according to claim 1 , wherein the object detection arrangement is a one of smartphone and a tablet computer.

8. The object detection arrangement according to claim 1 , wherein the object detection arrangement is an optical see-through device.

9. The object detection arrangement according to claim 1 , wherein the object detection arrangement is arranged to be used in image retrieval, industrial use, robotic vision and/or video surveillance.

10. The object detection arrangement according to claim 2 , wherein at least one of the plurality of neural network layers exhibiting temporal dynamic behavior is a recurrent neural network (RNN).

11. The object detection arrangement according to claim 2 , wherein one image data stream comprises images at least partially overlapping images in another image data stream.

12. The object detection arrangement according to claim 2 , further comprising a plurality of image capturing devices, wherein each image data stream originates from one of the image capturing devices.

13. A method for object detection in an object detection arrangement, the method comprising:

receiving a plurality of image data streams, each of the plurality of image data streams originating from one of a plurality of image capturing devices;

performing local feature extraction on each of the received plurality of images to provide a plurality of feature data streams;

performing a common feature extraction at least by applying a plurality of neural network layers exhibiting temporal dynamic behavior to the plurality of feature data streams, the plurality of neural network layers being arranged such that one of the plurality of neural network layers receives input from another of the plurality of neural network layers not immediately preceding the one of the plurality of neural network layers, the common feature extraction being based on the plurality of feature data streams providing to provide a common feature data stream for object detection; and

configuring at least one of the plurality of image capturing devices to perform the local feature extraction on a corresponding image data stream.

14. A non-transitory computer-readable storage medium storing computer instructions that when loaded into and executed by a controller of an object detection arrangement causes the object detection arrangement to perform a method, the method comprising:

receiving a plurality of image data streams, each of the plurality of image data streams originating from one of a plurality of image capturing devices;

performing local feature extraction on each of the received plurality of images to provide a plurality of feature data streams;

performing a common feature extraction at least by applying a plurality of neural network layers exhibiting temporal dynamic behavior to the plurality of feature data streams, the plurality of neural network layers being arranged such that one of the plurality of neural network layers receives input from another of the plurality of neural network layers not immediately preceding the one of the plurality of neural network layers, the common feature extraction being based on the plurality of feature data streams to provide a common feature data stream for object detection; and

configuring at least one of the plurality of image capturing devices to perform the local feature extraction on a corresponding image data stream.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 7, 2022
From: BASTANI, SAEED; CALTENCO, HECTOR; KALANTARI, ASHKAN; LI, YUN
To: TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
Reel/Frame 059534/0229 →
Continuity (1)
Related Publication 20220284689A1 · Sep 8, 2022
References Cited (29)
US 10055853B1 · Fisher et al. · 2018 [cited by applicant]
US 10198655B2 · Hotson et al. · 2019 [cited by applicant]
US 20030215141A1 · Zakrzewski et al. · 2003 [cited by applicant]
US 20060083423A1 · Brown et al. · 2006 [cited by applicant]
US 20150304634A1 · Karvounis · 2015 [cited by applicant]
US 20160275375A1 · Kant · 2016 [cited by applicant]
US 20170206415A1 · Redden · 2017 [cited by applicant]
US 20170220854A1 · Yang · 2017 [cited by examiner]
US 20180005054A1 · Yu et al. · 2018 [cited by applicant]
US 20190025773A1 · Yang et al. · 2019 [cited by applicant]
US 20190065910A1 · Wang et al. · 2019 [cited by applicant]
US 20190087975A1 · Versace et al. · 2019 [cited by applicant]
US 20190138855A1 · Sohn · 2019 [cited by examiner]
CN 105512680A · 2016 [cited by applicant]
CN 108038420A · 2018 [cited by applicant]
CN 108629368A · 2018 [cited by applicant]
GB 2549836A · 2017 [cited by applicant]
KR 20180125885 · 2018 [cited by applicant]
WO 2018088170A1 · 2018 [cited by applicant]
Lee, Daehyun, Jongmin Lee, and Kee-Eung Kim. “Multi-view automatic lip-reading using neural network.” Computer Vision—ACCV 2016 Workshops: ACCV 2016 International Workshops, Taipei, Taiwan, Nov. 20-24, 2016, Revised Sel… [cited by examiner]
Donahue, Jeff, et al. “Long-term Recurrent Convolutional Networks for Visual Recognition and Description.” arXiv preprint arXiv: 1411.4389v4 (2016). (Year: 2016). [cited by examiner]
Yu, Sheng, et al. “A novel recurrent hybrid network for feature fusion in action recognition.” Journal of Visual Communication and Image Representation 49 (2017): 192-203. (Year: 2017). [cited by examiner]
International Search Report and Written Opinion dated May 20, 2020 for Application No. PCT/EP2019/069244 filed Jul. 17, 2019, consisting of 10 pages. [cited by applicant]
Donahue et al. “Long-term Recurrent Convolutional Networks for Visual Recognition and Description” California Univ Berkeley Elec Eng Sci Dept; May 31, 2016, consisting of 14 pages. [cited by applicant]
Daehyun Lee et al. “Multi-view Automatic Lip-Reading Using Neural Network” School of Computing, KAIST, Daejeon, Korea, 2017, consisting of 13 pages. [cited by applicant]
Kavi et al. “Multiview fusion for activity recognition using deep neural networks”; Journal of Electronic Imaging; 2016, consisting of 9 pages. [cited by applicant]
Tao et al. “Moving object recognition using multi-view three-dimensional convolutional neural networks”; Neural Comput &Applic; CrossMark; 2017, consisting of 9 pages. [cited by applicant]
Liu et al. “Mobile Video Object Detection with Temporally-Aware Feature Maps”; arXiv:1711.06368v2; Mar. 28, 2018; consisting of 10 pages. [cited by applicant]
Zhang et al. “ML-LocNet: Improving Object Localization with Multi-view Learning Network”; National University of Singapore; ECCV 2018, consisting of 16 pages. [cited by applicant]