IP Library › Granted Patent US 12,738,046
Granted Patent B2
US 12,738,046 · App. 18/639,763 · Granted Sep 15, 2026

Localization by a neural network

Inventors: Csaba Mate Jozsa (Budapest, HU); Gábor Sörös (Budapest, HU); Krisztián Zsolt Varga (Budapest, HU); Lóránt Farkas (Budapest, HU)
Assignee: Nokia Solutions and Networks Oy
G06V10/82G06V10/757G06V10/7715
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,738,046
App. No.
18/639,763
Granted
Sep 15, 2026
Kind
B2
Abstract

Encoding a first 3D point cloud of a first coordinate system into a first encoded map comprising first feature center points and first feature vectors; encoding a second 3D point cloud of a second coordinate system into a second encoded map comprising second feature center points and second feature vectors; adapting the first input feature vectors based on the first input feature vectors and the second input feature vectors, to obtain a first joint map; adapting the second input feature vectors based on the first input feature vectors and the second input feature vectors to obtain a second joint map; checking whether a correlation condition is fulfilled; extracting the coordinates of the first joint feature center points and the second joint feature center points if the correlation condition is fulfilled; calculating a transformation between the first coordinate system and the second coordinate system based on the pairs of extracted coordinates.

Claims (94)

1 . Apparatus comprising:

one or more processors, and memory storing instructions that, when executed by the one or more processors, cause the apparatus to perform:

inputting a first 3D point cloud of a first modality and a first grid comprising one or more first cells into a first encoder layer, wherein coordinates of points of the first 3D point cloud are indicated in a first coordinate system;

encoding, by the first encoder layer, the first 3D point cloud of the first modality into a first encoded map of the first modality, wherein the first encoded map of the first modality comprises, for each of the first cells of the first grid, a respective first feature center point and a respective first feature vector, and the coordinates of the first feature center points are indicated in the first coordinate system;

inputting a first encoded input map into a first matching descriptor layer of a first hierarchical layer, wherein the first encoded input map is based on the first encoded map of the first modality and comprises, for each of the first cells, a respective first joint feature center point and a respective first input feature vector, wherein the coordinates of the first joint feature center points are indicated in the first coordinate system;

inputting a second 3D point cloud of the first modality and a second grid comprising one or more second cells into a second encoder layer, wherein the coordinates of points of the second 3D point cloud are indicated in a second coordinate system;

encoding, by the second encoder layer, the second 3D point cloud of the first modality into a second encoded map of the first modality, wherein the second encoded map of the first modality comprises, for each of the second cells of the second grid, a respective second feature center point and a respective second feature vector, and the coordinates of the second feature center points are indicated in the second coordinate system;

inputting a second encoded input map into the first matching descriptor layer, wherein the second encoded input map is based on the second encoded map of the first modality and comprises, for each of the second cells, a respective second joint feature center point and a respective second input feature vector, wherein the coordinates of the second joint feature center points are indicated in the second coordinate system;

adapting, by the first matching descriptor layer, each of the first input feature vectors based on the first input feature vectors, with their first feature center points, and the second input feature vectors, with their second feature center points, to obtain a first joint map, wherein the first joint map comprises, for each of the first joint feature center points, a respective first joint feature vector;

adapting, by the first matching descriptor layer, each of the second input feature vectors based on the first input feature vectors, with their first feature center points, and the second input feature vectors, with their second feature center points, to obtain a second joint map, wherein the second joint map comprises, for each of the second joint feature center points, a respective second joint feature vector;

calculating, by a first optimal matching layer, for each of the first cells and each of the second cells, a similarity between the first joint feature vector of the respective first cell and the second joint feature vector of the respective second cell;

calculating, by the first optimal matching layer, for each of the first cells and each of the second cells, a first point correlation between the respective first cell and the respective second cell based on the similarities between the first joint feature vector of the respective first cell and the second joint feature vectors of the second cells and based on the similarities between the second joint feature vector of the respective second cell and the first joint feature vectors of the first cells;

checking, for each of the first cells and each of the second cells, whether at least one of one or more first correlation conditions for the respective first cell and the respective second cell is fulfilled, wherein the one or more first correlation conditions comprise:

the first point correlation between the respective first cell and the respective second cell is larger than a first correlation threshold of the first hierarchical layer; or

the first point correlation between the respective first cell and the respective second cell is among the largest k values of the first point correlations, and k is a fixed value;

extracting, for each of the first cells and each of the second cells, the coordinates of the first joint feature center point of the respective first cell in the first coordinate system and the coordinates of the second joint feature center point of the respective second cell in the second coordinate system to obtain a respective pair of extracted coordinates of the first hierarchical layer in response to checking that the at least one of the one or more first correlation conditions for the respective first cell and the respective second cell is fulfilled;

calculating an estimated transformation between the first coordinate system and the second coordinate system based on the pairs of extracted coordinates of the first hierarchical layer; and

wherein the first encoder layer, the second encoder layer, and the first matching descriptor layer of the first hierarchical layer are configured by respective first parameters, and the instructions, when executed by the one or more processors, further cause the apparatus to perform:

training of the apparatus for improved localization in different environments by determining the first parameters such that a deviation between the estimated transformation and a known transformation between the first coordinate system and the second coordinate system becomes less than a transformation estimation threshold and such that the first point correlations are increased.

2 . The apparatus according to claim 1 , wherein the first modality comprises one of photographic images, LIDAR images, and ultrasound images.

3 . The apparatus according to claim 1 , wherein the first encoder layer is the same as the second encoder layer.

4 . The apparatus according to claim 1 , wherein the instructions, when executed by the one or more processors, further cause the apparatus to perform at least one of the following:

generating the first 3D point cloud based on a sequence of first images of the first modality, wherein a respective location in the first coordinate system is associated to each of the first images; or

generating the second 3D point cloud based on a sequence of second images of the first modality, wherein a respective location in the second coordinate system is associated to each of the second images.

5 . The apparatus according to claim 1 , wherein the first encoded input map is the first encoded map, the first joint feature center points are the first feature center points, the second encoded input map is the second encoded map, and the second joint feature center points are the second feature center points.

6 . The apparatus according to claim 1 , wherein the instructions, when executed by the one or more processors, further cause the apparatus to perform:

inputting a third 3D point cloud of a second modality and the first grid into a third encoder layer, wherein the coordinates of points of the third 3D point cloud are indicated in the first coordinate system, and the second modality is different from the first modality; encoding, by the third encoder layer, the third 3D point cloud of the second modality into a third encoded map of the second modality, wherein the third encoded map of the second modality comprises, for each of the first cells of the first grid, a respective third feature center point and a respective third feature vector, and the coordinates of the third feature center points are indicated in the first coordinate system;

inputting the third encoded map into a first neural fusion layer;

inputting the first encoded map into the first neural fusion layer;

generating, by the first neural fusion layer, based on the first encoded map and the third encoded map, the first encoded input map, wherein, for each of the first cells, the respective first joint feature center point is based on the first feature center point of the respective first cell and the third feature center point of the respective first cell, and the respective first input feature vector is based on the first feature vector and the first feature center point of the respective first cell adapted based on the first feature vectors with their first feature center points and the third feature vectors, with their third feature center points, and based on the third feature vector and the third feature center point of the respective first cell adapted by the first feature vectors, with their first feature center points, and the third feature vectors, with their third feature center points;

inputting a fourth 3D point cloud of the second modality and the second grid into a fourth encoder layer, wherein the coordinates of points of the fourth 3D point cloud are indicated in the second coordinate system;

encoding, by the fourth encoder layer, the fourth 3D point cloud of the second modality into a fourth encoded map of the second modality, wherein the fourth encoded map of the second modality comprises, for each of the second cells of the second grid, a respective fourth feature center point and a respective fourth feature vector, and the coordinates of the fourth feature center points are indicated in the second coordinate system;

inputting the fourth encoded map into a second neural fusion layer;

inputting the second encoded map into the second neural fusion layer;

generating, by the second neural fusion layer, based on the second encoded map and the fourth encoded map, the second encoded input map, wherein, for each of the second cells, the respective second joint feature center point is based on the second feature center point of the respective second cell and the fourth feature center point of the respective second cell, and the respective second input feature vector is based on the second feature vector and the second feature center point of the respective second cell adapted by the second feature vectors, with their second feature center points, and the fourth feature vectors, with their fourth feature center points, and based on the fourth feature vector and the fourth feature center point of the respective second cell adapted by the second feature vectors, with their second feature center points, and the fourth feature vectors, with their fourth feature center points.

7 . The apparatus according to claim 6 , wherein the second modality comprises at least one of photographic images, RF fingerprints, LIDAR images, or ultrasound images.

8 . The apparatus according to claim 6 , wherein the first encoder layer, the second encoder layer, and the first matching descriptor layer of the first hierarchical layer are configured by respective first parameters, and the instructions, when executed by the one or more processors, further cause the apparatus to perform training of the apparatus to determine the first parameters such that a deviation between the estimated transformation and a known transformation between the first coordinate system and the second coordinate system becomes less than a transformation estimation threshold and such that the first point correlations are increased, and

wherein the third encoder layer, the fourth encoder layer, the first neural fusion layer, and the second neural fusion layer are configured by respective second parameters, and the instructions, when executed by the one or more processors, further cause the apparatus to perform the training of the apparatus to determine the first and second parameters such that the deviation between the estimated transformation and the known transformation between the first coordinate system and the second coordinate system becomes less than the transformation estimation threshold and such that the first point correlations are increased.

9 . The apparatus according to claim 6 , wherein the first encoded input map is the first encoded map, the first joint feature center points are the first feature center points, the second encoded input map is the second encoded map, and the second joint feature center points are the second feature center points, and wherein

the first images are taken by a first sensor of a first sensor device;

the second images are taken by a second sensor of a second sensor device; and the instructions, when executed by the one or more processors, further cause the apparatus to perform conducting a plurality of first measurements of the second modality with a third sensor of the first sensor device;

conducting a plurality of second measurements of the second modality with a fourth sensor of the second sensor device;

generating the third 3D point cloud of the second modality based on the plurality of first measurements, the locations associated to the first images, and a calibration between the first sensor of the first sensor device and the third sensor of the first sensor device;

generating the fourth 3D point cloud of the second modality based on the plurality of second measurements, the locations associated to the second images, and a calibration between the second sensor of the second sensor device and the fourth sensor of the second sensor device.

10 . The apparatus according to claim 6 , wherein at least one of the following:

the third encoder layer is the same as the fourth encoder layer; or

the first neural fusion layer is the same as the second neural fusion layer.

11 . The apparatus according to claim 1 , wherein the instructions, when executed by the one or more processors, further cause the apparatus to perform;

inputting the first encoded map of the first modality and a third grid comprising one or more third cells into a fifth encoder layer, wherein each of the third cells comprises respective two or more of the first cells, and the first grid comprises two or more of the first cells;

encoding, by the fifth encoder layer, the first encoded map of the first modality into a fifth encoded map of the first modality, wherein the fifth encoded map of the first modality comprises, for each of the third cells of the third grid, a respective fifth feature center point and a respective fifth feature vector, and the coordinates of the fifth feature center points are indicated in the first coordinate system;

inputting a third encoded input map into a second matching descriptor layer of a second hierarchical layer, wherein the third encoded input map is based on the fifth encoded map of the first modality and comprises, for each of the third cells, a respective third joint feature center point and a respective third input feature vector, wherein the coordinates of the third joint feature center points are indicated in the first coordinate system;

inputting the second encoded map of the first modality and a fourth grid comprising one or more fourth cells into a sixth encoder layer, wherein each of the fourth cells comprises respective two or more of the second cells, and the second grid comprises two or more of the second cells;

encoding, by the sixth encoder layer, the second encoded map of the first modality into a sixth encoded map of the first modality, wherein the sixth encoded map of the first modality comprises, for each of the fourth cells of the fourth grid, a respective sixth feature center point and a respective sixth feature vector, and the coordinates of the sixth feature center points are indicated in the second coordinate system;

inputting a fourth encoded input map into the second matching descriptor layer, wherein the fourth encoded input map is based on the sixth encoded map of the first modality and comprises, for each of the fourth cells, a respective fourth joint feature center point and a respective fourth input feature vector, wherein the coordinates of the fourth joint feature center points are indicated in the second coordinate system;

adapting, by the second matching descriptor layer, each of the third input feature vectors based on the third input feature vectors, with their third joint feature center points, and the fourth input feature vectors, with their fourth joint feature center points, to obtain a third joint map, wherein the third joint map comprises, for each of the third joint feature center points, a respective third joint feature vector;

adapting, by the second matching descriptor layer, each of the fourth input feature vectors based on the third input feature vectors, with their third joint feature center points, and the fourth input feature vectors, with their fourth joint feature center points, to obtain a fourth joint map, wherein the fourth joint map comprises, for each of the fourth joint feature center points, a respective fourth joint feature vector;

calculating, by a second optimal matching layer, for each of the third cells and each of the fourth cells, the similarity between the third joint feature vector of the respective third cell and the fourth joint feature vector of the respective fourth cell;

calculating, by the second optimal matching layer, for each of the third cells and each of the fourth cells, a second point correlation between the respective third cell and the respective fourth cell based on the similarities between the third joint feature vector of the respective third cell and the fourth joint feature vectors of the fourth cells and based on the similarities between the fourth joint feature vector of the respective fourth cell and the third joint feature vectors of the third cells;

checking, for each of the third cells and each of the fourth cells, whether at least one of one or more second correlation conditions for the respective third cell and the respective fourth cell is fulfilled, wherein the one or more second correlation conditions comprise:

the second point correlation between the respective third cell and the respective fourth cell is larger than a second correlation threshold of the second hierarchical layer; or

the second point correlation between the respective third cell and the respective fourth cell is among the largest 1 values of the second point correlations, and 1 is a fixed value.

12 . The apparatus according to claim 11 , wherein the instructions, when executed by the one or more processors, further cause the apparatus to perform inhibiting the first matching descriptor layer to check, for each of the first cells and each of the second cells, whether the first point correlation between the first joint feature vector of the respective first cell and the second joint feature vector of the respective second cell is larger than the first correlation threshold of the first hierarchical layer in response to checking that the at least one of the one or more second correlation conditions for the respective third cell and the respective fourth cell is fulfilled.

13 . A method comprising:

inputting a first 3D point cloud of a first modality and a first grid comprising one or more first cells into a first encoder layer, wherein coordinates of points of the first 3D point cloud are indicated in a first coordinate system;

encoding, by the first encoder layer, the first 3D point cloud of the first modality into a first encoded map of the first modality, wherein the first encoded map of the first modality comprises, for each of the first cells of the first grid, a respective first feature center point and a respective first feature vector, and the coordinates of the first feature center points are indicated in the first coordinate system;

inputting a first encoded input map into a first matching descriptor layer of a first hierarchical layer, wherein the first encoded input map is based on the first encoded map of the first modality and comprises, for each of the first cells, a respective first joint feature center point and a respective first input feature vector, wherein the coordinates of the first joint feature center points are indicated in the first coordinate system;

inputting a second 3D point cloud of the first modality and a second grid comprising one or more second cells into a second encoder layer, wherein the coordinates of points of the second 3D point cloud are indicated in a second coordinate system;

encoding, by the second encoder layer, the second 3D point cloud of the first modality into a second encoded map of the first modality, wherein the second encoded map of the first modality comprises, for each of the second cells of the second grid, a respective second feature center point and a respective second feature vector, and the coordinates of the second feature center points are indicated in the second coordinate system;

inputting a second encoded input map into the first matching descriptor layer, wherein the second encoded input map is based on the second encoded map of the first modality and comprises, for each of the second cells, a respective second joint feature center point and a respective second input feature vector, wherein the coordinates of the second joint feature center points are indicated in the second coordinate system;

adapting, by the first matching descriptor layer, each of the first input feature vectors based on the first input feature vectors, with their first feature center points, and the second input feature vectors, with their second feature center points, to obtain a first joint map, wherein the first joint map comprises, for each of the first joint feature center points, a respective first joint feature vector;

adapting, by the first matching descriptor layer, each of the second input feature vectors based on the first input feature vectors, with their first feature center points, and the second input feature vectors, with their second feature center points, to obtain a second joint map, wherein the second joint map comprises, for each of the second joint feature center points, a respective second joint feature vector;

calculating, by a first optimal matching layer, for each of the first cells and each of the second cells, a similarity between the first joint feature vector of the respective first cell and the second joint feature vector of the respective second cell;

calculating, by the first optimal matching layer, for each of the first cells and each of the second cells, a first point correlation between the respective first cell and the respective second cell based on the similarities between the first joint feature vector of the respective first cell and the second joint feature vectors of the second cells and based on the similarities between the second joint feature vector of the respective second cell and the first joint feature vectors of the first cells;

checking, for each of the first cells and each of the second cells, whether at least one of one or more first correlation conditions for the respective first cell and the respective second cell is fulfilled, wherein the one or more first correlation conditions comprise:

the first point correlation between the respective first cell and the respective second cell is larger than a first correlation threshold of the first hierarchical layer; or

the first point correlation between the respective first cell and the respective second cell is among the largest k values of the first point correlations, and k is a fixed value;

extracting, for each of the first cells and each of the second cells, the coordinates of the first joint feature center point of the respective first cell in the first coordinate system and the coordinates of the second joint feature center point of the respective second cell in the second coordinate system to obtain a respective pair of extracted coordinates of the first hierarchical layer in response to checking that the at least one of the one or more first correlation conditions for the respective first cell and the respective second cell is fulfilled;

calculating an estimated transformation between the first coordinate system and the second coordinate system based on the pairs of extracted coordinates of the first hierarchical layer.

14 . A non-transitory computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform at least the following:

inputting a first 3D point cloud of a first modality and a first grid comprising one or more first cells into a first encoder layer, wherein coordinates of points of the first 3D point cloud are indicated in a first coordinate system;

encoding, by the first encoder layer, the first 3D point cloud of the first modality into a first encoded map of the first modality, wherein the first encoded map of the first modality comprises, for each of the first cells of the first grid, a respective first feature center point and a respective first feature vector, and the coordinates of the first feature center points are indicated in the first coordinate system;

inputting a first encoded input map into a first matching descriptor layer of a first hierarchical layer, wherein the first encoded input map is based on the first encoded map of the first modality and comprises, for each of the first cells, a respective first joint feature center point and a respective first input feature vector, wherein the coordinates of the first joint feature center points are indicated in the first coordinate system;

inputting a second 3D point cloud of the first modality and a second grid comprising one or more second cells into a second encoder layer, wherein the coordinates of points of the second 3D point cloud are indicated in a second coordinate system;

encoding, by the second encoder layer, the second 3D point cloud of the first modality into a second encoded map of the first modality, wherein the second encoded map of the first modality comprises, for each of the second cells of the second grid, a respective second feature center point and a respective second feature vector, and the coordinates of the second feature center points are indicated in the second coordinate system;

inputting a second encoded input map into the first matching descriptor layer, wherein the second encoded input map is based on the second encoded map of the first modality and comprises, for each of the second cells, a respective second joint feature center point and a respective second input feature vector, wherein the coordinates of the second joint feature center points are indicated in the second coordinate system;

adapting, by the first matching descriptor layer, each of the first input feature vectors based on the first input feature vectors, with their first feature center points, and the second input feature vectors, with their second feature center points, to obtain a first joint map, wherein the first joint map comprises, for each of the first joint feature center points, a respective first joint feature vector;

adapting, by the first matching descriptor layer, each of the second input feature vectors based on the first input feature vectors, with their first feature center points, and the second input feature vectors, with their second feature center points, to obtain a second joint map, wherein the second joint map comprises, for each of the second joint feature center points, a respective second joint feature vector;

calculating, by a first optimal matching layer, for each of the first cells and each of the second cells, a similarity between the first joint feature vector of the respective first cell and the second joint feature vector of the respective second cell;

calculating, by the first optimal matching layer, for each of the first cells and each of the second cells, a first point correlation between the respective first cell and the respective second cell based on the similarities between the first joint feature vector of the respective first cell and the second joint feature vectors of the second cells and based on the similarities between the second joint feature vector of the respective second cell and the first joint feature vectors of the first cells;

checking, for each of the first cells and each of the second cells, whether at least one of one or more first correlation conditions for the respective first cell and the respective second cell is fulfilled, wherein the one or more first correlation conditions comprise:

the first point correlation between the respective first cell and the respective second cell is larger than a first correlation threshold of the first hierarchical layer; or

the first point correlation between the respective first cell and the respective second cell is among the largest k values of the first point correlations, and k is a fixed value;

extracting, for each of the first cells and each of the second cells, the coordinates of the first joint feature center point of the respective first cell in the first coordinate system and the coordinates of the second joint feature center point of the respective second cell in the second coordinate system to obtain a respective pair of extracted coordinates of the first hierarchical layer in response to checking that the at least one of the one or more first correlation conditions for the respective first cell and the respective second cell is fulfilled;

calculating an estimated transformation between the first coordinate system and the second coordinate system based on the pairs of extracted coordinates of the first hierarchical layer.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2024
From: MATE JOZSA, CSABA; SÖRÖS, GÁBOR; ZSOLT VARGA, KRISZTIÁN; FARKAS, LÓRÁNT
To: NOKIA SOLUTIONS AND NETWORKS KFT.
Reel/Frame 069064/0764 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2024
From: NOKIA SOLUTIONS AND NETWORKS KFT.
To: NOKIA SOLUTIONS AND NETWORKS OY
Reel/Frame 069064/0774 →
Priority Claims (1)
FI 20235463 · Apr 26, 2023 · national
Continuity (1)
Related Publication 20240362905A1 · Oct 31, 2024
References Cited (63)
US 9031283B2 · Arth et al. · 2015 [cited by applicant]
US 9668146B2 · Lau · 2017 [cited by applicant]
US 9763035B2 · Valaee et al. · 2017 [cited by applicant]
US 9934587B2 · Senthamil · 2018 [cited by applicant]
US 10834532B2 · Fuerst et al. · 2020 [cited by applicant]
US 10841556B2 · Casas · 2020 [cited by examiner]
US 11393179B2 · Fleischman et al. · 2022 [cited by applicant]
US 11675068B2 · Jiang · 2023 [cited by examiner]
US 20140133741A1 · Wang · 2014 [cited by examiner]
US 20150119086A1 · Mirowski et al. · 2015 [cited by applicant]
US 20160183050A1 · Yiu et al. · 2016 [cited by applicant]
US 20180033160A1 · Ishigami · 2018 [cited by applicant]
US 20200015047A1 · Song et al. · 2020 [cited by applicant]
US 20200058156A1 · Tran · 2020 [cited by examiner]
US 20200082567A1 · Liu · 2020 [cited by applicant]
US 20200357143A1 · Chiu et al. · 2020 [cited by applicant]
US 20210017587A1 · Cai · 2021 [cited by examiner]
US 20210063200A1 · Kroepfl et al. · 2021 [cited by applicant]
US 20210112427A1 · Shveki et al. · 2021 [cited by applicant]
US 20210150252A1 · Sarlin · 2021 [cited by examiner]
US 20210373161A1 · Lu et al. · 2021 [cited by applicant]
US 20220155079A1 · Chidlovskii et al. · 2022 [cited by applicant]
US 20220237446A1 · Lei · 2022 [cited by examiner]
US 20220335638A1 · Kar et al. · 2022 [cited by applicant]
JP 2021536634A · 2021 [cited by applicant]
WO 2014162044A1 · 2014 [cited by applicant]
WO 2015048434A1 · 2015 [cited by applicant]
WO 2020194079A1 · 2020 [cited by applicant]
Office action received for corresponding Japanese Patent Application No. 2024-069433, dated Apr. 24, 2025, 3 pages of office action and 4 pages of translation available. [cited by applicant]
Kanazaki et al., “Transformating Topological Maps between Heterogeneous Multi-Agents”, IEICE Technical Report, vol. 104, No. 636, Jan. 27, 2005, 10 pages. [cited by applicant]
Extended European Search Report received for corresponding European Patent Application No. 24171494.8, dated Sep. 25, 2024, 5 pages. [cited by applicant]
Xue et al., “Research on 3D LiDAR Point Cloud Matching Localization Method”, 5th World Conference on Mechanical Engineering and Intelligent Manufacturing (WCMEIM), Nov. 18-20, 2022, pp. 878-881. [cited by applicant]
Zhang et al., “Robust LiDAR Localization on an HD Vector Map without a Separate Localization Layer”, IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Sep. 27-Oct. 1, 2021, pp. 5536-5543. [cited by applicant]
“Niantic and 6D.ai: Building The Next Tech Platform”, 6D.ai, Retrieved on Apr. 21, 2024, Webpage available at : https://www.6d.ai/. [cited by applicant]
Friis, “A Note on a Simple Transmission Formula”, Proceedings of the I.R.E and Waves and Electrons, May 1946, pp. 254-256. [cited by applicant]
Sinkhorn et al., “Concerning Nonnegative Matrices and Doubly Stochastic Matrices”, Pacific Journal of Mathematics, vol. 21, No. 02, Dec. 1967, 9 pages. [cited by applicant]
Besl et al., “A Method for registration of 3-D shapes”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 14, No. 02, Feb. 1992, pp. 239-256. [cited by applicant]
Irschara et al., “From Structure-from-Motion Point Clouds to Fast Location Recognition”, IEEE Conference on Computer Vision and Pattern Recognition, Jun. 20-25, 2009, pp. 2599-2606. [cited by applicant]
Kendall et al., “PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization”, Proceedings of the IEEE International Conference on Computer Vision (ICCV), Dec. 7-13, 2015, pp. 2938-2946. [cited by applicant]
Ioffe et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift”; Proceedings of the 32nd International Conference on Machine Learning, vol. 37, Jul. 7-9, 2015, 9 pages. [cited by applicant]
Ba et al., “Layer Normalization”, arxiv, Jul. 21, 2016, pp. 1-14. [cited by applicant]
Krizhevsky et al., “ImageNet Classification with Deep Convolutional Neural Networks”, Advances in Neural Information Processing Systems 25 (NIPS 2012), Dec. 3-6, 2012, pp. 1-9. [cited by applicant]
Vaswani et al., “Attention is all you need”, Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS 2017), Dec. 4-9, 2017, pp. 1-11. [cited by applicant]
Wu et al., “Group Normalization”, Proceedings of the European Conference on Computer Vision (ECCV), Sep. 8-14, 2018, pp. 1-17. [cited by applicant]
Sarlin et al., “SuperGlue: Learning Feature Matching with Graph Neural Networks”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 13-19, 2020, pp. 4938-4947. [cited by applicant]
Hashemifar et al., “Augmenting visual SLAM with Wi-Fi sensing for indoor applications”, Autonomous Robots, vol. 43, Jul. 18, 2019, pp. 2245-2260. [cited by applicant]
Thomas et al., “KPConv: Flexible and Deformable Convolution for Point Clouds”, Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 27-Nov. 2, 2019, pp. 6411-6420. [cited by applicant]
Sarlin et al., “From Coarse to Fine: Robust Hierarchical Localization at Large Scale”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 15-20, 2019, pp. 12716-12725. [cited by applicant]
Ayyalasomayajula et al., “Deep Learning based Wireless Localization for Indoor Navigation”, Proceedings of the 26th Annual International Conference on Mobile Computing and Networking, Sep. 21-25, 2020, 14 pages. [cited by applicant]
Oertel et al., “Augmenting Visual Place Recognition With Structural Cues”, IEEE Robotics and Automation Letters, vol. 05, No. 04, Oct. 2020, pp. 5534-5541. [cited by applicant]
Sun et al., “WiFi Based Fingerprinting Positioning Based on Seq2seq Model”, Sensors, vol. 20, No. 13, Jul. 5, 2020, pp. 1-19. [cited by applicant]
Brachmann et al., “Visual Camera Re-Localization From RGB and RGB-D Images Using DSAC”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, No. 09, Sep. 2022, pp. 5847-5865. [cited by applicant]
Wang et al., “PointLoc: Deep Pose Regressor for LiDAR Point Cloud Localization”, IEEE Sensors Journal, vol. 22, No. 01, Jan. 1, 2022, pp. 959-968. [cited by applicant]
Shu et al., “3D Point Cloud-Based Indoor Mobile Robot in 6-DoF Pose Localization Using a Wi-Fi-Aided Localization System”, IEEE Access, vol. 09, 2021, pp. 38636-38648. [cited by applicant]
Zanjani et al., “Modality-Agnostic Topology Aware Localization”, 35th Conference on Neural Information Processing Systems (NeurIPS 2021), Dec. 6-14, 2021, pp. 1-12. [cited by applicant]
Yu et al., “Multi-Modal Recurrent Fusion for Indoor Localization”, IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 23-27, 2022, pp. 5083-5087. [cited by applicant]
Qin et al., “Geometric Transformer for Fast and Robust Point Cloud Registration”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 19-24, 2022, pp. 11143-11152. [cited by applicant]
Seidel et al., “914 MHz Path Loss Prediction Models for Indoor Wireless Communications in Multifloored Buildings”, IEEE Transactions on Antennas and Propagation, vol. 40, No. 02, Feb. 1992, pp. 207-217. [cited by applicant]
Sarlin et al., “LaMAR: Benchmarking Localization and Mapping for Augmented Reality”, European Conference on Computer Vision (ECCV), Oct. 23-27, 2022, pp. 1-30. [cited by applicant]
Nabati et al., “Using Synthetic Data to Enhance the Accuracy of Fingerprint-Based Localization: A Deep Learning Approach”, IEEE Sensors Letters, vol. 04, No. 04, Apr. 2020, 4 pages. [cited by applicant]
Song et al., “CNNLoc: Deep-Learning Based Indoor Localization with WIFI Fingerprinting”, IEEE SmartWorld, Ubiquitous Intelligence & Computing, Advanced & Trusted Computing, Scalable Computing & Communications, Cloud & B… [cited by applicant]
Fischer et al., “StickyPillars: Robust and Efficient Feature Matching on Point Clouds Using Graph Neural Networks”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 19-25, 2… [cited by applicant]
Office action received for corresponding Finnish Patent Application No. 20235463, dated Jan. 18, 2024, 14 pages. [cited by applicant]