IP Library › Granted Patent US 12,444,103
Granted Patent B2
US 12,444,103 · App. 17/964,670 · Granted Oct 14, 2025

System, method, and apparatus for machine learning-based map generation

Inventor: Ofer Melnik (Weehawken, NJ)
Assignee: HERE Global B.V.
G06T11/206G06T3/4046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,444,103
App. No.
17/964,670
Granted
Oct 14, 2025
Kind
B2
Abstract

An approach is provided for machine learning-based map generation. The approach involves, for example, receiving an encoder output for an image depicting a geographic area, wherein the encoder output comprises an encoding for each tile of a plurality of tiles of the image, and wherein the encoding represents data associated with a location of each tile. The approach also involves using a machine learning decoder to determine a window over the encoder output comprising a tile of the plurality of tiles and one or more neighboring tiles and to process the encoding associated with the tile and the one or more neighboring tiles in the window to generate a map representation for the location of the tile and providing the map representation as an output.

Claims (40)

1. A method of machine learning-based map generation comprising:

receiving an encoder output for an image depicting a geographic area, wherein the encoder output comprises an encoding for each tile of a plurality of tiles of the image, and wherein the encoding represents data associated with a location of each tile;

using a machine learning decoder to determine a window over the encoder output comprising a tile of the plurality of tiles and one or more neighboring tiles and to process the encoding associated with the tile and the one or more neighboring tiles in the window to generate a map representation for the location of the tile; and

providing the map representation as an output.

2. The method of claim 1 , further comprising:

using the machine learning decoder to move the window to another tile of the plurality of tiles, wherein the moved window is over the another tile and one or more other neighboring tiles, and to process the encoding associated with the another tile and the one or more other neighboring tiles in the moved window to generate another map representation for the another tile.

3. The method of claim 1 , wherein the one or more neighboring tiles are immediate neighbors to the tile.

4. The method of claim 1 , further comprising:

determining the one or more neighboring tiles based on one or more cartographic relationships to the tile.

5. The method of claim 1 , wherein the encoding for each tile is associated with a probability that the encoding represents a true encoding for the each tile, and wherein the generating of the map representation by the machine learning decoder is further based on the probability.

6. The method of claim 1 , wherein the encoding for each tile is based on an encoding codebook.

7. The method of claim 1 , wherein the machine learning decoder is extracted from a trained encoder-decoder stack and used independently from a machine learning encoder of the encoder-decoder stack.

8. The method of claim 1 , wherein a hidden layer between the machine learning encoder and the machine learning decoder of the encoder-decoder stack is used as the representational layer of the encoding.

9. The method of claim 1 , wherein the map representation is a raster representation.

10. The method of claim 9 , further comprising:

processing the raster representation using a neural network comprising at least one convolutional center of mass (CCOM) layer to generate a vector representation.

11. The method of claim 10 , wherein the CCOM layer is a function that returns a coordinate for each kernel location corresponding to a feature of the raster representation and a weight at the coordinate, and wherein the weight represents an intensity of the feature at each kernel location.

12. The method of claim 9 , wherein each kernel of the CCOM layer overlaps.

13. The method of claim 1 , wherein the machine learning decoder comprises one or more convolutional layers that upscale the encoder output to a target image resolution.

14. An apparatus for machine learning-based map generation comprising:

at least one processor; and

at least one memory including computer program code for one or more programs,

the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to perform at least the following:

receive an encoder output for an image depicting a geographic area, wherein the encoder output comprises an encoding for each tile of a plurality of tiles of the image, and wherein the encoding represents data associated with a location of each tile;

use a machine learning decoder to determine a window over the encoder output comprising a tile of the plurality of tiles and one or more neighboring tiles and to process the encoding associated with the tile and the one or more neighboring tiles in the window to generate a map representation for the location of the tile; and

provide the map representation as an output.

15. The apparatus of claim 14 , wherein the apparatus is further caused to:

use the machine learning decoder to move the window to another tile of the plurality of tiles, wherein the moved window is over the another tile and one or more other neighboring tiles and to process the encoding associated with the another tile and the one or more other neighboring tiles in the moved window to generate another map representation for the another tile.

16. The apparatus of claim 14 , wherein the map representation is a raster representation, and wherein the apparatus is further caused to:

processing the raster representation using a neural network comprising at least one convolutional center of mass (CCOM) layer to generate a vector representation.

17. The apparatus of claim 16 , wherein the CCOM layer is a function that returns a coordinate for each kernel location corresponding to a feature of the raster representation and a weight at the coordinate, and wherein the weight represents an intensity of the feature at each kernel location.

18. A non-transitory computer-readable storage medium for machine learning-based map generation, carrying one or more sequences of one or more instructions which, when executed by one or more processors, cause an apparatus to at least perform the following steps:

receiving an encoder output for an image depicting a geographic area, wherein the encoder output comprises an encoding for each tile of a plurality of tiles of the image, and wherein the encoding represents data associated with a location of each tile;

using a machine learning decoder to determine a window over the encoder output comprising a tile of the plurality of tiles and one or more neighboring tiles and to process the encoding associated with the tile and the one or more neighboring tiles in the window to generate a map representation for the location of the tile; and

providing the map representation as an output.

19. The non-transitory computer-readable storage medium of claim 18 , wherein the apparatus is caused to further perform:

using the machine learning decoder to move the window to another tile of the plurality of tiles, wherein the moved window is over the another tile and one or more other neighboring tiles and to process the encoding associated with the another tile and the one or more other neighboring tiles in the moved window to generate another map representation for the another tile.

20. The non-transitory computer-readable storage medium of claim 18 , wherein the map representation is a raster representation, and wherein the apparatus is caused to further perform:

processing the raster representation using a neural network comprising at least one convolutional center of mass (CCOM) layer to generate a vector representation,

wherein the CCOM layer is a function that returns a coordinate for each kernel location corresponding to a feature of the raster representation and a weight at the coordinate, and wherein the weight represents an intensity of the feature at each kernel location.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 13, 2022
From: MELNIK, OFER
To: HERE GLOBAL B.V.
Reel/Frame 061416/0749 →
Continuity (1)
Related Publication 20240127504A1 · Apr 18, 2024
References Cited (25)
US 20200041276A1 · Chakravarty et al. · 2020 [cited by applicant]
US 20200111238A1 · Covell · 2020 [cited by examiner]
US 20210319046A1 · Bandyopadhyay · 2021 [cited by examiner]
US 20210377566A1 · Iguchi et al. · 2021 [cited by applicant]
US 20220012579A1 · Asama · 2022 [cited by examiner]
US 20220101639A1 · Shugurov et al. · 2022 [cited by applicant]
US 20220196415A1 · Sameer · 2022 [cited by examiner]
US 20220357176A1 · Yin · 2022 [cited by examiner]
US 20230304826A1 · Zhang · 2023 [cited by examiner]
CN 114445630A · 2022 [cited by applicant]
K. Bruhwiler, P. Khandelwal, D. Rammer, S. Armstrong, S. L. Pallickara and S. Pallickara, “Lightweight, Embeddings Based Storage and Model Construction Over Satellite Data Collections,” 2020 IEEE International Conferenc… [cited by examiner]
Minnen, David, et al. “Spatially Adaptive Image Compression Using a Tiled Deep Network.” arXiv.Org, Feb. 7, 2018, arxiv.org/abs/1802.02629 (Year: 2018). [cited by examiner]
Mattheuwsen, Lukas, and Maarten Vergauwen. “Manhole Cover Detection on Rasterized Mobile Mapping Point Cloud Data Using Transfer Learned Fully Convolutional Neural Networks.” MDPI, Multidisciplinary Digital Publishing I… [cited by examiner]
Lee, J. “A Combinatorial Data Model for Representing Topological Relations among 3D Geographical Features in micro-Spatial Environments: International Journal of Geographical Information Science: vol. 19, No. 10.” Tandf… [cited by examiner]
Lee et. al (Year: 2007). [cited by examiner]
Bruhwiler et. al (Year: 2020). [cited by examiner]
Mattheuwsen et. al (Year: 2020). [cited by examiner]
Minnen et. al (Year: 2018). [cited by examiner]
Mukherjee et al., “Predicting vehicle behaviour using automotive radar and recurrent neural networks”, 2021, pp. 1-14. [cited by applicant]
Carlier et al., “DeepSVG: A Hierarchical Generative”, Generative Network for Vector Graphics Animation, 2020, 19 pages. [cited by applicant]
Jadon, “A survey of loss functions for semantic segmentation”, 2020, 2020 IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB), 6 pages. [cited by applicant]
Li et al., “Differentiable vector graphics rasterization for editing and learning”, ACM Transactions on Graphics, vol. 39, Issue 6, Dec. 2020 Article No. 193, 15 pages. [cited by applicant]
Lopes et al., “A Learned Representation for Scalable Vector Graphics”, Apr. 4, 2019, pp. 1-13. [cited by applicant]
Nash et al., “PolyGen: An Autoregressive Generative Model of 3D Meshes”, 2020, 10 pages. [cited by applicant]
Reddy et al., “Im2Vec: Synthesizing Vector Graphics without Vector Supervision”, Apr. 1, 2021, 10 pages. [cited by applicant]