Cross-view image geo-localization
View Patent ↗CNN-based methods for cross-view image geo-localization rely on polar transform and fail to model global correlation. A pure transformer-based approach (TransGeo) is described to address these limitations from a different perspective. TransGeo takes full advantage of the strengths of the transformer related to global information modeling and explicit position information encoding. The claimed invention further leverages transformer input's flexibility and discloses an attention-guided non-uniform cropping method so that uninformative image patches are removed with a negligible drop in performance to reduce computation cost. The saved computation can be reallocated to increase resolution only for informative patches, resulting in performance improvement with no additional computation cost. This “attend and zoom-in” strategy is highly similar to human behavior when observing images. Remarkably, TransGeo achieves state-of-the-art results on both urban and rural datasets, with significantly less computation cost than CNN-based methods. It does not rely on polar transform and provides faster methods.
1 . A cross-view image geo-localization method comprising:
electronically performing with an information processor each of,
a first stage operation for
acquiring ground-view images and aerial-view images of a geographical position, the aerial-view images are at a first resolution;
establishing a first training set using each of the ground-view images and its corresponding ground-truth aerial image;
training a ground-view image transformer-encoder with the first training set to produce ground-view image transformer/encoder weights;
training a first aerial-view image transformer-encoder with the first training set to produce a first set of aerial-view image encoder weights;
a second stage operation for
generating, using the first set of aerial-view image encoder weights after completion of the first stage training, a spatial attention map identifying high-saliency regions of the aerial-view images at the first resolution;
accessing aerial-view images at a second resolution, the second resolution being higher than the first resolution;
performing non-uniform spatial cropping of the aerial-view images at the second resolution based on the spatial attention map generated from the first set of aerial-view image encoder weights;
establishing a second training set comprising the non-uniformly cropped higher-resolution aerial-view images; and
training a second aerial-view image transformer-encoder with the second training set, wherein the first aerial-view image transformer-encoder and the second aerial-view image transformer-encoder are distinct training instances.
2 . The method of claim 1 , wherein the training the ground-view image transformer-encoder further includes training with a first set of class tokens to integrate classification information.
3 . The method of claim 2 , wherein the training the first aerial-view image transformer-encoder further includes training with a second set of class tokens to integrate classification information.
4 . The method of claim 3 , wherein the generating, using the first set of aerial-view image encoder weights after completion of the first stage training, the spatial attention map identifying high-saliency regions of the aerial-view images at the first resolution includes the second set of class tokens.
5 . The method of claim 3 , wherein the training the second aerial-view image transformer-encoder further includes training with a third set of class tokens to integrate classification information.
6 . The method of claim 1 , wherein the first stage operation is independent of polar transforms.
7 . The method of claim 2 , wherein the second stage operation is independent of polar transforms.
8 . The method of claim 1 , wherein the first stage operation is without data augmentation.
9 . The method of claim 8 , wherein the second stage operation is without data augmentation.
10 . The method of claim 1 , wherein the aerial images at the first resolution are a down-sampled version of the aerial images at the second resolution.
11 . The method of claim 1 , wherein the aerial images at the second resolution are an up-sampled version of the aerial images at the first resolution.
12 . A system for cross-view image geo-localization, the system comprising
memory;
at least one processor operatively coupled to the memory for performing each of a first stage operation for
acquiring ground-view images and aerial-view images of a geographical position, the aerial-view images are at a first resolution;
establishing a first training set using each of the ground-view images and its corresponding ground-truth aerial image;
training a ground-view image transformer-encoder with the first training set to produce ground-view image transformer/encoder weights;
training a first aerial-view image transformer-encoder with the first training set to produce a first set of aerial-view image encoder weights;
a second stage operation for
generating, using the first set of aerial-view image encoder weights after completion of the first stage training, a spatial attention map identifying high-saliency regions of the aerial-view images at the first resolution;
accessing aerial-view images at a second resolution, the second resolution being higher than the first resolution;
performing non-uniform spatial cropping of the aerial-view images at the second resolution based on the spatial attention map generated from the first set of aerial-view image encoder weights;
establishing a second training set comprising the non-uniformly cropped higher-resolution aerial-view images; and
training a second aerial-view image transformer-encoder with the second training set, wherein the first aerial-view image transformer-encoder and the second aerial-view image transformer-encoder are distinct training instances.
13 . The system of claim 12 , wherein the training the ground-view image transformer-encoder further includes training with a first set of class tokens to integrate classification information.
14 . The system of claim 13 , wherein the training the first aerial-view image transformer-encoder further includes training with a second set of class tokens to integrate classification information.
15 . The system of claim 14 , wherein the generating, using the first set of aerial-view image encoder weights after completion of the first stage training, the spatial attention map identifying high-saliency regions of the aerial-view images at the first resolution includes the second set of class tokens.
16 . The system of claim 14 , wherein the training the second aerial-view image transformer-encoder further includes training with a third set of class tokens to integrate classification information.
17 . The system of claim 12 , wherein the first stage operation is independent of polar transforms and the second stage operation is independent of polar transforms.
18 . The system of claim 12 , wherein the first stage operation is without data augmentation and the second stage operation is without data augmentation.
19 . The system of claim 12 , wherein the aerial images at the first resolution are a down-sampled version of the aerial images at the second resolution.
20 . He system of claim 12 , wherein the aerial images at the second resolution are an up-sampled version of the aerial images at the first resolution.