IP Library › Granted Patent US 12,743,744
Granted Patent B2
US 12,743,744 · App. 18/414,626 · Granted Sep 22, 2026

Cross-view image geo-localization

Inventors: Sijie Zhu (Santa Clara, CA); Chen Chen (Orlando, FL); Mubarak Shah (Winter Park, FL)
Assignee: University of Central Ponios Research: Foundation, Inc
G06T3/4053G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,743,744
App. No.
18/414,626
Granted
Sep 22, 2026
Kind
B2
Abstract

CNN-based methods for cross-view image geo-localization rely on polar transform and fail to model global correlation. A pure transformer-based approach (TransGeo) is described to address these limitations from a different perspective. TransGeo takes full advantage of the strengths of the transformer related to global information modeling and explicit position information encoding. The claimed invention further leverages transformer input's flexibility and discloses an attention-guided non-uniform cropping method so that uninformative image patches are removed with a negligible drop in performance to reduce computation cost. The saved computation can be reallocated to increase resolution only for informative patches, resulting in performance improvement with no additional computation cost. This “attend and zoom-in” strategy is highly similar to human behavior when observing images. Remarkably, TransGeo achieves state-of-the-art results on both urban and rural datasets, with significantly less computation cost than CNN-based methods. It does not rely on polar transform and provides faster methods.

Claims (44)

1 . A cross-view image geo-localization method comprising:

electronically performing with an information processor each of,

a first stage operation for

acquiring ground-view images and aerial-view images of a geographical position, the aerial-view images are at a first resolution;

establishing a first training set using each of the ground-view images and its corresponding ground-truth aerial image;

training a ground-view image transformer-encoder with the first training set to produce ground-view image transformer/encoder weights;

training a first aerial-view image transformer-encoder with the first training set to produce a first set of aerial-view image encoder weights;

a second stage operation for

generating, using the first set of aerial-view image encoder weights after completion of the first stage training, a spatial attention map identifying high-saliency regions of the aerial-view images at the first resolution;

accessing aerial-view images at a second resolution, the second resolution being higher than the first resolution;

performing non-uniform spatial cropping of the aerial-view images at the second resolution based on the spatial attention map generated from the first set of aerial-view image encoder weights;

establishing a second training set comprising the non-uniformly cropped higher-resolution aerial-view images; and

training a second aerial-view image transformer-encoder with the second training set, wherein the first aerial-view image transformer-encoder and the second aerial-view image transformer-encoder are distinct training instances.

2 . The method of claim 1 , wherein the training the ground-view image transformer-encoder further includes training with a first set of class tokens to integrate classification information.

3 . The method of claim 2 , wherein the training the first aerial-view image transformer-encoder further includes training with a second set of class tokens to integrate classification information.

4 . The method of claim 3 , wherein the generating, using the first set of aerial-view image encoder weights after completion of the first stage training, the spatial attention map identifying high-saliency regions of the aerial-view images at the first resolution includes the second set of class tokens.

5 . The method of claim 3 , wherein the training the second aerial-view image transformer-encoder further includes training with a third set of class tokens to integrate classification information.

6 . The method of claim 1 , wherein the first stage operation is independent of polar transforms.

7 . The method of claim 2 , wherein the second stage operation is independent of polar transforms.

8 . The method of claim 1 , wherein the first stage operation is without data augmentation.

9 . The method of claim 8 , wherein the second stage operation is without data augmentation.

10 . The method of claim 1 , wherein the aerial images at the first resolution are a down-sampled version of the aerial images at the second resolution.

11 . The method of claim 1 , wherein the aerial images at the second resolution are an up-sampled version of the aerial images at the first resolution.

12 . A system for cross-view image geo-localization, the system comprising

memory;

at least one processor operatively coupled to the memory for performing each of a first stage operation for

acquiring ground-view images and aerial-view images of a geographical position, the aerial-view images are at a first resolution;

establishing a first training set using each of the ground-view images and its corresponding ground-truth aerial image;

training a ground-view image transformer-encoder with the first training set to produce ground-view image transformer/encoder weights;

training a first aerial-view image transformer-encoder with the first training set to produce a first set of aerial-view image encoder weights;

a second stage operation for

generating, using the first set of aerial-view image encoder weights after completion of the first stage training, a spatial attention map identifying high-saliency regions of the aerial-view images at the first resolution;

accessing aerial-view images at a second resolution, the second resolution being higher than the first resolution;

performing non-uniform spatial cropping of the aerial-view images at the second resolution based on the spatial attention map generated from the first set of aerial-view image encoder weights;

establishing a second training set comprising the non-uniformly cropped higher-resolution aerial-view images; and

training a second aerial-view image transformer-encoder with the second training set, wherein the first aerial-view image transformer-encoder and the second aerial-view image transformer-encoder are distinct training instances.

13 . The system of claim 12 , wherein the training the ground-view image transformer-encoder further includes training with a first set of class tokens to integrate classification information.

14 . The system of claim 13 , wherein the training the first aerial-view image transformer-encoder further includes training with a second set of class tokens to integrate classification information.

15 . The system of claim 14 , wherein the generating, using the first set of aerial-view image encoder weights after completion of the first stage training, the spatial attention map identifying high-saliency regions of the aerial-view images at the first resolution includes the second set of class tokens.

16 . The system of claim 14 , wherein the training the second aerial-view image transformer-encoder further includes training with a third set of class tokens to integrate classification information.

17 . The system of claim 12 , wherein the first stage operation is independent of polar transforms and the second stage operation is independent of polar transforms.

18 . The system of claim 12 , wherein the first stage operation is without data augmentation and the second stage operation is without data augmentation.

19 . The system of claim 12 , wherein the aerial images at the first resolution are a down-sampled version of the aerial images at the second resolution.

20 . He system of claim 12 , wherein the aerial images at the second resolution are an up-sampled version of the aerial images at the first resolution.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 17, 2024
From: ZHU, SIJIE; CHEN, CHEN; SHAH, MUBARAK
To: UNIVERSITY OF CENTRAL FLORIDA RESEARCH FOUNDATION, INC.
Reel/Frame 068006/0424 →
Continuity (2)
Provisional Application 63488548 · Mar 6, 2023
Related Publication 20240303770A1 · Sep 12, 2024
References Cited (42)
US 20230290135A1 · Zhou · 2023 [cited by examiner]
WO WO2023168613A1 · 2023 [cited by examiner]
Zhu, Y., Yang, H., Lu, Y., & Huang, Q. (2023). Simple, Effective and General: A New Backbone for Cross-view Image Geo-localization. ArXiv, abs/2302.01572. (Year: 2023). [cited by examiner]
Zhu, S., Shah, M., & Chen, C. (2022). TransGeo: Transformer Is All You Need for Cross-view Image Geo-localization. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 1152-1161. (Year: 2022). [cited by examiner]
Zhu, Yingying et al. “Simple, Effective and General: A New Backbone for Cross-view Image Geo-localization.” ArXiv abs/2302.01572 (2023): n. pag. (Year: 2023). [cited by examiner]
Zhu, Sijie et al. “TransGeo: Transformer Is All You Need for Cross-view Image Geo-localization.” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022): 1152-1161. (Year: 2022). [cited by examiner]
Eli Brosh, Matan Friedmann, Ilan Kadar, Lev Yitzhak Lavy, Elad Levi, Shmuel Rippa, Yair Lempert, Bruno Fernandez-Ruiz, Roei Herzig, and Trevor Darrell. Accurate visual localization for automotive applications. In Procee… [cited by applicant]
Sudong Cai, Yulan Guo, Salman Khan, Jiwei Hu, and Gongjian Wen. Ground-to-aerial image geo-localization with a hard exemplar reweighting triplet loss. In Proceedings of the IEEE International Conference on Computer Visi… [cited by applicant]
Xiangning Chen, Cho-Jui Hsieh, and Boqing Gong. When vision transformers outperform resnets without pretraining or strong data augmentations. arXiv preprint arXiv:2106.01548, 2021. [cited by applicant]
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248-255. Ieee, 2009. [cited by applicant]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pretraining of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. [cited by applicant]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transfo… [cited by applicant]
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Advances in neural information processing systems, pp. 267… [cited by applicant]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770-778, 2016. [cited by applicant]
Sixing Hu, Mengdan Feng, Rang MH Nguyen, and Gim Hee Lee. Cvm-net: Cross-view matching network for image-based ground-to-aerial geo-localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Reco… [cited by applicant]
Jungmin Kwon, Jeongseop Kim, Hyunseo Park, and In Kwon Choi. Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks. arXiv preprint arXiv:2102.11600, 2021. [cited by applicant]
Tsung-Yi Lin, Serge Belongie, and James Hays. Cross-view image geolocalization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 891-898, 2013. [cited by applicant]
Liu and Hongdong Li. Lending orientation to neural networks for cross-view geo-localization. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5624-5633, 2019. [cited by applicant]
Ang Li, Huiyi Hu, Piotr Mirowski, and Mehrdad Farajtabar. Cross-view policy learning for street navigation. In Proceedings of the IEEE International Conference on Computer Vision, pp. 8100-8109, 2019. [cited by applicant]
Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016. [cited by applicant]
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. [cited by applicant]
Piotr Mirowski, Matt Grimes, Mateusz Malinowski, Karl Moritz Hermann, Keith Anderson, Denis Teplyashin, Karen Simonyan, Andrew Zisserman, Raia Hadsell, et al. Learning to navigate in cities without a map. In Advances in… [cited by applicant]
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning libra… [cited by applicant]
Krishna Regmi and Mubarak Shah. Bridging the domain gap for ground-to-aerial image matching. In Proceedings of the IEEE International Conference on Computer Vision, pp. 470-479, 2019. [cited by applicant]
Florian Schroff, Dmitry Kalenichenko, and James Philbin. Facenet: A unified embedding for face recognition and clustering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 815-823, 2… [cited by applicant]
Yujiao Shi, Liu, Xin Yu, and Hongdong Li. Spatial-aware feature aggregation for image based cross-view geo-localization. In Advances in Neural Information Processing Systems, pp. 10090-10100, 2019. [cited by applicant]
Yujiao Shi, Xin Yu, Dylan Campbell, and Hongdong Li. Where am i looking at? joint location and orientation estimation by cross-view matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco… [cited by applicant]
Bin Sun, Chen, Yingying Zhu, and Jianmin Jiang. Geocapsnet: Ground to aerial view image geo-localization using capsule network. In 2019 IEEE International Conference on Multimedia and Expo (ICME), pp. 742-747. IEEE, 201… [cited by applicant]
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recogni… [cited by applicant]
Yicong Tian, Chen, and Mubarak Shah. Cross-view image matching for geo-localization in urban environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3608-3616, 2017. [cited by applicant]
Aysim Toker, Qunjie Zhou, Maxim Maximov, and Laura Leal-Taixé. Coming down to earth: Satellite-to-street view synthesis for geo-localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco… [cited by applicant]
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. In International Conference on Machine … [cited by applicant]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pp. 5998-6008… [cited by applicant]
Nam N Vo and James Hays. Localizing and orienting street views using overhead imagery. In European conference on computer vision, pp. 494-509. Springer, 2016. [cited by applicant]
Scott Workman, Richard Souvenir, and Nathan Jacobs. Wide-area image geolocalization with aerial reference imagery. In Proceedings of the IEEE International Conference on Computer Vision, pp. 3961-3969, 2015. [cited by applicant]
Hongji Yang, Xiufan Lu, and Yingying Zhu. Cross-view geo-localization with layer-to-layer transformer. Advances in Neural Information Processing Systems, 34, 2021. [cited by applicant]
Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF Internationa… [cited by applicant]
Amir Roshan Zamir and Mubarak Shah. Accurate image localization based on google maps street view. In European Conference on Computer Vision, pp. 255-268. Springer, 2010. [cited by applicant]
Menghua Zhai, Zachary Bessinger, Scott Workman, and Nathan Jacobs. Predicting ground-level scene layout from aerial imagery. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 867-875,… [cited by applicant]
Sijie Zhu, Taojiannan Yang, and Chen. Revisiting street-to-aerial view image geo-localization and orientation estimation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 7… [cited by applicant]
Sijie Zhu, Taojiannan Yang, and Chen. Vigor: Cross-view image geo-localization beyond one-to-one retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3640-3649, 2021. [cited by applicant]
https://developers.google.com/maps/documentation/maps-static/intro. [cited by applicant]