IP Library › Granted Patent US 12,197,498
Granted Patent B2
US 12,197,498 · App. 17/662,260 · Granted Jan 14, 2025

Method, apparatus, and system for determining pose

Inventors: Jinxue Liu (Shenzhen, CN); Jun Cao (Shenzhen, CN); Ruihua Li (Shenzhen, CN); Jiange Ge (Shenzhen, CN); Yuanlin Chen (Shenzhen, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06F16/587G06F16/532G06T7/0002G06T7/74G06V10/24G06V10/74G06V20/63G06T2207/30168G06T2207/30184
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,197,498
App. No.
17/662,260
Granted
Jan 14, 2025
Kind
B2
Abstract

A method, an apparatus, and a medium for determining a pose are provided. The method includes: receiving a query image sent by a terminal and N text fields included in the query image; determining a candidate reference image based on a prestored correspondence between a reference image and a text field and based on the N text fields; determining an initial pose of the terminal at the first location based on the query image and the candidate reference image; and sending the initial pose to the terminal. According to the method, a candidate reference image is queried based on a text field, and an initial pose of a terminal is obtained based on a candidate reference image with higher accuracy. Therefore, the obtained pose is more accurate.

Claims (63)

1. A method for determining a pose, wherein the method comprises:

receiving a query image sent by a terminal and N text fields comprised in the query image, wherein N is greater than or equal to 1, wherein the query image is obtained based on an image captured by the terminal at a first location, and wherein a scene at the first location comprises a scene in the query image;

determining at least one candidate reference image based on a prestored correspondence between a reference image and a text field and based on the N text fields;

determining a target reference image in the at least one candidate reference image, wherein the scene at the first location comprises a scene in the target reference image;

determining an initial pose of the terminal at the first location based on the query image and the target reference image; and

sending the initial pose to the terminal.

2. The method according to claim 1 , wherein the determining of the initial pose of the terminal at the first location based on the query image and the target reference image comprises:

determining a 2D-2D correspondence between the query image and the target reference image; and

determining the initial pose of the terminal at the first location based on the 2D-2D correspondence and a preset 2D-3D correspondence of the target reference image.

3. The method according to claim 1 , wherein the determining of the target reference image in the at least one candidate reference image comprises:

determining an image similarity between each candidate reference image and the query image; and

determining a candidate reference image whose image similarity is greater than or equal to a preset similarity threshold as the target reference image.

4. The method according to claim 1 , wherein the determining of the target reference image in the at least one candidate reference image comprises:

obtaining a global image feature of each candidate reference image;

determining a global image feature of the query image;

determining a distance between the global image feature of each candidate reference image and the global image feature of the query image; and

determining a candidate reference image whose distance is less than or equal to a preset distance threshold as the target reference image.

5. The method according to claim 1 , wherein the determining of the at least one candidate reference image based on the prestored correspondence between the reference image and the text field and based on the N text fields comprises:

inputting the N text fields comprised in the query image into a pre-trained text classifier to obtain a text type of each text field comprised in the query image;

determining a text field whose text type is a preset salient type; and

searching, based on the prestored correspondence between the reference image and the text field, for a candidate reference image corresponding to the text field of the salient type.

6. An apparatus for determining a pose, wherein the apparatus comprises:

one or more processors;

a receiving module, configured by the one or more processors to receive a query image sent by a terminal and N text fields comprised in the query image, wherein N is greater than or equal to 1, wherein the query image is obtained based on an image captured by the terminal at a first location, and wherein a scene at the first location comprises a scene in the query image;

a determining module, configured by the one or more processors to determine at least one candidate reference image based on a prestored correspondence between a reference image and a text field and based on the N text fields,

determine a target reference image in at least one candidate reference image, wherein the scene at the first location comprises a scene in the target reference image, and determine an initial pose of the terminal at the first location based on the query image and the target reference image; and

a sending module, configured by the one or more processors to send the initial pose to the terminal.

7. The apparatus according to claim 6 , wherein the determining module is configured to:

determine a 2D-2D correspondence between the query image and the target reference image; and

determine the initial pose of the terminal at the first location based on the 2D-2D correspondence and a preset 2D-3D correspondence of the target reference image.

8. The apparatus according to claim 6 , wherein the determining module is configured to:

determine an image similarity between each candidate reference image and the query image; and

determine a candidate reference image whose image similarity is greater than or equal to a preset similarity threshold as the target reference image.

9. The apparatus according to claim 6 , wherein the determining module is configured to:

obtain a global image feature of each candidate reference image;

determine a global image feature of the query image;

determine a distance between the global image feature of each candidate reference image and the global image feature of the query image; and

determine a candidate reference image whose distance is less than or equal to a preset distance threshold as the target reference image.

10. The apparatus according to claim 6 , wherein the determining module is configured to:

input the N text fields comprised in the query image into a pre-trained text classifier to obtain a text type of each text field comprised in the query image;

determine a text field whose text type is a preset salient type; and

search, based on the prestored correspondence between a reference image and a text field, for a candidate reference image corresponding to the text field of the salient type.

11. A non-transitory computer readable medium storing program instructions, which, when executed by one or more processors, causes the one or more processors to perform operations for determining a pose, the operations comprising:

receiving a query image sent by a terminal and N text fields comprised in the query image, wherein N is greater than or equal to 1, wherein the query image is obtained based on an image captured by the terminal at a first location, and wherein a scene at the first location comprises a scene in the query image;

determining at least one candidate reference image based on a prestored correspondence a reference image and a text field and based on the N text fields;

determining a target reference image in the at least one candidate reference image, wherein the scene at the first location comprises a scene in the target reference image;

determining an initial pose of the terminal at the first location based on the query image and the target reference image; and

sending the initial pose to the terminal.

12. The non-transitory computer readable medium according to claim 11 , wherein the determining of the initial pose of the terminal at the first location based on the query image and the target reference image comprises:

determining a 2D-2D correspondence between the query image and the target reference image; and

determining the initial pose of the terminal at the first location based on the 2D-2D correspondence and a preset 2D-3D correspondence of the target reference image.

13. The non-transitory computer readable medium according to claim 11 , wherein the determining of the target reference image in the at least one candidate reference image comprises:

determining an image similarity between each candidate reference image and the query image; and

determining a candidate reference image whose image similarity is greater than or equal to a preset similarity threshold as the target reference image.

14. The non-transitory computer readable medium according to claim 11 , wherein the determining of the target reference image in the candidate reference image comprises:

obtaining a global image feature of each candidate reference image;

determining a global image feature of the query image;

determining a distance between the global image feature of each candidate reference image and the global image feature of the query image; and

determining a candidate reference image whose distance is less than or equal to a preset distance threshold as the target reference image.

15. The non-transitory computer readable medium according to claim 11 , wherein the determining of the at least one candidate reference image based on the prestored correspondence between the reference image and the text field and based on the N text fields comprises:

inputting the N text fields comprised in the query image into a pre-trained text classifier to obtain a text type of each text field comprised in the query image;

determining a text field whose text type is a preset salient type; and

searching, based on the prestored correspondence between the reference image and the text field, for a candidate reference image corresponding to the text field of the salient type.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2023
From: LIU, JINXUE; CAO, JUN; LI, RUIHUA; GE, JIANGE; CHEN, YUANLIN
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 064133/0159 →
Priority Claims (2)
CN 201911089900.7 · Nov 8, 2019 · national
CN 202010124987.3 · Feb 27, 2020 · national
Continuity (2)
Continuation PCTCN2020100475 · Jul 6, 2020
Related Publication 20220262035A1 · Aug 18, 2022
References Cited (14)
US 9080882B2 · Gupta et al. · 2015 [cited by applicant]
US 20130157682A1 · Ling · 2013 [cited by applicant]
US 20130243250A1 · France et al. · 2013 [cited by applicant]
US 20170046580A1 · Lu · 2017 [cited by examiner]
US 20180033162A1 · Chang · 2018 [cited by examiner]
US 20190370998A1 · Ciecko · 2019 [cited by examiner]
US 20220262035A1 · Liu · 2022 [cited by examiner]
CN 103489002A · 2014 [cited by applicant]
CN 104422439A · 2015 [cited by applicant]
CN 106304335A · 2017 [cited by applicant]
CN 109919157A · 2019 [cited by applicant]
ITU-T H.264(Jun. 2019),Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Advanced video coding for generic audiovisual services, total 836 pages. [cited by applicant]
Torsten Sattler et al, Ef cient and Effective Prioritized Matching for Large-Scale Image-Based Localization, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, No. 9, Sep. 2017, 13 pages. [cited by applicant]
Paul-Edouard Sarlin et al, From Coarse to Fine: Robust Hierarchical Localization at Large Scale, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, 10 pages. [cited by applicant]