IP Library › Granted Patent US 12,651,371
Granted Patent B2
US 12,651,371 · App. 18/372,477 · Granted Jun 9, 2026

Visual positioning method, storage medium and electronic device

Inventors: Yuhao Zhou (Dongguan, CN); Jijunnan Li (Dongguan, CN); Yandong Guo (Dongguan, CN)
Assignee: GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP., LTD.
G06T7/73G06T7/33G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,371
App. No.
18/372,477
Granted
Jun 9, 2026
Kind
B2
Abstract

Provided are a visual positioning method, a non-transitory computer-readable storage medium and an electronic device. Surface normal vectors of a current image frame is obtained. A first transformation parameter between the current image frame and a reference image frame is determined, by projecting the surface normal vectors to a Manhattan coordinate system. A matching operation between feature points of the current image frame and feature points of the reference image frame is performed, and a second transformation parameter between the current image frame and the reference image frame is determined based on a matching result. A target transformation parameter is obtained, based on the first transformation parameter and the second transformation parameter. A visual localization result corresponding to the current image frame is output, based on the target transformation parameter.

Claims (84)

1 . A visual localization method, comprising:

obtaining surface normal vectors of a current image frame;

determining a first transformation parameter between the current image frame and a reference image frame, by projecting the surface normal vectors to a Manhattan coordinate system;

performing a matching operation between feature points of the current image frame and feature points of the reference image frame, and determining, based on a matching result, a second transformation parameter between the current image frame and the reference image frame;

obtaining a target transformation parameter, based on the first transformation parameter and the second transformation parameter; and

outputting, based on the target transformation parameter, a visual localization result corresponding to the current image frame, comprising:

determining a first pose corresponding to the current image frame, based on the target transformation parameter and a pose corresponding to the reference image frame;

projecting, based on the first pose, a three-dimensional point cloud of a target scene to a plane of the current image frame, and obtaining projection points corresponding to the target scene, the target scene being a scene for which the current image frame and the reference image frame are captured;

performing a matching operation between the feature points of the current image frame and the projection points, and determining, based on matching point pairs of the feature points of the current image frame and the projection points, a second pose corresponding to the current image frame; and

outputting the second pose as the visual localization result corresponding to the current image frame.

2 . The method as claimed in claim 1 , wherein obtaining the surface normal vectors of the current image frame, comprises:

obtaining the surface normal vectors of the current image frame, by processing the current image frame through a trained surface-normal-vector estimation network.

3 . The method as claimed in claim 2 , wherein the surface-normal-vector estimation network comprises an encoding sub-network, a decoding sub-network and a convolutional sub-network, and obtaining the surface normal vectors of the current image frame by processing the current image frame through the trained surface-normal-vector estimation network, comprises:

obtaining a down-sampled intermediate image and a down-sampled target image, by down-sampling the current image frame through the encoding sub-network;

obtaining an up-sampled target image, by up-sampling the down-sampled target image and performing a concatenation operation on the down-sampled target image after undergoing the up-sampling and the down-sampled intermediate image, through the decoding sub-network; and

obtaining the surface normal vectors, by performing a convolution operation on the up-sampled target image through the convolutional sub-network.

4 . The method as claimed in claim 1 , wherein determining the first transformation parameter between the current image frame and the reference image frame by projecting the surface normal vectors to the Manhattan coordinate system, comprises

mapping, based on a transformation parameter for transformation from a camera coordinate system corresponding to the reference image frame to the Manhattan coordinate system, the surface normal vectors to the Manhattan coordinate system;

determining, based on an offset of the surface normal vectors in the Manhattan coordinate system, a transformation parameter for transformation from a camera coordinate system corresponding to the current image frame to the Manhattan coordinate system; and

determining the first transformation parameter between the current image frame and the reference image frame, based on the transformation parameter for transformation from the camera coordinate system corresponding to the reference image frame to the Manhattan coordinate system and the transformation parameter for transformation from the camera coordinate system corresponding to the current image frame to the Manhattan coordinate system.

5 . The method as claimed in claim 4 , wherein mapping, based on the transformation parameter for transformation from the camera coordinate system corresponding to the reference image frame to the Manhattan coordinate system, the surface normal vectors to the Manhattan coordinate system, comprises:

mapping, based on the transformation parameter for transformation from the camera coordinate system corresponding to the reference image frame to the Manhattan coordinate system, three-dimensional coordinates of the surface normal vectors in the camera coordinate system corresponding to the reference image frame to three-dimensional coordinates in the Manhattan coordinate system; and

mapping the three-dimensional coordinates of the surface normal vectors in the Manhattan coordinate system to two-dimensional coordinates of the surface normal vectors on a tangent plane of each axis of the Manhattan coordinate system.

6 . The method as claimed in claim 5 , wherein determining, based on the offset of the surface normal vectors in the Manhattan coordinate system, the transformation parameter for transformation from the camera coordinate system corresponding to the current image frame to the Manhattan coordinate system, comprises:

clustering the two-dimensional coordinates of the surface normal vectors on the tangent plane, and determining, based on a cluster center, an offset of the surface normal vectors on the tangent plane;

mapping two-dimensional coordinates of the offset on the tangent plane into three-dimensional coordinates in the Manhattan coordinate system; and

determining the transformation parameter for transformation from the camera coordinate system corresponding to the current image frame to the Manhattan coordinate system, based on the transformation parameter for transformation from the camera coordinate system corresponding to the reference image frame to the Manhattan coordinate system and the three-dimensional coordinates of the offset in the Manhattan coordinate system.

7 . The method as claimed in claim 6 , wherein mapping the two-dimensional coordinates of the offset on the tangent plane into the three-dimensional coordinates in the Manhattan coordinate system, comprises:

obtaining the three-dimensional coordinates of the offset in the Manhattan coordinate system, by mapping the two-dimensional coordinates of the offset on the tangent plane to a unit sphere of the Manhattan coordinate system through exponential mapping.

8 . The method as claimed in claim 4 , wherein the transformation parameter for transformation from the camera coordinate system corresponding to the reference image frame to the Manhattan coordinate system comprises: a relative rotation matrix between the camera coordinate system corresponding to the reference image frame and the Manhattan coordinate system.

9 . The method as claimed in claim 1 , wherein performing the matching operation between the feature points of the current image frame and the feature points of the reference image frame, comprises:

obtaining first matching information, by performing the matching operation from the feature points of the reference image frame to the feature points of the current image frame;

obtaining second matching information, by performing the matching operation from the feature points of the current image frame to the feature points of the reference image frame; and

obtaining the matching result, based on the first matching information and the second matching information.

10 . The method as claimed in claim 9 , wherein obtaining the matching result based on the first matching information and the second matching information, comprises:

taking, as the matching result, an intersection or a union of the first matching information and the second matching information.

11 . The method as claimed in claim 9 , wherein performing the matching operation between the feature points of the current image frame and the feature points of the reference image frame, further comprises:

removing, based on a geometric constraint on the current image frame and the reference image frame, false matching point pairs of the current image frame and the reference image frame, from the matching result.

12 . The method as claimed in claim 1 , wherein the first transformation parameter comprises a first rotation matrix, the second transformation parameter comprises a second rotation matrix, and obtaining the target transformation parameter based on the first transformation parameter and the second transformation parameter, comprises:

establishing a loss function, based on a deviation between first rotation matrix and the second rotation matrix; and

adjusting the second rotation matrix iteratively to reduce a value of the loss function until the loss function converges, and determining the adjusted second rotation matrix as a rotation matrix of the target transformation parameter.

13 . The method as claimed in claim 1 , wherein performing the matching operation between the feature points of the current image frame and the projection points, comprises:

obtaining third matching information, by performing the matching operation from the projection points to the feature points of the current image frame;

obtaining fourth matching information, by performing the matching operation from the feature points of the current image frame to the projection points; and

obtaining, based on the third matching information and the fourth matching information, the matching point pairs of the feature points of the current image frame and the projection points.

14 . The method as claimed in claim 13 , wherein determining, based on the matching point pairs of the feature points of the current image frame and the projection points, the second pose corresponding to the current image frame, comprises:

obtaining, by replacing the projection points in the matching point pairs with three-dimensional points in the three-dimensional point cloud, a matching relationship between the feature points of the current image frame and the three-dimensional points, and solving the second pose based on the matching relationship.

15 . The method as claimed in claim 1 , wherein obtaining the surface normal vectors of the current image frame, comprises:

obtaining the surface normal vector of each pixel point in the current image frame.

16 . The method as claimed in claim 15 , wherein determining the first transformation parameter between the current image frame and the reference image frame by projecting the surface normal vectors to the Manhattan coordinate system, comprises:

determining the first transformation parameter between the current image frame and the reference image frame, by projecting the surface normal vector of each pixel point to the Manhattan coordinate system.

17 . A non-transitory computer-readable storage medium storing a computer program thereon, wherein the computer program, when being executed by a processor, causes a visual localization method to be implemented, and the method comprises:

obtaining a surface normal vector of a current image frame;

projecting the surface normal vector to a Manhattan coordinate system, and determining, based on an offset of the projection of the surface normal vector on the Manhattan coordinate system, a first transformation parameter between the current image frame and a reference image frame;

performing a matching operation between feature points of the current image frame and feature points of the reference image frame, and determining, based on a matching result, a second transformation parameter between the current image frame and the reference image frame;

obtaining a target transformation parameter, based on the first transformation parameter and the second transformation parameter; and

outputting, based on the target transformation parameter, a visual localization result corresponding to the current image frame, comprising:

determining a first pose corresponding to the current image frame, based on the target transformation parameter and a pose corresponding to the reference image frame;

projecting, based on the first pose, a three-dimensional point cloud of a target scene to a plane of the current image frame, and obtaining projection points corresponding to the target scene, the target scene being a scene for which the current image frame and the reference image frame are captured;

performing a matching operation between the feature points of the current image frame and the projection points, and determining, based on matching point pairs of the feature points of the current image frame and the projection points, a second pose corresponding to the current image frame; and

outputting the second pose as the visual localization result corresponding to the current image frame.

18 . An electronic device, comprising:

a processor; and

a memory, configured to store executable instructions for the processor,

wherein the processor is configured to execute the executable instructions to implement a visual localization method comprising:

obtaining surface normal vectors of a current image frame;

determining, based on projections of the surface normal vectors on a Manhattan coordinate system, a first transformation parameter between the current image frame and a reference image frame;

determining a second transformation parameter between the current image frame and the reference image frame, by performing a matching operation between feature points of the current image frame and feature points of the reference image frame;

obtaining a target transformation parameter, based on the first transformation parameter and the second transformation parameter; and

outputting, based on the target transformation parameter, a visual localization result corresponding to the current image frame, comprising:

determining a first pose corresponding to the current image frame, based on the target transformation parameter and a pose corresponding to the reference image frame;

projecting, based on the first pose, a three-dimensional point cloud of a target scene to a plane of the current image frame, and obtaining projection points corresponding to the target scene, the target scene being a scene for which the current image frame and the reference image frame are captured;

performing a matching operation between the feature points of the current image frame and the projection points, and determining, based on matching point pairs of the feature points of the current image frame and the projection points, a second pose corresponding to the current image frame; and

outputting the second pose as the visual localization result corresponding to the current image frame.

19 . The electronic device as claimed in claim 18 , wherein determining, based on the projections of the surface normal vectors on the Manhattan coordinate system, the first transformation parameter between the current image frame and the reference image frame, comprises:

mapping, based on a transformation parameter for transformation from a camera coordinate system corresponding to the reference image frame to the Manhattan coordinate system, the surface normal vectors to the Manhattan coordinate system;

determining, based on an offset of the surface normal vectors in the Manhattan coordinate system, a transformation parameter for transformation from a camera coordinate system corresponding to the current image frame to the Manhattan coordinate system; and

determining the first transformation parameter between the current image frame and the reference image frame, based on the transformation parameter for transformation from the camera coordinate system corresponding to the reference image frame to the Manhattan coordinate system and the transformation parameter for transformation from the camera coordinate system corresponding to the current image frame to the Manhattan coordinate system.

20 . The electronic device as claimed in claim 18 , wherein performing the matching operation between the feature points of the current image frame and the projection points, comprises:

obtaining third matching information, by performing the matching operation from the projection points to the feature points of the current image frame;

obtaining fourth matching information, by performing the matching operation from the feature points of the current image frame to the projection points; and

obtaining, based on the third matching information and the fourth matching information, the matching point pairs of the feature points of the current image frame and the projection points; and

wherein determining, based on the matching point pairs of the feature points of the current image frame and the projection points, the second pose corresponding to the current image frame, comprises:

obtaining, by replacing the projection points in the matching point pairs with three-dimensional points in the three-dimensional point cloud, a matching relationship between the feature points of the current image frame and the three-dimensional points, and solving the second pose based on the matching relationship.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2023
From: ZHOU, YUHAO; LI, JIJUNNAN; GUO, YANDONG
To: GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP., LTD.
Reel/Frame 065013/0392 →
Priority Claims (1)
CN 202110336267.8 · Mar 29, 2021 · national
Continuity (2)
Continuation PCTCN2022078435 · Feb 28, 2022
Related Publication 20240029297A1 · Jan 25, 2024
References Cited (28)
US 20170343356A1 · Roumeliotis et al. · 2017 [cited by applicant]
US 20200167943A1 · Kim · 2020 [cited by examiner]
US 20230360262A1 · He · 2023 [cited by examiner]
CN 103543434A · 2014 [cited by applicant]
CN 106910210A · 2017 [cited by applicant]
CN 107229055A · 2017 [cited by applicant]
CN 107292956A · 2017 [cited by applicant]
CN 107292965A · 2017 [cited by applicant]
CN 108717712A · 2018 [cited by applicant]
CN 109544677A · 2019 [cited by applicant]
CN 109814572A · 2019 [cited by applicant]
CN 109974693A · 2019 [cited by applicant]
CN 110322500A · 2019 [cited by applicant]
CN 110335316A · 2019 [cited by applicant]
CN 111768489A · 2020 [cited by applicant]
CN 111784776A · 2020 [cited by applicant]
CN 111967481A · 2020 [cited by applicant]
Pyojin Kim “Low-Drift Visual Odometry for Indoor Robotics” Dissertation—Seoul National University (Year: 2019). [cited by examiner]
Ning, Ruixin “Visual Positioning and Mapping of Indoor Robotics” Master Thesis—Hangzhou Dianzi University, (Year: 2020). [cited by examiner]
Ning, Ruixin “Visual Positioning and Mapping of Indoor Robotics” Master Thesis—Hangzhou Dianzi University, (Year: 2021). [cited by examiner]
Gao, Ge, et al. “6d object pose regression via supervised learning on point clouds.” IEEE (Year: 2020). [cited by examiner]
CNIPA, Office Action issued for CN Application No. 202110336267.8, Apr. 26, 2023. [cited by applicant]
CNIPA, First Office Action for CN Application No. 202110336267.8, Dec. 20, 2022. [cited by applicant]
Ning Ruixin, Indoor robot visual localization and mapping, China Outstanding Master's Degree Thesis Full Text Database, Feb. 15, 2021, ISSN:1674-0246, Chapter 3. [cited by applicant]
Structure-SLAM: Low-Drift Monocular SLAM in Indoor Environments, Yanyan Li, Aug. 5, 2020. [cited by applicant]
Monocular SLAM methods in structured environments, Salt Particles, Aug. 24, 2020. [cited by applicant]
Divide and Conquer: Efficient Density-Based Tracking of 3D Sensors in Manhattan Worlds, Sep. 4, 2017, Yi Zhou. [cited by applicant]
WIPO, International Search Report for PCT Application No. PCT/CN2022/078435, Feb. 28, 2022. [cited by applicant]