IP Library › Granted Patent US 12,423,853
Granted Patent B2
US 12,423,853 · App. 17/976,367 · Granted Sep 23, 2025

Method and apparatus for correcting perspective of road image

Inventors: Junjie Cai (Beijing, CN); Kai Zhong (Beijing, CN); Jianzhong Yang (Beijing, CN); Deguo Xia (Beijing, CN); Tongbin Zhang (Beijing, CN); Zhen Lu (Beijing, CN)
Assignee: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
G06T7/70G06V10/44G06V20/588G06T2207/30256
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,423,853
App. No.
17/976,367
Granted
Sep 23, 2025
Kind
B2
Abstract

A method and apparatus for processing an image. The method may include: acquiring a top view of a road; identifying a position of a lane line from the top view; cutting the top view into at least two areas, and determining, according to the position of the lane line in each area, a width of a lane in the each area and an average width of the lane in the top view; calculating a first perspective correction matrix by optimizing a first loss function, the first loss function being used to represent a difference between the width of the lane in the each area and the average width of the lane in the top view; and performing a lateral correction on the top view through the first perspective correction matrix to obtain a first corrected image.

Claims (90)

1. A method for processing an image, comprising:

acquiring a top view of a road;

identifying a position of a lane line from the top view;

cutting the top view into at least two areas, and determining, according to the position of the lane line in each area, a width of a lane in the each area and an average width of the lane in the top view;

calculating a first perspective correction matrix by optimizing a first loss function, the first loss function being used to represent a difference between the width of the lane in the each area and the average width of the lane in the top view; and

performing a lateral correction on the top view through the first perspective correction matrix to obtain a first corrected image.

2. The method according to claim 1 , further comprising:

identifying a dashed lane line from the top view;

determining a length of the dashed lane line in the each area and an average length of the dashed lane line in the top view;

calculating a second perspective correction matrix by optimizing a second loss function, the second loss function being used to represent a difference between the length of the dashed lane line in the each area and the average length of the dashed lane line in the top view; and

performing a longitudinal correction on the first corrected image through the second perspective correction matrix to obtain a second corrected image.

3. The method according to claim 1 , wherein the identifying the position of the lane line from the top view comprises:

identifying a pixel and type of the lane line from the top view through a semantic segmentation model, to generate a semantic segmentation image; and

extracting the position of the lane line from the semantic segmentation image.

4. The method according to claim 3 , wherein the extracting the position of the lane line from the semantic segmentation image comprises:

transforming the semantic segmentation image into a binary image;

performing a contour detection on the binary image to obtain a rectangular contour;

splitting the rectangular contour into a plurality of segments along a direction of a long side of the rectangular contour, and performing the contour detection on each segment again to generate a plurality of sub-contours;

extracting a center line of each sub-contour rectangle as a linear vector of the lane line; and

fitting the linear vector of the lane line through a quadratic curve, and predicting and supplementing a missing part of the lane line.

5. The method according to claim 1 , wherein the determining, according to the position of the lane line in the each area, the width of the lane in the each area and the average width of the lane in the top view comprises:

performing a near neighbor search on the identified lane line, and pairing each two lane lines to obtain a matching pair set, wherein each matching pair corresponds to one lane;

calculating, for each lane, a lane width of a middle position of the each area as a width of the lane in the each area; and

calculating the average width of the lane in the top view based on the width of the lane in the each area.

6. The method according to claim 2 , wherein the determining the length of the dashed lane line in the each area and the average length of the dashed lane line in the top view comprises:

using, for the each area, a length of a complete dashed lane line in middle of the area as the length of the dashed lane line in the area; and

calculating the average length of the dashed lane line based on the length of the dashed lane line in the each area.

7. The method according to claim 1 , wherein the acquiring the top view of the road comprises:

acquiring a panoramic view of the road; and

transforming the panoramic view into the top view through a perspective projection method.

8. An electronic device, comprising:

at least one processor; and

a memory, communicatively connected to the at least one processor,

wherein the memory stores an instruction executable by the at least one processor, and the instruction is executed by the at least one processor, to enable the at least one processor to perform operations, the operations comprising:

acquiring a top view of a road;

identifying a position of a lane line from the top view;

cutting the top view into at least two areas, and determining, according to the position of the lane line in each area, a width of a lane in the each area and an average width of the lane in the top view;

calculating a first perspective correction matrix by optimizing a first loss function, the first loss function being used to represent a difference between the width of the lane in the each area and the average width of the lane in the top view; and

performing a lateral correction on the top view through the first perspective correction matrix to obtain a first corrected image.

9. The electronic device according to claim 8 , further comprising:

identifying a dashed lane line from the top view;

determining a length of the dashed lane line in the each area and an average length of the dashed lane line in the top view;

calculating a second perspective correction matrix by optimizing a second loss function, the second loss function being used to represent a difference between the length of the dashed lane line in the each area and the average length of the dashed lane line in the top view; and

performing a longitudinal correction on the first corrected image through the second perspective correction matrix to obtain a second corrected image.

10. The electronic device according to claim 8 , wherein the identifying the position of the lane line from the top view comprises:

identifying a pixel and type of the lane line from the top view through a semantic segmentation model, to generate a semantic segmentation image; and

extracting the position of the lane line from the semantic segmentation image.

11. The electronic device according to claim 10 , wherein the extracting the position of the lane line from the semantic segmentation image comprises:

transforming the semantic segmentation image into a binary image;

performing a contour detection on the binary image to obtain a rectangular contour;

splitting the rectangular contour into a plurality of segments along a direction of a long side of the rectangular contour, and performing the contour detection on each segment again to generate a plurality of sub-contours;

extracting a center line of each sub-contour rectangle as a linear vector of the lane line; and

fitting the linear vector of the lane line through a quadratic curve, and predicting and supplementing a missing part of the lane line.

12. The electronic device according to claim 8 , wherein the determining, according to the position of the lane line in the each area, the width of the lane in the each area and the average width of the lane in the top view comprises:

performing a near neighbor search on the identified lane line, and pairing each two lane lines to obtain a matching pair set, wherein each matching pair corresponds to one lane;

calculating, for each lane, a lane width of a middle position of the each area as a width of the lane in the each area; and

calculating the average width of the lane in the top view based on the width of the lane in the each area.

13. The electronic device according to claim 9 , wherein the determining the length of the dashed lane line in the each area and the average length of the dashed lane line in the top view comprises:

using, for the each area, a length of a complete dashed lane line in middle of the area as the length of the dashed lane line in the area; and

calculating the average length of the dashed lane line based on the length of the dashed lane line in the each area.

14. The electronic device according to claim 8 , wherein the acquiring the top view of the road comprises:

acquiring a panoramic view of the road; and

transforming the panoramic view into the top view through a perspective projection electronic device.

15. A non-transitory computer readable storage medium, storing a computer instruction, wherein the computer instruction, when executed by a processor, causes the processor to perform operations, the operations comprising:

acquiring a top view of a road;

identifying a position of a lane line from the top view;

cutting the top view into at least two areas, and determining, according to the position of the lane line in each area, a width of a lane in the each area and an average width of the lane in the top view;

calculating a first perspective correction matrix by optimizing a first loss function, the first loss function being used to represent a difference between the width of the lane in the each area and the average width of the lane in the top view; and

performing a lateral correction on the top view through the first perspective correction matrix to obtain a first corrected image.

16. The non-transitory computer readable storage medium according to claim 15 , further comprising:

identifying a dashed lane line from the top view;

determining a length of the dashed lane line in the each area and an average length of the dashed lane line in the top view;

calculating a second perspective correction matrix by optimizing a second loss function, the second loss function being used to represent a difference between the length of the dashed lane line in the each area and the average length of the dashed lane line in the top view; and

performing a longitudinal correction on the first corrected image through the second perspective correction matrix to obtain a second corrected image.

17. The non-transitory computer readable storage medium according to claim 15 , wherein the identifying the position of the lane line from the top view comprises:

identifying a pixel and type of the lane line from the top view through a semantic segmentation model, to generate a semantic segmentation image; and

extracting the position of the lane line from the semantic segmentation image.

18. The non-transitory computer readable storage medium according to claim 17 , wherein the extracting the position of the lane line from the semantic segmentation image comprises:

transforming the semantic segmentation image into a binary image;

performing a contour detection on the binary image to obtain a rectangular contour;

splitting the rectangular contour into a plurality of segments along a direction of a long side of the rectangular contour, and performing the contour detection on each segment again to generate a plurality of sub-contours;

extracting a center line of each sub-contour rectangle as a linear vector of the lane line; and

fitting the linear vector of the lane line through a quadratic curve, and predicting and supplementing a missing part of the lane line.

19. The non-transitory computer readable storage medium according to claim 15 , wherein the determining, according to the position of the lane line in the each area, the width of the lane in the each area and the average width of the lane in the top view comprises:

performing a near neighbor search on the identified lane line, and pairing each two lane lines to obtain a matching pair set, wherein each matching pair corresponds to one lane;

calculating, for each lane, a lane width of a middle position of the each area as a width of the lane in the each area; and

calculating the average width of the lane in the top view based on the width of the lane in the each area.

20. The non-transitory computer readable storage medium according to claim 16 , wherein the determining the length of the dashed lane line in the each area and the average length of the dashed lane line in the top view comprises:

using, for the each area, a length of a complete dashed lane line in middle of the area as the length of the dashed lane line in the area; and

calculating the average length of the dashed lane line based on the length of the dashed lane line in the each area.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2022
From: CAI, JUNJIE; ZHONG, KAI; YANG, JIANZHONG; XIA, DEGUO; ZHANG, TONGBIN; LU, ZHEN
To: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
Reel/Frame 061649/0430 →
Priority Claims (1)
CN 202111568762.8 · Dec 21, 2021 · national
Continuity (1)
Related Publication 20230052842A1 · Feb 16, 2023
References Cited (20)
US 20120283895A1 · Noda · 2012 [cited by applicant]
US 20150278611A1 · Chi · 2015 [cited by examiner]
US 20210248392A1 · Zaheer · 2021 [cited by examiner]
CN 103035005B · 2015 [cited by applicant]
CN 106682563A · 2017 [cited by applicant]
CN 107424196A · 2017 [cited by applicant]
CN 108805823A · 2018 [cited by applicant]
CN 109255316A · 2019 [cited by applicant]
CN 111428538A · 2020 [cited by examiner]
CN 111652937A · 2020 [cited by applicant]
CN 110163039B · 2020 [cited by examiner]
CN 111996883B · 2021 [cited by examiner]
KR 2006084745A · 2006 [cited by examiner]
KR 101637535B1 · 2016 [cited by applicant]
KR 20180022277A · 2018 [cited by applicant]
WO 2020087322A1 · 2020 [cited by applicant]
Zhang et al., “Lane line recognition based on improved 2D-gamma function and variable threshold Canny algorithm under complex environment,” Measurement and Control, 2020. [cited by applicant]
Zheng et al., “Lane recognition technique of prediction and inverse projection based on Kalman,” Computer Engineering and Design, 2009. [cited by applicant]
Guosheng et al., “Laneline semantic segmentation algorithm based on convolutional neural network,” Journal of Electronic Measurement and Instrumentation, vol. 32, No. 7, 2018. [cited by applicant]
Chinese Notice of Allowance for Chinese Application No. 114267027, dated Mar. 13, 2025, 5 pages. [cited by applicant]