IP Library › Granted Patent US 12,548,177
Granted Patent B2
US 12,548,177 · App. 18/205,424 · Granted Feb 10, 2026

Method for training depth estimation model, training apparatus, and electronic device applying the method

Inventors: Chin-Pin Kuo (New Taipei, TW); Tsung-Wei Liu (New Taipei, TW)
Assignee: HON HAI PRECISION INDUSTRY CO., LTD.
G06T7/55G06T7/74G06V10/54G06V10/761
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,177
App. No.
18/205,424
Granted
Feb 10, 2026
Kind
B2
Abstract

A method for training a depth estimation model comprise acquires a first image and a second image being inputted into the depth estimation model. The depth estimation model outputs a first depth image. A posture conversion relationship between the first image and the second image is extracted by a posture estimation model. A restored image is generated based on the first depth image, the posture conversion relationship, and pre-obtained camera parameters. A similarity between the restored image and the first image is calculated to obtain a two-dimension loss image. A first similarity of pixel points of each weak texture region are determined based on the two-dimension loss image. A ratio of the first similarity for adjusting the parameters of the depth estimation model is decreased and a first loss value is obtained. A training apparatus and an electronic device applying the method are also disclosed.

Claims (77)

1 . A method for training a depth estimation model, being applicable in an electronic device; the electronic device comprises a storage medium with computer programs and a processor; the processor executes the computer programs to implement the following steps:

acquiring a first image and a second image, the first image and the second image being images viewed from different viewing angles;

inputting the first image and the second image into the depth estimation model, and outputting a first depth image based on parameters of the depth estimation model;

inputting the first image and the second image into a posture estimation model, and extracting a posture conversion relationship between the first image and the second image;

generating a restored image based on the first depth image, the posture conversion relationship, and pre-obtained camera parameters;

calculating a similarity between the restored image and the first image, and obtaining a two-dimension loss image;

extracting weak texture regions in the first image;

determining a first similarity of pixel points of each of the weak texture regions based on the two-dimension loss image;

decreasing a ratio of the first similarity, and obtaining a first loss value; and

adjusting the parameters of the depth estimation model based on the first loss value.

2 . The method of claim 1 , wherein the method further comprises:

extracting non-weak texture regions in the first image;

determining a second similarity of pixel points of each of the extracted non-weak texture regions based on the two-dimension loss image;

increasing a ratio of the second similarity, and obtaining a second loss value; and

further adjusting the parameters of the depth estimation model based on the second loss value.

3 . The method of claim 2 , wherein the steps of increasing a ratio of the second similarity, and obtaining a second loss value comprises:

acquiring an enlarged scale of the second similarity; and

adjusting the second similarity based on the enlarged scale, and obtaining the second loss value.

4 . The method of claim 2 , wherein the step of extracting non-weak texture regions in the first image comprises:

acquiring information of color and brightness of the first image;

dividing the first image into regions based on the information of color and brightness;

calculating gradient information of the first image; and

selecting the regions based on the gradient information, wherein a gradient average value of the regions is outside a predetermined range, as the non-weak texture regions.

5 . The method of claim 2 , wherein the non-weak texture regions comprise regions of object edges.

6 . The method of claim 1 , wherein the steps of decreasing a ratio of the first similarity, and obtaining a first loss value comprises:

acquiring a reduced scale of the first similarity; and

adjusting the first similarity based on the reduced scale, and obtaining the first loss value.

7 . The method of claim 1 , wherein the step of extracting weak texture regions in the first image comprises:

acquiring information of color and brightness of the first image;

dividing the first image into regions based on the information of color and brightness;

calculating gradient information of the first image; and

selecting the regions based on the gradient information, wherein a gradient average value of the regions is in a predetermined range, as the weak texture regions.

8 . A training apparatus comprises a storage medium and at least one processor; the storage medium stores at least one command; the at least one commands is implemented by the at least one processor to execute functions; the storage medium comprising:

an acquiring module, configured to acquire a first image and a second image; the first image and the second image are images from different viewing angles;

a first inputting module, configured to input the first image and the second image into a depth estimation model, and output a first depth image based on parameters of the depth estimation model;

a second inputting module, configured to input the first image and the second image into a posture estimation model, and extract a posture conversion relationship between the first image and the second image;

a generating module, configured to generate a restored image based on the first depth image, the posture conversion relationship, and pre-obtained camera parameters;

a calculating module, configured to calculate a similarity between the restored image and the first image, and obtain a two-dimension loss image

a extracting module, configured to extract weak texture regions in the first image;

a determining module, configured to determine a first similarity of pixel points of each of the extracted weak texture regions based on the two-dimension loss image;

a decreasing module, configured to decrease a ratio of the first similarity, and obtain a first loss value; and

a adjusting module, configured to adjust the parameters of the depth estimation model based on the first loss value.

9 . An electronic device comprises:

a storage medium; and

a processor,

wherein the storage medium stores computer programs, and

the processor executes the computer programs to implement the following steps:

acquiring a first image and a second image; the first image and the second image are images from different viewing angles;

inputting the first image and the second image into the depth estimation model, and outputting a first depth image based on parameters of the depth estimation model;

inputting the first image and the second image into a posture estimation model, and extracting a posture conversion relationship between the first image and the second image;

generating a restored image based on the first depth image, the posture conversion relationship, and pre-obtained camera parameters;

calculating a similarity between the restored image and the first image, and obtaining a two-dimension loss image;

extracting weak texture regions in the first image;

determining a first similarity of pixel points of each of the extracted weak texture regions based on the two-dimension loss image;

decreasing a ratio of the first similarity, and obtaining a first loss value; and

adjusting the parameters of the depth estimation model based on the first loss value.

10 . The electronic device of claim 9 , wherein the processor further executes the computer programs to implement:

extracting non-weak texture regions in the first image;

determining a second similarity of pixel points of each of the extracted non-weak texture regions based on the two-dimension loss image;

increasing a ratio of the second similarity, and obtaining a second loss value; and

further adjusting the parameters of the depth estimation model based on the second loss value.

11 . The electronic device of claim 10 , wherein the step of increasing a ratio of the second similarity, and obtaining a second loss value comprises:

acquiring an enlarged scale of the second similarity; and

adjusting the second similarity based on the enlarged scale, and obtaining the second loss value.

12 . The electronic device of claim 9 , wherein the steps of decreasing a ratio of the first similarity, and obtaining a first loss value comprises:

acquiring a reduced scale of the first similarity; and

adjusting the first similarity based on the reduced scale, and obtaining the first loss value.

13 . The electronic device of claim 9 , wherein the steps of extracting weak texture regions in the first image comprises:

acquiring information of color and brightness of the first image;

dividing the first image into regions based on the information of color and brightness;

calculating gradient information of the first image; and

selecting the regions based on the gradient information, wherein a gradient average value of the regions is in a predetermined range, as the weak texture regions.

14 . The electronic device of claim 9 , wherein the step of extracting non-weak texture regions in the first image comprises:

acquiring information of color and brightness of the first image;

dividing the first image into regions based on the information of color and brightness;

calculating gradient information of the first image; and

selecting the regions based on the gradient information, wherein a gradient average value of the regions is outside a predetermined range, as the non-weak-texture regions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 2, 2023
From: KUO, CHIN-PIN; LIU, TSUNG-WEI
To: HON HAI PRECISION INDUSTRY CO., LTD.
Reel/Frame 063846/0854 →
Priority Claims (1)
CN 202210624025.3 · Jun 2, 2022 · national
Continuity (1)
Related Publication 20230394693A1 · Dec 7, 2023
References Cited (8)
US 20210312650A1 · Ye et al. · 2021 [cited by applicant]
US 20210398302A1 · Guizilini · 2021 [cited by examiner]
US 20220148207A1 · Varekamp et al. · 2022 [cited by applicant]
CN 103177451A · 2013 [cited by applicant]
CN 112561978A · 2021 [cited by examiner]
CN 112819875A · 2021 [cited by applicant]
TW 202101374A · 2021 [cited by applicant]
IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, No. 4, Depth Estimation Using a Self-Supervised Network Based on Cross-Layer Feature Fusion and the Quadtree Constraint, Tian et al., Apr. 2022. (… [cited by examiner]