IP Library Granted Patent US 12,307,690
Granted Patent B2
US 12,307,690 · App. 17/866,354 · Granted May 20, 2025

Multi-modal image alignment method and system

Inventors: Jay Huang (Tainan, TW); Hian-Kun Tenn (Tainan, TW); Wen-Hung Ting (Tainan, TW); Chia-Chang Li (Pingtung, TW)
Assignee: INDUSTRIAL TECHNOLOGY RESEARCH INSTITUTE
G06T7/33G06T7/60G06T7/80G05B2219/35063G06T2200/04G06T2207/10028Y02T10/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,307,690
App. No.
17/866,354
Granted
May 20, 2025
Kind
B2
Abstract

A multi-modal image alignment method includes obtaining first points corresponding to a center vertex of a calibration object and second point groups corresponding to side vertices of the calibration object from two-dimensional images, obtaining third points corresponding to the center vertex from three-dimensional images, performing first optimizing computation using a first coordinate system associated with the two-dimensional images, the first points and the third points to obtain a first transformation matrix, processing on the three-dimensional images using the first transformation matrix to generate firstly-transformed images respectively, performing second optimizing computation using the firstly-transformed images, the first points and the second point groups to obtain a second transformation matrix, and transforming an image to be processed from a second coordinate system associated with the three-dimensional images to the first coordinate system using the first transformation matrix and the second transformation matrix.

Claims (44)

1. A multi-modal image alignment system, comprising:

a calibration object, comprising:

a main body having a center vertex and a plurality of side vertices; and

a plurality of indicators respectively disposed on the center vertex and the plurality of side vertices;

a 2D image capturing device having a first 3D coordinate system, and configured to generate a plurality of 2D images associated with the calibration object;

a 3D image capturing device having a second 3D coordinate system, and configured to generate a plurality of 3D images associated with the calibration object; and

a processing device connected to the 2D image capturing device and the 3D image capturing device, and configured to obtain a coordinate transformation matrix according to the plurality of 2D images and the plurality of 3D images, and transform an image to be processed to the second 3D coordinate system or the first 3D coordinate system by using the coordinate transformation matrix.

2. The multi-modal image alignment system according to claim 1 , further comprising:

another 2D image capturing device connected to the processing device, having a third 3D coordinate system, and configured to generate a plurality of second 2D images associated with the calibration object;

wherein the processing device is further configured to obtain another coordinate transformation matrix according to the plurality of second 2D images and the plurality of 3D images, and perform transformation between the first 3D coordinate system and the third 3D coordinate system using the two coordinate transformation matrices.

3. The multi-modal image alignment system according to claim 1 , wherein the plurality of indicators have different colors or temperatures.

4. The multi-modal image alignment system according to claim 1 , wherein to obtain the coordinate transformation matrix, the processing device is further configured to:

obtain a plurality of first points corresponding to the center vertex and a plurality of second point groups corresponding to the plurality of side vertices from the plurality of 2D images;

obtain a plurality of third points corresponding to the center vertex from the plurality of 3D images;

perform first optimizing computation based on the first 3D coordinate system by using an initial first transformation matrix, the plurality of first points and the plurality of third points to obtain an optimized first transformation matrix;

process on the plurality of 3D images by using the optimized first transformation matrix to generate a plurality of firstly-transformed images respectively; and

perform second optimizing computation based on the first 3D coordinate system by using the plurality of firstly-transformed images, the plurality of first points, the plurality of second point groups and a predetermined specification parameter set of the calibration object to obtain an optimized second transformation matrix;

wherein the coordinate transformation matrix comprises the optimized first transformation matrix and the optimized second transformation matrix.

5. The multi-modal image alignment system according to claim 4 , wherein the plurality of first points and the plurality of third points respectively correspond to a plurality of calibration positions where the calibration object is placed, and to perform the first optimizing computation, the processing device is further configured to:

perform a distance calculation on the first point and the third point corresponding to each of the plurality of calibration positions to obtain a plurality of calculation results respectively corresponding to the plurality of calibration positions, wherein the distance calculation comprises:

obtaining a line connecting an origin of the first 3D coordinate system and the first point corresponding to each of the plurality of calibration positions;

transforming the third point corresponding to each of the plurality of calibration positions by using the initial first transformation matrix; and

calculating a distance between the third point transformed by the initial first transformation matrix and the line; and

iteratively adjusting the initial first transformation matrix based on a convergence function, and regarding the iteratively-adjusted initial first transformation matrix as the optimized first transformation matrix, wherein the convergence function is to minimize a sum of the plurality of calculation results.

6. The multi-modal image alignment system according to claim 4 , wherein to perform the second optimizing computation, the processing device is further configured to:

process on the plurality of firstly-transformed images by using an initial second transformation matrix to respectively generate a plurality of secondly-transformed images;

obtain a plurality of fourth point groups respectively from the plurality of secondly-transformed images according to the plurality of first points and the plurality of second point groups;

obtain a plurality of estimated specification parameter sets according to the plurality of fourth point groups respectively; and

iteratively adjust the initial second transformation matrix based on a convergence function, and regard the iteratively-adjusted initial second transformation matrix as the optimized second transformation matrix, wherein the convergence function is to minimize a difference level between the plurality of estimated specification parameter sets and the predetermined specification parameter set of the calibration object.

7. The multi-modal image alignment system according to claim 6 , wherein the plurality of secondly-transformed images, the plurality of first points, the plurality of second point groups respectively correspond to a plurality of calibration positions where the calibration object is placed, and to obtain the plurality of fourth point groups, the processing device is further configured to:

perform steps on a first point, a second point group and a secondly-transformed image corresponding to each of the plurality of calibration positions, wherein the steps comprise:

projecting the first point onto the secondly-transformed image to obtain a fifth point corresponding to each of the plurality of calibration positions; and

projecting a plurality of points in the second point groups onto the secondly-transformed image to obtain a plurality of sixth points corresponding to each of the plurality of calibration positions;

wherein the fifth point and the plurality of sixth points form one of the plurality of fourth point groups.

8. The multi-modal image alignment system according to claim 7 , wherein to obtain the plurality of estimated specification parameter sets, the processing device is further configured to:

perform specification parameter estimation on each of the plurality of fourth point groups to obtain one of the plurality of estimated specification parameter sets, wherein the specification parameter estimation comprises:

obtaining a plurality of connecting lines between the fifth point and the plurality of sixth points; and

calculating a plurality of estimated lengths of the plurality of connecting lines and a plurality of estimated angles between the plurality of connecting lines;

wherein the predetermined specification parameter set comprises a plurality of predetermined side lengths and a plurality of predetermined angles of the calibration object, the difference level between the plurality of estimated specification parameter sets and the predetermined specification parameter set is a weighted sum of a first value and a second value, the first value is a sum of differences respectively between the plurality of estimated lengths and the plurality of predetermined side lengths, and the second value is a sum of differences respectively between the plurality of estimated angles and the plurality of predetermined angles.

9. The multi-modal image alignment system according to claim 4 , wherein to obtaining the plurality of third points, the processing device is further configured to:

regard each of the plurality of 3D images as a target image;

obtain three planes in the target image, wherein the planes are adjacent to each other and three normal vectors of the three planes are perpendicular to each other; and

obtain a point of intersection of the planes as one of the plurality of third points.

10. The multi-modal image alignment system according to claim 1 , wherein the first 3D coordinate system is a camera coordinate system of a 2D image capturing device, the processing device is further configured to transform the image to be processed from the first 3D coordinate system to an image plane coordinate system of the 2D image capturing device using a parameter matrix of the 2D image capturing device, and the parameter matrix comprises a local length parameter and a projection center parameter of the 2D image capturing device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 19, 2022
From: HUANG, JAY; TENN, HIAN-KUN; TING, WEN-HUNG; LI, CHIA-CHANG
To: INDUSTRIAL TECHNOLOGY RESEARCH INSTITUTE
Reel/Frame 060547/0857 →
Priority Claims (1)
TW 110146318 · Dec 10, 2021 · national
Continuity (2)
Provisional Application 63234924 · Aug 19, 2021
Related Publication 20230055649A1 · Feb 23, 2023
References Cited (41)
US 5764786A · Kuwashima · 1998 [cited by examiner]
US 7789562B2 · Strobel · 2010 [cited by examiner]
US 9270974B2 · Zhang et al. · 2016 [cited by applicant]
US 9325969B2 · Takemoto et al. · 2016 [cited by applicant]
US 10129490B2 · Beall · 2018 [cited by applicant]
US 10445898B2 · Liu et al. · 2019 [cited by applicant]
US 10796403B2 · Choi et al. · 2020 [cited by applicant]
US 11030763B1 · Srivastava · 2021 [cited by examiner]
US 20050075585A1 · Kim · 2005 [cited by examiner]
US 20090296893A1 · Strobel · 2009 [cited by examiner]
US 20190279399A1 · Yasunaga et al. · 2019 [cited by applicant]
US 20210174529A1 · Srivastava · 2021 [cited by examiner]
US 20210209750A1 · Aponte et al. · 2021 [cited by applicant]
US 20210243369A1 · Dal Mutto et al. · 2021 [cited by applicant]
CN 106846461A · 2017 [cited by applicant]
CN 107564089A · 2018 [cited by examiner]
CN 207218846U · 2018 [cited by examiner]
CN 110648367A · 2020 [cited by applicant]
CN 114545377A · 2022 [cited by examiner]
JP 2002328012A · 2002 [cited by examiner]
JP 2002328013A · 2002 [cited by examiner]
JP 2002328014A · 2002 [cited by examiner]
KR 20110006360A · 2011 [cited by examiner]
KR 20140049361A · 2014 [cited by examiner]
WO WO2019163211A1 · 2019 [cited by examiner]
WO WO2019163212A1 · 2019 [cited by examiner]
Walch et al., “A combined calibration of 2D and 3D sensors a novel calibration for laser triangulation sensors based on point correspondences,” 2014 International Conference on Computer Vision Theory and Applications (V… [cited by examiner]
CN 107564089 A (machine translation) (Year: 2018). [cited by examiner]
CN 114545377 A (machine translation) (Year: 2022). [cited by examiner]
CN 207218846 U (machine translation) (Year: 2018). [cited by examiner]
JP 2002328012 A (machine translation) (Year: 2002). [cited by examiner]
JP 2002328013 A (machine translation) (Year: 2002). [cited by examiner]
JP 2002328014 A (machine translation) (Year: 2002). [cited by examiner]
KR 20110006360 A (machine translation) (Year: 2011). [cited by examiner]
KR 20140049361 A (machine translation) (Year: 2014). [cited by examiner]
WO 2019163211 A1 (machine translation) (Year: 2019). [cited by examiner]
WO 2019163212 A1 (machine translation) (Year: 2019). [cited by examiner]
Rangel et al. “3D Thermal Imaging: Fusion of Thermography and Depth Cameras” University of Kassel, Faculty of Mechanical Engineering, Measurement and Control Department, Jan. 2014. [cited by applicant]
Vidas et al. “3D Thermal Mapping of Building Interiors using an RGB-D and Thermal Camera” 2013 IEEE International Conference on Robotics and Automation (ICRA) May 6-10, 2013. [cited by applicant]
Mouats et al. “Thermal Stereo Odometry for UAVs” IEEE Sensors Journal, vol. 15, No. 11, Nov. 2015. [cited by applicant]
TW Office Action in TW application No. 110146318 dated May 5, 2022. [cited by applicant]