IP Library › Granted Patent US 12,423,826
Granted Patent B2
US 12,423,826 · App. 18/732,924 · Granted Sep 23, 2025

3D semantic segmentation method and computer program recorded on recording medium to execute the same

Inventors: Jae Geun Yoon (Seoul, KR); Seung Jin Oh (Seongnam-si, KR); Kwang Ho Song (Seoul, KR); Jiyeon Jeon (Incheon, KR)
Assignee: INFINIQ CO., LTD.
G06T7/12G01S17/89G06T5/30G06T7/11G06V10/771G06V10/82G06T7/174G06T2207/10028G06T2207/20081G06T2207/20084G06T2207/30261
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,423,826
App. No.
18/732,924
Granted
Sep 23, 2025
Kind
B2
Abstract

3D semantic segmentation methods and devices for performing 3D semantic segmentation on the basis of fusion data obtained through sensor fusion of cameras and LiDAR are described. According to one embodiment, a method comprises receiving an image photographed by a camera and point cloud data acquired from a LiDAR, by a learning data generation device, generating a projection image expressing the point cloud data in polar coordinates of a size equal to those of the image, by the learning data generation device, and inputting the image and the projection image into an artificial intelligence (AI) machine learned in advance to estimate a 2D segment map and a 3D segment map having dimensions as high as a number of types of classes to be predicted, by the learning data generation device.

Claims (189)

1. A 3D semantic segmentation method comprising steps of:

receiving an image photographed by a camera and point cloud data acquired from a LiDAR sensor;

generating, by a learning data generation device, a projection image expressing the point cloud data in polar coordinates of a size equal to those of the image, wherein the generating step includes generating the projection image through a multiplication operation of calibration matrix information between the LiDAR sensor and the camera and coordinates of the point cloud data, and generating an image and the projection image having a same height and width by truncating a preset area from the projection image and by equally truncating the image;

inputting the image and the projection image into an artificial intelligence (AI) machine, learned in advance to estimate a 2D segment map and a 3D segment map having dimensions as high as a number of types of classes to be predicted; and

learning artificial intelligence before the 2D segment map and the 3D segment map are estimated based on a synthesis loss function that simultaneously calculates and sums loss values for estimating the 2D segment map and the 3D segment map,

wherein the synthesis loss function is expressed as shown in a following equation:

L

total

=

L

3

⁢

D

(

p

⁢

r

⁢

e

⁢

d

3

⁢

D

,

label

3

⁢

D

)

+

L

2

⁢

D

(

p

⁢

r

⁢

e

⁢

d

2

⁢

D

,

label

2

⁢

D

)

here, L 2D denotes a first loss value for estimating the 2D segment map, L 3D denotes a second loss value for estimating the 3D segment map, pred 2D denotes a predicted answer for estimating the 2D segment map, pred 3D denotes a predicted answer for estimating the 3D segment map, label 2D denotes a first correct answer value for estimating the 2D segment map, and label 3D denotes a second correct answer value for estimating the 3D segment map,

wherein the learning step includes a step of setting pixels neighboring as much as a preset distance from each point included in the first correct answer value with a same label,

wherein the first correct answer value and the second correct answer value used to calculate each of the loss values are sparse data as the first correct answer value and the second correct answer value are generated in a method of assigning a 3D correct value of the point cloud data to 2D projective points projected through the multiplication operation of calibration matrix information,

wherein the artificial intelligence (AI) machine includes a neural network for images and a neural network for projection images symmetrical to each other, and wherein each of the neural network for images and the neural network for projection images includes:

an encoder including three contextual blocks and four residual blocks (res blocks) for learning a structure and context information of the image and the projection image;

a decoder including four dilation blocks (up blocks) for dilating data output from the encoder, and an output layer for outputting the 2D segment map and the 3D segment map; and

an attention fusion module including eight attention fusion blocks for fusing feature maps output from the three contextual blocks, the four residual blocks, and the four dilation blocks, and

wherein the AI machine generates two feature maps with emphasized features of important locations by finding first locations and reflection rates of important features from a feature map of an image through a spatial attention module, multiplying the first locations and the reflection rates with a feature map of the projection image and the feature map of the image and connecting the feature map of the projection image and the feature map of the image to original feature maps through a residual path.

2. The method according to claim 1 , wherein the first loss value and the second loss value are calculated through a following equation:

L

2

⁢

D

=

L

Focal

(

pred

2

⁢

D

,

label

2

⁢

D

)

+

L

Dice

(

pred

2

⁢

D

,

label

2

⁢

D

)

⁢

L

3

⁢

D

=

L

Focal

(

pred

3

⁢

D

,

label

3

⁢

D

)

+

L

Lovasz

(

pred

3

⁢

D

,

label

3

⁢

D

)

here, pred 2D denotes a predicted answer for estimating the 2D segment map, pred 3D denotes a predicted answer for estimating the 3D segment map, label 2D denotes the first correct answer value, and label 3D denotes the second correct answer value.

3. The method according to claim 1 , wherein the encoder sequentially generates feature maps of ½, ¼, ⅛, and 1/16 times of the size of the image and the projection image, and transfers the feature maps to the four dilation blocks of the decoder, the decoder sequentially restores the feature maps received from the encoder in sizes of ⅛, ¼, ½, and 1, and the four dilation blocks include a pixel shuffle layer for dilating or reducing the feature maps received from the encoder, a dilated convolution layer for learning features of dilated feature maps, and a concatenation layer for concatenating the dilated feature maps with the feature maps transferred from the four residual blocks of the encoder through a residual connection.

4. The method according to claim 3 , wherein the eight attention fusion blocks are arranged between a plurality of residual blocks and a plurality of dilation blocks, excluding the contextual block, to infer features of the projection image having an insufficient amount of information about shapes, structures, and boundaries of objects based on features of an image having color information.

5. A non-transitory computer-readable medium, comprising a set of instructions stored therein, that when executed by a processor, causes the processor to perform a 3D semantic segmentation method comprising:

receiving an image photographed by a camera and point cloud data acquired from a LiDAR sensor;

generating a projection image expressing the point cloud data in polar coordinates of a size equal to those of the image, wherein the generating step includes generating the projection image through a multiplication operation of calibration matrix information between the LiDAR sensor and the camera and coordinates of the point cloud data, and generating an image and the projection image having a same height and width by truncating a preset area from the projection image and by equally truncating the image;

inputting the image and the projection image into an artificial intelligence (AI) machine, learned in advance to estimate a 2D segment map and a 3D segment map, each having dimensions as high as a number of types of classes to be predicted; and

learning artificial intelligence before the 2D segment map and the 3D segment map are estimated based on a synthesis loss function that simultaneously calculates and sums loss values for estimating the 2D segment map and the 3D segment map,

wherein the synthesis loss function is expressed as shown in a following equation:

L

total

=

L

3

⁢

D

(

p

⁢

r

⁢

e

⁢

d

3

⁢

D

,

label

3

⁢

D

)

+

L

2

⁢

D

(

p

⁢

r

⁢

e

⁢

d

2

⁢

D

,

label

2

⁢

D

)

here, L 2D denotes a first loss value for estimating the 2D segment map, L 3D denotes a second loss value for estimating the 3D segment map, pred 2D denotes a predicted answer for estimating the 2D segment map, pred 3D denotes a predicted answer for estimating the 3D segment map, label 2D denotes a first correct answer value for estimating the 2D segment map, and label 3D denotes a second correct answer value for estimating the 3D segment map,

wherein the learning step includes a step of setting pixels neighboring as much as a preset distance from each point included in the first correct answer value with a same label,

wherein the first correct answer value and the second correct answer value used to calculate each of the loss values are sparse data as the first correct answer value and the second correct answer value are generated in a method of assigning a 3D correct value of the point cloud data to 2D projective points projected through the multiplication operation of calibration matrix information,

wherein the artificial intelligence (AI) machine includes a neural network for images and a neural network for projection images symmetrical to each other, and wherein each of the neural network for images and the neural network for projection images includes:

an encoder including three contextual blocks and four residual blocks (res blocks) for learning a structure and context information of the image and the projection image;

a decoder including four dilation blocks (up blocks) for dilating data output from the encoder, and an output layer for outputting the 2D segment map and the 3D segment map; and

an attention fusion module including eight attention fusion blocks for fusing feature maps output from the three contextual blocks, the four residual blocks, and the four dilation blocks, and

wherein the AI machine generates two feature maps with emphasized features of important locations by finding first locations and reflection rates of important features from a feature map of an image through a spatial attention module, multiplying the first locations and the reflection rates with a feature map of the projection image and the feature map of the image and connecting the feature map of the projection image and the feature map of the image to original feature maps through a residual path.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 4, 2024
From: YOON, JAE GEUN; OH, SEUNG JIN; SONG, KWANG HO; JEON, JIYEON
To: INFINIQ CO., LTD.
Reel/Frame 067612/0129 →
Priority Claims (1)
KR 10-2023-0086542 · Jul 4, 2023 · national
Continuity (1)
Related Publication 20250014187A1 · Jan 9, 2025
References Cited (13)
US 20210201145A1 · Pham · 2021 [cited by examiner]
US 20230081913A1 · Tsai · 2023 [cited by examiner]
US 20230267615A1 · Agia · 2023 [cited by examiner]
KR 102073873B1 · 2020 [cited by applicant]
KR 1020210036244A · 2021 [cited by applicant]
He et al., “Multimodal Fusion and Data Augmentation for 3D Semantic Segmentation”, Dec. 1, 2022, ICROS, the 22nd International Conference on Control, Automation and Systems (ICCAS 2022), p. 1143-1148. (Year: 2022). [cited by examiner]
Zhao et al., “LIF-Seg: LiDAR and Camera Image Fusion for 3D LiDAR Semantic Segmentation”, May 17, 2023, IEEE, IEEE Transactions on Multimedia, vol. 26, p. 1158-1168. (Year: 2023). [cited by examiner]
Alnaggar et al., “Multi Projection Fusion for Real-time Semantic Segmentation of 3D LiDAR Point Clouds”, Jan. 2021, IEEE, 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), p. 1799-1808. (Year: 2021). [cited by examiner]
Aksoy et al. “SalsaNet: Fast Road and Vehicle Segmentation in LiDAR Point Clouds for Autonomous Driving”, Oct. 23, 2020, IEEE, 2020 IEEE Intelligent Vehicles Symposium (IV), p. 926-932. (Year: 2020). [cited by examiner]
Zhuang et al., “Perception-Aware Multi-Sensor Fusion for 3D LiDAR Semantic Segmentation”, Oct. 2021, IEEE, 2021 IEEE/CVF International Conference on Computer Vision (ICCV), p. 16260-16270. (Year: 2021). [cited by examiner]
Candan et al., “U-Net-based RGB and LiDAR image fusion for road segmentation”, Jan. 2023, Springer, Signal, Image and Video Processing (2023), vol. 17, p. 2837-2843. (Year: 2023). [cited by examiner]
El-Madawi et al., “RGB and LiDAR fusion based 3D Semantic Segmentation for Autonomous Driving”, Oct. 2019, IEEE, 2019 IEEE Intelligent Transportation Systems Conference (ITSC), p. 7-12. (Year: 2019). [cited by examiner]
Jaegeun Yoon et al., “TwinAMFNet : Twin Attention-based Multi-modal Fusion Network for 3D Semantic Segmentation”, Journal of KIISE, Sep. 2023, pp. 784-794, vol. 50, No. 9, doi: 10.5626/JOK.2023.50.9.784. [cited by applicant]