IP Library › Granted Patent US 12,586,362
Granted Patent B2
US 12,586,362 · App. 17/987,060 · Granted Mar 24, 2026

Method and apparatus with multi-modal feature fusion

Inventors: Hao Wang (Beijing, CN); Weiming Li (Beijing, CN); Qiang Wang (Beijing, CN); Jiyeon Kim (Suwon-si, KR); Hyun Sung Chang (Suwon-si, KR); Sunghoon Hong (Suwon-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G06V10/806G06T7/50G06T7/62G06T7/70G06T17/00G06T2207/10024G06T2207/10028G06T2207/20016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,362
App. No.
17/987,060
Granted
Mar 24, 2026
Kind
B2
Abstract

A method, apparatus, electronic device, and non-transitory computer-readable storage medium with multi-modal feature fusion are provided. The method includes generating three-dimensional (3D) feature information and two-dimensional (2D) feature information based on a color image and a depth image, generating fused feature information by fusing the 3D feature information and the 2D feature information based on an attention mechanism, and generating predicted image information by performing image processing based on the fused feature information.

Claims (30)

1 . A processor-implemented method performed by a computing apparatus, the method comprising:

generating three-dimensional (3D) feature information and two-dimensional (2D) feature information based on a color image and a depth image;

generating fused feature information by fusing the 3D feature information and the 2D feature information based on an attention mechanism; and

performing at least one of estimating a six-dimensional (6D) pose of an object, estimating a size of the object, reconstructing a shape of the object, or segmenting the object based on the fused feature information,

wherein the attention mechanism includes a self-attention mechanism and cross-attention mechanism,

wherein the fused feature information is generated by fusing the 3D feature information of at least one scale and the 2D feature information of at least one scale, and

wherein the generating of the fused feature information comprises, for the 3D feature information of one scale and the 2D feature information of one scale, generating fused feature information of a current scale by performing a feature fusion on 3D feature information of the current scale and 2D feature information of the current scale based on the attention mechanism, the 3D feature information of the current scale being determined based on fused feature information of a previous scale and 3D feature information of the previous scale, and the 2D feature information of the current scale being determined based on 2D feature information of the previous scale.

2 . The method of claim 1 , wherein the generating of the fused feature information comprises:

acquiring point cloud voxel feature information and/or voxel position feature information based on the 3D feature information;

generating first image voxel feature information based on the 2D feature information; and

performing the generating of the fused feature information based on the point cloud voxel feature information, the voxel position feature information, and/or the first image voxel feature information, based on the attention mechanism.

3 . The method of claim 2 , wherein the performing of the generating of the fused feature information based on the point cloud voxel feature information, the voxel position feature information, and/or the first image voxel feature information, based on the attention mechanism comprises one of:

generating the fused feature information by fusing features using the cross-attention mechanism that is dependent on the first image voxel feature information and feature information output by the self-attention mechanism, that is dependent on the voxel position feature information, the point cloud voxel feature information, and the first image voxel feature information; or

generating the fused feature information by fusing features using the cross-attention mechanism that is dependent on the first image voxel feature information and another feature information output by the self-attention mechanism, that is dependent on the point cloud voxel feature information.

4 . A non-transitory computer-readable storage medium storing instructions that, when executed in one or more processors of the computing apparatus, configure the one or more processors to perform the method of claim 1 .

5 . An apparatus comprising:

one or more processors comprising processing circuitry;

memory comprising one or more storage media storing instructions that, when executed by the one or more processors individually or collectively, cause the apparatus to:

generate three-dimensional (3D) feature information based on a depth image;

generate two-dimensional (2D) feature information based on a color image;

fuse the 3D feature information and the 2D feature information using an attention mechanism to generate fused feature information; and

perform at least one of an estimation of a six-dimensional (6D) pose of an object, an estimation of a size of the object, a reconstruction of a shape of the object, or a segmenting of the object based on the fused feature information,

wherein the attention mechanism includes a self-attention mechanism and cross-attention mechanism,

wherein the fused feature information is generated through a fusing of the 3D feature information of at least one scale and the 2D feature information of at least one scale, and

wherein the generation of the fused feature information comprises, for the 3D feature information of one scale and the 2D feature information of one scale, generation of fused feature information of a current scale by performing a feature fusion on 3D feature information of the current scale and 2D feature information of the current scale based on the attention mechanism, the 3D feature information of the current scale being determined based on fused feature information of a previous scale and 3D feature information of the previous scale, and the 2D feature information of the current scale being determined based on 2D feature information of the previous scale.

6 . The apparatus of claim 5 , wherein the instructions, when executed by the one or more processors individually or collectively, further cause the apparatus to:

generate point cloud voxel feature information and/or voxel position feature information based on the 3D feature information;

generate first image voxel feature information based on the 2D feature information; and

perform the fusing of the 3D feature information and the 2D feature information based on the point cloud voxel feature information, the voxel position feature information, and/or the first image voxel feature information, based on the attention mechanism.

7 . The apparatus of claim 5 , wherein the apparatus is an AR device that further comprises one or more cameras configured to respectively capture the depth image and the color image, and one or more displays to display AR image information based on the predicted image information.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 9, 2023
From: HONG, SUNGHOON
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 062314/0579 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2022
From: WANG, HAO; LI, WEIMING; WANG, QIANG; KIM, JIYEON; CHANG, HYUN SUNG
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 061769/0740 →
Priority Claims (2)
CN 202111348242.6 · Nov 15, 2021 · national
KR 10-2022-0111206 · Sep 2, 2022 · national
Continuity (1)
Related Publication 20230154170A1 · May 18, 2023
References Cited (24)
US 11921824B1 · Hester · 2024 [cited by examiner]
US 20200363815A1 · Mousavian et al. · 2020 [cited by applicant]
US 20220164597A1 · Ye · 2022 [cited by examiner]
US 20220410381A1 · Stoppi · 2022 [cited by examiner]
US 20230035475A1 · Cheng · 2023 [cited by examiner]
US 20240246240A1 · Teoh · 2024 [cited by examiner]
CN 111899301A · 2020 [cited by applicant]
CN 112767466A · 2021 [cited by examiner]
CN 113012122A · 2021 [cited by applicant]
“Liu, Attentive Cross-Modal Fusion Network for RGB-D Saliency Detection, 2021, IEEE, vol. 23, pp. 967-981” (Year: 2021). [cited by examiner]
Zhao, Shanshan, et al. “Adaptive Context-Aware Multi-Modal Network for Depth Completion.” IEEE Transactions on Image Processing 30, arXiv:2008.10833v1 [cs.CV] Aug. 25, 2020, (15 pages in English). [cited by applicant]
Srivastava, Siddharth, et al. “Self Attention Guided Depth Completion using RGB and Sparse LiDAR Point Clouds.” 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, Sep. 2021, (8 pages … [cited by applicant]
Liu, Di, et al. “Attentive Cross-Modal Fusion Network for RGB-D Saliency Detection.” IEEE Transactions on Multimedia 23 (2021)(Published Apr. 2020): 967-981. [cited by applicant]
Tan, Xun, et al. “MBDF-Net: Multi-Branch Deep Fusion Network for 3D Object Detection.” Proceedings of the 1st International Workshop on Multimedia Computing for Urban Data. Oct. 2021, (9 pages in English). [cited by applicant]
Sun, Yangjie, et al. “Deep Multimodal Fusion Network for Semantic Segmentation Using Remote Sensing Image and LiDAR Data.” IEEE Transactions on Geoscience and Remote Sensing 60, 2022 (Published Sep. 2021), (18 pages in … [cited by applicant]
Extended European search report issued on Feb. 22, 2023, in counterpart European Patent Application No. 22207194.6 (10 pages in English). [cited by applicant]
Li, Yi, et al. “DeepIM: Deep iterative matching for 6D pose estimation.” [cited by applicant]
Chen, Wei, et al. “FS-NET: Fast shape-based network for category-level 6D object pose estimation with decoupled rotation mechanism.” [cited by applicant]
He, Yisheng, et al. “FFB6D: A full flow bidirectional fusion network for 6D pose estimation.” [cited by applicant]
Prakash, Aditya, et al. “Multi-modal fusion transformer for end-to-end autonomous driving.” [cited by applicant]
Wang, Yue, et al. “Dynamic graph cnn for learning on point clouds.” [cited by applicant]
Hu, Haotian, et al. “VA-GCN: A vector attention graph convolution network for learning on point clouds.” arXiv preprint arXiv:2106.00227v1 (2021). pp 1-12. [cited by applicant]
Carion, Nicolas, et al. “End-to-end object detection with transformers.” arXiv:2005.12872v3 (2020). pp 1-26. [cited by applicant]
Vaswani, Ashish, et al. “Attention is all you need.” [cited by applicant]