IP Library › Granted Patent US 12,254,681
Granted Patent B2
US 12,254,681 · App. 17/903,393 · Granted Mar 18, 2025

Multi-modal test-time adaptation

Inventors: Yi-Hsuan Tsai (Santa Clara, CA); Bingbing Zhuang (San Jose, CA); Samuel Schulter (New York, NY); Buyu Liu (Cupertino, CA); Sparsh Garg (San Jose, CA); Ramin Moslemi (Pleasanton, CA); Inkyu Shin (Daejeon, KR)
Assignee: NEC Corporation
G06V10/811G01S17/89G06V10/776
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,681
App. No.
17/903,393
Granted
Mar 18, 2025
Kind
B2
Abstract

Systems and methods are provided for multi-modal test-time adaptation. The method includes inputting a digital image into a pre-trained Camera Intra-modal Pseudo-label Generator, and inputting a point cloud set into a pre-trained Lidar Intra-modal Pseudo-label Generator. The method further includes applying a fast 2-dimension (2D) model, and a slow 2D model, to the inputted digital image to apply pseudo-labels, and applying a fast 3-dimension (3D) model, and a slow 3D model, to the inputted point cloud set to apply pseudo-labels. The method further includes fusing pseudo-label predictions from the fast models and the slow models through an Inter-modal Pseudo-label Refinement module to obtain robust pseudo labels, and measuring a prediction consistency for the pseudo-labels. The method further includes selecting confident pseudo-labels from the robust pseudo labels and measured prediction consistencies to form a final cross-modal pseudo-label set as a self-training signal, and updating batch parameters utilizing the self-training signal.

Claims (120)

1. A method for multi-modal test-time adaptation, comprising:

inputting a digital image into a pre-trained Camera Intra-modal Pseudo-label Generator (C-Intra-PG);

inputting a Lidar point cloud set into a pre-trained Lidar Intra-modal Pseudo-label Generator (L-Intra-PG);

applying a fast 2-dimension (2D) model, F 2D , and a slow 2D model, S 2D , to the inputted digital image to apply pseudo-labels to the digital image;

applying a fast 3-dimension (3D) model, F 3D , and a slow 3D model, S 3D , to the inputted Lidar point cloud set to apply pseudo-labels to the Lidar point cloud set;

fusing pseudo-label predictions from the fast (F 2D , F 3D ) models and the slow (S 2D , S 3D ) models through Inter-modal Pseudo-label Refinement (Inter-PR) module to obtain robust pseudo labels;

measuring a prediction consistency for each of the digital image pseudo-labels and Lidar pseudo-labels separately;

selecting confident pseudo-labels from the robust pseudo labels and measured prediction consistencies to form a final cross-modal pseudo-label set as a self-training signal; and

updating batch parameters of the Camera Intra-modal Pseudo-label Generator and Lidar Intra-modal Pseudo-label Generator utilizing the self-training signal.

2. The method as recited in claim 1 , further comprising training segmentation models for the C-Intra-PG and the L-Intra-PG including the fast and the slow models with source domain data.

3. The method as recited in claim 2 , wherein the digital image and the Lidar point cloud set are multi-modal test-time target data that the C-Intra-PG and the L-Intra-PG see once.

4. The method as recited in claim 3 , wherein the digital image is captured at test-time by a camera and the Lidar point cloud set is captured at test-time by a LiDAR sensor.

5. The method as recited in claim 4 , wherein consistency is measured as:

cons M =Sim(G M (x t M ), S M (x t M )), where SIM(−) is the inverse of KL divergence, and M is the 2D or 3D modality.

6. The method as recited in claim 5 , wherein the pseudo-labels are obtained by:

y

^

t

M

=

arg

⁢

max

k

∈

K

⁢

p

⁢

(

x

t

M

)

(

k

)

.

7. The method as recited in claim 6 , wherein the robust pseudo-labels are selected based on a hard select criteria or a soft select criteria.

8. The method as recited in claim 7 , wherein pseudo labels with a maximum consistency measure over the two modalities is below a threshold is ignored.

9. A computer system for multi-modal test-time adaptation, comprising:

one or more processors;

a display screen coupled to the one or more processors through a bus; and

memory coupled to the one or more processors through the bus, wherein the memory includes a multi-modal test-time adaptation tool configured to:

receive a digital image into a pre-trained Camera Intra-modal Pseudo-label Generator (C-Intra-PG);

receive a Lidar point cloud set into a pre-trained Lidar Intra-modal Pseudo-label Generator (L-Intra-PG);

apply a fast 2-dimension (2D) model, F 2D , and a slow 2D model, S 2D , to the inputted digital image to apply pseudo-labels to the digital image;

apply a fast 3-dimension (3D) model, F 3D , and a slow 3D model, S 3D , to the inputted Lidar point cloud set to apply pseudo-labels to the Lidar point cloud set;

fuse pseudo-label predictions from the fast (F 2D , F 3D ) models and the slow (S 2D , S 3D ) models through Inter-modal Pseudo-label Refinement (Inter-PR) module to obtain robust pseudo labels;

measure a prediction consistency for each of the digital image pseudo-labels and Lidar pseudo-labels separately;

select confident pseudo-labels from the robust pseudo labels and measured prediction consistencies to form a final cross-modal pseudo-label set as a self-training signal; and

update batch parameters of the Camera Intra-modal Pseudo-label Generator and Lidar Intra-modal Pseudo-label Generator utilizing the self-training signal.

10. The computer system as recited in claim 9 , wherein the multi-modal test-time adaptation tool is further configured to train segmentation models for the C-Intra-PG and the L-Intra-PG including the fast and the slow models with source domain data.

11. The computer system as recited in claim 10 , wherein the digital image and the Lidar point cloud set are multi-modal test-time target data that the C-Intra-PG and the L-Intra-PG see once.

12. The computer system as recited in claim 11 , wherein the digital image is captured at test-time by a camera and the Lidar point cloud set is captured at test-time by a LIDAR sensor.

13. The computer system as recited in claim 12 , wherein consistency is measured as:

cons M =Sim(G M (x t M ), S M (x t M )), where SIM(−) is the inverse of KL divergence, and M is the 2D or 3D modality.

14. The computer system as recited in claim 13 , wherein the pseudo-labels are obtained by:

y

^

t

M

=

arg

⁢

max

k

∈

K

⁢

p

⁢

(

x

t

M

)

(

k

)

.

15. The computer system as recited in claim 14 , wherein the robust pseudo-labels are selected based on a hard select criteria or a soft select criteria.

16. The computer system as recited in claim 15 , wherein pseudo labels with a maximum consistency measure over the two modalities below a threshold is ignored.

17. A non-transitory computer readable storage medium comprising a computer readable program for multi-modal test-time adaptation, wherein the computer readable program when executed on a computer causes the computer to perform the steps of:

receiving a digital image into a pre-trained Camera Intra-modal Pseudo-label Generator (C-Intra-PG);

receiving a Lidar point cloud set into a pre-trained Lidar Intra-modal Pseudo-label Generator (L-Intra-PG);

applying a fast 2-dimension (2D) model, F 2D , and a slow 2D model, S 2D , to the inputted digital image to apply pseudo-labels to the digital image;

applying a fast 3-dimension (3D) model, F 3D , and a slow 3D model, S 3D , to the inputted Lidar point cloud set to apply pseudo-labels to the Lidar point cloud set;

fusing pseudo-label predictions from the fast (F 2D , F 3D ) models and the slow (S 2D , S 3D ) models through Inter-modal Pseudo-label Refinement (Inter-PR) module to obtain robust pseudo labels;

measuring a prediction consistency for each of the digital image pseudo-labels and Lidar pseudo-labels separately;

selecting confident pseudo-labels from the robust pseudo labels and measured prediction consistencies to form a final cross-modal pseudo-label set as a self-training signal; and

updating batch parameters of the Camera Intra-modal Pseudo-label Generator and Lidar Intra-modal Pseudo-label Generator utilizing the self-training signal.

18. The non-transitory computer readable storage medium as recited in claim 17 , wherein the digital image and the Lidar point cloud set are multi-modal test-time target data that the C-Intra-PG and the L-Intra-PG see once.

19. The non-transitory computer readable storage medium as recited in claim 18 , wherein consistency is measured as:

cons M =Sim(G M (x t M ), S M (x t M )), where SIM(−) is the inverse of KL divergence, and M is the 2D or 3D modality.

20. The non-transitory computer readable storage medium as recited in claim 19 , wherein the pseudo-labels are obtained by:

y

^

t

M

=

arg

⁢

max

k

∈

K

⁢

p

⁢

(

x

t

M

)

(

k

)

,

and the robust pseudo-labels are selected based on a hard select criteria or a soft select criteria.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 9, 2025
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 069802/0620 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2022
From: TSAI, YI-HSUAN; ZHUANG, BINGBING; SCHULTER, SAMUEL; LIU, BUYU; GARG, SPARSH; MOSLEMI, RAMIN; SHIN, INKYU
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 060997/0189 →
Continuity (3)
Provisional Application 63279715 · Nov 16, 2021
Provisional Application 63241137 · Sep 7, 2021
Related Publication 20230081913A1 · Mar 16, 2023
References Cited (56)
US 11120276B1 · Zhang · 2021 [cited by examiner]
US 11703566B2 · Cohen · 2023 [cited by examiner]
US 20200219264A1 · Brunner · 2020 [cited by examiner]
US 20210012166A1 · Braley · 2021 [cited by examiner]
US 20230213643A1 · Hwang · 2023 [cited by examiner]
CN 117572457A · 2024 [cited by examiner]
EP 3686776A1 · 2020 [cited by examiner]
WO WO2021007100A1 · 2021 [cited by examiner]
Sparse-to-dense Feature Matching: Intra and Inter domain Cross-modal Learning in Domain Adaptation for 3D Semantic Segmentation, Duo Peng et al., arXiv, Aug. 2021, pp. 1-10 (Year: 2021). [cited by examiner]
Self-Supervised Person Detection in 2D Range Data using a Calibrated Camera, Dan Jia et al., IEEE, 2021, pp. 13301-13307 (Year: 2021). [cited by examiner]
Lidar-Camera Semi-Supervised Learning for Semantic Segmentation, Luca Caltagirone et al., MDPI, 2021, pp. 1-13 (Year: 2021). [cited by examiner]
Complete & Label: A Domain Adaptation Approach to Semantic Segmentation of LiDAR Point Clouds, Li Yi et al., CVF,2021, pp. 15363-15373 (Year: 2021). [cited by examiner]
Tent: Fully Test-Time Adaptation by Entropy Minimization, Dequan Wang et al., arXiv, Mar. 2021, pp. 1-15 (Year: 2021). [cited by examiner]
Test-Time Training with Self-Supervision for Generalization under Distribution Shifts, Yu Sun et al., PMLR, 2020, pp. 1-20 (Year: 2020). [cited by examiner]
Two-phase Pseudo Label Densification for Self-training based Domain Adaptation, Inkya Shin et al., arXiv, 2020, pp. 1-17 (Year: 2020). [cited by examiner]
Pseudo-labeling for Scalable 3D Object Detection, Benjamin Caine et al., arXiv, Mar. 2021, pp. 1-16 (Year: 2021). [cited by examiner]
RGB and LiDAR fusion based 3D Semantic Segmentation for Autonomous Driving, Khaled El Mandawi et al., arXiv, Jul. 2019, pp. 1-7 (Year: 2019). [cited by examiner]
XMUDA: Cross-Modal Unsupervised Domain Adaptation for 3D Semantic Segmentation, Maximilian Jaritz et al., arXiv, Mar. 2020, pp. 1-12 (Year: 2020). [cited by examiner]
Jaritz, M., Vu, T. H., Charette, R. D., Wirbel, E., & Pérez, P. (Jun. 13, 2020). xmuda: Cross-modal unsupervised domain adaptation for 3d semantic segmentation. In Proceedings of the IEEE/CVF conference on computer visi… [cited by applicant]
Bayoudh, K., Knani, R., Hamdaoui, F., & Mtibaa, A. (Jun. 10, 2021). A survey on deep multimodal learning for computer vision: advances, trends, applications, and datasets. The Visual Computer, 38(8), 2939-2970. [cited by applicant]
Behley, J., Garbade, M., Milioto, A., Quenzel, J., Behnke, S., Stachniss, C., & Gall, J. (Jun. 15, 2019). Semantickitti: A dataset for semantic scene understanding of lidar sequences. In Proceedings of the IEEE/CVF Inte… [cited by applicant]
Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., . . . & Beijbom, O. (Jun. 13, 2020). nuscenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF conference on computer vis… [cited by applicant]
Chi, Z., Wang, Y., Yu, Y., & Tang, J. (Jun. 20, 2021). Test-time fast adaptation for dynamic scene deblurring via meta- auxiliary learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogn… [cited by applicant]
Choy, C., Gwak, J., & Savarese, S. (Jun. 15, 2019). 4d spatio-temporal convnets: Minkowski convolutional neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 3075-30… [cited by applicant]
Duerr, F., Weigel, H., Machlisch, M., & Beyerer, J. (Nov. 9, 2020). Iterative deep fusion for 3D semantic segmentation. In 2020 Fourth IEEE International Conference on Robotic Computing (IRC) (pp. 391-397). IEEE. [cited by applicant]
Geyer, J., Kassahun, Y., Mahmudi, M., Ricou, X., Durgesh, R., Chung, A. S., . . . & Schuberth, P. (Apr. 14, 2020), A2d2: Audi autonomous driving dataset. arXiv preprint arXiv:2004.06320. [cited by applicant]
Graham, B., Engeicke, M., & Van Der Maaten, L. (Jun. 18, 2018). 3d semantic segmentation with submanifold sparse convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (… [cited by applicant]
Graham, B., & van der Maaten, L. (Jun. 5, 2017). Submanifold sparse convolutional networks. arXiv preprint arXiv:1706.01307. [cited by applicant]
He, K., Zhang, X., Ren, S., & Sun, J. (Jun. 26, 2016). Deep residual learning for image recognition. CVPR. 2016. arXM preprint arXiv:1512.03385. [cited by applicant]
Iandola, F. N., Han, S., Moskewicz, M. W., Ashraf, K., Dally, W. J., & Keutzer, K. (Feb. 24, 2016). SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and< 0.5 MB model size. arXiv preprint arXiv:1602.07360. [cited by applicant]
Ioffe, S., & Szegedy, C. (2015, June 1). Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning (pp. 448-456). PMLR. [cited by applicant]
Kim, D., Tsai, Y. H., Zhuang, B., Yu, X., Sclaroff, S., Saenko, K., & Chandraker, M. (Jun. 20, 2021). Learning cross-modal contrastive features for video domain adaptation. In Proceedings of the IEEE/CVF International C… [cited by applicant]
Li, Y., Hao, M., Di, Z., Gundavarapu, N. B., & Wang, X. (Dec. 6, 2021). Test-time personalization with a transformer for human pose estimation. Advances in Neural Information Processing Systems, 34, 2583-2597. [cited by applicant]
El Madawi, K., Rashed, H., El Sallab, A., Nasr, O., Kamel, H., & Yogamani, S. (Oct. 27, 2019). Rgb and lidar fusion based 3d semantic segmentation for autonomous driving. In 2019 IEEE Intelligent Transportation Systems … [cited by applicant]
Meyer, G. P., Charland, J., Hegde, D., Laddha, A., & Vallespi-Gonzalez, C. (Jun. 15, 2019). Sensor fusion for joint 3d object detection and semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vi… [cited by applicant]
Peng, D., Lei, Y., Li, W., Zhang, P., & Guo, Y. (Jun. 20, 2021). Sparse-to-dense feature matching: Intra and inter domain cross-modal learning in domain adaptation for 3d semantic segmentation. In Proceedings of the IEE… [cited by applicant]
Prabhu, V., Khare, S., Kartik, D., & Hoffman, J. (Jul. 21, 2021). S4t: Source-free domain adaptation for semantic segmentation via self-supervised selective self-training. arXiv preprint arXiv:2107.10140. [cited by applicant]
Rist, C. B., Enzweiler, M., & Gavrila, D. M. (Jun. 9, 2019). Cross-sensor deep domain adaptation for LiDAR detection and segmentation. In 2019 IEEE Intelligent Vehicles Symposium (IV) (pp. 1535-1542). IEEE. [cited by applicant]
Ronneberger, O., Fischer, P., & Brox, T. (Oct. 5, 2015). U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention (pp. 23… [cited by applicant]
Ros, G., Sellart, L., Materzynska, J., Vazquez, D., & Lopez, A. M. (Jun. 26, 2016). The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes. In Proceedings of the IEEE confe… [cited by applicant]
Saleh, K., Abobakr, A., Attia, M., Iskander, J., Nahavandi, D., Hossny, M., & Nahvandi, S. (Jun. 15, 2019). Domain adaptation for vehicle detection from bird's eye view LIDAR point cloud data. In Proceedings of the IEEE… [cited by applicant]
Shin, I., Woo, S., Pan, F., & Kweon, I. S. (Aug. 23, 2020). Two-phase pseudo label densification for self-training based domain adaptation. In European conference on computer vision (pp. 532-548). Springer, Cham. [cited by applicant]
Sun, Y., Wang, X., Liu, Z., Miller, J., Efros, A., & Hardt, M. (Nov. 21, 2020). Test-time training with self-supervision for generalization under distribution shifts. In International conference on machine learning (pp.… [cited by applicant]
Tang, H., Liu, Z., Zhao, S., Lin, Y., Lin, J., Wang, H., & Han, S. (Aug. 23, 2020). Searching efficient 3d architectures with sparse point-voxel convolution. In European conference on computer vision (pp. 685-702). Spri… [cited by applicant]
Thomas, H., Qi, C. R., Deschaud, J. E., Marcotegui, B., Goulette, F., & Guibas, L. J. (Jun. 15, 2019). Kpconv: Flexible and deformable convolution for point clouds. In Proceedings of the IEEE/CVF international conferenc… [cited by applicant]
Tsai, Y. H., Hung, W. C., Schulter, S., Sohn, K., Yang, M. H., & Chandraker, M. (Jun. 18, 2018). Learning to adapt structured output space for semantic segmentation. In Proceedings of the IEEE conference on computer vis… [cited by applicant]
Van der Maaten, L., & Hinton, G. (Nov. 1, 2008). Visualizing data using t-SNE. Journal of machine learning research, 9(11). [cited by applicant]
Vu, T. H., Jain, H., Bucher, M., Cord, M., & Pérez, P. (Jun. 15, 2019). Advent: Adversarial entropy minimization for domain adaptation in semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Visi… [cited by applicant]
Wang, D., Shelhamer, E., Liu, S., Olshausen, B., & Darrell, T. (Jun. 18, 2020). Tent: Fully test-time adaptation by entropy minimization. arXiv preprint arXiv:2006.10726. [cited by applicant]
Wang, Y., Shi, T., Yun, P., Tai, L., & Liu, M. (Jul. 17, 2018). Pointseg: Real-time semantic segmentation based on 3d lidar point cloud. arXiv preprint arXiv:1807.06288. [cited by applicant]
Wu, B., Wan, A., Yue, X., & Keutzer, K. (May 21, 2018). Squeezeseg: Convolutional neural nets with recurrent crf for real-time road-object segmentation from 3d lidar point cloud. In 2018 IEEE International Conference on… [cited by applicant]
Wu, B., Zhou, X., Zhao, S., Yue, X., & Keutzer, K. (May 20, 2019). Squeezesegv2: Improved model structure and unsupervised domain adaptation for road-object segmentation from a lidar point cloud. In 2019 International C… [cited by applicant]
Xu, J., Zhang, R., Dou, J., Zhu, Y., Sun, J., & Pu, S. (Jun. 20, 2021). Rpvnet: A deep and efficient range-point-voxel fusion network for lidar point cloud segmentation. In Proceedings of the IEEE/CVF International Conf… [cited by applicant]
Yi, L., Gong, B., & Funkhouser, T. (Jun. 20, 2021). Complete & label: A domain adaptation approach to semantic segmentation of lidar point clouds. In Proceedings of the IEEE/CVF conference on computer vision and pattern… [cited by applicant]
Zou, Y., Yu, Z., Kumar, B. V. K., & Wang, J. (Sep. 8, 2018). Unsupervised domain adaptation for semantic segmentation via class-balanced self-training. In Proceedings of the European conference on computer vision (ECCV)… [cited by applicant]
Zou, Y., Yu, Z., Liu, X., Kumar, B. V. K., & Wang, J. (Jun. 15, 2019). Confidence regularized self-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 5982-5991). [cited by applicant]