IP Library Granted Patent US 12,705,752
Granted Patent B2
US 12,705,752 · App. 18/097,405 · Granted Aug 11, 2026

Systems and methods for continuous adaptation of semantic image segmentation model

Inventors: Riccardo Volpi (Grenoble, FR); Gabriela Csurka Khedari (Crolles, FR); Diane Larlus (La Tronche, FR)
Assignee: NAVER CORPORATION
G06T7/12G06V10/82G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,705,752
App. No.
18/097,405
Granted
Aug 11, 2026
Kind
B2
Abstract

A semantic image segmentation (SIS) system includes: a neural network module trained to generate semantic image segmentation maps based on input images, the semantic image segmentation maps grouping pixels of the input images under respective class labels, respectively; a minimum entropy module configured to, at a first time, determine first minimum entropies of pixels, respectively, in the semantic image segmentation maps generated for a received image and N images received before the received image, where N is an integer greater than or equal to 1; and an adaptation module configured to selectively adjust parameters of the neural network module based on optimization of a loss function that minimizes the first minimum entropies.

Claims (42)

1 . A semantic image segmentation (SIS) system, comprising:

a neural network module trained to generate semantic image segmentation maps based on input images, the semantic image segmentation maps grouping pixels of the input images under respective class labels, respectively;

a minimum entropy module configured to, at a first time, determine first minimum entropies of pixels, respectively, in the semantic image segmentation maps generated for a received image and N images received before the received image,

where N is an integer greater than or equal to 1; and

an adaptation module configured to selectively adjust parameters of the neural network module based on optimization of a loss function that minimizes the first minimum entropies.

2 . The SIS system of claim 1 wherein the adaptation module is configured to adjust batch norm parameters of the neural network module based on the optimization of a loss function that minimizes the first minimum entropies.

3 . The SIS system of claim 2 wherein the batch norm parameters include β and γ of each layer of the neural network module.

4 . The SIS system of claim 1 wherein the neural network module includes a ResNet-50 convolutional neural network.

5 . The SIS system of claim 1 wherein the neural network module includes a visual network having the Transformer architecture.

6 . The SIS system of claim 1 further comprising a buffer module configured to store the N images received before the received image.

7 . The SIS system of claim 6 wherein the received image and the N images received before the received image are captured consecutively in time.

8 . The SIS system of claim 7 wherein the received image and the N images received before the received image are captured non-consecutively in time.

9 . The SIS system of claim 1 wherein:

the minimum entropy module is configured to, at a second time after the first time, determine second minimum entropies of the pixels, respectively, in the semantic image segmentation maps generated for a second received image and N images received before the second received image; and

the adaptation module is configured to selectively adjust the parameters of the neural network module based on the second minimum entropies.

10 . The SIS system of claim 9 wherein the received image, the second received image, and the N images form a continuous video stream.

11 . The SIS system of claim 1 wherein the neural network module is further configured to, after the adjustment of the parameters, determine a semantic image segmentation map based on the received image.

12 . A robot comprising:

a camera;

the SIS system of claim 1 , wherein the received image is captured using the camera; and

a control module configured to actuate an actuator of the robot based on one of the semantic image segmentation maps from the neural network module.

13 . The system of claim 1 wherein the neural network module is configured to receive the received image from a camera.

14 . The system of claim 1 wherein the neural network module is configured to receive the received image from a video stored in memory.

15 . A semantic image segmentation (SIS) method, comprising:

by a neural network module, generating semantic image segmentation maps based on input images, the semantic image segmentation maps grouping pixels of the input images under respective class labels, respectively;

at a first time, determining first minimum entropies of pixels, respectively, in the semantic image segmentation maps generated for a received image and N images received before the received image,

where N is an integer greater than or equal to 1; and

selectively adjusting parameters of the neural network module based on optimization of a loss function that minimizes the first minimum entropies.

16 . The SIS method of claim 15 wherein the selectively adjusting includes adjusting batch norm parameters of the neural network module based on the optimization of a loss function that minimizes the first minimum entropies.

17 . The SIS method of claim 16 wherein the batch norm parameters include β and γ of each layer of the neural network module.

18 . The SIS method of claim 15 wherein the neural network module includes one of:

a ResNet-50 convolutional neural network; and

a visual network having the Transformer architecture.

19 . The SIS method of claim 15 further comprising storing the N images received before the received image in a buffer module.

20 . The SIS method of claim 15 wherein one of:

the received image and the N images received before the received image are captured consecutively in time; and

the received image and the N images received before the received image are captured non-consecutively in time.

21 . A semantic image segmentation (SIS) method, comprising:

by a neural network module, generating semantic image segmentation maps based on input images that group pixels of the input images under respective class labels, respectively;

at a first time, determining first minimum entropies of pixels, respectively, in the semantic image segmentation maps generated for a received image and N images received before the received image,

where N is an integer greater than or equal to 1; and

selectively adjusting the respective class labels of the semantic image segmentation map of the received image based on a function that minimizes the first minimum entropies of the pixels.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2024
From: NAVER LABS CORPORATION
To: NAVER CORPORATION
Reel/Frame 068820/0495 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 16, 2023
From: VOLPI, RICCARDO; CSURKA KHEDARI, GABRIELA; LARLUS, DIANE
To: NAVER CORPORATION; NAVER LABS CORPORATION
Reel/Frame 062384/0603 →
Continuity (1)
Related Publication 20240242357A1 · Jul 18, 2024
References Cited (22)
US 10452978B2 · Shazeer et al. · 2019 [cited by applicant]
Liu, Zhizheng, et al. “Unsupervised Continual Semantic Adaptation through Neural Rendering.” arXiv e-prints (2022): arXiv-2211. (Year: 2022). [cited by examiner]
Vu, Tuan-Hung, et al. “Advent: Adversarial entropy minimization for domain adaptation in semantic segmentation.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019. (Year: 2019). [cited by examiner]
Yang, Tao, et al. “Test-time batch normalization.” arXiv preprint arXiv:2205.10210 (2022). (Year: 2022). [cited by examiner]
Xie, Enze, et al. “SegFormer: Simple and efficient design for semantic segmentation with transformers.” Advances in neural information processing systems 34 (2021): 12077-12090. (Year: 2021). [cited by examiner]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vi… [cited by applicant]
Chen, L.-C., Papandreou, G., Kokkinos, I., Murphy, K., and Yuille, A. (2017). Deeplab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs. IEEE Transactions on Pattern… [cited by applicant]
Hayes, T. and Kanan, C. (2020). Lifelong machine learning with deep streaming linear discriminant analysis. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Workshops. [cited by applicant]
Hayes, T., Cahill, N. D., and Kanan, C. (2019). Memory efficient experience replay for streaming learning. In International Conference on Robotics and Automation (ICRA). [cited by applicant]
Hayes, T., Kafle, K., Shrestha, R., Acharya, M., and Kanan, C. (2020). Remind your neural network to prevent catastrophic forgetting. In European Conference on Computer Vision (ECCV). [cited by applicant]
Parisi, G., Kemker, R., Part, J., Kanan, C., and Wermter, S. (2019). Continual Lifelong Learning with Neural Networks: A Review. Neural Networks, 113:57-71. [cited by applicant]
Riccardo Volpi, Dlane Larlus, Gabriela Csurka: “Entropy minimization strategies over a Buffer of samples for continual adaptation over a stream of unlabeled samples.” Jul. 2022. [cited by applicant]
Richter, S., Vineet, V., Roth, S., and Vladlen, K. (2016). Playing for Data: Ground Truth from Computer Games. In European Conference on Computer Vision (ECCV). [cited by applicant]
Ros, G., Sellart, L., Materzynska, J., V´azquez, D., and Lopez, A. (2016). The SYNTHIA Dataset: a Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes. In IEEE Conference on Computer Vision and… [cited by applicant]
Sakaridis, C., Dai, D., and Van Gool, L. (2021). ACDC: The Adverse Conditions Dataset with Correspondences for Semantic Driving Scene Understanding. arXiv:2104.13395. [cited by applicant]
Schneider, S., Rusak, E., Eck, L., Bringmann, O., Brendel, W., and Bethge, M. (2020). Improving robustness against common corruptions by covariate shift adaptation. In Neural Information Processing Systems (NeurIPS). [cited by applicant]
Sun, Y., Wang, X., Liu, Z., Miller, J., Efros, A., and Hardt, M. (2020). Test-time training with self-supervision for generalization under distribution shifts. In International Conference on Machine Learning (ICML). [cited by applicant]
Volpi, R., de Jorge, P., Larlus, D., and Csurka, G. (2022). On the road to online adaptation for semantic image segmentation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). [cited by applicant]
Wang, D., Shelhamer, E., Liu, S., Olshausen, B., and Darrell, T. (2021). Tent: Fully test-time adaptation by entropy minimization. In International Conference on Learning Representations (ICLR). [cited by applicant]
Zhang, M., Levine, S., and Finn, C. (2021). Memo: Test time robustness via adaptation and augmentation. arXiv:2110.09506. [cited by applicant]
Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. (2017). Pyramid Scene Parsing Network. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). [cited by applicant]
Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., and Torralba, A. (2017). Scene parsing through ade20k dataset. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). [cited by applicant]