IP Library Granted Patent US 12,561,962
Granted Patent B2
US 12,561,962 · App. 18/339,075 · Granted Feb 24, 2026

Systems and methods for multimodal fusion of missing and unpaired image and tabular data for defect classification

Inventors: Qisen Cheng (San Jose, CA); Shuhui Qu (San Jose, CA); Kaushik Balakrishnan (San Jose, CA); Janghwan Lee (San Jose, CA)
Assignee: Samsung Display Co., Ltd.
G06V10/803G06T7/0004G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,962
App. No.
18/339,075
Granted
Feb 24, 2026
Kind
B2
Abstract

A method may include providing a data set including rows of data. The rows of data may include at least one row of unpaired modality including a first modality, and at least one row of paired modality may include both the first modality and a second modality. The method may further include imputing, by a modality-specific encoder, the at least one row of unpaired modality by interpolating embeddings from the second modality of the paired modality; training, in a latent space, the modality-specific encoder based on the imputation for unimodal prediction and bimodal prediction; and generating a confidence value for the unimodal prediction and the bimodal prediction.

Claims (34)

1 . A method, comprising:

providing a first data set comprising rows of data, the rows of data comprising at least one row of unpaired modality comprising a first modality, and at least one row of paired modality comprising both the first modality and a second modality;

imputing, by a modality-specific encoder, the at least one row of unpaired modality by interpolating embeddings from the second modality of the paired modality based on one or more prior embeddings of the second modality of the paired modality;

based on the imputing of the at least one row of the unpaired modality and based on a second set of data, training, in a latent space, the modality-specific encoder for a unimodal prediction of a missing modality and a bimodal prediction of the missing modality; and

generating a confidence value for the unimodal prediction and the bimodal prediction.

2 . The method of claim 1 , wherein the generating the confidence value comprises computing Shapley-based explanations for the unimodal prediction and the bimodal prediction.

3 . The method of claim 2 , wherein the computing the Shapley-based explanations comprises comparing an impact of the unimodal prediction and an impact of the bimodal prediction with a predetermined threshold.

4 . The method of claim 3 , further comprising selecting either the unimodal prediction or the bimodal prediction based on the generated confidence value.

5 . The method of claim 1 , wherein the second modality is a missing modality from the unpaired modality,

wherein the one or more prior embeddings of the second modality of the paired modality comprise K priors of the second modality, and

wherein the interpolating the embeddings comprises selecting the K priors of the second modality, the K priors being second modality embeddings of K samples of the at least one row of paired modality having closest embeddings of observed modality.

6 . The method of claim 5 , further comprising computing a weighted sum of the K priors by taking a cross-attention between the K samples and the K priors.

7 . The method of claim 1 , wherein the first modality corresponds to an image modality and the second modality corresponds to a tabular modality.

8 . The method of claim 1 , wherein the first modality correspond to a tabular modality and the second modality corresponds to an image modality.

9 . The method of claim 1 , further comprising training the modality-specific encoder for image modality by performing Vision Transformer.

10 . The method of claim 1 , further comprising training the modality-specific encoder for tabular data modality by performing Feature-Tokenizer Transformer.

11 . A system, comprising:

a memory; and

a processor configured to execute instructions stored in the memory to perform operations comprising:

providing a first data set comprising rows of data, the rows of data comprising at least one row of unpaired modality comprising a first modality, and at least one row of paired modality comprising both the first modality and a second modality;

imputing, by a modality-specific encoder, the at least one row of unpaired modality by interpolating embeddings from the second modality of the paired modality based on one or more prior embeddings of the second modality of the paired modality;

based on the imputing of the at least one row of the unpaired modality and based on a second set of data, training, in a latent space, the modality-specific encoder for a unimodal prediction of a missing modality and a bimodal prediction of the missing modality; and

generating a confidence value for the unimodal prediction and the bimodal prediction.

12 . The system of claim 11 , wherein the generating the confidence value comprises computing Shapley-based explanations for the unimodal prediction and the bimodal prediction.

13 . The system of claim 12 , wherein the computing the Shapley-based explanations comprises comparing an impact of the unimodal prediction and an impact of the bimodal prediction with a predetermined threshold.

14 . The system of claim 13 , wherein the operations further comprise selecting either the unimodal prediction or the bimodal prediction based on the generated confidence value.

15 . The system of claim 11 , wherein the second modality is a missing modality from the unpaired modality,

wherein the one or more prior embeddings of the second modality of the paired modality comprise K priors of the second modality, and

wherein the interpolating the embeddings comprises selecting the K priors of the second modality, the K priors being second modality embeddings of K samples of the at least one row of paired modality having closest embeddings of observed modality.

16 . The system of claim 15 , wherein the operations further comprise computing a weighted sum of the K priors by taking a cross-attention between the K samples and the K priors.

17 . The system of claim 11 , wherein the first modality corresponds to an image modality and the second modality corresponds to a tabular modality.

18 . The system of claim 11 , wherein the first modality correspond to a tabular modality and the second modality corresponds to an image modality.

19 . The system of claim 11 , wherein the operations further comprise training the modality-specific encoder for image modality by performing Vision Transformer.

20 . The system of claim 11 , wherein the operations further comprise training the modality-specific encoder for tabular data modality by performing Feature-Tokenizer Transformer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 21, 2023
From: CHENG, QISEN; QU, SHUHUI; BALAKRISHNAN, KAUSHIK; LEE, JANGHWAN
To: SAMSUNG DISPLAY CO., LTD.
Reel/Frame 064021/0625 →
Continuity (2)
Provisional Application 63452638 · Mar 16, 2023
Related Publication 20240312193A1 · Sep 19, 2024
References Cited (30)
US 8843423B2 · Chu et al. · 2014 [cited by applicant]
US 11488694B2 · Malone et al. · 2022 [cited by applicant]
US 11921824B1 · Hester · 2024 [cited by examiner]
US 20200234086A1 · Taha · 2020 [cited by examiner]
US 20200372369A1 · Gong · 2020 [cited by examiner]
US 20230045548A1 · Yakut · 2023 [cited by examiner]
US 20230117247A1 · Xiao · 2023 [cited by examiner]
US 20230255564A1 · Pascual-Leone · 2023 [cited by examiner]
US 20240134937A1 · Ni · 2024 [cited by examiner]
US 20240297957A1 · Bakunov · 2024 [cited by examiner]
US 20240312193A1 · Cheng · 2024 [cited by examiner]
US 20240374136A1 · Sørensen · 2024 [cited by examiner]
US 20250166746A1 · Balazard · 2025 [cited by examiner]
CN 113706558A · 2021 [cited by applicant]
Lee, Mihee et al, Private-Shared Disentangled Multimodal VAE for Learning of Latent Representations, Dec. 23, 2020 (Year: 2020). [cited by examiner]
EPO Extended European Search Report issued in corresponding EP Application No. 24163661.2, dated Jul. 31, 2024, 11 pages. [cited by applicant]
Gorishniy, Y. et al., “Revisiting Deep Learning Models for Tabular Data,” arXiv:2106.11959v2, Nov. 10, 2021, 25 pages, www.arXiv.org. [cited by applicant]
Han, H. et al., “SSGD: A Smartphone Screen Glass Dataset for Defect Detection,” arXiv:2303.06673v1, Mar. 12, 2023, 5 pages, www.arXiv.org. [cited by applicant]
Lee, M. et al., “Explainable AI for domain experts: a post Hoc analysis of deep learning for defect classification of TFT-LCD panels,” Journal of Intelligent Manufacturing, vol. 33, Mar. 2021, pp. 1747-1759. [cited by applicant]
Liang, P. et al., “Foundations & Recent Trends in Multimodal Machine Learning: Principles, Challenges, & Open Questions,” arXiv:2209.03430v1, Sep. 7, 2022, 65 pages, www.arXiv.org. [cited by applicant]
Ma, M. et al., “Are Multimodal Transformers Robust to Missing Modality?”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 18156-18165. [cited by applicant]
Bahari, Dara et al., “SCARF: Self-Supervised Contrastive Learning Using Random Feature Corruption,” Published as a conference paper at ICLR 2022, 24 pages. [cited by applicant]
Gorishniy, Y., et al., “Revisiting Deep Learning Models for Tabular Data,” 35 [cited by applicant]
He, K., et al., “Masked Autoencoders Are Scalable Vision Learners,” IEEE Xplore, 10 pages. [cited by applicant]
Jain, D.K., et al., “Employing Co-Learning to Evaluate the Explainability of Multimodal Sentiment Analysis,” IEEE Transaction on Computational Social Systems, Dec. 8, 2022, 8 pages. [cited by applicant]
Jethani, N., et al., “FastSHAP: Real-Time Shapely Value Estimation,” Published as a conference paper at ICLR 2022, 23 pages. [cited by applicant]
Joshi, G., et al., “A Review on Explainability in Multimodal Deep Neural Nets,” IEEE Access, Apr. 26, 2021, 22 pages. [cited by applicant]
Ma, M., et al., “SMIL: Multimodal Learning with Severely Missing Modality,” The Thirty-Fifth AAAI Conference on Artificial Intelligence, (AAAI-21) (www.aaai.org), 2019, 9 pages. [cited by applicant]
Ma, M., et al., “Are Multimodal Transformers Robust to Missing Modality?” IEEE Xplore, 10 pages. [cited by applicant]
Yoon, J., et al., “VIME: Extending the Success of Self- and Semi-supervised Learning to Tabular Domain,” 34 [cited by applicant]