IP Library › Granted Patent US 12,579,800
Granted Patent B2
US 12,579,800 · App. 18/308,767 · Granted Mar 17, 2026

Multi-modal automated evaluation for improved accessibility

Inventors: Iam Palatnik de Sousa (Rio de Janeiro, BR); Shary Beshara (Cairo, EG)
Assignee: Dell Products L.P.
G06V10/82G06F40/117G06V10/764
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,800
App. No.
18/308,767
Granted
Mar 17, 2026
Kind
B2
Abstract

One example method includes a machine-learning (ML) model receiving a first input that includes images that have been extracted from a web page and a second input that includes alt-texts that have been extracted from the web page. The alt-texts describe the images. The ML model converts the images into a first embedding representation and converts the alt-texts into a second embedding representation. Based on the first and second embedding representations, a similarity score between the images and the alt-texts is calculated. The similarity score specifies how accurately each of the alt-texts describe the images. The one of the alt-texts having the highest similarity score is then selected.

Claims (50)

1 . A method, comprising: receiving, at a machine-learning (ML) model, a first input including one or more images that have been extracted from a web page and a second input including one or more alt-texts that have been extracted from the web page, the one or more alt-texts describing the one or more images;

converting, by the ML model the one or more images into a first embedding representation;

converting, by the ML model the one or more alt-texts into a second embedding representation;

calculating, based on the first and second embedding representations, a similarity score between the one or more images and the one or more alt-texts, the similarity score specifying how accurately each of the alt-texts describe the one or more images;

calculating a complexity score for each of the one or more alt-texts based on one or more user defined complexity rules;

combining the complexity score and the similarity score for each of the one or more alt-texts to generate a final score; and

selecting the one of the one or more alt-texts having a highest final score.

2 . The method of claim 1 , wherein selecting the one of the one or more alt-texts having the highest similarity score comprises:

comparing the similarity score of each one or more alt-texts with a predefined threshold; and

selecting the one of the one or more alt-texts having the highest similarity score from the alt-texts having a similarity score higher than the predefined threshold.

3 . The method of claim 2 , further comprising:

ranking the alt-texts having a similarity score higher than the predefined threshold based on their similarity scores; and

selecting the one of the one or more alt-texts having a highest ranking.

4 . The method of claim 2 , further comprising:

providing a notification when none of the one or more alt-texts have a similarity score higher than the predefined threshold.

5 . The method of claim 1 , wherein the ML model is a Contrastive Language Image Pretraining (CLIP) model.

6 . The method of claim 1 , wherein calculating, based on the first and second embedding representations comprises:

combining the first embedding representation and the second embedding representation to generate a muti-modal embedding matrix; and

using the muti-modal embedding matrix in the similarity score calculation.

7 . The method of claim 1 , wherein the similarity score is a cosine similarity score.

8 . The method of claim 1 , wherein selecting the one of the one or more alt-texts having the highest final score comprises:

comparing the final score of each of the one or more alt-texts with a predefined threshold; and

selecting the one of the one or more alt-texts having the highest final score from the alt-texts having a final score higher than the predefined threshold.

9 . The method of claim 8 , further comprising:

providing a notification when none of the one or more alt-texts have a final score higher than the predefined threshold.

10 . A non-transitory storage medium having stored therein instructions that are executable by one or more hardware processors to perform operations comprising: receiving, at a machine-learning (ML) model, a first input including one or more images that have been extracted from a web page and a second input including one or more alt-texts that have been extracted from the web page, the one or more alt-texts describing the one or more images;

converting, by the ML models the one or more images into a first embedding representation;

converting, by the ML model the one or more alt-texts into a second embedding representation;

calculating, based on the first and second embedding representations, a similarity score between the one or more images and the one or more alt-texts, the similarity score specifying how accurately each of the alt-texts describe the one or more images;

calculating a complexity score for each of the one or more alt-texts based on one or more user defined complexity rules;

combining the complexity score and the similarity score for each of the one or more alt-texts to generate a final score; and

selecting the one of the one or more alt-texts having a highest final score.

11 . The non-transitory storage medium of claim 10 , wherein selecting the one of the one or more alt-texts having the highest similarity score comprises:

comparing the similarity score of each one or more alt-texts with a predefined threshold; and

selecting the one of the one or more alt-texts having the highest similarity score from the alt-texts having a similarity score higher than the predefined threshold.

12 . The non-transitory storage medium of claim 11 , further comprising:

ranking the alt-texts having a similarity score higher than the predefined threshold based on their similarity scores; and

selecting the one of the one or more alt-texts having a highest ranking.

13 . The non-transitory storage medium of claim 11 , further comprising:

providing a notification when none of the one or more alt-texts have a similarity score higher than the predefined threshold.

14 . The non-transitory storage medium of claim 10 , wherein the ML model is a Contrastive Language Image Pretraining (CLIP) model.

15 . The non-transitory storage medium of claim 10 , wherein calculating, based on the first and second embedding representations comprises:

combining the first embedding representation and the second embedding representation to generate a muti-modal embedding matrix; and

using the muti-modal embedding matrix in the similarity score calculation.

16 . The non-transitory storage medium of claim 10 , wherein the similarity score is a cosine similarity score.

17 . The non-transitory storage medium of claim 10 , wherein selecting the one of the one or more alt-texts having the highest final score comprises:

comparing the final score of each of the one or more alt-texts with a predefined threshold; and

selecting the one of the one or more alt-texts having the highest final score from the alt-texts having a final score higher than the predefined threshold.

18 . The non-transitory storage medium of claim 17 , further comprising:

providing a notification when none of the one or more alt-texts have a final score higher than the predefined threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2023
From: PALATNIK DE SOUSA, IAM; BESHARA, SHARY
To: DELL PRODUCTS L.P.
Reel/Frame 063475/0486 →
Continuity (1)
Related Publication 20240362903A1 · Oct 31, 2024
References Cited (10)
US 20130024441A1 · Sun · 2013 [cited by examiner]
US 20210073617A1 · Bazzani · 2021 [cited by examiner]
US 20220067506A1 · Li · 2022 [cited by examiner]
US 20230153522A1 · Cho · 2023 [cited by examiner]
US 20230326178A1 · Marri · 2023 [cited by examiner]
US 20240046616A1 · Kim · 2024 [cited by examiner]
US 20240256597A1 · Zhang · 2024 [cited by examiner]
US 20250173613A1 · Oktay · 2025 [cited by examiner]
[1] Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J. and Krueger, G., Jul. 2021. Learning transferable visual models from natural language supervision… [cited by applicant]
https://equalizedigital.com/accessibility-checker/low-quality-alternative-text (accessed Jan. 2023). [cited by applicant]