IP Library › Granted Patent US 12,456,182
Granted Patent B2
US 12,456,182 · App. 18/184,911 · Granted Oct 28, 2025

Anomaly detection using masked auto-encoder

Inventors: Eliyahu Schwartz (Haifa, IL); Leonid Karlinsky (Acton, MA); Sivan Harary (Manof, IL); Assaf Arbelle (Lehvot Haviva, IL)
Assignee: International Business Machines Corporation
G06T7/0002G06N3/0455G06T2207/10024G06T2207/20021G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,456,182
App. No.
18/184,911
Granted
Oct 28, 2025
Kind
B2
Abstract

An example system includes a processor that can randomly mask tokens using different masks to generate different subsets of masked tokens. The processor can process the different sets of masked tokens via a pretrained masked auto-encoder (MAE) encoder to output intermediate representations. The processor can process the intermediate representations via a pretrained MAE decoder to output reconstructed images. The processor can further compare input image with the output reconstructed images to generate an anomaly score.

Claims (34)

1 . A system, comprising a processor to:

randomly mask tokens using different masks to generate different subsets of masked tokens;

process the different subsets of masked tokens via a pretrained masked auto-encoder (MAE) encoder to output intermediate representations;

process the intermediate representations via a pretrained MAE decoder to output reconstructed images; and

compare an input image with the output reconstructed images to generate an anomaly score, wherein generating the anomaly score comprises channel-wise filtering squared error maps with a Gaussian kernel to remove noise.

2 . The system of claim 1 , wherein the processor is to split the input image into non-overlapping patches and flatten the patches into the tokens.

3 . The system of claim 1 , wherein the MAE encoder and the MAE decoder are pretrained on a large collection of general images.

4 . The system of claim 1 , wherein the processor is to fine-tune the pretrained MAE encoder and the pretrained MAE decoder, process the different subsets of masked tokens via the fine-tuned MAE encoder to output the intermediate representations, and process the intermediate representations via the fine-tuned MAE decoder to output the reconstructed images.

5 . The system of claim 1 , wherein generating the anomaly score further comprises summing the squared error maps over three color channels.

6 . The system of claim 1 , wherein generating the anomaly score comprises calculating a mean of a plurality of error maps to generate a single error map.

7 . The system of claim 1 , wherein the anomaly score comprises an image-level anomaly score, and wherein generating the image-level anomaly score comprises calculating a max error of pixel-level anomaly scores of an error map.

8 . The system of claim 1 , wherein the anomaly score comprises a pixel-level anomaly score.

9 . A computer-implemented method, comprising:

randomly masking, via a processor, tokens using different masks to generate different subsets of masked tokens;

processing, via the processor, the different subsets of masked tokens via a pretrained masked auto-encoder (MAE) encoder to output intermediate representations;

processing, via the processor, the intermediate representations via a pretrained MAE decoder to output reconstructed images;

fine-tuning, via the processor, the pretrained MAE encoder and the pretrained MAE decoder, and processing the different subsets of masked tokens via the fine-tuned MAE encoder to output the intermediate representations; and

comparing, via the processor, an input image with the output reconstructed images to generate an anomaly score.

10 . The computer-implemented method of claim 9 , further comprising splitting, via the processor, the input image into non-overlapping patches and flatten the patches into the tokens.

11 . The computer-implemented method of claim 9 , further comprising processing the intermediate representations via the fine-tuned MAE decoder to output the reconstructed images.

12 . The computer-implemented method of claim 11 , wherein fine-tuning the pretrained MAE encoder and the pretrained MAE decoder comprises using an attention mechanism of a transformer to share information between reference tokens and query tokens.

13 . The computer-implemented method of claim 9 , wherein generating the anomaly score comprises channel-wise filtering squared error maps with a Gaussian kernel to remove noise and summing the squared error maps over three color channels.

14 . The computer-implemented method of claim 9 , wherein generating the anomaly score comprises calculating a mean of a plurality of error maps to generate a single error map.

15 . The computer-implemented method of claim 9 , wherein generating the anomaly score comprises calculating a max error of a single error map.

16 . The computer-implemented method of claim 9 , comprising detecting, via the processor, a foreign object in the input image from which the tokens are generated based on the anomaly score.

17 . A computer program product for generating anomaly scores, the computer program product comprising a computer-readable storage medium having program code embodied therewith, the program code executable by a processor to cause the processor to:

randomly mask tokens using different masks to generate different subsets of masked tokens;

process the different subsets of masked tokens via a pretrained masked auto-encoder (MAE) encoder to output intermediate representations;

process the intermediate representations via a pretrained MAE decoder to output reconstructed images;

fine-tune, via the processor, the pretrained MAE encoder and the pretrained MAE decoder by using an attention mechanism of a transformer to share information between reference tokens and query tokens; and

compare an input image with the output reconstructed images to generate an anomaly score.

18 . The computer program product of claim 17 , further comprising program code executable by the processor to split the input image into non-overlapping patches and flatten the patches into the tokens.

19 . The computer program product of claim 17 , further comprising program code executable by the processor to process the different subsets of masked tokens via the fine-tuned MAE encoder to output the intermediate representations, and process the intermediate representations via the fine-tuned MAE decoder to output the reconstructed images.

20 . The computer program product of claim 17 , further comprising program code executable by the processor to generate the anomaly score by channel-wise filtering squared error maps with a Gaussian kernel to remove noise.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 16, 2023
From: SCHWARTZ, ELIYAHU; KARLINSKY, LEONID; HARARY, SIVAN; ARBELLE, ASSAF
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 063002/0915 →
Continuity (1)
Related Publication 20240311987A1 · Sep 19, 2024
References Cited (34)
US 20220164626A1 · Bird · 2022 [cited by examiner]
US 20220335631A1 · Denney · 2022 [cited by examiner]
US 20230206067A1 · Wang · 2023 [cited by examiner]
US 20230410483A1 · Chen · 2023 [cited by examiner]
US 20240096072A1 · He · 2024 [cited by examiner]
US 20240169541A1 · Zhang · 2024 [cited by examiner]
US 20240221128A1 · Chen · 2024 [cited by examiner]
CN 112767331A · 2021 [cited by applicant]
CN 113487571A · 2021 [cited by applicant]
Chaoqin Huang et al. , “Self-Supervised Masking for Unsupervised Anomaly Detection and Localization,” May 13, 2022, Computer Vision and Pattern Recognition (cs.CV), pp. 1-11. [cited by examiner]
Adín Ramírez Rivera et al.,“Anomaly Detection Based on Zero-Shot Outlier Synthesis and Hierarchical Feature Distillation,” Jan. 5, 2022, IEEE Transactions on Neural Networks and Learning Systems, vol. 33, No. 1, Jan. 20… [cited by examiner]
Christoph Baur et al., “Autoencoders for unsupervised anomaly segmentation in brain MR images: A comparative study,” Jan. 2, 2021, Medical Image Analysis 69 (2021) , 101952, pp. 1-10. [cited by examiner]
Yatian Pang et al.,“Masked Autoencoders for Point Cloud Self-supervised Learning,” Nov. 11, 2022, Lecture Notes in Computer Science ((LNCS,vol. 13662)), pp. 604-616. [cited by examiner]
Yu Tian et al.,“Unsupervised Anomaly Detection in Medical Images with a Memory-Augmented Multi-level Cross-Attentional Masked Autoencoder,” 15th 2023, Lecture Notes in Computer Science ((LNCS,vol. 14349)), pp. 1-15. [cited by examiner]
Adín Ramírez Rivera et al., “Anomaly Detection based on Zero-Shot Outlier Synthesis and Hierarchical Feature Distillation”, Published in arxiv.org, Oct. 10, 2020, 16 pages. [cited by applicant]
Bauer, Alexander, “Self-Supervised Training with Autoencoders for Visual Anomaly Detection”, Published in arxiv.org, Jun. 28, 2022, 9 pages. [cited by applicant]
Chaoqin Huang et al. “Self-Supervised Masking for Unsupervised Anomaly Detection and Localization”, Journal of Latex Class Files, Jul. 2021, 15 pages. [cited by applicant]
Christoph Baur et al., “Autoencoders for Unsupervised Anomaly Segmentation in Brain MR Images: A Comparative Study”, Journal of Latex Class Files, vol. 14, No. 8, Aug. 2015, 16 pages. [cited by applicant]
Yatian Pang et al., “Masked Autoencoders for Point Cloud Self-supervised Learning”, arXIV, Mar. 2022, 18 pages. [cited by applicant]
Yu Tian et al., “Unsupervised Anomaly Detection in Medical Images with a Memory-augmented Muti-level Cross-attentional Masked Autoencoder”, Published in arXiv.org, Mar. 22, 2022, 11 pages. [cited by applicant]
Bergmann et al., “MVTec AD—A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection”, International Journal of Computer Vision, Jan. 6, 2021, pp. 1038-1059. [cited by applicant]
Cohen et al., “Sub-Image Anomaly Detection with Deep Pyramid Correspondences”, arXiv:2005.02357v3 [cs.CV], Feb. 3, 2021, 17 pages. [cited by applicant]
Defard et al., “PaDiM: a Patch Distribution Modeling Framework for Anomaly Detection and Localization”, arXiv:2011.08785v1 [cs.CV], Nov. 17, 2020, 7 pages. [cited by applicant]
Geoffrey E. Hinton, “Connectionist learning procedures”, Artificial Intelligence, Sep. 1989, pp. 185-234. [cited by applicant]
Goodfellow et al., “Generative adversarial nets”, arXiv:1406.2661v1 [stat.ML], Jun. 10, 2014, 9 pages. [cited by applicant]
He et al., “Masked autoencoders are scalable vision learners”, arXiv:2111.06377v3 [cs.CV], Dec. 19, 2021, 14 pages. [cited by applicant]
Huang et al., “Attribute restoration framework for anomaly detection”, arXiv:1911.10676v3 [cs.CV], Dec. 12, 2020, 14 pages. [cited by applicant]
Japkowicz et al., “A novelty detection approach to classification”, IJCAI'95: Proceedings of the 14th international joint conference on Artificial intelligence—vol. 1, Aug. 20, 1995, pp. 518-523. [cited by applicant]
Leo Tolstoy, “Anna karenina”, https://www.goodreads.com/book/show/15823480-anna-karenina, Jan. 1, 1878, 550 pages. [cited by applicant]
Roth et al., “Towards total recall in industrial anomaly detection”, arXiv:2106.08265v2 [cs.CV] , May 5, 2022, 18 pages. [cited by applicant]
Sakurada et al., “Anomaly Detection Using Autoencoders with Nonlinear Dimensionality Reduction”, MLSDA'14: Proceedings of the MLSDA 2014 2nd Workshop on Machine Learning for Sensory Data Analysis, Dec. 2, 2014, pp. 4-11. [cited by applicant]
Schlegl et al., “Unsupervised Anomaly Detection with Generative Adversarial Networks to Guide Marker Discovery”, arXiv:1703.05921v1 [cs.CV], Mar. 17, 2017, 12 pages. [cited by applicant]
Yan et al., “Learning semantic context from normal samples for unsupervised anomaly detection”, The Thirty-Fifth AAAI Conference on Artificial Intelligence (AAAI-21), May 2021, pp. 3110-3118. [cited by applicant]
Zavrtanik et al., “Reconstruction by inpainting for visual anomaly detection”, Pattern Recognition, Apr. 2021, 16 pages. [cited by applicant]