IP Library › Granted Patent US 12,456,187
Granted Patent B2
US 12,456,187 · App. 17/364,341 · Granted Oct 28, 2025

Pretraining framework for neural networks

Inventors: Xiaosong Wang (Rockville, MD); Ziyue Xu (Reston, VA); Lickkong Tam (Santa Clara, CA); Dong Yang (Pocatello, ID); Daguang Xu (Potomac, MD)
Assignee: NVIDIA Corporation
G06T7/0012G06N3/048G06T3/4046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,456,187
App. No.
17/364,341
Granted
Oct 28, 2025
Kind
B2
Abstract

Apparatuses, systems, and techniques to indicate an extent, to which text corresponds to one or more images. In at least one embodiment, an extent to which text corresponds to one or more images is indicated using one or more neural networks and used to train the one or more neural networks.

Claims (64)

1. One or more processors, comprising:

circuitry to train one or more neural networks to perform inference on one or more images based, at least in part, on an extent to which training text corresponds to one or more training images, the extent being determined by the one or more neural networks using inputs of paired text and image data and unpaired text and image data.

2. The one or more processors of claim 1 , wherein the circuitry is further to:

calculate features of the training text and the one or more training images;

process the features using at least one or more multiplication and pooling processes; and

use the one or more neural networks to calculate the extent based at least in part on the processed features.

3. The one or more processors of claim 2 , wherein the circuitry is to use at least the extent to train the one or more neural networks to perform one or more classification tasks.

4. The one or more processors of claim 2 , wherein the circuitry is to calculate the extent by at least inputting the processed features to one or more sigmoid functions.

5. The one or more processors of claim 2 , wherein the circuitry is to calculate the features of the training text and the one or more images using one or more encoders.

6. The one or more processors of claim 1 , wherein the unpaired text and image data comprises text that does not describe a first feature in the one or more images.

7. The one or more processors of claim 1 , wherein the one or more images include one or more medical images.

8. A system, comprising:

one or more computers having one or more processors to train one or more neural networks to perform inference on one or more images based, at least in part, on an extent to which training text corresponds to one or more training images, the extent being determined by the one or more neural networks using inputs of paired text and image data and unpaired text and image data.

9. The system of claim 8 , wherein the one or more processors are further to:

obtain a first set of features corresponding to the training text and a second set of features corresponding to the one or more training images; and

use the one or more neural networks to calculate a first value indicating the extent based at least in part on one or more sigmoid functions, the first set of features, and the second set of features.

10. The system of claim 9 , wherein the one or more processors are further to train the one or more neural networks to perform one or more similarity search tasks based at least in part on the first value.

11. The system of claim 9 , wherein the one or more processors are further to:

perform one or more scaling operations to generate a set of scaled images based at least in part on the one or more training images; and

calculate the second set of features based at least in part on the set of scaled images.

12. The system of claim 9 , wherein the one or more sigmoid functions include a logistic function.

13. The system of claim 8 , wherein the paired text and image data comprises text that describes a feature in the one or more images.

14. The system of claim 8 , wherein the training text is a medical text.

15. A method, comprising:

training one or more neural networks to perform inference on one or more images based, at least in part, on an extent to which training text corresponds to one or more training images, the extent being determined by the one or more neural networks using inputs of paired text and image data and unpaired text and image data.

16. The method of claim 15 , further comprising:

obtaining training data comprising a ground truth value, the training text, and the one or more training images;

causing the one or more neural networks to calculate a value indicating the extent to which the training text corresponds to the one or more training images; and

updating the one or more neural networks based at least in part on a difference between the value and the ground truth value.

17. The method of claim 16 , further comprising:

using the one or more neural networks to generate one or more image patches based at least in part on the one or more training images; and

updating the one or more neural networks based at least in part on the one or more image patches and the one or more training images.

18. The method of claim 17 , further comprising training the one or more neural networks for one or more image regeneration tasks based at least in part on the one or more image patches and the one or more training images.

19. The method of claim 16 , further comprising:

generating one or more embeddings based at least in part on the training text and the one or more training images;

generating one or more features based at least in part on the one or more embeddings; and

inputting the one or more features to the one or more neural networks.

20. The method of claim 16 , wherein the ground truth value is a binary value indicating whether the training text corresponds to the one or more training images.

21. The method of claim 15 , wherein the one or more neural networks are part of one or more transformer neural network models.

22. One or more processors, comprising:

circuitry to cause one or more neural networks to perform inference on one or more images based, at least in part, on an extent to which training text corresponds to one or more training images, the extent being determined by the one or more neural networks using inputs of paired text and image data and unpaired text and image data.

23. The one or more processors of claim 22 , wherein the circuitry is further to:

obtain training data; and

train the one or more neural networks using one or more loss functions based at least in part on the extent to which the training text corresponds to the one or more training images and a corresponding ground truth value of the training data.

24. The one or more processors of claim 23 , wherein the circuitry is further to cause the one or more neural networks to output an indication of the extent to which the training text corresponds to the one or more training images.

25. The one or more processors of claim 23 , wherein the one or more loss functions include one or more Binary Cross Entropy (BCE) loss functions.

26. The one or more processors of claim 23 , wherein the training data comprises the unpaired text and image data that comprises text that does not describe a first feature in the one or more images and paired the text and image data that comprises text that does describe a second feature in the one or more images.

27. The one or more processors of claim 22 , wherein the circuitry is further to:

use the one or more neural networks to generate one or more text segments based at least in part on the training text; and

train the one or more neural networks based at least in part on the one or more text segments and the training text.

28. A system, comprising:

one or more computers having one or more processors to cause one or more neural networks to perform inference on one or more images based, at least in part, on an extent to which training text corresponds to one or more training images, the extent being determined by the one or more neural networks using inputs of paired text and image data and unpaired text and image data.

29. The system of claim 28 , wherein the one or more processors are further to:

calculate a set of text features based at least in part on the training text;

calculate a set of image features based at least in part on the one or more training images;

perform one or more average pooling operations based at least in part on the set of text features and the set of image features; and

use the one or more neural networks and results of the one or more average pooling operations to calculate an indication of the extent to which the training text corresponds to the one or more training images.

30. The system of claim 29 , wherein the one or more processors are further to:

use the one or more neural networks and the set of text features to generate one or more text segments;

use the one or more neural networks and the set of image features to generate one or more image patches; and

train the one or more neural networks based at least in part on the one or more text segments and the one or more image patches.

31. The system of claim 30 , wherein the one or more processors are further to train the one or more neural networks for one or more text regeneration tasks based at least in part on the one or more text segments.

32. The system of claim 30 , wherein the one or more processors are further to generate the one or more image patches based at least in part on one or more scaled images generated from the one or more training images.

33. The system of claim 30 , wherein the one or more processors are further to use one or more multi-layer perceptrons (MLPs) to generate the one or more text segments and the one or more image patches.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2022
From: WANG, XIAOSONG; XU, ZIYUE; TAM, LICKKONG; YANG, DONG; XU, DAGUANG
To: NVIDIA CORPORATION
Reel/Frame 060869/0211 →
Continuity (1)
Related Publication 20230019211A1 · Jan 19, 2023
References Cited (82)
US 11734570B1 · Kurz · 2023 [cited by examiner]
US 20180108124A1 · Guo · 2018 [cited by examiner]
US 20200294654A1 · Harzig · 2020 [cited by examiner]
US 20210012150A1 · Liu et al. · 2021 [cited by applicant]
US 20210124979A1 · Andreotti · 2021 [cited by examiner]
US 20210150684A1 · Elmoznino · 2021 [cited by examiner]
US 20210358604A1 · Kearney · 2021 [cited by examiner]
US 20210365736A1 · Kearney · 2021 [cited by examiner]
US 20220051396A1 · Bar · 2022 [cited by examiner]
US 20220058340A1 · Aggarwal et al. · 2022 [cited by applicant]
US 20220180572A1 · Maheshwari · 2022 [cited by examiner]
US 20220261494A1 · Truong · 2022 [cited by examiner]
CN 109935294A · 2019 [cited by applicant]
CN 110647632A · 2020 [cited by applicant]
CN 113035311A · 2021 [cited by applicant]
CN 109783655B · 2022 [cited by examiner]
EP 3832594A1 · 2021 [cited by applicant]
JP 2019139713A · 2019 [cited by applicant]
JP 2020524355A · 2020 [cited by applicant]
JP 2020149682A · 2020 [cited by applicant]
WO 2021050776A1 · 2021 [cited by applicant]
Office Action for United Kingdom Application No. GB2209080.7, mailed Nov. 28, 2023, 3 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2209080.7, mailed Mar. 7, 2025, 3 pages. [cited by applicant]
Wang et al., “Self-supervised Image-text Pre-training With Mixed Data in Chest X-rays,” 2021, 10 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2209080.7, mailed May 21, 2024, 3 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2209080.7, mailed Dec. 3, 2024, 4 pages. [cited by applicant]
Avcibas et al., “Statistical Evaluation of Image Quality Measures,” Journal of Electronic Imaging, 11(2): 2002, 18 pages. [cited by applicant]
Brisimi et al., “Federated Learning of Predictive Models from Federated Electronic Health Records,” IJMI, Apr. 2018, 22 pages. [cited by applicant]
Bustos et al., “PadChest: A Large Chest X-ray Image Dataset with Multi-Label Annotated Reports,” Medical Image Analysis, 66:101797, 2020, 35 pages. [cited by applicant]
Cao et al., “Deep Cauchy Hashing for Hamming Space Retrieval,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 9 pages. [cited by applicant]
Chen et al., Generative pretraining from pixels. In International Conference on Machine Learning, PMLR, 2020, 13 pages. [cited by applicant]
Chen et al., “UNITER: Universal Image-TExt Representation Learning,” European Conference on Computer Vision, 2020, 17 pages. [cited by applicant]
Chen et al., “VisualGPT: Data-efficient Image Captioning by Balancing Visual Input and Linguistic Knowledge from Pretraining,” Apr. 17, 2021, 18 pages. [cited by applicant]
Cho et al., “Unifying Vision- and-Language Tasks via Text Generation,” May 23, 2021, 15 pages. [cited by applicant]
Demner-Fushman et al., “Preparing a Collection of Radiology Examinations for Distribution and Retrieval,” Journal of the American Medical Informatics Association, 23(2), 2015, 7 pages. [cited by applicant]
Deng et al., “Imagenet: A Large-Scale Hierarchical Image Database,” CVPR, 2009, 8 pages. [cited by applicant]
Desai et al., “VirTex: Learning Visual Representations from Textual Annotations,” Jun. 11, 2020, 27 pages. [cited by applicant]
Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” Oct. 11, 2018, 14 pages. [cited by applicant]
Dosovitskiy et al., “An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale,” Oct. 22, 2020, 21 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition,” CVPR, 2016, 9 pages. [cited by applicant]
He et al., “Mask R-CNN,” ICCV, 2017, 9 pages. [cited by applicant]
He et al., “Rethinking ImageNet Pre-training,” Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, 10 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Irvin et al., “CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, 2019, 8 pages. [cited by applicant]
Johnson et al., “MIMIC-CXR, A De-identified Publicly Available Database of Chest Radiographs with Free-text Reports,” Scientific Data, 6(1), 2019, 8 pages. [cited by applicant]
Kolesnikov et al., “Big Transfer (BiT): General Visual Representation Learning,” Dec. 24, 2019, 23 pages. [cited by applicant]
Konecny et al., “Federated Learning: Strategies for Improving Communication Efficiency,” Oct. 18, 2016, 5 pages. [cited by applicant]
Krishna et al., “Visual Genome: Connecting Language and Vision using Crowdsourced Dense Image Annotations,” Springer, 2016, 42 pages. [cited by applicant]
Lee et al., “Biobert: A Pre-trained Biomedical Language Representation Model for Biomedical Text Mining,” Bioinformatics, 36(4), 2020, 7 pages. [cited by applicant]
Li et al., “VisualBERT: A Simple and Performant Baseline for Vision and Language,” Aug. 9, 2019, 14 pages. [cited by applicant]
Lin et al., “Interbert: Vision-and-Language Interaction for Multi-Modal Pretraining,” Jun. 10, 2020, 15 pages. [cited by applicant]
Lin et al., “Microsoft COCO: Common Objects in Context,” European Conference on Computer Vision, Jul. 5, 2014, 14 pages. [cited by applicant]
Lu et al., “ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision- and-Language Tasks,” Neural Information Processing Systems, 2019, 11 pages. [cited by applicant]
Mikolov et al., “Efficient Estimation of Word Representations in Vector Space,” Sep. 7, 2013, 12 pages. [cited by applicant]
Mikolov et al., “Recurrent Neural Network Based Language Model,” in Eleventh annual Conference of the International Speech Communication Association, 2010, 4 pages. [cited by applicant]
Pennington et al., “GloVe: Global Vectors for Word Representation,” vol. 14, Jan. 2014, 12 pages. [cited by applicant]
Radford et al., “Improving Language Understanding by Generative Pre-Training,” 2018, 12 pages. [cited by applicant]
Radford et al., “Language Models are Unsupervised Multitask Learners,” OpenAI Blog 1(8), 2019, 24 pages. [cited by applicant]
Ren et al., “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” Advances in Neural Information Processing Systems, 2015, 9 pages. [cited by applicant]
Sariyildiz et al., “Learning Visual Representations with Caption Annotations,” Aug. 4, 2020, 28 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
Su et al., “VL-BERT: Pre-training of Generic Visuallinguistic Representations,” Nov. 22, 2019, 15 pages. [cited by applicant]
Sun et al., “Revisiting Unreasonable Effectiveness of Data in Deep Learning Era,” Jul. 2017, 10 pages. [cited by applicant]
Tan et al., “Lxmert: Learning Crossmodality Encoder Representations from Transformers,” Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 2019, 12 pages. [cited by applicant]
Vaswani et al., “Attention is All You Need,” Dec. 6, 2017, 15 pages. [cited by applicant]
Wang et al., “ChestX-ray8: Hospital-Scale Chest X-Ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases,” Proceedings of the IEEE Conference on Computer Vision and Pa… [cited by applicant]
Wang et al., “TieNet: Text-Image Embedding Network for Common Thorax Disease Classification and Reporting in Chest X-Rays,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 10 pages. [cited by applicant]
Wei et al., “Multi-Modality Cross Attention Network for Image and Sentence Matching,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, 10 pages. [cited by applicant]
Xu et al., “Show, Attend and Tell: Neural Image Caption Generation with Visual Attention,” ICML, 2015, 10 pages. [cited by applicant]
Yang et al., “Federated Semisupervised Learning for Covid Region Segmentation in Chest CT using Multi-National Data from China, Italy, Japan,” Medical image analysis, 2020, 19 pages. [cited by applicant]
Yang et al., “XLNet: Generalized Autoregressive Pretraining for Language Understanding,” Advances in Neural Information Processing Systems, 2019, 11 pages. [cited by applicant]
Zhang et al., “Contrastive Learning of Medical Visual Representations from Paired Images and Text,” 2020, 15 pages. [cited by applicant]
Zhou et al., “Learning Deep Features for Scene Recognition using Places Database,” Neural Information Processing Systems, 2014, 9 pages. [cited by applicant]
Endo et al., “Prediction of Regions in Images Corresponding to Text Phrases Using Image Attention with Tree LSTM,” IEICE Technical Report, 118(259): Oct. 2018, 8 pages. [cited by applicant]
Hotta et al., “Object Identification by Combining Neural Network and Logistic Function,” IEICE Technical Report, 101(713): Mar. 8, 2002, 12 pages. [cited by applicant]
Kato et al., “Prediction of Thermal Displacement of Machine Tool by Neural Network-Analysis of Valid Temperature Data-,” IEICE Technical Report, 101(534): Dec. 14, 2001, 10 pages. [cited by applicant]
Kotani et al., “On Data Compression Capability of Acoustic Signals by Neural Networks,” Transactions of the Society of Instrument and Control Engineers, 28(7): Jul. 31, 1992, 12 pages. [cited by applicant]
Office Action for Japanese Application No. 2022-088126, mailed Aug. 8, 2024, 9 pages. [cited by applicant]
Sasaki et al., “Fundamental Study on Pattern Recognition and Action Learning by Convolutional Neural Networks,” IEEJ Workshop Materials, Dec. 6, 2015, 7 pages. [cited by applicant]
Wen et al., “Dual Semantic Relationship Attention Network for Image-Text Matching,” International Joint Conference on Neural Networks, Jul. 19, 2020, 8 pages. [cited by applicant]
United Kingdom Combined Search and Examination Report for Application No. GB2209080.7, mailed Nov. 22, 2022, 6 pages. [cited by applicant]
Cited By (1)
US 12,688,057