IP Library Granted Patent US 12,367,261
Granted Patent B1
US 12,367,261 · App. 17/324,653 · Granted Jul 22, 2025

Extending supervision using machine learning

Inventors: Lickkong Tam (Santa Clara, CA); Ali Hatamizadeh (Los Angeles, CA); Kevin Frank Lu (Washington, DC); Yuhong Wen (McLean, VA); Daguang Xu (Potomac, MD); Riddhish Bhalodia (Salt Lake City, UT); Wenqi Li (London, GB); Can Zhao (Rockville, MD); Xiaosong Wang (Rockville, MD)
Assignee: NVIDIA Corporation
G06F18/2155G06F40/169G06N3/045G06N3/08G06V10/225G06V30/1448G06V2201/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,261
App. No.
17/324,653
Granted
Jul 22, 2025
Kind
B1
Abstract

Apparatuses, systems, and techniques to generate labeled training data. In at least one embodiment, labeled training images are generated from medial images annotated with natural language text.

Claims (44)

1. A processor, comprising:

one or more circuits to:

parse one or more natural language annotations corresponding to one or more non-fully labeled images to extract localizing language corresponding to respective locations of one or more features depicted in the one or more non-fully labeled images; and

use one or more neural networks to generate one or more fully labeled images to include one or more labels indicating respective locations of the one or more features depicted in the one or more non-fully labeled images, based, at least in part, on:

the extracted localizing language, and

the one or more non-fully labeled images.

2. The processor of claim 1 , wherein each image of the fully labeled images includes a natural language annotation and an associated indication of a location on the image.

3. The processor of claim 2 , wherein the natural language annotation includes a first portion that identifies a disease and a second portion that identifies a location.

4. The processor of claim 1 , wherein each image of the non-fully labeled images includes a natural language annotation.

5. The processor of claim 1 , wherein the one or more neural networks is trained using fully labeled training images that include natural language annotations and indications of locations on the fully labeled training images.

6. The processor of claim 5 , wherein the indications of locations are bounding boxes.

7. The processor of claim 1 , wherein the one or more fully labeled images are used to train one or more additional neural networks to recognize, in an image, a characteristic described in natural language annotations of the non-fully labeled images.

8. The processor of claim 7 , wherein:

the non-fully labeled images and the fully labeled images are medical images; and

the characteristic is a medical condition.

9. A computer system, comprising:

one or more processors and memory storing executable instructions that, as a result of being executed by the one or more processors, cause the computer system to:

parse one or more natural language annotations corresponding to one or more non-fully labeled images to extract localizing language corresponding to respective locations of one or more features depicted in the one or more non-fully labeled images; and

use one or more neural networks to generate one or more fully labeled images to include one or more labels indicating respective locations of the one or more features depicted in the one or more non-fully labeled images, based, at least in part, on:

the extracted localizing language, and

the one or more non-fully labeled images.

10. The computer system of claim 9 , wherein each image of the fully labeled images includes a natural language annotation and an associated indication of a location on the image.

11. The computer system of claim 10 , wherein the natural language annotation includes a first portion that identifies a disease and a second portion that identifies a location.

12. The computer system of claim 9 , wherein each image of the non-fully labeled images includes an annotation.

13. The computer system of claim 9 , wherein the one or more neural networks is trained using fully labeled training images that include natural language annotations and indications of locations on the fully labeled training images.

14. The computer system of claim 13 , wherein the indications of locations are geometric shapes that encompass a region of an image of the fully labeled training images.

15. The computer system of claim 9 , wherein the one or more fully labeled images are used to train one or more additional neural networks to recognize, in an image, a characteristic described in natural language annotations of the non-fully labeled images.

16. The computer system of claim 15 , wherein:

the non-fully labeled images and the fully labeled images are medical images; and

the characteristic is a medical condition.

17. A computer-implemented method, comprising:

parsing one or more natural language annotations corresponding to one or more non-fully labeled images to extract localizing language corresponding to respective locations of one or more features depicted in the one or more non-fully labeled images; and

using one or more neural networks to generate one or more fully labeled images to include one or more labels indicating respective locations of the one or more features depicted in the one or more non-fully labeled images, based, at least in part, on:

the extracted localizing language, and

the one or more non-fully labeled images.

18. The computer-implemented method of claim 17 , wherein each image of the fully labeled images includes a natural language annotation and an associated indication of a location on the image.

19. The computer-implemented method of claim 18 , wherein the natural language annotation includes a first portion that identifies a disease and a second portion that identifies a location.

20. The computer-implemented method of claim 17 , wherein each image of the non-fully labeled images includes a natural language annotation.

21. The computer-implemented method of claim 17 , wherein the one or more neural networks is trained using fully labeled training images that include natural language annotations and indications of locations on the fully labeled training images.

22. The computer-implemented method of claim 21 , wherein the indications of locations are bounding boxes.

23. The computer-implemented method of claim 17 , wherein the one or more fully labeled images are used to train one or more additional neural networks to recognize, in an image, a characteristic described in natural language annotations of the non-fully labeled images.

24. The computer-implemented method of claim 23 , wherein:

the non-fully labeled images and the fully labeled images are medical images; and

the characteristic is a medical condition.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2021
From: TAM, LICKKONG; HATAMIZADEH, ALI; LU, KEVIN FRANK; WEN, YUHONG; XU, DAGUANG; BHALODIA, RIDDHISH; LI, WENQI; ZHAO, CAN; WANG, XIAOSONG
To: NVIDIA CORPORATION
Reel/Frame 056866/0629 →
References Cited (76)
US 10650929B1 · Beck · 2020 [cited by examiner]
US 10922793B2 · Baek · 2021 [cited by examiner]
US 11322256B2 · Sati · 2022 [cited by examiner]
US 20190188848A1 · Madani · 2019 [cited by examiner]
US 20190213443A1 · Cunningham · 2019 [cited by examiner]
US 20190259493A1 · Xu et al. · 2019 [cited by applicant]
US 20190286880A1 · Jackson et al. · 2019 [cited by applicant]
US 20190317986A1 · Kobayashi · 2019 [cited by applicant]
US 20200093455A1 · Wang · 2020 [cited by applicant]
US 20200286405A1 · Buras et al. · 2020 [cited by applicant]
US 20200372301A1 · Kearney et al. · 2020 [cited by applicant]
US 20200388033A1 · Matlock et al. · 2020 [cited by applicant]
US 20210035015A1 · Edgar · 2021 [cited by examiner]
US 20210059758A1 · Avendi et al. · 2021 [cited by applicant]
US 20210065859A1 · Mckinney · 2021 [cited by examiner]
US 20210090250A1 · Soans et al. · 2021 [cited by applicant]
US 20210303931A1 · Wang · 2021 [cited by examiner]
CA 3110581A1 · 2021 [cited by applicant]
JP 2019185551 · 2019 [cited by applicant]
WO 2018009405A1 · 2018 [cited by applicant]
WO 2018101985A1 · 2018 [cited by applicant]
WO 2019084697A1 · 2019 [cited by applicant]
WO 2021050976A1 · 2021 [cited by applicant]
Luo, et al. (Computer English Translation of Chinese Patent No. CN111738454 B), pp. 1-17. (Year: 2020). [cited by examiner]
United Kingdom Combined Search and Examination Report for Application No. GB2207326.6, mailed Nov. 22, 2022, 6 pages. [cited by applicant]
Bergstra et al., “Algorithms for Hyper-Parameter Optimization,” Neural Information Processing Sysytems, 2011, 9 pages. [cited by applicant]
Brooks, “Coco Annotator,” retrieved from, github.com/jsbroks/coco-annotator/, 2019, 6 pages. [cited by applicant]
Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” Annual Conference of the North American Chapter of the Association for Computational Linguistics, 2019, 16 pages. [cited by applicant]
Ghesu et al., “Quantifying and Leveraging Classification Uncertainty for Chest Radiograph Assessment,” vol. 11769, 2019, 9 pages. [cited by applicant]
Gupta et al., “Contrastive Learning for Weakly Supervised Phrase Grounding,” Aug. 5, 2020, 18 pages. [cited by applicant]
Gutmann et al., “Noise-contrastive: A New Estimation Principle for Unormalized Statistical Models,” vol. 9, 2010, 8 pages. [cited by applicant]
Huang et al., “ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission,” 2019, 19 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Irvin et al., “CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, 2019, 8 pages. [cited by applicant]
Jia et al., “Scaling up Visual and Vision-Language Representation Learning with Noisy Text Supervision,” 2021, 11 pages. [cited by applicant]
Jocher et al., “Ultralytics/yolov5: v4.0—nn.SiLU() activations, Weights & Biases logging, PyTorch Hub Integration,” retrieved from https://github.com/ultralytics/yolov5/tree/v4.0, 2021, 8 pages. [cited by applicant]
Johnson et al., “CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning,” CVPR, 2017, 10 pages. [cited by applicant]
Johnson et al., “MIMIC-CXR, A De-identified Publicly Available Database of Chest Radiographs with Free-text Reports,” Scientific Data, 6(1), 2019, 8 pages. [cited by applicant]
Kashyap et al., “Looking in the Right Place for Anomalies: Explainable Al through Automatic Location Learning,” SBI, 2020, 6 pages. [cited by applicant]
Kazemzadeh et al., “ReferltGame: Referring to Objects in Photographs of Natural Scenes,” 2014, 12 games. [cited by applicant]
Kendall et al., What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision? In Advances in Neural Information Processing Systems, Oct. 5, 2017, 12 pages. [cited by applicant]
Kingma et al., “Adam: A Method for Stochastic Optimization,” ICLR, 2015, 15 pages. [cited by applicant]
Krizhevsky et al., “ImageNet Classification with Deep Convolutional Neural Networks,” In Advances in Neural Information Processing Systems, 2012, 9 pages. [cited by applicant]
Lecun et al., “Gradient-Based Learning Applied to Document Recognition,” Proceedings of the IEEE, 86(11): 1998, 47 pages. [cited by applicant]
Lee et al., “Biobert: A Pre-trained Biomedical Language Representation Model for Biomedical Text Mining,” Bioinformatics, 36(4), 2020, 7 pages. [cited by applicant]
Li et al., “Thoracic Disease Identification and Localization with Limited Supervision,” CVPR, 2019, 10 pages. [cited by applicant]
Lin et al., “Microsoft COCO: Common Objects in Context,” European Conference on Computer Vision, Jul. 5, 2014, 14 pages. [cited by applicant]
Loper et al., “NLTK: The Natural Language Toolkit,” 2002, 8 pages. [cited by applicant]
Lu et al., “ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks,” Neural Information Processing Systems, 2019, 11 pages. [cited by applicant]
Manning et al., “The Stanford CoreNLP Natural Language Processing Toolkit,” Association for Computational Linguistics, 2014, 6 pages. [cited by applicant]
Moradi et al., “Bimodal Network Architectures for Automatic Generation of Image Annotation from Text,” MICCAI, 2018, 8 pages. [cited by applicant]
Nguyen et al., “VinDr-CXR: An Open Dataset of Chest X-rays with Radiologist's Annotations,” 2020, 10 pages. [cited by applicant]
Pan et al., “Tackling the Radiological Society of North America Pneumonia Detection Challenge,” Am J Roentgenol 213(3): Sep. 2019, 7 pages. [cited by applicant]
Perez et al., “FiLM: Visual Reasoning with a General Conditioning Layer,” AAAI Conference on Artificial Intelligence, 2018, 10 pages. [cited by applicant]
Perez, “Retrospective for FiLM: Visual Reasoning with a General Conditioning Layer,” retrieved from https://ml-retrospectives.github.io/neurips2019/accepted_retrospectives/2019/film/, Nov. 27, 2019, 6 pages. [cited by applicant]
Radford et al., “Learning Transferable Visual Models from Natural Language Supervision,” Feb. 26, 2021, 48 pages. [cited by applicant]
Rajpurkar et al., “CheXNet: Radiologist-Level Pneumonia Detection on Chest X-Rays with Deep Learning,” Dec. 25, 2017, 7 pages. [cited by applicant]
Redmon et al., “YOLOv3: An Incremental Improvement”, Apr. 8, 2018, 6 pages. [cited by applicant]
Redmon et al., “You Only Look Once: Unified, Real-Time Object Detection, ” 2015, pp. 1-9. [cited by applicant]
Shin et al., “BioMegatron: Larger Biomedical Domain Language Model,” Empirical Methods in Natural Language Processing, 2020, 7 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
Tam et al., “Transformer Query-Target Knowledge Discovery (TEND): Drug Discovery from CORD-19,” 2020, 10 pages. [cited by applicant]
Tam et al., “Weakly Supervised One-Stage Vision and Language Disease Detection using Large Scale Pneumonia and Pneumothorax Studies,” MICCAI. vol. 12264, 2020, 11 pages. [cited by applicant]
Vaswani et al., “Attention is All You Need,” Dec. 6, 2017, 15 pages. [cited by applicant]
Wang et al., “ChestX-ray8: Hospital-Scale Chest X-Ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases,” Proceedings of the IEEE Conference on Computer Vision and Pa… [cited by applicant]
Wang et al., “Frustratingly Simple Few-Shot Object Detection,” Mar. 16, 2020, 12 pages. [cited by applicant]
Wang et al., “Score-CAM: Score-Weighted Visual Explanations for Convolutional Neural Networks,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020, 9 pages. [cited by applicant]
Yan et al., “Learning from Multiple Datasets with Heterogeneous and Partial Labels for Universal Lesion Detection in CT,” 2020, 12 pages. [cited by applicant]
Yang et al., “A Fast and Accurate One-Stage Approach to Visual Grounding,” ICCV, 2019, 11 pages. [cited by applicant]
Yang et al., “Improving One-Stage Visual Grounding by Recursive Sub-Query Construction,” ECCV, vol. 12359, 2020, 21 pages. [cited by applicant]
Zhang et al., “Contrastive Learning of Medical Visual Representations from Paired Images and Text,” 2020, 15 pages. [cited by applicant]
Zhang et al., “Stanza : A Python Natural Language Processing Toolkit for Many Human Languages,” 2020, 8 pages. [cited by applicant]
Notice of Preliminary Rejection from Korean Patent Application No. 10-2022-0058189, dated Sep. 27, 2024, pp. 1-10 (includes English translation). [cited by applicant]
Office Action from Japanese Patent Application No. 2022-072977, mailed Aug. 27, 2024, pp. 1-6 (English translation included). [cited by applicant]
Kelvin Xu, et al., “Show, Attend and tell: Neural Image Caption Generation with Visual Attention,” arXiv, Apr. 19, 2026, pp. 1022, https://arxiv.org/pdf/1502.03044v3. [cited by applicant]
Cited By (2)
US 12,587,903 US 12,639,965