IP Library › Granted Patent US 12,505,905
Granted Patent B2
US 12,505,905 · App. 18/952,233 · Granted Dec 23, 2025

Method and system for the computer-aided processing of medical images

Inventors: Jeffrey Chang (Berkeley, CA); Doktor Gurson (Berkeley, CA); Joseph Zachary Allen (Berkeley, CA); Christopher Johnson (Berkeley, CA)
Assignee: RAD AI, Inc.
G16H15/00G06F40/40G06T7/0012G06V10/77G06T2207/10116G06V2201/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,905
App. No.
18/952,233
Granted
Dec 23, 2025
Kind
B2
Abstract

Methods and systems described enable automatic generation of a significant portion of or all of a clinical report (e.g., radiology report), using multimodal models trained on image and language data. Methods described can transform unstructured language and image information into findings, as well as an accurate and comprehensive clinical report, in a designated style (e.g., writing style). The methods and systems described thus significantly improve performance in generation and processing of clinical reports, in relation to time saved per clinical shift, dictation effort, medical billing, and other performance factors.

Claims (33)

1. A method of using a multimodal model for generation of radiology outputs, the method comprising:

at a computing system, receiving a request from a radiologist to retrieve a report associated with a session with a patient; and

returning the report with a level of completion above a threshold level of completion, wherein the report is returned in a writing style of the radiologist, wherein returning the report comprises:

training the multimodal model, comprising a language-aligned vision encoder and an adapter between the language-aligned vision encoder and a fixed large language model (LLM), wherein the adapter is structured to pass information between the language-aligned image encoder and the LLM with an attention mechanism, to generate a trained multimodal model;

re-training the multimodal model to generate a re-trained multimodal model, wherein re-training the multimodal model comprises training only the adapter and fixing all other components of the multimodal model;

transforming a set of images generated during the session into a set of image representations;

returning a set of radiology outputs upon processing the set of image representations with the re-trained multimodal model, wherein the multimodal model is structured to process image-based inputs and language-based inputs;

transforming the set of radiology outputs into the report; and

transmitting the report to the radiologist.

2. The method of claim 1 , wherein training the multimodal model comprises training the multimodal model in a first stage of training and in a second stage of training.

3. The method of claim 2 , wherein the first stage of training involves a contrastive learning with language-image pre-training operation, with a contrastive loss function, structured to drive image representations and associated reports closer in a high-dimension space, and structured to drive mismatched image representations and unrelated reports, to generate a first trained representation of the multimodal model.

4. The method of claim 3 , wherein the second stage of training involves training the first trained representation of the multimodal model with bootstrapping language-image pre-training architecture applied to a Q-former comprising a shared embedding space.

5. The method of claim 1 , wherein returning the set of radiology outputs comprises detecting an anomaly captured in the set of images.

6. The method of claim 5 , further comprising determining that the anomaly is associated with a clinical indication, wherein the anomaly comprises one of a global anomaly and a local anomaly.

7. The method of claim 6 , further comprising retrieving a set of candidate actions to perform based upon the clinical indication.

8. The method of claim 7 , wherein the clinical indication comprises a hemorrhage, and wherein detecting the anomaly is performed with at least 70% sensitivity and 80% specificity.

9. The method of claim 7 , further comprising executing an action of the set of candidate actions, wherein the action comprises executing a critical results workflow with administration of care for treating the clinical indication.

10. The method of claim 1 , wherein the writing style comprises a word choice component and a length of an impressions section of the report.

11. The method of claim 1 , wherein receiving the request and returning the report is performed within a duration of 5 minutes, and wherein the threshold level of completion is 80%.

12. A method of using a multimodal model for generation of radiology outputs, the method comprising:

training the multimodal model, comprising a language-aligned vision encoder and an adapter between the language-aligned vision encoder and a fixed large language model (LLM), wherein the adapter is structured to pass information between the language-aligned image encoder and the LLM with an attention mechanism, to generate a trained multimodal model;

re-training the multimodal model to generate a re-trained multimodal model, wherein re-training the multimodal model comprises training only the adapter and fixing all other components of the multimodal model;

transforming a set of images generated during a session with a patient, into a set of image representations;

returning a set of radiology outputs upon processing the set of image representations with the re-trained multimodal model, wherein the multimodal model is structured to process image-based inputs and language-based inputs;

detecting an anomaly captured in the set of images, using the trained multimodal model;

determining that the anomaly is associated with a clinical indication;

retrieving a set of candidate actions to perform based upon the clinical indication;

transforming the set of radiology outputs into a report comprising a style and level of completion above a threshold level of completion;

transmitting the report to the radiologist; and

executing an action of the set of candidate actions, wherein the action comprises administering care according to a critical results workflow corresponding to the clinical indication.

13. The method of claim 12 , wherein training the multimodal model involves a contrastive learning with language-image pre-training operation, with a contrastive loss function, to generate a first trained representation of the multimodal model.

14. The method of claim 12 , wherein the style comprises a word choice component and a length of an impressions section of the report.

15. The method of claim 12 , wherein transmitting the report in response to a request by a radiologist is performed within a duration of 2 minutes, and wherein the threshold level of completion is 90%.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2025
From: CHANG, JEFFREY; GURSON, DOKTOR; ALLEN, JOSEPH ZACHARY; JOHNSON, CHRISTOPHER
To: RAD AI, INC.
Reel/Frame 070423/0805 →
Continuity (2)
Provisional Application 63602098 · Nov 22, 2023
Related Publication 20250166764A1 · May 22, 2025
References Cited (108)
US 10127662B1 · Reicher et al. · 2018 [cited by applicant]
US 10140421B1 · Bernard et al. · 2018 [cited by applicant]
US 10395772B1 · Lucas et al. · 2019 [cited by applicant]
US 10667794B2 · Beymer et al. · 2020 [cited by applicant]
US 11064902B2 · Wallace et al. · 2021 [cited by applicant]
US 11069432B2 · Guo et al. · 2021 [cited by applicant]
US 11164045B2 · Paik et al. · 2021 [cited by applicant]
US 11263749B1 · Purushottam et al. · 2022 [cited by applicant]
US 11342055B2 · Chang et al. · 2022 [cited by applicant]
US 11380432B2 · Glottmann · 2022 [cited by examiner]
US 11403786B2 · Farri et al. · 2022 [cited by applicant]
US 11475358B2 · Edgar et al. · 2022 [cited by applicant]
US 11610667B2 · Gurson et al. · 2023 [cited by applicant]
US 11615890B2 · Chang et al. · 2023 [cited by applicant]
US 11803759B2 · Hu et al. · 2023 [cited by applicant]
US 11961624B2 · Smurro · 2024 [cited by applicant]
US 12032658B2 · Kiraly et al. · 2024 [cited by applicant]
US 20030154085A1 · Kelley · 2003 [cited by applicant]
US 20070078679A1 · Rose · 2007 [cited by applicant]
US 20100138239A1 · Reicher et al. · 2010 [cited by applicant]
US 20130290031A1 · Kay et al. · 2013 [cited by applicant]
US 20140379378A1 · Cohen-Solal et al. · 2014 [cited by applicant]
US 20150112725A1 · Ryan · 2015 [cited by applicant]
US 20150348229A1 · Aguirre-Valencia et al. · 2015 [cited by applicant]
US 20160364862A1 · Reicher et al. · 2016 [cited by applicant]
US 20180060533A1 · Reicher et al. · 2018 [cited by applicant]
US 20180322254A1 · Smurro · 2018 [cited by applicant]
US 20180330828A1 · Hayter · 2018 [cited by applicant]
US 20190021677A1 · Grbic et al. · 2019 [cited by applicant]
US 20190122073A1 · Ozdemir et al. · 2019 [cited by applicant]
US 20190139218A1 · Song et al. · 2019 [cited by applicant]
US 20190156947A1 · Nakamura et al. · 2019 [cited by applicant]
US 20190231288A1 · Profio et al. · 2019 [cited by applicant]
US 20190362835A1 · Sreenivasan et al. · 2019 [cited by applicant]
US 20200226481A1 · Sim et al. · 2020 [cited by applicant]
US 20200311613A1 · Ma et al. · 2020 [cited by applicant]
US 20200334416A1 · Vianu et al. · 2020 [cited by applicant]
US 20200342967A1 · Bronkalla et al. · 2020 [cited by applicant]
US 20200380675A1 · Golden et al. · 2020 [cited by applicant]
US 20210005297A1 · Oez · 2021 [cited by applicant]
US 20210035015A1 · Edgar et al. · 2021 [cited by applicant]
US 20210090694A1 · Colley et al. · 2021 [cited by applicant]
US 20210110912A1 · Mukherjee · 2021 [cited by applicant]
US 20210216822A1 · Paik et al. · 2021 [cited by applicant]
US 20210327596A1 · Tahmasebi Maraghoosh et al. · 2021 [cited by applicant]
US 20210334462A1 · Kukreja et al. · 2021 [cited by applicant]
US 20220020495A1 · Sadeghi · 2022 [cited by applicant]
US 20220051771A1 · Lyman et al. · 2022 [cited by applicant]
US 20220059200A1 · Rahbar et al. · 2022 [cited by applicant]
US 20220246258A1 · Chang et al. · 2022 [cited by applicant]
US 20220254464A1 · Pinto · 2022 [cited by applicant]
US 20230051982A1 · Sasidharan et al. · 2023 [cited by applicant]
US 20230079929A1 · Bradski et al. · 2023 [cited by applicant]
US 20230145535A1 · Hatamizadeh et al. · 2023 [cited by applicant]
US 20230197276A1 · Chang · 2023 [cited by examiner]
US 20230289529A1 · Alikaniotis et al. · 2023 [cited by applicant]
US 20240266074A1 · Smurro · 2024 [cited by applicant]
US 20240347156A1 · Paulett · 2024 [cited by examiner]
CN 111414464A · 2020 [cited by applicant]
CN 116364227A · 2023 [cited by examiner]
CN 117558439A · 2024 [cited by applicant]
CN 117995344A · 2024 [cited by applicant]
CN 118070227A · 2024 [cited by applicant]
EP 3246836A1 · 2017 [cited by applicant]
WO 0239415A2 · 2002 [cited by applicant]
WO 2019025601A1 · 2019 [cited by applicant]
WO 2020043673A1 · 2020 [cited by applicant]
WO 2020214683A1 · 2020 [cited by applicant]
WO 2022162167A1 · 2022 [cited by applicant]
WO 2022212771A2 · 2022 [cited by applicant]
Li, J., Li, D., Savarese, S., & Hoi, S. (Jul. 2023). Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning (pp. 19730-197… [cited by examiner]
Machine translation of CN 116364227 obtained from google patents (Year: 2023). [cited by examiner]
“Improving and automating lung cancer screening”, Nuance, Data Sheet, Nov. 2020, https://www.nuance.com/asset/en_us/collateral/healthcare/data-sheet/ds-powerscribe-lung-cancer-screening-en-us.pdf. [cited by applicant]
“Lung Cancer Orchestrator”, Philips, first downloaded May 19, 2023, https://www.usa.philips.com/healthcare/product/HC841017/lung-cancer-orchestrator. [cited by applicant]
“Lung Cancer Screening”, Eon Health, first downloaded May 19, 2023, https://eonhealth.com/lung-cancer-screening-software/. [cited by applicant]
“Merge Announces Release of Fusion PACS MX(TM) 3.0 Featuring Integrated Digital Mammography, and Release of Merge Mammo(TM) 7.10”, Business Wire, May 19, 2008, p. NA. ProQuest. Web. Aug. 29, 2023. (Year: 2008). [cited by applicant]
“PowerScribe One for radiology reporting, Next-generation radiology reporting”, Nuance, https://www.nuance.com/healthcare/diagnostics-solutions/workflow-radiology-reporting/powerscribe-one.html#, first downloaded Mar. 1… [cited by applicant]
“Radiology reports designed for patients”, Scanslated, first downloaded May 19, 2023, https://scanslated.com. [cited by applicant]
Agarwal, Sheela, et al., “Beyond the impression: How AI-driven clinical intelligence transforms the radiology experience”, Radiology Business Journal sponsored webinar, Nov. 16, 2022. [cited by applicant]
Chang, Jeffrey, et al., “Method and System for the Computer-Aided Processing of Medical Images”, U.S. Appl. No. 18/952,233, filed Nov. 19, 2024. [cited by applicant]
Chang, Jeffrey, et al., “Method and System for the Computer-Assisted Implementation of Radiology Recommendations”, U.S. Appl. No. 18/215,354, filed Jun. 28, 2023. [cited by applicant]
Chang, Jeffrey, et al., “System and Method for Automatically Displaying Information at a Radiologist Dashboard”, U.S. Appl. No. 18/952,147, filed Nov. 19, 2024. [cited by applicant]
Dai, Ning, et al., “Style transformer: Unpaired text style transfer without disentangled latent representation”, Ithaca: Cornell University Library, arXiv.org. (Year: 2019). [cited by applicant]
Deng, Mingkai, “RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning”, Machine Learning, Carnegie Mellon University, published Feb. 24, 2023. [cited by applicant]
Ebesu, Travis Akira, “Deep learning for recommender systems”, (Order No. 13900137). Available from ProQuest Dissertations & Theses Global. (2293976827). (Year: 2019). [cited by applicant]
Garud, Hrishkesh Deepak, et al., “Transforming human pose forecasting”, (Order No. 27814956). Available from ProQuest Dissertations & Theses Global. (2399247743). (Year: 2019). [cited by applicant]
Jettaku, Amarin , et al., “Relation extraction between bacteria and biotopes from biomedical texts with X attention mechanisms and domain-specific contextual representations”, BMC Bioinformatics, 20, 1-17, (2019). [cited by applicant]
Koncel-Kedziorski, Rik, et al., “Understanding and generating multi-sentence texts”, (Order No. 13814316). Available from ProQuest Dissertations & Theses Global. (2305944561). (Year: 2019). [cited by applicant]
Lambert, Nathan, et al., “Illustrating Reinforcement Learning from Human Feedback (RLHF)”, Hugging Face, https://huggingface.co/blog/rlhf, published Dec. 9, 2022. [cited by applicant]
Lou, Robert, et al., “Automated detection of radiology reports that require follow-up imaging using natural language processing feature engineering and machine learning classification”, Journal of digital imaging, 33(1)… [cited by applicant]
Lu, Edward, et al., “LORA: Low-Rank Adaptation of Large Lan-Guage Models”, arXiv:2106.09685, https://doi.org/10.48550/arXiv.2106.09685, Jun. 17, 2021. [cited by applicant]
Malhotra, Tanya, “Exploring The Differences Between ChatGPT/GPT-4 and Traditional Language Models: The Impact of Reinforcement Learning from Human Feedback (RLHF)”, MarkTestPost, https://www.marktechpost.com/2023/03/21/… [cited by applicant]
Nandhakumar, Nidhin , et al., “Clinically Significant Information Extraction from Radiology Report”, DocEng '17: Proceedings of the 2017 ACM Symposium on Document Engineering, Aug. 2017, pp. 153-162. [cited by applicant]
Paulett, John, et al., “System and Method for Radiology Reporting”, U.S. Appl. No. 18/638,368, filed Apr. 17, 2024. [cited by applicant]
Sanjabi, Nima, “Abstractive text summarization with attention-based mechanism”, (Projecte Final de Master Oficial). UPC, Facultat d'Informatica de Barcelona. (Year: 2018). [cited by applicant]
Sean, Xiao, “Fine-tuning LLMs Made Easy with LoRA and Generative AI-Stable Diffusion LoRA”, Medium, https://xiaosean5408.medium.com/fine-tuning-llms-made-easy-with-lora-and-generative-ai-stable-diffusion-lora-39ff27480f… [cited by applicant]
Song, Huan, “Data-driven representation learning in multimodal feature fusion”, (Order No. 10838232). Available from ProQuest Dissertations & Theses Global. (2094858110). (Year: 2018). [cited by applicant]
Van Veen, Dave, et al., “RadAdapt: Radiology Report Summarization via Lightweight Domain Adaptation of Large Language Models”, arXiv:2305.01146, https://doi.org/10.48550/arXiv.2305.01146, May 2, 2023. [cited by applicant]
Witteveen, Sam , “Building a Summarization System with LangChain and GPT-3—Part 2”, https://www.youtube.com/watch?v=d-yeHDLgKHw, Mar. 11, 2023. [cited by applicant]
Xu, Shawn, et al., “ELIXR: Towards a general purpose X-ray artificial intelligence system through alignment of large language models and radiology vision encoders”, arXiv:2308.01317 [cs.CV], Aug. 2, 2023. [cited by applicant]
Xue, Y. , et al., “Multimodal Recurrent Model with Attention for Automated Radiology Report Generation”, Medical Image Computing and Computer Assisted Intervention—MICCAI 2018. MICCAI 2018. Lecture Notes in Computer Sci… [cited by applicant]
Yao, Shunyu , et al., “ReAct: Synergizing Reasoning and Acting in Language Models”, https://react-lm.github.io, Oct. 6, 2022. [cited by applicant]
Zech, John, et al., “Natural Language-based Machine Learning Models for the Annotation of Clinical Radiology Reports”, Radiology Reports. Jan. 30, 2018 (Jan. 30, 2018). [cited by applicant]
Zhang, Yuhao , et al., “Learning to Summarize Radiology Findings”, Oct. 8, 2018 (Oct. 8, 2018). 1-20 [retrieved on Nov. 18, 2020]. [cited by applicant]
Chang, et al., “Method and System for the Computer-Assisted Implementation of Radiology Recommendations”, U.S. Appl. No. 19/174,282, filed Apr. 9, 2025. [cited by applicant]
Li, Junnan , et al., “BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models”, arXiv:2301.12597v1, Jan. 30, 2023. [cited by applicant]
Chadha, Aman, “Exploring the Art of Artificial Intelligence One Concept at a Time”, aman.ai, https://aman.ai/, first downloaded Jun. 7, 2023. [cited by applicant]
Li, et al., “BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models”, arXiv:2301.12597v1, Jan. 30, 2023. [cited by applicant]