IP Library › Granted Patent US 12,626,433
Granted Patent B2
US 12,626,433 · App. 18/323,029 · Granted May 12, 2026

Generating supplemental text and image content in multimodal digital content items via machine learning

Inventors: Anant Shankhdhar (New Delhi, IN); Samyak Sanjay Mehta (Mumbai, IN); Shreya Singh (Ahmedabad, IN); K V Vikram (Coimbatore, IN); Tripti Shukla (Lucknow, IN); Srikrishna Karanam (Bangalore, IN); Balaji Vasan Srinivasan (Bangalore, IN); Vishwa Vinay (Bangalore, IN); Niyati Himanshu Chhaya (Bangalore, IN)
Assignee: Adobe Inc.
G06T11/60G06F16/5866G06F40/211G06V30/418
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,433
App. No.
18/323,029
Granted
May 12, 2026
Kind
B2
Abstract

The present disclosure relates to systems, non-transitory computer-readable media, and methods for expanding a digital document including a sequence of informational data via supplemental multimodal digital content. In particular, the system expands digital documents with multimodal granular details to dynamically integrate supplemental in-depth information to the digital document. For example, in response to a selection of a specific portion of a digital document, the system generates expanded multimodal content (e.g., text and image content) for the selected portion of the digital document from external text and image sources. Indeed, the system uses existing content from the digital document to select images and combine the selected images with text into image-text pairs that are textually and visually consistent with the digital document. Moreover, the system expands the digital document by inserting the image-text pairs in connection with the selected portion of the digital document.

Claims (68)

1 . A computer-implemented method comprising:

generating, by at least one processor and for a selected content item from a plurality of content items of a selected digital document, a plurality of extracted text content items by extracting text from one or more digital documents of a plurality of ranked digital documents of a digital document repository, wherein the plurality of extracted text content items correspond to supplemental textual content comprising granular detail related to the selected content item;

retrieving, by the at least one processor and from an image repository, a plurality of selected digital images based on the plurality of extracted text content items; and

modifying, by the at least one processor, the selected digital document by inserting image-text pairs comprising the plurality of extracted text content items and the plurality of selected digital images in connection with the selected content item.

2 . The computer-implemented method of claim 1 , further comprising detecting the selected content item from the plurality of content items by detecting a selection of an image-text pair within the selected digital document.

3 . The computer-implemented method of claim 1 , wherein determining the plurality of ranked digital documents from the digital document repository comprises selecting one or more digital documents from the digital document repository based on a similarity of textual content within the one or more digital documents to the selected content item.

4 . The computer-implemented method of claim 1 , wherein generating the plurality of extracted text content items comprises:

generating text content dependency graphs for the extracted text from the one or more digital documents;

generating a content item dependency graph for the selected content item; and

selecting a subset of the extracted text from the one or more digital documents based on the text content dependency graphs and the content item dependency graph.

5 . The computer-implemented method of claim 4 , wherein modifying the selected digital document comprises:

determining a text content order for the plurality of extracted text content items based on an order of the extracted text in the one or more digital documents; and

inserting the plurality of extracted text content items and the plurality of selected digital images in the selected digital document based on the text content order.

6 . The computer-implemented method of claim 1 , wherein modifying the selected digital document comprises replacing the selected content item with the image-text pairs.

7 . The computer-implemented method of claim 1 , wherein retrieving the plurality of selected digital images comprises:

parsing the plurality of extracted text content items to obtain key phrases associated with the selected content item; and

retrieving, from the image repository, the plurality of selected digital images based on the key phrases.

8 . The computer-implemented method of claim 7 , wherein:

parsing the plurality of extracted text content items to obtain the key phrases comprises:

extracting a plurality of keywords from the plurality of extracted text content items; and

generating a set of queries comprising the key phrases based on the plurality of keywords or one or more combinations of the plurality of keywords; and

retrieving the plurality of selected digital images comprises performing digital image searches based on the set of queries comprising the key phrases.

9 . The computer-implemented method of claim 1 , wherein determining the plurality of selected digital images comprises:

extracting first image features from the plurality of selected digital images;

extracting second image features from one or more digital images in the selected digital document; and

selecting the plurality of selected digital images based on the first image features and the second image features.

10 . A system comprising:

one or more memory devices comprising a digital document; and

one or more processors configured to cause the system to:

determine in response to an indication of a selected content item from a plurality of ordered content items of a selected digital document, a plurality of ranked digital documents from a digital document repository;

generate, for the selected content item and utilizing a natural language processing model, a plurality of text content items by comparing text of the selected content item to text extracted from one or more documents of the plurality of ranked digital documents, wherein the text extracted from the one or more documents corresponds to supplemental textual content comprising granular detail related to the selected content item;

select, from an image repository, a plurality of selected digital images based on one or more queries generated from the plurality of text content items; and

modify the selected digital document by inserting digital content comprising the plurality of text content items and the plurality of selected digital images into the plurality of ordered content items of the selected digital document.

11 . The system of claim 10 , where the one or more processors are further configured to cause the system to determine the indication of the selected content item from the plurality of ordered content items by detecting a selection of an image-text pair within the selected digital document.

12 . The system of claim 10 , where the one or more processors are further configured to cause the system to generate the plurality of text content items by:

generating text content dependency graphs for the extracted text from the plurality of ranked digital documents;

generating a content item dependency graph for the selected content item; and

selecting a subset of the extracted text from the plurality of ranked digital documents based on the text content dependency graphs and the content item dependency graph.

13 . The system of claim 12 , where the one or more processors are further configured to cause the system to modify the selected digital document by:

determining a text content order for the plurality of text content items based on an order of the extracted text in the plurality of ranked digital documents; and

inserting the plurality of text content items and the plurality of selected digital images in the selected digital document based on the text content order.

14 . The system of claim 10 , where the one or more processors are further configured to cause the system to modify the selected digital document by inserting digital content comprising the plurality of text content items and the plurality of selected digital images into the plurality of ordered content items adjacent to the selected content item.

15 . The system of claim 10 , where the one or more processors are further configured to cause the system to retrieve the plurality of selected digital images by:

parsing the plurality of text content items to obtain key phrases associated with the selected content item; and

retrieving, from the image repository, the plurality of selected digital images based on the one or more queries comprising the key phrases.

16 . The system of claim 10 , where the one or more processors are further configured to cause the system to:

determining, in response to an indication of a second selected content item from the plurality of ordered content items of the digital document, a second plurality of ranked digital documents from the digital document repository;

generating, for the selected content item, a second plurality of text content items by extracting text from one or more documents of the second plurality of ranked digital documents;

retrieving, from the image repository, a second plurality of selected digital images based on the plurality of text content items; and

modifying, the selected digital document by inserting second digital content comprising the second plurality of text content items and the second plurality of selected digital images in connection with the second selected content item.

17 . A non-transitory computer readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising:

determining, in response to an indication of a selected content item from a plurality of content items of a selected digital document, a plurality of ranked digital documents from a digital document repository;

generating, for the selected content item, a plurality of text content items by extracting text from one or more documents of the plurality of ranked digital documents, wherein the text from the one or more documents corresponds to supplemental textual content comprising granular detail related to the selected content item;

retrieving, from an image repository, a plurality of selected digital images based on the plurality of text content items; and

modifying, the selected digital document by inserting digital content comprising the plurality of text content items and the plurality of selected digital images in connection with the selected content item.

18 . The non-transitory computer readable medium of claim 17 , wherein determining the plurality of selected digital images comprises:

extracting first image features from the plurality of selected digital images;

extracting second image features from one or more digital images in the selected digital document; and

selecting the plurality of selected digital images based on the first image features and the second image features.

19 . The non-transitory computer readable medium of claim 17 , wherein retrieving the plurality of selected digital images comprises:

parsing the plurality of text content items to obtain key phrases associated with the selected content item;

generating an abbreviated text content item from the plurality of text content items by removing stop-words from the plurality of text content items; and

retrieving the plurality of selected digital images comprises performing digital image searches based on queries comprising the key phrases and the abbreviated text content item.

20 . The non-transitory computer readable medium of claim 17 , wherein the executable instructions cause the processing device to perform operations further comprising:

determining, in response to an indication of a second selected content item from the plurality of content items of the selected digital document, a second plurality of ranked digital documents from the digital document repository;

generating, for the selected content item, a second plurality of text content items by extracting text from one or more documents of the second plurality of ranked digital documents;

retrieving, from the image repository, a second plurality of selected digital images based on the plurality of text content items; and

modifying, the selected digital document by inserting second digital content comprising the second plurality of text content items and the second plurality of selected digital images in connection with the second selected content item.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE SPELLING OF THE FIRST ASSIGNOR FROM ANANT SHANKHDAR PREVIOUSLY RECORDED ON REEL 063751 FRAME 0727. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jul 10, 2023
From: SHANKHDHAR, ANANT; MEHTA, SAMYAK SANJAY; SINGH, SHREYA; VIKRAM, K V; SHUKLA, TRIPTI; KARANAM, SRIKRISHNA; SRINIVASAN, BALAJI VASAN; VINAY, VISHWA; CHHAYA, NIYATI HIMANSHU
To: ADOBE INC.
Reel/Frame 065023/0062 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 24, 2023
From: SHANKHDAR, ANANT; MEHTA, SAMYAK SANJAY; SINGH, SHREYA; VIKRAM, K V; SHUKLA, TRIPTI; KARANAM, SRIKRISHNA; SRINIVASAN, BALAJI VASAN; VINAY, VISHWA; CHHAYA, NIYATI HIMANSHU
To: ADOBE INC.
Reel/Frame 063751/0727 →
Continuity (1)
Related Publication 20240394942A1 · Nov 28, 2024
References Cited (24)
US 20080320384A1 · Nagarajan · 2008 [cited by examiner]
US 20120163707A1 · Baker · 2012 [cited by examiner]
US 20190065579A1 · Daher · 2019 [cited by examiner]
US 20210342399A1 · Sisto · 2021 [cited by examiner]
US 20220398230A1 · Sinha · 2022 [cited by examiner]
US 20250061290A1 · Gardner · 2025 [cited by examiner]
“Welcome to Faiss Documentation”, webpage <https://faiss.ai/>, 3 pages, retrieved from internet on Sep. 11, 2023. [cited by applicant]
Apache Lucene, “Welcom to Apache Lucene”, webpage <https://lucene.apache.org/>, 2 pages, 2011, retreived from the internet on Sep. 11, 2023. [cited by applicant]
Cai, Guanyu, et al. “Ask & confirm: active detail enriching for cross-modal retrieval with partial query.” Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021. [cited by applicant]
Castro, Santiago, et al. “Fill-in-the-blank as a challenging video understanding evaluation framework.” arXiv preprint arXiv:2104.04182 (2021). [cited by applicant]
Chen, Danqi, and Christopher D. Manning. “A fast and accurate dependency parser using neural networks.” Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP). 2014. [cited by applicant]
Dr. Ruth Colvin Clark and Dr. Richard E. Mayer. “E-Learning and the Science of Instruction”. via Google Books. 3 pages, Feb. 2016, < http://bit.ly/RichardMayer>. [cited by applicant]
Guthrie, John T., Stan Bennett, and Shelley Weber. “Processing procedural documents: A cognitive model for following written directions.” Educational Psychology Review 3.3 (1991): 249-265. [cited by applicant]
Haveliwala, Taher. Efficient computation of PageRank. Stanford, 1999. [cited by applicant]
Kirsch, I. S., and A. Jungeblut. “Literacy: Profiles of America's young adults (NAEP Rep. No. 16-PL-02).” New Jersey: Princeton University, National Assessment of Educational Progress of Educational Testing Service (198… [cited by applicant]
Koupaee, Mahnaz, and William Yang Wang. “Wikihow: A large scale text summarization dataset.” arXiv preprint arXiv:1810.09305 (2018). [cited by applicant]
Lerner, Paul, et al. “ViQuAE, a Dataset for Knowledge-based Visual Question Answering about Named Entities.” Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieva… [cited by applicant]
Li, Yangguang, et al. “Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm.”arXiv preprint arXiv:2110.05208 (2021). [cited by applicant]
Lin, Jimmy, et al. “Pyserini: An easy-to-use python toolkit to support replicable ir research with sparse and dense representations.” arXiv preprint arXiv:2102.10073 (2021). [cited by applicant]
Manning, Christopher D., et al. “The Stanford CoreNLP natural language processing toolkit.” Proceedings of 52nd annual meeting of the association for computational linguistics: system demonstrations. 2014. [cited by applicant]
Pang, Liang & Lan, Yanyan & Guo, Jiafeng & Xu, Jun & Xu, Jingfang & Cheng, Xueqi. (2017). DeepRank: A New Deep Architecture for Relevance Ranking in Information Retrieval. 10.1145/3132847.3132914. [cited by applicant]
Reddy, Mr D. Murahari, et al. “Dall-e: Creating images from text.” UGC Care Group I Journal 8.14 (2021): 71-75. [cited by applicant]
Tarau, Paul, and Eduardo Blanco. “Interactive text graph mining with a prolog-based dialog engine.” Theory and Practice of Logic Programming 21.2 (2021): 244-263. [cited by applicant]
Wielemaker, Jan. “An overview of the SWI-Prolog Programming Environment.” WLPE 3 (2003): 1-16. [cited by applicant]