IP Library › Granted Patent US 12,688,227
Granted Patent B2
US 12,688,227 · App. 18/403,808 · Granted Jul 21, 2026

Automatic generation of presentations from documents

Inventors: Himanshu Maheshwari (Bengaluru, IN); Aparna Garimella (Bengaluru, IN); Niyati Himanshu Chhaya (Bangalore, IN)
Assignee: ADOBE INC.
G06F16/4393G06F40/30G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,688,227
App. No.
18/403,808
Filed
Jan 4, 2024
Granted
Jul 21, 2026
Kind
B2
Art Unit
2653
USPC
704/9
Abstract

Embodiments of the present disclosure include extracting structured text from a source document. The structured text comprises a plurality of source sections. Some embodiments generate a semantic outline based on the structured text. In some examples, the semantic outline comprises a plurality of output headings. Some embodiments generate text content corresponding to each of the plurality of output headings. An image is selected from the source document for each of the plurality of output headings by computing a similarity score between the image and the text content. Then, an output document is generated based on the semantic outline, where the output document comprises a plurality of output sections corresponding to the plurality of output headings, respectively.

Claims (57)

1 . A method comprising:

extracting structured text from a source document, wherein the structured text comprises a plurality of source sections;

generating, using a language generation model comprising a transformer model, a semantic outline based on the structured text by performing a self-attention mechanism that combines information from the structured text with an input prompt, wherein the input prompt comprises instructions to a plurality of output headings for the semantic outline from the plurality of source sections, and wherein the self-attention mechanism is performed by an encoder or a decoder of the transformer model;

generating, using the language generation model, text content corresponding to each of the plurality of output headings;

selecting an image from the source document for each of the plurality of output headings by computing a similarity score between the image and the generated text content of a respective output heading of the plurality of output headings; and

generating an output document based on the semantic outline, wherein the output document comprises the selected image, the text content, and a plurality of output sections corresponding to the plurality of output headings, respectively.

2 . The method of claim 1 , wherein

the input prompt includes the instructions to generate the semantic outline to cover important sections of the plurality of source sections.

3 . The method of claim 1 , further comprising:

extracting the image from the source document; and

selecting the image to represent a heading of the plurality of output headings, wherein the output document includes the selected image in an output section corresponding to the heading.

4 . The method of claim 1 , further comprising:

identifying a plurality of images from the source document and a pre-determined selection factor; and

filtering the plurality of images based on the pre-determined selection factor to obtain a filtered set of images, wherein the filtered set of images includes the selected image.

5 . The method of claim 1 , wherein selecting the image comprises:

generating a multi-modal text embedding based on the text content;

generating a multi-modal image embedding based on the image; and

computing the similarity score by comparing the multi-modal text embedding and the multi-modal image embedding.

6 . The method of claim 1 , further comprising:

generating a synthesized image based on a heading of the plurality of output headings, wherein the output document includes the synthesized image in an output section corresponding to the heading.

7 . The method of claim 1 , wherein:

the semantic outline includes a plurality of sub-headings for a heading of the plurality of output headings.

8 . The method of claim 1 , wherein:

a number of the plurality of source sections is greater than a number of the plurality of output headings.

9 . The method of claim 1 , further comprising:

identifying a predetermined number of output headings in the output document, wherein the semantic outline is generated based on the predetermined number of output headings.

10 . The method of claim 1 , wherein generating the output document comprises:

obtaining a document template; and

inserting the plurality of output headings into the document template.

11 . The method of claim 1 , wherein:

the output document comprises a slide presentation including a plurality of slides corresponding to the plurality of output sections, respectively.

12 . The method of claim 1 , further comprising:

receiving user feedback on the output document; and

updating the output document based on the user feedback.

13 . An apparatus comprising:

at least one processor;

at least one memory including instructions executable by the at least one processor;

an extraction component comprising parameters stored in the at least one memory and configured to extract structured text from a source document, wherein the structured text comprises a plurality of source sections;

a language generation model comprising a transformer model and including parameters stored in the at least one memory, wherein the language generation model is configured to generate a semantic outline based on the structured text by performing a self-attention mechanism that combines information from the structured text with an input prompt, wherein the input prompt comprises instructions to select a plurality of output headings for the semantic outline from the plurality of source sections, wherein the self-attention mechanism is performed by an encoder or a decoder of the transformer model, and wherein the language generation model is configured to generate text content corresponding to each of the plurality of output headings;

an image selection component configured to select an image from the source document for each of the plurality of output headings by computing a similarity score between the image and the generated text content of a respective output heading of the plurality of output headings; and

a document generator comprising parameters stored in the at least one memory and configured to generate an output document based on the semantic outline, wherein the output document comprises the text content and a plurality of output sections corresponding to the plurality of output headings, respectively.

14 . The apparatus of claim 13 , wherein:

the image selection component selects the image to represent a heading of the plurality of output headings, wherein the output document includes the selected image in an output section corresponding to the heading.

15 . The apparatus of claim 13 , further comprising:

an image generator configured to generate a synthesized image based on a heading of the plurality of output headings, wherein the output document includes the synthesized image in an output section corresponding to the heading.

16 . The apparatus of claim 13 , further comprising:

a user interface configured to receive user feedback on the output document.

17 . A non-transitory computer readable medium storing code for natural language processing, the code comprising instructions executable by at least one processor to:

extracting structured text from a source document, wherein the structured text comprises a plurality of source sections;

generating, using a language generation model comprising a transformer model, a semantic outline based on the structured text by performing a self-attention mechanism that combines information from the structured text with an input prompt, wherein the input prompt comprises instructions to select a plurality of output headings for the semantic outline from the plurality of source sections, and wherein the self-attention mechanism is performed by an encoder or a decoder of the transformer model;

generating, using the language generation model, text content corresponding to each of the plurality of output headings;

selecting an image from the source document for each of the plurality of output headings by computing a similarity score between the image and the generated text content of a respective output heading of the plurality of output headings; and

generating an output document based on the semantic outline, wherein the output document comprises the selected image, the text content, and a plurality of output sections corresponding to the plurality of output headings, respectively.

18 . The non-transitory computer readable medium of claim 17 , wherein:

the input prompt includes the instructions to generate the semantic outline to cover important sections of the plurality of source sections.

19 . The non-transitory computer readable medium of claim 17 , the code further comprising instructions executable by the at least one processor to:

identify a predetermined number of output headings in the output document, wherein the semantic outline is generated based on the predetermined number of output headings.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2024
From: MAHESHWARI, HIMANSHU; GARIMELLA, APARNA; CHHAYA, NIYATI HIMANSHU
To: ADOBE INC.
Reel/Frame 066014/0932 →
Continuity (1)
Related Publication 20250225173A1 · Jul 10, 2025
References Cited (22)
US 10222942B1 · Zeiler · 2019 [cited by examiner]
US 11748577B1 · Aberle · 2023 [cited by examiner]
US 20180232340A1 · Lee · 2018 [cited by examiner]
US 20190095803A1 · Raskovic · 2019 [cited by examiner]
US 20210279269A1 · Verma · 2021 [cited by examiner]
US 20220092097A1 · Gupta · 2022 [cited by examiner]
US 20230343126A1 · Salacinski · 2023 [cited by examiner]
US 20230351102A1 · Tran · 2023 [cited by examiner]
US 20240013562A1 · Montero · 2024 [cited by examiner]
US 20240062019A1 · Aberle · 2024 [cited by examiner]
Tsu-Jui Fu, William Yang Wang, Daniel McDuff, Yale Song2, “DOC2PPT: Automatic Presentation Slides Generation from Scientific Documents”, UC Santa Barbara Microsoft Research, The Thirty-Sixth AAAI Conference on Artificia… [cited by examiner]
“Introducing ChatGPT”, Oct. 2023, found on the internet at https://openai.com/blog/chatgpt, 14 pages. [cited by applicant]
Masum, et al., “A Multi-Agent System for Building Automatic Multi-Modal Presentation of a Topic from World Wide Web Information”, In IEEE/WIC/ACM International Conference on Intelligent Agent Technology, (Sep. 2005), 4 … [cited by applicant]
Winters, et al., “Automatically Generating Engaging Presentation Slide Decks”, In International Conference on Computational Intelligence in Music, Sound, Art and Design, (Apr. 2019), https://doi.org/10.1007/978-3-030-16… [cited by applicant]
Hu, et al., “PPSGen: Learning-Based Presentation Slides Generation for Academic Papers”, in IEEE transactions on knowledge and data engineering, vol. 27, No. 4, Apr. 2015, pp. 1085-1097. [cited by applicant]
Bhandare, et al., “Automatic Era: Presentation slides from Academic Paper”, In 2016 International Conference on Automatic Control and Dynamic Optimization Techniques (ICACDOT), (Sep. 2016), pp. 809-814. [cited by applicant]
Syamili, et al., “Presentation Slides Generation from Scientific papers using Support Vector Regression”, In International Conference on Inventive Communication and Computational Technologies (ICICCT 2017), (Mar. 2017),… [cited by applicant]
Sefid, edt al., “Automatic Slide Generation for Scientific Papers”, In Third International Workshop on Capturing Scientific Knowledge co-located with the 10th International Conference on Knowledge Capture, Nov. 2019, fo… [cited by applicant]
Wang, et al., “Phrase-Based Presentation Slides Generation for Academic Papers”, In Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, vol. 31, No. 1, Feb. 2017, pp. 196-202. [cited by applicant]
Sun, et al., “D2S: Document-to-Slide Generation Via Query-Based Text Summarization”, arXiv preprint arXiv:2105.03664v1 [cs.CL] May 8, 2021, 14 pages. [cited by applicant]
“python-pptx 0.6.22”, Oct. 2023, found on the internet at https://pypi.org/project/python-pptx/, 9 pages. [cited by applicant]
“SlidesAI.io—Create Slides with AI”, Oct. 2023, found on the internet at https://workspace.google.com/marketplace/app/slidesaiio_create_slides_with_ai/904276957168, 1 page. [cited by applicant]