IP Library › Granted Patent US 12,400,384
Granted Patent B2
US 12,400,384 · App. 18/460,401 · Granted Aug 26, 2025

Reflowing documents to display semantically related content

Inventors: Christopher Tensmeyer (Fulton, MD); Fuxiao Liu (Baltimore, MD); Hao Tan (Santa Clara, CA); Ani Nenkova (Philadelphia, PA)
Assignee: Adobe Inc.
G06T11/60G06F3/04842G06F3/04845G06F2203/04803G06T2200/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,400,384
App. No.
18/460,401
Granted
Aug 26, 2025
Kind
B2
Abstract

Embodiments are disclosed for reflowing documents to display semantically related content. The method may include receiving a request to view a document that includes body text and one or more images. A trimodal document relationship model identifies relationships between segments of the body text and the one or more images. A linearized view of the document is generated based on the relationships and the linearized view is caused to be displayed on a user device.

Claims (44)

1. A method, comprising: receiving a request to view a document that includes body text, one or more images, and one or more associated captions;

identifying, using a trimodal document relationship model, relationships between segments of the body text and the one or more images, wherein the trimodal document relationship model: generates a contextual embedding for each segment of the body text, image, and associated caption, and

predicts at least one segment of the body text associated with each image from the one or more images based on a similarity score determined between a plurality of image-caption pairs and segments of the body text based on their contextual embeddings which encode a combination of image embeddings, text embeddings, segment embeddings, and position embeddings;

generating a linearized view of the document based on the relationships; and

causing the linearized view to be displayed on a user device.

2. The method of claim 1 , wherein identifying, using a trimodal document relationship model, relationships between segments of the body text and the one or more images, further comprises:

receiving, by the trimodal document relationship model, a plurality of segments of the body text, the one or more images, and one or more associated captions from the document.

3. The method of claim 2 , wherein the trimodal document relationship model includes a transformer encoder.

4. The method of claim 3 , wherein each segment embedding defines a segment type and each position embedding indicates a position of the segment of body text, image, or associated caption.

5. The method of claim 1 , wherein a segment of body text includes a section, a paragraph, or a sentence.

6. The method of claim 1 , wherein the linearized view is a linear presentation of the segments of the body text, wherein each segment of the body text determined to be associated with an image from the one or more images has an associated user interface element rendered in the linearized view.

7. The method of claim 6 , further comprising:

receiving a selection of a first user interface element associated with a first segment of the body text in the linearized view; and

causing an adjustable split screen to be displayed on the user device, wherein a first pane of the split screen displays a first image determined to be associated with the first segment, and a second pane of the split screen displays at least some of the first segment of the body text.

8. The method of claim 7 , wherein multiple images are associated with the first segment of the body text, and wherein the first screen of the split screen includes a second user interface element which, when selected, causes a different image from the multiple images to be displayed in the first screen.

9. The method of claim 7 , wherein the first pane and the second pane of the adjustable split screen are resizable using an interactive user interface element.

10. A non-transitory computer-readable medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising: receiving a request to view a document that includes body text, one or more images, and one or more associated captions;

identifying, using a trimodal document relationship model, relationships between segments of the body text and the one or more images, wherein the trimodal document relationship model: generates a contextual embedding for each segment of the body text, image, and associated caption, and

predicts at least one segment of the body text associated with each image from the one or more images based on a similarity score determined between a plurality of image-caption pairs and segments of the body text based on their contextual embeddings which encode a combination of image embeddings, text embeddings, segment embeddings, and position embeddings;

generating a linearized view of the document based on the relationships; and

causing the linearized view to be displayed on a user device.

11. The non-transitory computer-readable medium of claim 10 , wherein the operation of identifying, using a trimodal document relationship model, relationships between segments of the body text and the one or more images, further comprises:

receiving, by the trimodal document relationship model, a plurality of segments of the body text, the one or more images, and one or more associated captions from the document.

12. The non-transitory computer-readable medium of claim 11 , wherein the trimodal document relationship model includes a transformer encoder.

13. The non-transitory computer-readable medium of claim 12 , wherein each segment embedding defines a segment type and each position embedding indicates a position of the segment of body text, image, or associated caption.

14. The non-transitory computer-readable medium of claim 10 , wherein the linearized view is a linear presentation of the segments of the body text, wherein each segment of the body text determined to be associated with an image from the one or more images has an associated user interface element rendered in the linearized view.

15. The non-transitory computer-readable medium of claim 14 , wherein the operations further comprise:

receiving a selection of a first user interface element associated with a first segment of the body text in the linearized view; and

causing an adjustable split screen to be displayed on the user device, wherein a first pane of the split screen displays a first image determined to be associated with the first segment, and a second pane of the split screen displays at least some of the first segment of the body text.

16. The non-transitory computer-readable medium of claim 15 , wherein multiple images are associated with the first segment of the body text, and wherein the first screen of the split screen includes a second user interface element which, when selected, causes a different image from the multiple images to be displayed in the first screen.

17. The non-transitory computer-readable medium of claim 15 , wherein the first pane and the second pane of the adjustable split screen are resizable using an interactive user interface element.

18. A system, comprising:

a memory component; and

a processing device coupled to the memory component, the processing device to perform operations comprising:

receiving, by a trimodal document relationship model, a plurality of elements of a document, wherein the elements include segments of body text, images, and image captions;

generating, by a feature extractor of the trimodal document relationship model, an element embedding for each element of the document;

generating a segment embedding, indicating an element type, and a position embedding, indicating an element position within the document, for each element of the document;

combining each element embedding, segment embedding, and position to create a combined embedding for each element of the document;

generating, by a transformer encoder of the trimodal document relationship model, a contextual embedding for each element of the document corresponding to each combined embedding; and

determining semantic relationships between the segments of the body text and the images in the document using their contextual embeddings.

19. The system of claim 18 , wherein the operations further comprise:

generating a reflowed document based on the semantic relationships.

20. The system of claim 18 , wherein the operation of determining semantic relationships between the segments of the body text and the images in the document using their contextual embeddings further comprises:

determining a similarity score between a plurality of image-caption pairs and segment sentences based on their contextual embeddings, the similarity score indicating a likelihood that an image is associated with a segment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2023
From: TENSMEYER, CHRISTOPHER; LIU, FUXIAO; TAN, HAO; NENKOVA, ANI
To: ADOBE INC.
Reel/Frame 064820/0673 →
Continuity (1)
Related Publication 20250078350A1 · Mar 6, 2025
References Cited (10)
US 8539342B1 · Lewis · 2013 [cited by examiner]
US 9766782B2 · Migos · 2017 [cited by examiner]
US 11176310B2 · Agrawal · 2021 [cited by examiner]
US 20090210828A1 · Kahn · 2009 [cited by examiner]
US 20120096344A1 · Ho · 2012 [cited by examiner]
Devlin et al. “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, 2019, 16 pgs. Bert.pdf (Year: 2019). [cited by examiner]
Liu, F., et al., “Upgrading the Newsroom: An Automated Image Selection System for News Articles,” ACM Transactions on Multimedia Computing, Communications, and Applications, vol. 16, Issue 3, Article No. 81, Jul. 2020, … [cited by applicant]
Muraoka, M., et al., “Image Position Prediction in Multimodal Documents,” Proceedings of the 12th Language Resources and Evaluation Conference, May 2020, pp. 4265-4274. [cited by applicant]
Radford, A., et al., “Learning Transferable Visual Models From Natural Language Supervision,” Proceedings of the 38th International Conference on Machine Learning, PMLR 139, Feb. 2021, pp. 1-16. [cited by applicant]
Shibata, Y., et al., “Byte Pair Encoding: A Text Compression Scheme That Accelerates Pattern Matching,” Sep. 1999, pp. 1-13. [cited by applicant]