IP Library › Granted Patent US 12,340,166
Granted Patent B2
US 12,340,166 · App. 18/328,593 · Granted Jun 24, 2025

Document decomposition based on determined logical visual layering of document content

Inventors: Punit Singh (New Delhi, IN); Jayant Vaibhav Srivastava (Noida, IN); Ankit Bal (Noida, IN)
Assignee: Adobe Inc.
G06F40/169G06F40/197
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,340,166
App. No.
18/328,593
Granted
Jun 24, 2025
Kind
B2
Abstract

Techniques for document decomposition based on determined logical visual layering of document content. The techniques include iteratively identifying a plurality of logical visual layers of a document resulting in each logical visual layer being associated with one or more document content objects of the document. The one or more document content objects associated with each logical visual layer are annotated to be indicative of the associated logical visual layer. The document is then displayed with an indication of one or more of the annotated document objects.

Claims (81)

1. A method comprising:

iteratively identifying one or more document content objects of an electronic document using a visual foreground of the electronic document, wherein the visual foreground of the electronic document is based on a logical visual layer of a plurality of logical visual layers of an electronic document;

annotating the one or more document content objects associated with each logical visual layer of the plurality of logical visual layers to be indicative of the associated logical visual layer, wherein annotations for the one or more document content objects are stored as annotation metadata associated with the electronic document; and

causing the electronic document to be displayed with a visual indication of one or more of the annotated document content objects.

2. The method of claim 1 , further comprising:

during a first iteration of a plurality of iterations:

rendering the electronic document as a first document image,

determining a foreground of the electronic document using the first document image,

rendering a first document content object in a first foreground image,

determining a first rendering area of the first foreground image, and

determining that the first document content object corresponds to the first rendering area; and

during a second iteration of the plurality of iterations:

rendering a version of the electronic document without the first document content object as a second document image,

determining a foreground of the version of the electronic document without the first document content object using the second document image,

rendering a second document content object in a second foreground image,

determining a second rendering area of the second foreground image, and

determining that the second document content object corresponds to the second rendering area.

3. The method of claim 1 , wherein a document content object associated with a first logical visual layer of the plurality of logical visual layers partially overlaps a second document content object associated with a second logical visual layer of the plurality of logical visual layers.

4. The method of claim 1 , wherein a document content object of the electronic document comprises text or a raster image; and wherein the method further comprises: determining, during a first iteration of a plurality of iterations, that the document content object is in a foreground of the electronic document based on:

determining a rendering area of the document content object; and

determining how much of the rendering area is within a foreground mask.

5. The method of claim 1 , wherein a document content object of the electronic document comprises a plurality of vector objects; and wherein the method further comprises: determining, during a first iteration of a plurality of iterations, that the document content object is in a foreground of the electronic document based on determining how much of each vector object of the plurality of vector objects is within a foreground mask.

6. The method of claim 1 , wherein:

determining, during a first iteration of a plurality of iterations, that a document content object of the electronic document is in a foreground of the electronic document based on using a machine learning model to separate an image of the electronic document into the foreground.

7. The method of claim 1 , further comprising:

receiving a user selection of a document content object of the electronic document;

determining instructions or data of the electronic document for rendering the document content object;

receiving a user edit to the document content object; and

applying the user edit to the instructions or data of the electronic document for rendering the document content object.

8. The method of claim 1 , wherein a document content object of the electronic document comprises instructions or data for rendering text, a raster image, a vector graphics image, audio, video, or a link.

9. A non-transitory computer-readable medium storing executable instructions which, when executed by a processing device, cause the processing device to perform operations comprising:

receiving a first user input requesting to decompose an electronic document into constituent document content objects;

iteratively decomposing the electronic document into a set of document content objects using a visual foreground of the electronic document, wherein the visual foreground of the electronic document is based on a plurality of logical visual layers of the electronic document;

storing the set of document content objects as annotation metadata associated with the electronic document;

receiving a second user input selecting a particular document content object of the set of document content objects;

receiving a third user input requesting to perform an action on the particular document content object; and

performing the action on the particular document content object.

10. The non-transitory computer-readable medium of claim 9 , wherein the electronic document is iteratively decomposed into the set of document content objects based on:

during a first iteration of a plurality of iterations:

rendering the electronic document as a first document image,

determining a foreground of the electronic document using the first document image,

rendering a first document content object in a first foreground image,

determining a first rendering area of the first foreground image, and

determining that the first document content object corresponds to the first rendering area; and

during a second iteration of the plurality of iterations:

rendering a version of the electronic document without the first document content object as a second document image,

determining, using the second document image, a foreground of the version of the electronic document without the first document content object,

rendering a second document content object in a second foreground image,

determining a second rendering area of the second foreground image, and

determining that the second document content object corresponds to the second rendering area.

11. The non-transitory computer-readable medium of claim 9 , wherein a first document content object of the set of document content objects partially overlaps a second document content object of the set of document content objects in the plurality of logical visual layers of the electronic document.

12. The non-transitory computer-readable medium of claim 9 , wherein a document content object of the electronic document comprises text or a raster image; and wherein the electronic document is decomposed into the set of document content objects based on determining, during a first iteration of a plurality of iterations, the document content object is in a foreground of the electronic document based on determining a rendering area of the document content object, and determining how much of the rendering area is within a foreground mask.

13. The non-transitory computer-readable medium of claim 9 , wherein a document content object comprises a plurality of vector objects; and wherein the electronic document is decomposed into the set of document content objects based on determining, during an iteration of a plurality of iterations, a document content object is in a foreground of the electronic document based on determining how much of each vector object of the plurality of vector objects is within a foreground mask.

14. The non-transitory computer-readable medium of claim 9 , wherein iteratively decomposing the electronic document into a set of document content objects based on a plurality of logical visual layers of the electronic document further comprises:

iteratively identifying a plurality of logical visual layers of the electronic document, wherein each logical visual layer is associated with one or more document content objects of the set of document content objects; and

annotating the one or more document content objects associated with each logical visual layer of the plurality of logical visual layers to be indicative of the associated logical visual layer.

15. A system comprising:

one or more memory components; and

one or more processing devices coupled to the one or more memory components, the one or more processing devices to perform operations comprising:

iteratively identifying one or more document content objects of an electronic document using a visual foreground of the electronic document, wherein the visual foreground of the electronic document is based on a logical visual layer of a plurality of logical visual layers of an electronic document;

annotating the one or more document content objects associated with each logical visual layer of the plurality of logical visual layers to be indicative of the associated logical visual layer, wherein annotations for the one or more document content objects are stored as annotation metadata associated with the electronic document; and

causing the electronic document to be displayed with a visual indication of one or more of the annotated document content objects.

16. The system of claim 15 , the one or more processing devices to further perform operations comprising:

during a first iteration of a plurality of iterations:

rendering the electronic document as a first document image,

determining a foreground of the electronic document using the first document image,

rendering a first document content object in a first foreground image,

determining a first rendering area of the first foreground image, and

determining that the first document content object corresponds to the first rendering area; and

during a second iteration of the plurality of iterations:

rendering a version of the electronic document as a second document image, wherein the second document image does not include a rendering of the first document content object,

determining a foreground of the version of the electronic document using the second document image,

rendering a second document content object in a second foreground image,

determining a second rendering area of the second foreground image, and

determining that the second document content object corresponds to the second rendering area.

17. The system of claim 15 , wherein a first document content object partially overlaps a second document content object in a visual layering of the electronic document.

18. The system of claim 15 , wherein a document content object comprises text or a raster image; and wherein the one or more processing devices are to perform determining, during an iteration of a plurality of iterations, that the document content object is in a foreground of the electronic document based on:

determining a rendering area of the document content object; and

determining how much of the rendering area is within a foreground mask.

19. The system of claim 15 , wherein a document content object comprises a plurality of vector objects; and wherein the one or more processing devices are to perform determining, during an iteration of a plurality of iterations, that the document content object is in a foreground of the electronic document based on determining how much of each vector object of the plurality of vector objects is within a foreground mask.

20. The system of claim 15 , wherein the one or more processing devices are to perform determining, during an iteration of a plurality of iterations, that a document content object is in a foreground of the electronic document based on using a machine learning model to separate an image of the electronic document into the foreground.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 12, 2023
From: SINGH, PUNIT; SRIVASTAVA, JAYANT VAIBHAV; BAL, ANKIT
To: ADOBE INC.
Reel/Frame 063924/0756 →
Continuity (1)
Related Publication 20240403543A1 · Dec 5, 2024
References Cited (21)
US 7386789B2 · Chao · 2008 [cited by examiner]
US 9922007B1 · Jain · 2018 [cited by examiner]
US 11514282B1 · Moore · 2022 [cited by examiner]
US 11798246B2 · Lee · 2023 [cited by examiner]
US 20050180642A1 · Curry · 2005 [cited by examiner]
US 20080194325A1 · Komuta · 2008 [cited by examiner]
US 20150082148A1 · Lai · 2015 [cited by examiner]
US 20150199314A1 · Ratnakar · 2015 [cited by examiner]
US 20160124813A1 · Jain · 2016 [cited by examiner]
US 20160232204A1 · Zholudev · 2016 [cited by examiner]
US 20170124717A1 · Baruch · 2017 [cited by examiner]
US 20170212870A1 · Thomsen · 2017 [cited by examiner]
US 20200175095A1 · Morariu · 2020 [cited by examiner]
US 20210319166A1 · Lyu · 2021 [cited by examiner]
US 20210334664A1 · Li · 2021 [cited by examiner]
US 20220262009A1 · Yu · 2022 [cited by examiner]
US 20220314121A1 · Fukutsuka · 2022 [cited by examiner]
US 20230206955A1 · Cole · 2023 [cited by examiner]
US 20240144141A1 · Cella · 2024 [cited by examiner]
US 20250036254A1 · Wang · 2025 [cited by examiner]
Park et al., A Smart Communication System for Avatar Agents in Virtual Environment, 2008, IEEE, 7 pages. (Year: 2008). [cited by examiner]