IP Library › Granted Patent US 7,958,444
Granted Patent B2
US 7,958,444 · App. 11/453,609 · Granted Jun 7, 2011

Visualizing document annotations in the context of the source document

Assignee: Xerox Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,958,444
App. No.
11/453,609
Granted
Jun 7, 2011
Kind
B2
Abstract

In a document annotator ( 8 ), a document converter ( 12 ) is configured to convert a source document ( 10 ) with a layout to a deterministic format ( 14, 64 ) including content and layout metadata. At least one annotation pipeline ( 20, 22 ) is configured to generate document annotations respective to received content. A merger ( 36, 46 ) is configured to associate the generated document annotations with positional tags based on the layout metadata, which locate the document annotations in the layout. A document visualizer ( 58 ) is configured to render at least some content of the deterministic format and one or more selected annotations ( 60 ) in substantial conformance with the layout based on the layout metadata and the positional tags associated with the selected one or more annotations ( 60 ).

Claims (59)

1. A document annotator comprising:

a document converter configured to generate an initial representation of a source document, the initial representation including initial content of the source document and initial layout metadata indicative of layout of said content in the source document;

at least one annotation pipeline configured to process at least some of the initial content of the source document to generate document annotations;

a merger configured to assign positional tags of the initial layout metadata to the generated document annotations to locate the document annotations respective to the initial layout metadata;

catalog data storage configured to store at least the document annotations and their assigned positional tags; and

a document visualizer implemented in software run via a processor that, subsequent to the storing:

generates a retrieval representation of the source document, the retrieval representation including retrieval content of the source document that is identical with the initial content and retrieval layout metadata that is identical with the initial layout metadata, the generating of the retrieval representation not comprising retrieving the initial representation of the source document,

renders the retrieval content using the retrieval layout metadata, and

renders at least one of the document annotations in conjunction with the rendering of the retrieval content based on the retrieval layout metadata and the assigned positional tag of the at least one document annotation.

2. The document annotator as set forth in claim 1 , wherein the document converter is configured to generate the initial representation of the source document having an XML document format including the layout metadata.

3. The document annotator as set forth in claim 2 , wherein the document converter is configured to receive the source document in at least one of HTML, PDF, and a native word processing application format.

4. The document annotator as set forth in claim 1 , wherein the at least one annotation pipeline is configured to generate document annotations in conformance with a document ontology.

5. The document annotator as set forth in claim 4 , wherein the document ontology includes annotations indicative of at least one of (i) the author of the source document; (ii) the title of the source document; and (iii) at least one topic or subject of the source document.

6. The document annotator as set forth in claim 1 , wherein the content includes textual content, and the at least one annotation pipeline includes at least one semantic processing pipeline.

7. The document annotator as set forth in claim 1 , wherein the content includes image content, and the at least one annotation pipeline includes at least one image classifier pipeline.

8. The document annotator as set forth in claim 1 , wherein the at least one annotation pipeline includes at least one of:

an autonomous semantic annotation pipeline configured to autonomously generate semantic content annotations respective to received content, and

an image annotation pipeline including an image classifier configured to generate an image annotation comprising an image classification.

9. The apparatus as set forth in claim 3 , wherein:

the catalog data storage is configured to store the document annotations and their assigned positional tags and a pointer to the source document; and

the generating of the retrieval representation of the source document performed by the document visualizer includes retrieving the source document using the stored pointer.

10. A document annotation method comprising:

generating an initial representation of a source document, the initial representation including initial content of the source document and initial layout metadata indicative of layout of said content in the source document;

processing at least some of the initial content of the source document to generate document annotations;

assigning positional tags of the initial layout metadata to the generated document annotations to locate the document annotations respective to the initial layout metadata;

storing at least the document annotations and their assigned positional tags;

subsequent to the storing, generating a retrieval representation of the source document, the retrieval representation including retrieval content of the source document that is identical with the initial content and retrieval layout metadata that is identical with the initial layout metadata, the generating of the retrieval representation not comprising retrieving the initial representation of the source document;

rendering the retrieval content using the retrieval layout metadata; and

rendering at least one of the document annotations in conjunction with the rendering of the retrieval content based on the retrieval layout metadata and the assigned positional tag of the at least one document annotation.

11. The document annotation method as set forth in claim 10 , wherein the generating an initial representation comprises:

generating an XML representation of the source document including the initial content of the source document and the initial layout metadata represented as metadata of the XML representation.

12. The document annotation method as set forth in claim 10 , wherein the content includes image content, and the processing comprises:

performing image classification to generate image class document annotations.

13. The document annotation method as set forth in claim 10 , wherein the content includes textual content, and the processing comprises:

performing semantic processing to generate semantic document annotations.

14. The document annotation method as set forth in claim 10 , wherein the storing at least the document annotations and their assigned positional tags further includes:

storing the document annotations and their assigned positional tags and a pointer to the source document;

wherein the generating of the retrieval representation of the source document includes retrieving the source document using the stored pointer.

15. The document annotation method as set forth in claim 10 , wherein the processing at least some of the initial content of the source document to generate document annotations includes at least one of:

autonomously generating semantic content annotations respective to received content, and

applying an image classifier to an image of the initial content of the source document generate an image annotation for the image comprising an image classification.

16. An apparatus comprising:

a document converter configured to generate an initial representation of a source document, the initial representation including initial content of the source document and initial layout metadata indicative of layout of said content in the source document;

at least one annotation pipeline configured to process at least some of the initial content of the source document to generate document annotations;

a merger configured to assign positional tags of the initial layout metadata to the generated document annotations to locate the document annotations respective to the initial layout metadata;

catalog data storage configured to store at least the document annotations and their assigned positional tags; and

a document visualizer implemented in software run via a processor that, subsequent to the storing:

generates a retrieval representation of the source document, the retrieval representation including retrieval content of the source document that is identical with the initial content and retrieval layout metadata that is identical with the initial layout metadata,

renders the retrieval content using the retrieval layout metadata, and

renders at least one of the document annotations in conjunction with the rendering of the retrieval content based on the retrieval layout metadata and the assigned positional tag of the at least one document annotation.

17. The apparatus as set forth in claim 16 , wherein the content includes image content, and the processing performed by the at least one annotation pipeline comprises:

performing image classification to generate image class document annotations.

18. The apparatus as set forth in claim 16 , wherein the content includes textual content, and the processing performed by the at least one annotation pipeline comprises:

performing semantic processing to generate semantic document annotations.

19. The apparatus as set forth in claim 16 , wherein the at least one annotation pipeline includes at least one of:

an autonomous semantic annotation pipeline configured to autonomously generate semantic content annotations respective to received content, and

an image annotation pipeline including an image classifier configured to generate an image annotation comprising an image classification.

20. The apparatus as set forth in claim 16 , wherein the document converter is configured to generate the initial representation of the source document having an XML document format including the layout metadata.

21. The document annotator as set forth in claim 20 , wherein the document converter is configured to receive the source document in at least one of HTML, PDF, and a native word processing application format.

Assignments (8)
SECOND LIEN NOTES PATENT SECURITY AGREEMENT Recorded Jul 2, 2025
From: XEROX CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 071785/0550 →
FIRST LIEN NOTES PATENT SECURITY AGREEMENT Recorded Apr 11, 2025
From: XEROX CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 070824/0001 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS RECORDED AT RF 064760/0389 Recorded Feb 13, 2024
From: CITIBANK, N.A., AS COLLATERAL AGENT
To: XEROX CORPORATION
Reel/Frame 068261/0001 →
SECURITY INTEREST Recorded Nov 20, 2023
From: XEROX CORPORATION
To: JEFFERIES FINANCE LLC, AS COLLATERAL AGENT
Reel/Frame 065628/0019 →
SECURITY INTEREST Recorded Jun 22, 2023
From: XEROX CORPORATION
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 064760/0389 →
RELEASE OF SECURITY INTEREST IN PATENTS AT R/F 062740/0214 Recorded May 18, 2023
From: CITIBANK, N.A., AS AGENT
To: XEROX CORPORATION
Reel/Frame 063694/0122 →
SECURITY INTEREST Recorded Nov 10, 2022
From: XEROX CORPORATION
To: CITIBANK, N.A., AS AGENT
Reel/Frame 062740/0214 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2007
From: JACQUIN, THIERRY; CHANOD, JEAN-PIERRE
To: XEROX CORPORATION
Reel/Frame 019058/0049 →
Continuity (1)
Related Publication 20070294614A1 · Dec 20, 2007