IP Library Granted Patent US 8,275,781
Granted Patent B2
US 8,275,781 · App. 12/357,469 · Granted Sep 25, 2012

Processing documents by modification relation analysis and embedding related document information

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,275,781
App. No.
12/357,469
Granted
Sep 25, 2012
Kind
B2
Abstract

A document processing apparatus is provided that facilitates location of elements within a document to be modified. To this end, the document processing apparatus analyzes document data for modification relations in character strings or between character strings within the document data, and embeds attribute tags within the text document data representing the modification relations. An XML document having the embedded attribute tags is stored in a data storage area, and can subsequently be searched using the embedded tags as search keys.

Claims (65)

1. A document processing apparatus comprising:

an extracting unit that extracts text document data from first document data contained in a first file;

an analyzing unit that identifies a modification relation between a first character string and a second character string included in the text document data;

an attribute embedding unit that embeds an attribute in the text document data, the attribute representing the modification relation;

a document specifying unit that identifies the second character string as being a document-specifying character string that specifies second document data contained in a second file differing from the first file;

a document-identification unit that embeds a document tag in the text document data, the document tag tagging the second character string as the document-specifying character string;

a receiving unit that receives an input character string;

a determining unit that determines that the input character string matches the first character string, and, in response to determination that the input character string matches the first character string, identifies the document-specifying character string having the modification relation with the first character string based on the attribute and the document tag embedded in the text document data;

an identifying unit that identifies the second document data contained in the second file specified by the document-specifying character string in response to the determination that the input character string matches the first character string; and

a central processing unit configured to execute at least the determining unit.

2. The apparatus according to claim 1 , further comprising a document obtaining unit that obtains the second document data identified by the identifying unit.

3. The apparatus according to claim 2 , wherein the document obtaining unit further obtains third document data identified by a second document-specifying character string in the text document data, the second document-specifying character string having no modification relation with the first character string.

4. The apparatus according to claim 2 , further comprising:

a type determining unit that determines a type of the text document data, and embeds the type in the text document data; and

a display unit that displays the second document data obtained by the document obtaining unit while classifying in units of types embedded in the text document data.

5. The apparatus according to claim 1 , further comprising:

a candidate display unit that displays a candidate character string associated with the attribute;

a selection receiving unit that receives a selection of the candidate character string displayed by the candidate display unit; and

a search unit that searches for one or more documents including the candidate character string and retrieves the one or more documents.

6. The apparatus according to claim 5 , further comprising:

a candidate extracting unit that extracts a set of candidate character strings having a modification relation with the candidate character string from each of the one or more documents retrieved by the searching unit, wherein

the candidate display unit further displays the set of candidate character string for selection.

7. The apparatus according to claim 1 , further comprising:

a link-name-information embedding unit that embeds, within the text document data, link identification information indicating whether the second document data is indicated in the text document data, wherein

the identifying unit identifies the second text document data represented by the document-specifying character string based on the link identification information in response to a determination that the text document data includes the document-specifying character string.

8. The apparatus according to claim 1 , wherein the document specifying unit identifies the second character string as being the document-specifying character string based on at least one of a document name, the document tag, or a clause or phrase in the second document data.

9. The apparatus according to claim 1 , wherein the analyzing unit identifies the first character string and the second character string from among an actor, an object, or an action performed by the actor for the modification relation.

10. A document processing method comprising:

extracting text document data from first document data of a first file;

analyzing a modification relation between a first character string and a second character string included in the text document data;

embedding an attribute in the text document data, the attribute indicating the modification relation;

identifying the second character string as being a document-specifying character string that specifies second document data of a second file differing from the first file;

embedding a document tag in the text document data, the document tag identifying the document-specifying character string;

receiving an input character string;

determining that the input character string matches the first character string in the text document data;

in response to the determining, identifying the document-specifying character string having the modification relation with the first character string based on the attribute and the document tag; and

identifying, in response to the identifying the document-specifying character string, the second document data represented by the document-specifying character string.

11. The method according to claim 10 , further comprising obtaining the second document data in response to the identifying the second document data.

12. The method according to claim 11 , wherein the obtaining further includes obtaining third document data corresponding to a second document-specifying character string in the text document data, the second document-specifying character string having no modification relation with the first character string.

13. The method according to claim 11 , further comprising:

determining a type of the text document data, and embedding the type in the text document data; and

displaying the second document data classified according to type based on the type embedded in the text document data.

14. The method according to claim 10 , further comprising:

displaying, as a candidate, a candidate character string associated with the attribute;

receiving a selection of the candidate character string; and

searching for the candidate character string in the text document data.

15. The method according to claim 14 , further comprising:

extracting, as a set of selection candidates, a set of extracted character strings having a modification relation with the candidate character string, from each of plural pieces of the text document data retrieved by the searching; and

displaying the set of extracted character string as the set of selection candidates.

16. The method according to claim 10 , further comprising:

embedding link identification information in the text document data, wherein

the identifying the second document data includes identifying the second document data based on the link identification information, in response to determining that the embedded text document data includes the document-specifying character string.

17. The method according to claim 10 , further comprising identifying the document-specifying character string based on at least one of a document name, the document tag, or a clause or phrase in the second document data.

18. A computer program product having a non-transitory computer-readable medium including programmed instructions for processing text information, wherein the instructions, in response to execution, cause a computing system to perform operations, including:

extracting text document data from first document data of a first file;

identifying a modification relation between a first character string and a second character string included in the text document data;

embedding an attribute in the text document data, the attribute representing the modification relation;

identifying the second character string as a document-specifying character string that specifies second document data of a second file differing from the first file;

embedding a document tag in the text document data, the document tag identifying the document-specifying character string;

receiving an input character string;

determining that the input character string matches the first character string in the text document;

in response to the determining, identifying the document-specifying character string having the modification relation with the first character string based on the attribute and the document tag embedded in the text document data; and

in response to the identifying the document-specifying character string, identifying the second document data represented by the document-specifying character string.

19. The computer program product according to claim 18 , the operations further including obtaining the second document data in response to the identifying the second document data.

20. The computer program product according to claim 19 , wherein the obtaining further includes obtaining third document data identified by a second document-specifying character string in the text document data, the second document-specifying character string having no modification relation with the first character string.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY'S ADDRESS PREVIOUSLY RECORDED ON REEL 048547 FRAME 0187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNORS INTEREST. Recorded May 6, 2020
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 052595/0307 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ADD SECOND RECEIVING PARTY PREVIOUSLY RECORDED AT REEL: 48547 FRAME: 187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 13, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: KABUSHIKI KAISHA TOSHIBA; TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 050041/0054 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 048547/0187 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 13, 2009
From: FUME, KOSEI
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 022537/0921 →