IP Library Granted Patent US 9,164,964
Granted Patent B2
US 9,164,964 · App. 13/742,473 · Granted Oct 20, 2015

Context-aware text document analysis

Inventors: Hiroaki Kikuchi (Yokohama, JP); Masaki Komedani (Yokohama, JP); Takuma Murakami (Tokyo, JP); Fumihiko Terui (Tokyo, JP)
Assignee: International Business Machines Corporation
G06F17/21G06F17/274G06F17/278
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,164,964
App. No.
13/742,473
Granted
Oct 20, 2015
Kind
B2
Abstract

An analysis device for analyzing a text document is provided. The analysis device includes a context storage unit configured to store context information that shows a position of a character set of a predetermined context in the text document. The analysis device also includes an index storage unit configured to store index information that shows a position of a word in the text document, for each word of a plurality of words contained in the text document. An input unit is configured to input a target word. A position detection unit is configured to detect from the index information a position of the target word contained in the text document. A frequency detection unit is configured to detect an appearance frequency of the target word per each type of context in the text document based on the position of the target word and on the context information.

Claims (50)

1. An analysis device for analyzing a text document, the analysis device comprising:

a memory having computer readable instructions; and

a processor for executing the computer readable instructions, the computer readable instructions comprising:

reading the text document from an external source;

forming context information by analyzing the text document for sentences based on punctuation marks and contents of words or language in the sentences to determine a position of a character set of a predetermined context in the text document;

storing the context information to a context storage unit;

storing in an index storage unit, index information that shows a position of a word in the text document, for each word of a plurality of words contained in the text document;

inputting a target word;

detecting from the index information read from the index storage unit, a position of the target word contained in the text document;

detecting an appearance frequency of the target word per each type of context in the text document based on comparing the position of the target word with the position of the character set of the predetermined context from the context information stored in the context storage unit, and counting a quantity of context information occurrences for each type of context in the text document for the target word; and

outputting the appearance frequency for each type of context in the text document for the target word based on the position of the target word and the context information.

2. The analysis device according to claim 1 , further comprising:

analyzing the text document;

forming index information for each word of the plurality of words contained in the text document; and

storing the formed index information in the index storage unit.

3. The analysis device according to claim 1 , wherein the context information is formed for a plurality of predetermined contexts.

4. The analysis device according to claim 2 , wherein the context information is formed for a plurality of predetermined contexts.

5. The analysis device according to claim 3 , wherein the context information includes sections having appearing positions that overlap with each other.

6. A computer-implemented method for analyzing a text document, the computer-implemented method comprising:

reading the text document from an external source;

forming context information by analyzing the text document for sentences based on punctuation marks and contents of words or language in the sentences to determine a position of a character set of a predetermined context in the text document;

storing the context information to a context storage unit;

storing in an index storage unit, index information that shows a position of a word in the text document, for each word of a plurality of words contained in the text document;

inputting a target word;

detecting from the index information read from the index storage unit, a position of the target word contained in the text document; and

detecting an appearance frequency of the target word per each type of context in the text document based on comparing the position of the target word with the position of the character set of the predetermined context from the context information stored in the context storage unit, and counting a quantity of context information occurrences for each type of context in the text document for the target word; and

outputting the appearance frequency for each type of context in the text document for the target word based on the position of the target word and the context information.

7. The computer-implemented method according to claim 6 , further comprising:

analyzing the text document;

forming index information for each word of the plurality of words contained in the text document; and

storing the formed index information.

8. The computer-implemented method according to claim 6 , wherein the context information is formed for a plurality of predetermined contexts.

9. The computer-implemented method according to claim 7 , wherein the context information is formed for a plurality of predetermined contexts.

10. The computer-implemented method according to claim 8 , wherein the context information includes sections having appearing positions that overlap with each other.

11. A computer program product for analyzing a text document, the computer program product comprising a computer readable non-transitory storage medium having computer readable program code embodied therewith, the computer readable program code comprising computer readable program code configured for:

reading the text document from an external source;

forming context information by analyzing the text document for sentences based on punctuation marks and contents of words or language in the sentences to determine a position of a character set of a predetermined context in the text document;

storing the context information to a context storage unit;

storing in an index storage unit, index information that shows a position of a word in the text document, for each word of a plurality of words contained in the text document;

inputting a target word;

detecting from the index information read from the index storage unit, a position of the target word contained in the text document;

detecting an appearance frequency of the target word per each type of context in the text document based on comparing the position of the target word with the position of the character set of the predetermined context from the context information stored in the context storage unit, and counting a quantity of context information occurrences for each type of context in the text document for the target word; and

outputting the appearance frequency for each type of context in the text document for the target word based on the position of the target word and the context information.

12. The computer program product according to claim 11 , further comprising:

analyzing the text document;

forming index information for each word of the plurality of words contained in the text document; and

storing the formed index information.

13. The computer program product according to claim 11 , wherein the context information is formed for a plurality of predetermined contexts.

14. The computer program product according to claim 12 , wherein the context information is formed for a plurality of predetermined contexts.

15. The computer program product according to claim 13 , wherein the context information includes sections having appearing positions that overlap with each other.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 15, 2021
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: DOORDASH, INC.
Reel/Frame 057826/0939 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 16, 2013
From: KIKUCHI, HIROAKI; KOMEDANI, MASAKI; MURAKAMI, TAKUMA; TERUI, FUMIHIKO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 029637/0395 →
Priority Claims (1)
JP 2012-032067 · Feb 16, 2012 · national
Continuity (1)
Related Publication 20130218555A1 · Aug 22, 2013