IP Library Granted Patent US 9,753,905
Granted Patent B2
US 9,753,905 · App. 14/721,340 · Granted Sep 5, 2017

Generating a document structure using historical versions of a document

Inventors: Sheng Hua Bao (San Jose, CA); HongLei Guo (Beijing, CN); Zhili Guo (Beijing, CN); Davide Pasetto (Bedford Hills, NY); Wei Hong Qian (Beijing, CN); Zhong Su (Beijing, CN)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F17/2288G06F17/2211G06F17/2241
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,753,905
App. No.
14/721,340
Granted
Sep 5, 2017
Kind
B2
Abstract

A method and apparatus for generating a document structure. The method includes the steps of: aligning various parts in the first version and the second version in at least one pair of historical versions among a plurality of historical versions of a document; dividing the first version and the second version into a plurality of blocks on the basis of a Levenshtein distance between the aligned parts; evaluating a level of the block in the document structure on the basis of text features of the block among the plurality of blocks; and generating the document structure on the basis of a position of the block according to an evaluation result. An apparatus for generating a document structure is also provided. According to the present invention, document structures can be more conveniently and rapidly generated.

Claims (53)

1. A method for generating a document structure, comprising:

aligning various parts in the first version and the second version in at least one pair of historical versions among a plurality of historical versions of a document;

dividing the first version and the second version into a plurality of blocks on the basis of a Levenshtein distance between the aligned parts in the first version and the second version;

evaluating a level of the block in the document structure on the basis of text features of the block among the plurality of blocks; and

generating the document structure on the basis of a position of the block according to an evaluation result.

2. The method according to claim 1 , wherein dividing the first version and the second version into a plurality of blocks on the basis of a Levenshtein distance between the aligned parts in the first version and the second version comprises:

in response to the Levenshtein distance being greater than a predetermined threshold, identifying the parts as difference blocks; and

in response to the Levenshtein distance being less than or equal to a predetermined threshold, identifying the parts as match blocks.

3. The method according to claim 1 , wherein evaluating a level of the block in the document structure on the basis of text features of the block among the plurality of blocks comprises:

analyzing an association relationship on the basis of a predetermined rule so as to evaluate a probability score that a block among the plurality of blocks acts as a level in the document structure.

4. The method according to claim 3 , wherein the text features describe an association relationship between text in the block and text in at least one historical version among the plurality of historical versions.

5. The method according to claim 4 , wherein the generating the document structure on the basis of a position of the block according to an evaluation result comprises:

in response to the evaluation result meeting a predetermined condition, identifying the position of the block as a level which is associated with the predetermined condition in the document structure.

6. The method according to claim 5 , wherein the text features of the block comprise at least one of:

number of the block's occurrences in the plurality of historical versions;

size of the block;

positions of the block in pages of the plurality of historical versions;

number of differences in the block;

distribution style of differences in the block; and

positions of differences in the block.

7. The method according to claim 6 , wherein evaluating a level of the block in the document structure on the basis of text features of the block among the plurality of blocks comprises:

with respect to at least one aspect of the text features, calculating a score component; and

calculating the probability score of the block on the basis of the score component.

8. The method according claim 1 , wherein aligning various parts in the first version and the second version comprises:

aligning the various parts in a minimum unit of lines.

9. The method according to claim 1 , wherein the plurality of historical versions are plain text files.

10. An apparatus for generating a document structure, comprising:

an aligning module configured to align various parts in the first version and the second version in at least one pair of historical versions among a plurality of historical versions of a document;

a dividing module configured to divide the first version and the second version into a plurality of blocks on the basis of a Levenshtein distance between the aligned parts in the first version and the second version;

an evaluating module configured to evaluate a level of the block in the document structure on the basis of text features of the block among the plurality of blocks; and

a generating module configured to generate the document structure on the basis of a position of the block according to an evaluation result.

11. The apparatus according to claim 10 , wherein the dividing module comprises:

a first identifying module configured to, in response to the Levenshtein distance being greater than a predetermined threshold, identify the parts as difference blocks; and

a second identifying module configured to, in response to the Levenshtein distance being less than or equal to a predetermined threshold, identify the parts as match blocks.

12. The apparatus according to claim 10 , wherein the evaluating module comprises:

a probability evaluating module configured to analyze an association relationship on the basis of a predetermined rule so as to evaluate a probability score that a block among the plurality of blocks acts as a level in the document structure.

13. The apparatus according to claim 10 , wherein the text features describe an association relationship between text in the block and text in at least one historical version among the plurality of historical versions.

14. The apparatus according to claim 13 , wherein the generating module comprises:

a level generating module configured to, in response to the evaluation result meeting a predetermined condition, identify the position of the block as a level which is associated with the predetermined condition in the document structure.

15. The apparatus according to claim 14 , wherein the text features of the block comprise at least one of:

number of the block's occurrences in the plurality of historical versions;

size of the block;

positions of the block in pages of the plurality of historical versions;

number of differences in the block;

distribution style of differences in the block; and

positions of differences in the block.

16. The apparatus according to claim 15 , wherein the probability evaluating module comprises:

a first calculating module configured to, with respect to at least one respect of the text features, calculate a score component; and

a second calculating module configured to calculate the probability score of the block on the basis of the score component.

17. The apparatus according to claim 10 , wherein the aligning module comprises:

a line aligning module configured to align the various parts in a minimum unit of lines.

18. The apparatus according to claim 10 , wherein the plurality of historical versions are plain text files.

19. A computer readable non-transitory article of manufacture tangibly embodying computer readable instructions, which when executed, cause a computer to carry out the steps of the method according to claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2015
From: BAO, SHENG HUA; GUO, HONGLEI; GUO, ZHILI; PASETTO, DAVIDE; QIAN, WEI HONG; SU, ZHONG
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 036256/0367 →
Priority Claims (1)
CN 2014 1 0225252 · May 26, 2014 · national
Continuity (1)
Related Publication 20150339278A1 · Nov 26, 2015