IP Library › Granted Patent US 12,277,389
Granted Patent B2
US 12,277,389 · App. 17/315,447 · Granted Apr 15, 2025

Text mining based on document structure information extraction

Inventors: Tetsuya Nasukawa (Kawasaki, JP); Shoko Suzuki (Yokohama, JP); Daisuke Takuma (Toshima-ku, JP); Issei Yoshida (Setagaya-ku, JP)
Assignee: International Business Machines Corporation
G06F40/279G06F16/93G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,277,389
App. No.
17/315,447
Granted
Apr 15, 2025
Kind
B2
Abstract

Frequent sequences extracted from a set of documents according to a common rule are obtained. Based on comparing occurrence frequencies of various sequences, confidence of the first frequent sequence being a label expression representing a document part in a target document is evaluated. Keywords are extracted from the target document based on evaluation of the confidence.

Claims (135)

1. A method for mining text by a computer-based text mining system, comprising:

obtaining, by the text mining system, a first frequent sequence of characters from a set of documents, the set of documents having structured contents according to a common rule, wherein the first frequent sequence satisfies a condition of maximality and the satisfying of the condition of maximality comprises:

performing a first comparison, the first comparison comprising comparing a first occurrence frequency of the first frequent sequence to a second occurrence frequency of a second sequence, wherein the second sequence is longer than the first sequence and the second sequence contains the first sequence;

determining that the first frequent sequence includes a symbol and the symbol comprises formatting data for a target document;

decomposing, responsive to the determining the first frequent sequence includes a symbol, the first frequent sequence into a symbol part and a remaining part;

evaluating, by the text mining system and based on the comparing, a first confidence of the first frequent sequence being a label expression, wherein the label expression represents a document part in the target document, the evaluating the first confidence comprises:

calculating a primary confidence value for the first frequent sequence across the set of documents;

computing a likelihood of the symbol being contained in the first frequent sequence observed in the target document; and

adjusting the primary confidence value, resulting in a secondary confidence value for the first frequent sequence within the target document, wherein the adjusting the primary confidence value for the first frequent sequence to obtain the secondary confidence value is based on the likelihood of the symbol being contained in the first frequent sequence;

determining that the first confidence is above a confidence threshold;

extracting, in response to the determining the first confidence is above the confidence threshold and by the text mining system, one or more keywords from the target document based on the secondary confidence value of the first frequent sequence, wherein the extracting the one or more keywords further comprises:

applying keyword extraction to the target document, resulting in a set of keywords;

identifying that a first keyword included in the set of keywords overlaps with the first frequent sequence;

removing, based on the determining and on the identifying, the first keyword from the set of keywords; and

assigning a label relating to the first frequent sequence to a second keyword included in the set of keywords, the assigning based on positions in the target document where the first frequent sequence and the second keyword have appeared; and

outputting the one or more keywords and the secondary confidence value.

2. The method of claim 1 , wherein the obtaining the first frequent sequence further comprises filtering the first frequent sequence based on a filtering condition with respect to at least a number of documents containing the first frequent sequence.

3. The method of claim 1 , wherein the obtaining the first frequent sequence further comprises enumerating an additional sequence based on a predetermined enumeration rule.

4. The method of claim 1 , wherein obtaining a first frequent sequence further comprises:

concatenating characters of each of the set of documents into a character array;

enumerating, based on the character array, each of a set of enumerated sequences observed in the set of documents with an occurrence frequency of each enumerated sequence observed in the set of documents;

performing a second comparison, the second comparison including comparing a third occurrence frequency of a first enumerated sequence from the set of enumerated sequences to a fourth occurrence frequency of a longer enumerated sequence containing the first enumerated sequence; and

designating, based on the second comparison, the first enumerated sequence as the first frequent sequence.

5. The method of claim 1 , wherein the first confidence is evaluated based on:

the first occurrence frequency of the first frequent sequence;

a fifth occurrence frequency of a substring of the first frequent sequence;

a sixth occurrence frequency of a longer sequence containing the first frequent sequence; and

the first confidence is calculated by:

conf( s )=conf 1 ( s )·conf 2 ( s )

where:

c

⁢

o

⁢

n

⁢

f

1

(

s

)

=

freq

⁡

(

s

)

/

freq

⁡

(

x

)

⁢

and

⁢

conf

2

(

s

)

=

(

1

-

freq

⁡

(

y

)

b

·

freq

⁡

(

s

)

)

where conf(s) is the first confidence, the freq(s) is the first occurrence frequency, the freq (x) is the fifth occurrence frequency, the freq (y) is the sixth occurrence frequency, and the b is a weighting factor.

6. The method of claim 1 , wherein:

each document is written in a natural language; and

each frequent sequence is a frequently occurring character sequence.

7. The method of claim 1 , wherein the secondary confidence is increased in response to determining the symbol is included in the first frequent sequence in a target document.

8. The method of claim 1 , wherein the keyword comprises a key, a value, and a category label, the category label is associated with the first frequent sequence, and the category label is the remaining part of the first frequent sequence.

9. A computer based text mining system, comprising:

a memory; and

a processor coupled to the memory, the processor configured to cause the text mining system to:

obtain, by the text mining system, a first frequent sequence of characters from a set of documents, the set of documents having structured contents according to a common rule, wherein the first frequent sequence satisfies a condition of maximality and satisfying the condition of maximality comprises:

performing a first comparison, the first comparison including comparing a first occurrence frequency of the first frequent sequence to a second occurrence frequency of a second sequence, wherein the second sequence is longer than the first sequence and the second sequence contains the first sequence;

determine that the first frequent sequence includes a symbol and the symbol comprises formatting data for a target document;

decompose, responsive to the determining the first frequent sequence includes a symbol, the first frequent sequence into a symbol part and a remaining part;

evaluate, based on the comparing, a first confidence of the first frequent sequence being a label expression, wherein the label expression represents a document part in the target document, the evaluating the first confidence comprises:

calculating a primary confidence value for the first frequent sequence across the set of documents; and

computing a likelihood of the symbol being contained in the first frequent sequence observed in the target document; and

adjusting the primary confidence value, resulting in a secondary confidence value for the first frequent sequence within the target document, wherein the adjusting the primary confidence value for the first frequent sequence to obtain the secondary confidence value is based on the likelihood of the symbol being contained in the first frequent sequence;

determine that the first confidence is above a confidence threshold;

extract, in response to the determining the first confidence is above the confidence threshold and by the text mining system, one or more keywords from the target document based on the secondary confidence value of the first frequent sequence, wherein the extracting the one or more keywords further comprises:

apply keyword extraction to the target document, resulting in a set of keywords;

identify that a first keyword included in the set of keywords overlaps with the first frequent sequence;

remove, based on the determining and on the identifying, the first keyword from the set of keywords; and

assign a label relating to the first frequent sequence to a second keyword included in the set of keywords, the assigning based on positions in the target document where the first frequent sequence and the second keyword have appeared; and

output the one or more keywords and the secondary confidence value.

10. The computer system of claim 9 , wherein the obtaining the first frequent sequence further comprises:

filtering the first frequent sequence based on a filtering condition with respect to at least a number of documents containing the first frequent sequence; and

enumerating an additional sequence based on a predetermined enumeration rule.

11. The computer system of claim 9 , wherein obtaining a first frequent sequence comprises:

concatenate characters of each of the set of documents into a character array;

enumerate, based on the character array, each of a set of enumerated sequences observed in the set of documents with an occurrence frequency of each enumerated sequence observed in the set of documents;

perform a second comparison, the second comparison including comparing a third occurrence frequency of a first enumerated sequence from the set of enumerated sequences to a fourth occurrence frequency of a longer enumerated sequence containing the first enumerated sequence; and

designate, based on the second comparison, the first enumerated sequence as the first frequent sequence.

12. A computer program product for a text mining system, comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to:

obtain, by the text mining system, a first frequent sequence of characters from a set of documents, the set of documents having structured contents according to a common rule, wherein the first frequent sequence satisfies a condition of maximality and satisfying the condition of maximality comprises:

performing a first comparison, the first comparison including comparing a first occurrence frequency of the first frequent sequence to a second occurrence frequency of a second sequence, wherein the second sequence is longer than the first sequence and the second sequence contains the first sequence;

determine that the first frequent sequence includes a symbol and the symbol comprises formatting data for a target document;

decompose, responsive to the determining the first frequent sequence includes a symbol, the first frequent sequence into a symbol part and a remaining part;

evaluate, by the text mining system and based on the comparing, a first confidence of the first frequent sequence being a label, wherein the label represents a document part in the target document, the evaluating the first confidence comprises:

calculating a primary confidence value for the first frequent sequence across the set of documents;

computing a likelihood of the symbol being contained in the first frequent sequence observed in the target document; and

adjusting the primary confidence value, resulting in a secondary confidence value for the first frequent sequence within the target document, wherein the adjusting the primary confidence value for the first frequent sequence to obtain the secondary confidence value is based on the likelihood of the symbol being contained in the first frequent sequence;

determine that the first confidence is above a confidence threshold;

extract, in response to the determining the first confidence is above the confidence threshold and by the text mining system, one or more keywords from the target document based on the secondary confidence value of the first frequent sequence, wherein the extracting the one or more keywords further comprises:

apply keyword extraction to the target document, resulting in a set of keywords;

identify that a first keyword included in the set of keywords overlaps with the first frequent sequence;

remove, based on the determining and on the identifying, the first keyword from the set of keywords; and

assign a label relating to the first frequent sequence to a second keyword included in the set of keywords, the assigning based on positions in the target document where the first frequent sequence and the second keyword have appeared; and

perform, based on the extracting the one or more keywords, a search of the set of documents; and

output the one or more keywords and the secondary confidence value.

13. The computer program product of claim 12 , wherein the obtaining the first frequent sequence further comprises:

filtering the first frequent sequence based on a filtering condition with respect to at least a number of documents containing the first frequent sequence; and

enumerating an additional sequence based on a predetermined enumeration rule.

14. The computer program product of claim 12 , wherein obtaining a first frequent sequence comprises:

concatenate characters of each of the set of documents into a character array;

enumerate, based on the character array, each of a set of enumerated sequences observed in the set of documents with an occurrence frequency of each enumerated sequence observed in the set of documents;

perform a second comparison, the second comparison including comparing a third occurrence frequency of a first enumerated sequence from the set of enumerated sequences to a fourth occurrence frequency of a longer enumerated sequence containing the first enumerated sequence; and

designate, based on the second comparison, the first enumerated sequence as the first frequent sequence.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 10, 2021
From: NASUKAWA, TETSUYA; SUZUKI, SHOKO; TAKUMA, DAISUKE; YOSHIDA, ISSEI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 056182/0584 →
Continuity (1)
Related Publication 20220358287A1 · Nov 10, 2022
References Cited (21)
US 8892422B1 · Kumar · 2014 [cited by examiner]
US 9348920B1 · Kesin · 2016 [cited by examiner]
US 20020116398A1 · Sugaya et al. · 2002 [cited by applicant]
US 20040024739A1 · Copperman · 2004 [cited by examiner]
US 20120278060A1 · Cancedda · 2012 [cited by examiner]
US 20140019438A1 · Le Chevalier · 2014 [cited by examiner]
US 20150161521A1 · Shah · 2015 [cited by examiner]
US 20170132205A1 · Novitskiy · 2017 [cited by examiner]
US 20200320170A1 · Bellert · 2020 [cited by examiner]
US 20210081613A1 · Begun · 2021 [cited by examiner]
JP 2002183175A · 2002 [cited by applicant]
JP 2007128224A · 2007 [cited by applicant]
JP 2009271772A · 2009 [cited by applicant]
JP 2010224823A · 2010 [cited by applicant]
Denny, Joshua C., et al. “Evaluation of a method to identify and categorize section headers in clinical documents.” Journal of the American Medical Informatics Association 16.6 (2009): pp. 806-815. (Year: 2009). [cited by examiner]
Pomares-Quimbaya, et al. “Current approaches to identify sections within clinical narratives from electronic health records: a systematic review.” BMC medical research methodology 19 (2019): pp. 1-20. (Year: 2019). [cited by examiner]
Wei, Mengxi, et al. “Robust layout-aware IE for visually rich documents with pre-trained language models.” Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval.… [cited by examiner]
“Python Strings are Arrays”, available at https://web.archive.org/web/20210509023147/https://www.w3schools.com/python/gloss_python_strings_are_arrays.asp, archived on May 9, 2021) (Year: 2021). [cited by examiner]
“Python Slice Strings”, available at https://web.archive.org/web/20230508083439/https://www.w3schools.com/python/gloss_python_string_slice.asp (archived on May 8, 2021) (Year: 2021). [cited by examiner]
Rajman, Martin, et al. “Text mining: natural language techniques and text mining applications.” Data Mining and Reverse Engineering: Searching for semantics. IFIP Seventh Conference on Database Semantics (DS-7) Oct. 7-1… [cited by examiner]
Matsuo et al., “Keyword Extraction From a Single Document Using Word Co-Occurrence Statistical Information,” World Scientific, International Journal on Artificial Intelligence Tools, vol. 13, No. 1, 2004, 14 pages. [cited by applicant]