IP Library › Granted Patent US 10,387,781
Granted Patent B2
US 10,387,781 · App. 14/747,011 · Granted Aug 20, 2019

Information processing using primary and secondary keyword groups

Inventors: Risa Kawanaka (Tokyo, JP); Issei Yoshida (Tokyo, JP)
Assignee: International Business Machines Corporation
G06N5/022G06F16/35G06N5/003G06N5/048
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,387,781
App. No.
14/747,011
Granted
Aug 20, 2019
Kind
B2
Abstract

An information processing device includes a keyword acquiring unit configured to acquire a plurality of primary keyword and secondary keyword groups; a classifying unit configured to classify each of the plurality of secondary keywords by a plurality of topics; an estimating unit configured to estimate whether or not each primary keyword in the plurality of groups is a related keyword related to any topic having a classified secondary keyword or a mixed keyword unrelated to any of the topics; and an assigning unit configured to preferentially assign a primary keyword estimated to be a related keyword to a topic having a classified secondary keyword in the same group, and assigning a primary keyword estimated to be a mixed keyword to any of all the topics given for classification.

Claims (17)

1. An information processing method executed by a computer, the method comprising:

acquiring a plurality of primary keyword and secondary keyword groups;

acquiring primary document information including more than one primary keyword, and secondary document information including more than one secondary keyword created by a first user;

classifying each of the plurality of secondary keywords by a plurality of topics based on a topic model, the topics being units for grouping a plurality of keywords having a threshold probability of appearing together in document information;

estimating whether each primary keyword in the plurality of groups is a related keyword related to any topic having a classified secondary keyword or a mixed keyword unrelated to any of the topics on the basis of primary topics having assigned primary keywords and secondary topics having assigned secondary keywords, the estimating comprising:

acquiring a primary topic of the primary keyword;

calculating the extent of the topic match being the proportion of secondary topics assigned to one or more secondary keywords that are the same as the primary topic;

acquiring the mixed keyword ratio or the ratio of primary keywords estimated to be mixed keywords among primary keywords included in primary document information from all users; and

calculating the mixed keyword probability of the primary keyword being a mixed keyword on the basis of the extent of the topic match and the mixed keyword ratio; and

assigning a primary keyword estimated to be a related keyword to a topic having a classified secondary keyword in the same group, and assigning a primary keyword estimated to be a mixed keyword to any of the topics for classification instead of a dedicated mixed keyword topic such that the primary keyword can later be assigned to a topic if determined to be a related keyword for a different user than the first user.

2. The method of claim 1 , further comprising:

determining a topic to assign a primary keyword estimated to be a related keyword on the basis of the proportion of secondary keywords classified by each topic in the same group; and

determining a topic to assign a primary keyword estimated to be a mixed keyword irrespective of the proportions.

3. The method of claim 1 , further comprising determining whether a primary keyword is a related keyword or a mixed keyword on the basis of the mixed keyword probability.

4. The method of claim 1 , further comprising calculating the likelihood Ψ kt of the t th primary keyword t, wherein t is an integer >1, appearing in the k th topic, wherein k is a predetermined integer >1, in all sets of primary document information from a user.

5. The method of claim 4 , further comprising generating the probability θ dk of the k th topic k in each set of secondary document information d, wherein d>1 but less than the total number of sets of secondary document information being generated.

6. The method of claim 5 , further comprising calculating a primary keyword generation probability P(t|d, D) of a primary keyword t being assigned to a single set of secondary document information d by totaling the θ dk Ψ kt of each topic k in the set of secondary document data d.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2015
From: KAWANAKA, RISA; YOSHIDA, ISSEI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 035882/0783 →
Priority Claims (1)
JP 2014-067158 · Mar 27, 2014 · national
Continuity (2)
Continuation 14635210 · Mar 2, 2015
Related Publication 20150286930A1 · Oct 8, 2015