IP Library Granted Patent US 9,904,677
Granted Patent B2
US 9,904,677 · App. 14/837,197 · Granted Feb 27, 2018

Data processing device for contextual analysis and method for constructing script model

Inventor: Shinichiro Hamada (Kanagawa, JP)
Assignees: KABUSHIKI KAISHA TOSHIBA; TOSHIBA SOLUTIONS CORPORATION
G06F17/28G06F17/27G06F17/2785G06K9/00469
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,904,677
App. No.
14/837,197
Granted
Feb 27, 2018
Kind
B2
Abstract

According to an embodiment, a data processing device includes an extractor, a generator, and a constructor. The extractor is configured to extract, from a document having been subjected to predicate argument structure analysis and anaphora resolution, an element sequence including elements each being a combination of predicate having a shared argument and case type information of the shared argument, together with the shared argument. The generator is configured to produce case example data expressed by a feature vector for each attention element which is one of the elements. The feature vector includes feature value(s) about a sub-sequence having the attention element and feature value(s) about a sequence of the shared argument corresponding to the sub-sequence. The constructor is configured to construct a script model for estimating the elements each following antecedent context by performing machine learning based on a discriminative model using the case example data.

Claims (17)

1. A data processing device comprising:

processing circuitry configured to function as:

an extractor configured to extract, from a document having been subjected to predicate argument structure analysis and anaphora resolution, an element sequence in which a plurality of elements are arranged in order of appearances of predicates in the document, each of the elements being a combination of a predicate having a shared argument and case type information indicating a type of a case of the shared argument, together with the shared argument;

a case example generator configured to produce case example data for training expressed by a feature vector for each attention element, the attention element being one of the elements included in the extracted element sequence, the feature vector including one or more feature values about a sub-sequence having the attention element as a last element of the sub-sequence in the extracted element sequence and one or more feature values about a sequence of the shared argument corresponding to the sub-sequence; and

a model constructor configured to train a script model using the produced case example data and employing a discriminative model training technique; and

a text analyzer configured to (1) analyze a target document, (2) predict unknown elements in the target document using the trained script model, and (3) output the prediction results of the unknown elements.

2. The data processing device according to claim 1 , wherein the case example generator is configured to produce, for each attention element, the case example data for training expressed by the feature vector further including one or more feature values about a combination obtained by a logical product of the shared argument and the sub-sequence in which a part of the elements is replaced with a wild card.

3. The data processing device according to claim 1 , wherein the one or more feature values about the sequence of the shared argument are one or more feature values that discriminate the shared argument using at least one of a surface or a normalized surface, information about a grammatical category, information about a semantic category, and information about a named entity type.

4. The data processing device according to claim 1 , wherein the sub-sequence includes a unigram sequence having only the attention element as a single element.

5. The data processing device according to claim 1 , wherein word sense identification information that identifies a word sense of the predicate is added to the predicate included in each of the elements.

6. A script model construction method that is implemented by a data processing device, the script model construction method comprising:

extracting, from a document having been subjected to predicate argument structure analysis and anaphora resolution, an element sequence in which a plurality of elements are arranged in order of appearances of predicates in the document, each of the elements being a combination of a predicate having a shared argument and case type information indicating a type of a case of the shared argument, together with the shared argument;

producing case example data for training expressed by a feature vector for each attention element, the attention element being one of the elements included in the extracted element sequence, the feature vector including one or more feature values about a sub-sequence having the attention element as a last element of the sub-sequence in the extracted element sequence and one or more feature values about a sequence of the shared argument corresponding to the sub-sequence;

training a script model using the produced case example data and employing a discriminative model training technique;

analyzing a target document;

predicting unknown elements in the target document using the trained script model; and

outputting the prediction results of the unknown elements.

Assignments (3)
CHANGE OF NAME Recorded Feb 8, 2021
From: TOSHIBA SOLUTIONS CORPORATION
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 055185/0728 →
CHANGE OF NAME Recorded Mar 8, 2019
From: TOSHIBA SOLUTIONS CORPORATION
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 048547/0215 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2015
From: HAMADA, SHINICHIRO
To: KABUSHIKI KAISHA TOSHIBA; TOSHIBA SOLUTIONS CORPORATION
Reel/Frame 036437/0314 →
Continuity (2)
Continuation PCTJP2013055477 · Feb 28, 2013
Related Publication 20160012040A1 · Jan 14, 2016