IP Library Granted Patent US 10,417,337
Granted Patent B2
US 10,417,337 · App. 15/253,548 · Granted Sep 17, 2019

Devices, systems, and methods for resolving named entities

Inventors: Dariusz T. Dusberger (Irvine, CA); Quentin Dietz (Irvine, CA)
Assignee: Canon Kabushiki Kaisha
G06F17/278G06F16/313G06F16/35G06F17/277
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,417,337
App. No.
15/253,548
Granted
Sep 17, 2019
Kind
B2
Abstract

An information processing apparatus to select a token from a document to describe a field of interest includes an obtaining unit, a determining unit, a clustering unit, and a selecting unit. The obtaining unit obtains a list of tokens output from extractors that received the document as an input. Each output token has an extractor score assigned to by an extractor. The determining unit determines, as a word frequency value, a frequency of each word in the list of tokens, determines a token score for each token in the list of tokens, and determines a distance between each token in the list of tokens. The clustering unit clusters each token in the list of tokens into a plurality of groups. The selecting unit selects a token with a group of the plurality of groups to describe the field of interest in the document.

Claims (24)

1. A method to be performed by an information processing apparatus to select a token from a document to describe a field of interest in the document, the method comprising:

obtaining a list of tokens output from a plurality of extractors that received the document as an input, wherein each output token has an extractor score assigned by an extractor of the plurality of extractors;

determining, as a word frequency value, a frequency of each word in the list of tokens, a token score for each token in the list of tokens, and a distance between each token in the list of tokens;

clustering each token in the list of tokens into a plurality of groups;

determining a cluster score of each group, wherein the cluster score of a group is determined by dividing a token score for a first token in the group by one plus a distance between the first token and a center token of the group; and

selecting, based on the determined cluster score of each group, a token with a group of the plurality of groups to describe the field of interest in the document.

2. The method according to claim 1 , wherein determining the token score of a first token includes multiplying an extractor score by a sum of word frequency values for the first token.

3. The method according to claim 1 , wherein determining the distance between two tokens includes taking word sequence/order into account and dividing a sum of word frequency values of words that are different between the two tokens by a sum of word frequency values of words that are common between the two tokens.

4. The method according to claim 1 , wherein determining the distance between two tokens includes multiplying a constant value and a sum of word frequency values of words that are different between the two tokens.

5. The method according to claim 1 , wherein clustering each token in the list of tokens includes determining whether a compactness of a group in the plurality of groups is less than a predetermined compactness threshold.

6. The method according to claim 5 , wherein, in a case where the compactness of a first group is greater than the predetermined compactness threshold, clustering includes dividing the first group into two groups.

7. A non-transitory computer-readable storage medium storing a program to cause an information processing apparatus to perform a method to select a token from a document to describe a field of interest in the document, the method comprising:

obtaining a list of tokens output from a plurality of extractors that received the document as an input, wherein each output token has an extractor score assigned by an extractor of the plurality of extractors;

determining, as a word frequency value, a frequency of each word in the list of tokens, a token score for each token in the list of tokens, and a distance between each token in the list of tokens;

clustering each token in the list of tokens into a plurality of groups;

determining a cluster score of each group, wherein the cluster score of a group is determined by dividing a token score for a first token in the group by one plus a distance between the first token and a center token of the group; and

selecting, based on the determined cluster score of each group, a token with a group of the plurality of groups to describe the field of interest in the document.

8. An information processing apparatus to select a token from a document to describe a field of interest in the document, the information processing apparatus comprising:

an obtaining unit configured to obtain a list of tokens output from a plurality of extractors that received the document as an input, wherein each output token has an extractor score assigned by an extractor of the plurality of extractors;

a determining unit configured to determine, as a word frequency value, a frequency of each word in the list of tokens, a token score for each token in the list of tokens, and a distance between each token in the list of tokens;

a clustering unit configured to cluster each token in the list of tokens into a plurality of groups;

a cluster score determining unit configured to determine a cluster score of each group, wherein the cluster score of a group is determined by dividing a token score for a first token in the group by one plus a distance between the first token and a center token of the group;

a selecting unit configured to select, based on the determined cluster score of each group, a token with a group of the plurality of groups to describe the field of interest in the document; and

at least one processor coupled to a memory, wherein the at least one processor implements the obtaining unit, the determining unit, the clustering unit, and the selecting unit.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 31, 2016
From: DUSBERGER, DARIUSZ T.; DIETZ, QUENTIN
To: CANON KABUSHIKI KAISHA
Reel/Frame 039606/0868 →
Continuity (2)
Provisional Application 62213535 · Sep 2, 2015
Related Publication 20170060837A1 · Mar 2, 2017