IP Library Granted Patent US 11,188,598
Granted Patent B2
US 11,188,598 · App. 16/893,077 · Granted Nov 30, 2021

Document data processing apparatus and non-transitory computer readable medium

Inventor: Tadafumi Kawaguchi (Kanagawa, JP)
Assignee: FUJIFILM Business Innovation Corp.
G06F16/93G06F16/3334G06F16/9035G06F16/9038G06F16/90344G06F17/175
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,188,598
App. No.
16/893,077
Granted
Nov 30, 2021
Kind
B2
Abstract

A document data processing apparatus includes a memory and a processor. The memory stores a distributed-representation set including multiple distributed representations corresponding to multiple pieces of data. The processor is configured to modify the distributed-representation set on the basis of multiple data pairs and multiple scores corresponding to the data pairs. The data pairs are subjected to learning. The processor is configured to modify the distributed-representation set in such a manner that, for each of the data pairs, a value indicating a relationship in a modified distributed-representation pair corresponding to the data pair comes close to a score corresponding to the data pair.

Claims (40)

1. A document data processing apparatus comprising:

a memory that stores a distributed-representation set including a plurality of distributed representations corresponding to a plurality of pieces of data; and

a processor configured to

modify the distributed-representation set on a basis of a plurality of data pairs and a plurality of scores corresponding to the plurality of data pairs, the plurality of data pairs being subjected to learning,

wherein the processor is configured to modify the distributed-representation set in such a manner that, for each data pair of the plurality of data pairs, a value indicating a relationship in a modified distributed-representation pair corresponding to the data pair comes close to a score corresponding to the data pair.

2. The document data processing apparatus according to claim 1 ,

wherein the processor is configured to modify the distributed-representation set in such a manner that a loss calculated by using a loss function is minimized, and

wherein the loss function involves calculation in which, for each data pair, the value indicating the relationship is subtracted from the score.

3. The document data processing apparatus according to claim 2 ,

wherein the value indicating the relationship is an inner product of two distributed representations included in the modified distributed-representation pair, and

wherein the score is a target value compared with the inner product.

4. The document data processing apparatus according to claim 1 ,

wherein the processor is configured to calculate the score for each data pair on a basis of a plurality of sub-scores defined for the data pair.

5. The document data processing apparatus according to claim 4 ,

wherein the processor is configured to calculate the score through weighted addition of the plurality of sub-scores.

6. The document data processing apparatus according to claim 5 ,

wherein the processor is configured to change a weight list on a basis of a user instruction or a document category, the weight list being used in the weighted addition.

7. The document data processing apparatus according to claim 2 ,

wherein the loss function further involves calculation, for each data pair, using a modified distributed-representation pair corresponding to a negative data pair specified by the data pair.

8. The document data processing apparatus according to claim 1 ,

wherein the plurality of data pairs are a plurality of data pairs with scores, the plurality of data pairs being subjected to learning,

wherein the memory stores a table including the plurality of data pairs with scores, and

wherein the processor is configured to refer to the table.

9. The document data processing apparatus according to claim 1 ,

wherein the plurality of data pairs are a plurality of data pairs with sub-score sets, the plurality of data pairs being subjected to learning,

wherein the memory stores a table including the plurality of data pairs with sub-score sets, and

wherein the processor is configured to refer to the table.

10. The document data processing apparatus according to claim 1 ,

wherein the processor is configured to specify the plurality of data pairs on a basis of a plurality of queries which are input in searching documents.

11. The document data processing apparatus according to claim 1 ,

wherein the processor is configured to

generate a recommendation list on a basis of the distributed-representation set, the recommendation list including one or more pieces of related data, the related data being related to data which is input by a user, and

present the recommendation list to a user.

12. A non-transitory computer readable medium storing a program causing a computer to execute a process for processing document data, the process comprising:

referring to a score for each data pair subjected to learning, the score corresponding to the data pair; and

modifying a distributed-representation table including a plurality of distributed representations corresponding to a plurality of pieces of data, the modification being performed in such a manner that, for each data pair subjected to learning, a value indicating a relationship in a modified distributed-representation pair corresponding to the data pair comes close to the score corresponding to the data pair.

13. A document data processing apparatus comprising:

means for storing a distributed-representation set including a plurality of distributed representations corresponding to a plurality of pieces of data; and

means for modifying the distributed-representation set on a basis of a plurality of data pairs and a plurality of scores corresponding to the plurality of data pairs, the plurality of data pairs being subjected to learning,

wherein the distributed-representation set is modified in such a manner that, for each data pair of the plurality of data pairs, a value indicating a relationship in a modified distributed-representation pair corresponding to the data pair comes close to a score corresponding to the data pair.

Assignments (2)
CHANGE OF NAME Recorded Apr 28, 2021
From: FUJI XEROX CO., LTD.
To: FUJIFILM BUSINESS INNOVATION CORP.
Reel/Frame 056078/0098 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 4, 2020
From: KAWAGUCHI, TADAFUMI
To: FUJI XEROX CO., LTD.
Reel/Frame 052843/0411 →
Priority Claims (1)
JP JP2019-226129 · Dec 16, 2019 · national
Continuity (1)
Related Publication 20210182345A1 · Jun 17, 2021