IP Library › Granted Patent US 10,360,302
Granted Patent B2
US 10,360,302 · App. 15/706,580 · Granted Jul 23, 2019

Visual comparison of documents using latent semantic differences

Inventor: Robert G. Farrell (Cornwall, NY)
Assignee: International Business Machines Corporation
G06F17/2785G06F16/338G06F16/3323G06F16/3344G06F16/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,360,302
App. No.
15/706,580
Granted
Jul 23, 2019
Kind
B2
Abstract

A method, computer system, and a computer program product for comparing documents using latent semantic differences is provided. The present invention may include receiving documents from a user. The present invention may also include extracting linguistic units associated with the received documents. The present invention may then include building latent semantic dimensions based on the extracted linguistic units. The present invention may then include weighting the extracted linguistic units utilizing the built latent semantic dimensions. The present invention may then include determining latent semantic differences between the received documents based on weighted linguistic units. The present invention may also include mapping the weighted linguistic units to a scaled visual feature. The present invention may further include generating a visualization to the user of the received documents based on the determined latent semantic differences and the scaled visual feature.

Claims (96)

1. A method for comparing documents using latent semantic differences, the method comprising:

receiving a plurality of documents from a user;

extracting a plurality of linguistic units associated with the received plurality of documents,

wherein a canonical unit for each linguistic unit from the extracted plurality of linguistic units is determined,

wherein that at least one variation of each linguistic unit from the extracted plurality of linguistic units are present is determined,

wherein a number of variations of each linguistic unit from the extracted plurality of linguistic units by utilizing a dictionary is determined,

wherein the determined number of variations of each linguistic unit is tracked by utilizing a set of tables,

wherein at least one start position and at least one end position with each linguistic unit from the extracted plurality of linguistic units is stored,

wherein each linguistic unit from the extracted plurality of linguistic units includes a plurality of words in a contiguous sequence;

building a plurality of latent semantic dimensions based on the extracted plurality of linguistic units;

weighting the extracted plurality of linguistic units utilizing the built plurality of latent semantic dimensions;

determining a plurality of latent semantic differences between the received plurality of documents based on weighted plurality of linguistic units;

mapping the weighted plurality of linguistic units to a scaled visual feature; and

generating a visualization to the user of the received plurality of documents based on the determined plurality of latent semantic differences and the scaled visual feature,

wherein a plurality of mark-ups is added to the mapped plurality of linguistic units,

wherein at least one value associated with at least one dimension of the mapped plurality of linguistic units from each of the received plurality of documents is correlated with at least one value associated with a hue, a saturation and a lightness based on the determined plurality of latent semantic differences associated with the at least one dimension of the mapped plurality of linguistic units,

wherein the at least one value associated with the hue, saturation and lightness is translated into a hexadecimal code.

2. The method of claim 1 , wherein building the plurality of latent semantic dimensions based on the extracted plurality of linguistic units, further comprises:

generating a plurality of reduced dimensions to a percentage weight associated with the extracted plurality of linguistic units using deep learning, wherein the deep learning comprises a combination of weights of the extracted plurality of linguistic units associated with the received plurality of documents.

3. The method of claim 1 , wherein generating the visualization to the user of the received plurality of documents based on the presented plurality of latent semantic differences and the scaled visual feature, further comprises:

subtracting at least one latent semantic dimension from one of the built plurality of latent semantic dimensions associated with another one of the received plurality of documents; and

generating a visual representation of the mapped plurality of latent semantic differences associated with the received plurality of documents.

4. The method of claim 3 , further comprising:

determining the user selected to visualize a plurality of latent semantic similarities between the received plurality of documents;

receiving, from the user, a threshold value to define the determined plurality of latent semantic similarities between the received plurality of documents;

comparing the determined threshold value with each of the determined plurality of latent semantic similarities; and

generating a visual representation of the mapped plurality of latent semantic similarities associated with the received plurality of documents.

5. The method of claim 1 , wherein receiving the plurality of documents from the user, further comprises:

uploading the plurality of documents onto a cloud storage by the user; and

selecting, by the user, the plurality of documents from the cloud storage for comparison.

6. The method of claim 1 , wherein mapping the weighted plurality of linguistic units to the scaled visual feature, further comprises:

assigning a hue to each of the latent semantic dimensions of the built plurality of latent semantic dimensions; and

determining a range of lightness and saturation for each assigned hue associated with each latent semantic dimension of the built plurality of latent semantic dimensions.

7. A computer system for comparing documents using latent semantic differences, comprising:

one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage medium, and program instructions stored on at least one of the one or more tangible storage medium for execution by at least one of the one or more processors via at least one of the one or more memories, wherein the computer system is capable of performing a method comprising:

receiving a plurality of documents from a user;

extracting a plurality of linguistic units associated with the received plurality of documents,

wherein a canonical unit for each linguistic unit from the extracted plurality of linguistic units is determined,

wherein that at least one variation of each linguistic unit from the extracted plurality of linguistic units are present is determined,

wherein a number of variations of each linguistic unit from the extracted plurality of linguistic units by utilizing a dictionary is determined,

wherein the determined number of variations of each linguistic unit is tracked by utilizing a set of tables,

wherein at least one start position and at least one end position with each linguistic unit from the extracted plurality of linguistic units is stored,

wherein each linguistic unit from the extracted plurality of linguistic units includes a plurality of words in a contiguous sequence;

building a plurality of latent semantic dimensions based on the extracted plurality of linguistic units;

weighting the extracted plurality of linguistic units utilizing the built plurality of latent semantic dimensions;

determining a plurality of latent semantic differences between the received plurality of documents based on weighted plurality of linguistic units;

mapping the weighted plurality of linguistic units to a scaled visual feature; and

generating a visualization to the user of the received plurality of documents based on the determined plurality of latent semantic differences and the scaled visual feature,

wherein a plurality of mark-ups is added to the mapped plurality of linguistic units,

wherein at least one value associated with at least one dimension of the mapped plurality of linguistic units from each of the received plurality of documents is correlated with at least one value associated with a hue, a saturation and a lightness based on the determined plurality of latent semantic differences associated with the at least one dimension of the mapped plurality of linguistic units,

wherein the at least one value associated with the hue, saturation and lightness is translated into a hexadecimal code.

8. The computer system of claim 7 , wherein building the plurality of latent semantic dimensions based on the extracted plurality of linguistic units, further comprises:

generating a plurality of reduced dimensions to a percentage weight associated with the extracted plurality of linguistic units using deep learning, wherein the deep learning comprises a combination of weights of the extracted plurality of linguistic units associated with the received plurality of documents.

9. The computer system of claim 7 , wherein generating the visualization to the user of the received plurality of documents based on the presented plurality of latent semantic differences and the scaled visual feature, further comprises:

subtracting at least one latent semantic dimension from one of the built plurality of latent semantic dimensions associated with another one of the received plurality of documents; and

generating a visual representation of the mapped plurality of latent semantic differences associated with the received plurality of documents.

10. The computer system of claim 9 , further comprising:

determining the user selected to visualize a plurality of latent semantic similarities between the received plurality of documents;

receiving, from the user, a threshold value to define the determined plurality of latent semantic similarities between the received plurality of documents;

comparing the determined threshold value with each of the determined plurality of latent semantic similarities; and

generating a visual representation of the mapped plurality of latent semantic similarities associated with the received plurality of documents.

11. The computer system of claim 7 , wherein receiving the plurality of documents from the user, further comprises:

uploading the plurality of documents onto a cloud storage by the user; and

selecting, by the user, the plurality of documents from the cloud storage for comparison.

12. The computer system of claim 7 , wherein mapping the weighted plurality of linguistic units to the scaled visual feature, further comprises:

assigning a hue to each of the latent semantic dimensions of the built plurality of latent semantic dimensions; and

determining a range of lightness and saturation for each assigned hue associated with each latent semantic dimension of the built plurality of latent semantic dimensions.

13. A computer program product for comparing documents using latent semantic differences, comprising:

one or more computer-readable storage media and program instructions stored on at least one of the one or more non-transitory storage media, the program instructions executable by a processor to cause the processor to perform a method comprising:

receiving a plurality of documents from a user;

extracting a plurality of linguistic units associated with the received plurality of documents,

wherein a canonical unit for each linguistic unit from the extracted plurality of linguistic units is determined,

wherein that at least one variation of each linguistic unit from the extracted plurality of linguistic units are present is determined,

wherein a number of variations of each linguistic unit from the extracted plurality of linguistic units by utilizing a dictionary is determined,

wherein the determined number of variations of each linguistic unit is tracked by utilizing a set of tables,

wherein at least one start position and at least one end position with each linguistic unit from the extracted plurality of linguistic units is stored,

wherein each linguistic unit from the extracted plurality of linguistic units includes a plurality of words in a contiguous sequence;

building a plurality of latent semantic dimensions based on the extracted plurality of linguistic units;

weighting the extracted plurality of linguistic units utilizing the built plurality of latent semantic dimensions;

determining a plurality of latent semantic differences between the received plurality of documents based on weighted plurality of linguistic units;

mapping the weighted plurality of linguistic units to a scaled visual feature; and

generating a visualization to the user of the received plurality of documents based on the determined plurality of latent semantic differences and the scaled visual feature,

wherein a plurality of mark-ups is added to the mapped plurality of linguistic units,

wherein at least one value associated with at least one dimension of the mapped plurality of linguistic units from each of the received plurality of documents is correlated with at least one value associated with a hue, a saturation and a lightness based on the determined plurality of latent semantic differences associated with the at least one dimension of the mapped plurality of linguistic units,

wherein the at least one value associated with the hue, saturation and lightness is translated into a hexadecimal code.

14. The computer program product of claim 13 , wherein building the plurality of latent semantic dimensions based on the extracted plurality of linguistic units, further comprises:

generating a plurality of reduced dimensions to a percentage weight associated with the extracted plurality of linguistic units using deep learning, wherein the deep learning comprises a combination of weights of the extracted plurality of linguistic units associated with the received plurality of documents.

15. The computer program product of claim 13 , wherein generating the visualization to the user of the received plurality of documents based on the presented plurality of latent semantic differences and the scaled visual feature, further comprises:

subtracting at least one latent semantic dimension from one of the built plurality of latent semantic dimensions associated with another one of the received plurality of documents; and

generating a visual representation of the mapped plurality of latent semantic differences associated with the received plurality of documents.

16. The computer program product of claim 13 , wherein receiving the plurality of documents from the user, further comprises:

uploading the plurality of documents onto a cloud storage by the user; and

selecting, by the user, the plurality of documents from the cloud storage for comparison.

17. The computer program product of claim 13 , wherein mapping the weighted plurality of linguistic units to the scaled visual feature, further comprises:

assigning a hue to each of the latent semantic dimensions of the built plurality of latent semantic dimensions; and

determining a range of lightness and saturation for each assigned hue associated with each latent semantic dimension of the built plurality of latent semantic dimensions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2017
From: FARRELL, ROBERT G.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 043607/0658 →
Continuity (1)
Related Publication 20190087409A1 · Mar 21, 2019