IP Library Granted Patent US 10,755,183
Granted Patent B1
US 10,755,183 · App. 15/416,611 · Granted Aug 25, 2020

Building training data and similarity relations for semantic space

Inventors: Eugene Livshitz (San Mateo, CA); Alexander Pashintsev (Cupertino, CA); Boris Gorbatov (Sunnyvale, CA)
Assignee: EVERNOTE CORPORATION
G06N5/04G06F16/334G06F40/205G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,755,183
App. No.
15/416,611
Granted
Aug 25, 2020
Kind
B1
Abstract

Selecting data from a source text corpus for training a semantic data analysis system includes selecting an item of the text corpus, validating the item, extracting at least one section of the item, determining a length of each of the at least one section of the item, and subdividing each of the sections having a length greater than a predetermined amount into a plurality of fragments that are deemed to be similar. The predetermined amount may be approximately twice a size of a fragment. A fragment may have approximately 100 words or between 40 and 60 words. Fragments from different items may be deemed to be dissimilar. Sections having a length less than the predetermined amount may be ignored. Validating the item may include parsing editorial notes and other accompanying data. The source text corpus may be Wikipedia. The item may be an article.

Claims (33)

1. A method of selecting data from a source text corpus for training a semantic data analysis system, comprising:

selecting an item of the text corpus, wherein the item includes a plurality of sections;

validating the item;

extracting a first section and a second section of the plurality of sections, wherein the first section and the second section are distinct;

determining a length of each of the first section and the second section of the plurality of sections;

based on the length of the first section being greater than a predetermined amount, subdividing the first section into a first plurality of fragments, wherein each fragment of the first plurality of fragments is deemed to be similar to each other; and

based on the length of the second section being greater than the predetermined amount, subdividing the second section into a second plurality of fragments, wherein each fragment of the second plurality of fragments is deemed to be similar to each other, and deemed to be of undefined similarity with regard to the first plurality of fragments.

2. The method of claim 1 , wherein the predetermined amount is approximately twice a size of a fragment.

3. The method of claim 2 , wherein a fragment has approximately 100 words.

4. The method of claim 2 , wherein a fragment has between 40 and 60 words.

5. The method of claim 1 , wherein fragments from different items are deemed to be dissimilar.

6. The method of claim 1 , wherein sections of the plurality of section having a length less than the predetermined amount are ignored.

7. The method of claim 1 , wherein validating the item includes parsing editorial notes and other accompanying data.

8. The method of claim 1 , wherein the source text corpus is Wikipedia.

9. The method of claim 8 , wherein the item is an article.

10. A nontransitory computer readable medium containing software that selects data from a source text corpus for training a semantic data analysis system, the software comprising:

executable code that selects an item of the text corpus, wherein the item includes a plurality of sections;

validating the item;

executable code that validates the item;

executable code that extracts a first section and a second section of the plurality of sections, wherein the first section and the second section are distinct;

executable code that determines a length of each of the first section and the second section of the plurality of sections;

executable code that based on the length of the first section being greater than a predetermined amount, subdivides the first section into a first plurality of fragments, wherein each fragment of the first plurality of fragments is deemed to be similar to each other; and

executable code that based on the length of the second section being greater than the predetermined amount, subdivides the second section into a second plurality of fragments, wherein each fragment of the second plurality of fragments is deemed to be similar to each other, and deemed to be of undefined similarity with regard to the first plurality of fragments.

11. The nontransitory computer readable medium of claim 10 , wherein the predetermined amount is approximately twice a size of a fragment.

12. The nontransitory computer readable medium of claim 11 , wherein a fragment has approximately 100 words.

13. The nontransitory computer readable medium of claim 11 , wherein a fragment has between 40 and 60 words.

14. The nontransitory computer readable medium of claim 10 , wherein fragments from different items are deemed to be dissimilar.

15. The nontransitory computer readable medium of claim 10 , wherein sections of the plurality of sections having a length less than the predetermined amount are ignored.

16. The nontransitory computer readable medium of claim 10 , wherein validating the item includes parsing editorial notes and other accompanying data.

17. The nontransitory computer readable medium of claim 10 , wherein the source text corpus is Wikipedia.

18. The nontransitory computer readable medium of claim 17 , wherein the item is an article.

19. The method of claim 1 , further comprising building a training set based on the first plurality of fragments and the second plurality of fragments, wherein the training set is used to train the semantic data analysis system.

20. The method of claim 19 , wherein the training set includes one or more pluralities of fragments from another item of the text corpus, wherein the one or more pluralities of fragments include respective similarity determinations.

Assignments (8)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 12, 2024
From: EVERNOTE CORPORATION
To: BENDING SPOONS S.P.A.
Reel/Frame 066288/0195 →
RELEASE OF SECURITY INTEREST Recorded Mar 17, 2023
From: MUFG BANK, LTD.
To: EVERNOTE CORPORATION
Reel/Frame 063116/0260 →
RELEASE OF SECURITY INTEREST Recorded Oct 8, 2021
From: EAST WEST BANK
To: EVERNOTE CORPORATION
Reel/Frame 057852/0078 →
SECURITY INTEREST Recorded Oct 6, 2021
From: EVERNOTE CORPORATION
To: MUFG UNION BANK, N.A.
Reel/Frame 057722/0876 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT TERMINATION AT R/F 048513/ 0003 Recorded Oct 22, 2020
From: HERCULES CAPITAL, INC.
To: EVERNOTE CORPORATION
Reel/Frame 054178/0892 →
SECURITY INTEREST Recorded Oct 19, 2020
From: EVERNOTE CORPORATION
To: EAST WEST BANK
Reel/Frame 054113/0876 →
SECURITY INTEREST Recorded Mar 5, 2019
From: EVERNOTE CORPORATION; EVERNOTE GMBH
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 048513/0003 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 3, 2017
From: LIVSHITZ, EUGENE; PASHINTSEV, ALEXANDER; GORBATOV, BORIS
To: EVERNOTE CORPORATION
Reel/Frame 044282/0179 →