IP Library Granted Patent US 9,460,080
Granted Patent B2
US 9,460,080 · App. 15/142,288 · Granted Oct 4, 2016

Modifying a tokenizer based on pseudo data for natural language processing

Inventors: Bing Zhao (Sunnyvale, CA); Ethan Zhang (San Jose, CA)
Assignee: LinkedIn Corporation
G06F17/277G06F17/278G06F17/2735
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,460,080
App. No.
15/142,288
Granted
Oct 4, 2016
Kind
B2
Abstract

Techniques for training a tokenizer (or word segmenter) are provided. In one technique, a tokenizer tokenizes a token string to identify individual tokens or words. A language model is generated based on the identified tokens or words. A vocabulary about an entity, such as a person or company, is identified. The vocabulary may be online data that refers to the entity, such as a news article or a profile page of a member of a social network. Some of the tokens in the vocabulary may be weighted higher than others. The language model accepts the weighted vocabulary as input and generates pseudo sentences. Alternatively, regular expressions are used to generate the pseudo sentences. The pseudo sentences are used to train the tokenizer.

Claims (50)

1. A method comprising:

identifying, in a profile of an entity in a social network, a name of the entity;

automatically generating, based on the name of the entity, one or more sentences;

based on the one or more sentences, training a tokenizer that is configured to identify tokens within a text string;

wherein the method is performed by one or more computing devices.

2. The method of claim 1 , wherein automatically generating the one or more sentences comprises automatically generating the one or more sentences using one or more regular expressions.

3. The method of claim 2 , further comprising:

selecting a particular regular expression from among a plurality of regular expressions;

wherein the one or more regular expressions includes the particular regular expression and are fewer than the plurality of regular expressions.

4. The method of claim 3 , further comprising:

prior to generating the one or more sentences, storing type data in association with the name of the entity;

wherein selecting the particular regular expression is performed based on the type data.

5. The method of claim 4 , wherein the type data indicates a type of the name or a type of the entity.

6. The method of claim 1 , wherein automatically generating the one or more sentences comprises automatically generating the one or more sentences based on context data.

7. The method of claim 6 , wherein the context data include data from the profile of the entity.

8. The method of claim 1 , wherein automatically generating the one or more sentences comprises automatically generating the one or more sentences to be segmented.

9. The method of claim 8 , wherein:

generating the one or more sentences comprises, for a sentence, of the one or more sentences, that comprises a plurality of tokens, labeling each token of the plurality of tokens;

wherein a label of a token in the plurality of tokens indicates that the token is a beginning of a word.

10. The method of claim 1 , wherein:

automatically generating the one or more sentences comprises automatically generating a plurality of sentences;

the method further comprising:

performing an analysis of each sentence of the plurality of sentences;

based on the analysis of a particular sentence in the plurality of sentences, filtering the particular sentence, wherein the particular sentence is not used to train the tokenizer.

11. The method of claim 1 , wherein the name is of an organization.

12. A system comprising:

one or more processors;

one or more storage media storing instructions which, when executed by the one or more processors, cause:

identifying, in a profile of an entity in a social network, a name of the entity;

automatically generating, based on the name of the entity, one or more sentences;

based on the one or more sentences, training a tokenizer that is configured to identify tokens within a text string.

13. The system of claim 12 , wherein automatically generating the one or more sentences comprises automatically generating the one or more sentences using one or more regular expressions.

14. The system of claim 13 , wherein the instructions, when executed by the one or more processors, further cause:

selecting a particular regular expression from among a plurality of regular expressions;

wherein the one or more regular expressions includes the particular regular expression and are fewer than the plurality of regular expressions.

15. The system of claim 14 , wherein the instructions, when executed by the one or more processors, further cause:

prior to generating the one or more sentences, storing type data in association with the name of the entity;

wherein selecting the particular regular expression is performed based on the type data.

16. The system of claim 15 , wherein the type data indicates a type of the name or a type of the entity.

17. The system of claim 12 , wherein automatically generating the one or more sentences comprises automatically generating the one or more sentences based on context data.

18. The system of claim 17 , wherein the context data include data from the profile of the entity.

19. The system of claim 12 , wherein automatically generating the one or more sentences comprises automatically generating the one or more sentences to be segmented.

20. The system of claim 19 , wherein:

generating the one or more sentences comprises, for a sentence, of the one or more sentences, that comprises a plurality of tokens, labeling each token of the plurality of tokens;

wherein a label of a token in the plurality of tokens indicates that the token is a beginning of a word.

21. The system of claim 12 , wherein:

automatically generating the one or more sentences comprises automatically generating a plurality of sentences;

the instructions, when executed by the one or more processors, further cause:

performing an analysis of each sentence of the plurality of sentences;

based on the analysis of a particular sentence in the plurality of sentences, filtering the particular sentence, wherein the particular sentence is not used to train the tokenizer.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2017
From: LINKEDIN CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 044746/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 29, 2016
From: ZHAO, BING; ZHANG, ETHAN
To: LINKEDIN CORPORATION
Reel/Frame 038583/0723 →
Continuity (2)
Continuation 14611816 · Feb 2, 2015
Related Publication 20160246776A1 · Aug 25, 2016