IP Library › Granted Patent US 11,227,004
Granted Patent B2
US 11,227,004 · App. 16/735,165 · Granted Jan 18, 2022

Semantic category classification

Inventor: Mingkuan Liu (San Jose, CA)
Assignee: eBay Inc.
G06F16/358G06F40/00G06F40/30G06N3/0454G06N20/00G06Q20/10G06Q40/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,227,004
App. No.
16/735,165
Filed
Jan 6, 2020
Granted
Jan 18, 2022
Kind
B2
Art Unit
2167
USPC
707/723
Abstract

In accordance with an example embodiment, large scale category classification based on sequence semantic embedding and parallel learning is described. In one example, one or more closest matches are identified by comparison between (i) a publication semantic vector that corresponds to at least part of the publication, the publication semantic vector based on a first machine-learned model that projects the at least part of the publication into a semantic vector space, and (ii) a plurality of category vectors corresponding to respective categories from a plurality of categories.

Claims (381)

1. A method performed via hardware processing circuitry, comprising:

receiving a request from a user device to add a publication to a publication corpus;

generating, based on the request, a publication semantic vector in a semantic vector space based on sequence sematic embedding of at least a portion of the publication;

comparing the publication semantic vector to each of a plurality of category vectors, each category vector resulting from a projection of a respective category of publications into the semantic vector space;

ranking a plurality of categories based on the comparison;

further ranking a portion of the ranked categories based on an expected perplexity of each of the categories in the portion and a perplexity of the at least the portion of the publication, the expected perplexity of each of the categories in the portion based on separate perplexities of each sentence of publications included in the respective category; and

causing display, on the user device, an output indicative of the further ranking.

2. The method of claim 1 , further comprising determining a distance between the publication semantic vector and each of the plurality of category vectors in the semantic vector space, wherein the ranking is based on the determined distances.

3. The method of claim 2 , wherein the ranking comprises assigning a highest ranking to a category having a respective category vector that is closest to the semantic vector of the publication in the semantic vector space of any of the plurality of categories, and the method further comprises selecting the portion from a highest ranked subset of the plurality of categories.

4. The method of claim 1 , further comprising hashing words of the portion of the publication, wherein the generation of the publication semantic vector is based on the hashed words.

5. The method of claim 1 , wherein the generating of the publication semantic vector is further based on a first machine learning model, and the first machine-learned model is trained at one or more of a sub-word level and a character level.

6. The method of claim 1 , wherein the separate perplexities of each sentence are determined by:

PP

⁡

(

S

)

=

P

⁡

(

w

1

⁢

⁢

…

⁢

⁢

w

N

)

-

1

/

N

=

∏

i

=

1

N

⁢

⁢

1

P

⁡

(

w

i

|

w

1

⁢

⁢

…

⁢

⁢

w

i

-

1

)

N

where:

N is the number of words in the sentence, and

{w 1 , w 2 , . . . , w N } are individual words in the sentence.

7. The method of claim 1 , wherein an expected perplexity value of a category is based on a mean perplexity value determined according to:

Mean_PP

⁢

(

C

)

=

Mean_PP

⁢

(

S

1

⁢

⁢

…

⁢

⁢

S

M

)

=

∑

PP

⁡

(

S

i

)

M

where:

Mean_PP(C) is a mean perplexity of a category of the plurality of categories,

S 1 . . . S M are sentences included in the category,

M is the number of sentences in the category.

8. The method of claim 1 , wherein the further ranking of the portion is further based on a standard deviation of the expected perplexity of each of the categories in the portion.

9. The method of claim 8 , further comprising determining the standard deviation of the expected perplexity according to:

STD_PP

⁢

(

C

)

=

STD_PP

⁢

(

S

1

⁢

⁢

…

⁢

⁢

S

M

)

=

∑

(

PP

⁡

(

S

i

)

-

Mean_PP

⁢

(

C

)

)

2

M

-

1

where:

STD_PP(C) is a standard deviation of a category C.

10. A system, comprising:

hardware processing circuitry;

one or more hardware memories storing instructions that when executed configure hardware processing circuitry to perform operations comprising:

receiving a request from a user device to add a publication to a publication corpus;

generating, based on the request, a publication semantic vector in a semantic vector space based on sequence sematic embedding of at least a portion of the publication;

comparing the publication semantic vector to each of a plurality of category vectors, each category vector resulting from a projection of a respective category of publications into the semantic vector space;

ranking a plurality of categories based on the comparison;

further ranking a portion of the ranked categories based on an expected perplexity of each of the categories in the portion and a perplexity of the at least the portion of the publication, the expected perplexity of each of the categories in the portion based on separate perplexities of each sentence of publications included in the respective category; and

causing display, on the user device, an output indicative of the further ranking.

11. The system of claim 10 , further comprising determining a distance between the publication semantic vector and each of the plurality of category vectors in the semantic vector space, wherein the ranking is based on the determined distances.

12. The system of claim 11 , wherein the ranking comprises assigning a highest ranking to a category having a respective category vector that is closest to the semantic vector of the publication in the semantic vector space of any of the plurality of categories, and the operations further comprising selecting the portion from a highest ranked subset of the plurality of categories.

13. The system of claim 10 , the operations further comprising hashing words of the portion of the publication, wherein the generation of the publication semantic vector is based on the hashed words.

14. The system of claim 10 , wherein the generating of the publication semantic vector is further based on a first machine learning model, and the first machine-learned model is trained at one or more of a sub-word level and a character level.

15. The system of claim 10 , wherein the separate perplexities of each sentence are determined by:

PP

⁡

(

S

)

=

P

⁡

(

w

1

⁢

⁢

…

⁢

⁢

w

N

)

-

1

/

N

=

∏

i

=

1

N

⁢

⁢

1

P

⁡

(

w

i

|

w

1

⁢

⁢

…

⁢

⁢

w

i

-

1

)

N

where:

N is the number of words in the sentence, and

{w 1 , w 2 , . . . , w N } are individual words in the sentence.

16. The system of claim 10 , wherein an expected perplexity value of a category is based on a mean perplexity value determined according to:

Mean_PP

⁢

(

C

)

=

Mean_PP

⁢

(

S

1

⁢

⁢

…

⁢

⁢

S

M

)

=

∑

PP

⁡

(

S

i

)

M

where:

Mean_PP(C) is a mean perplexity of a category of the plurality of categories,

S 1 . . . S M are sentences included in the category, and

M is the number of sentences in the category.

17. The system of claim 10 , wherein the further ranking of the portion is further based on a standard deviation of the expected perplexity of each of the categories in the portion.

18. The system of claim 17 , further comprising determining the standard deviation of the expected perplexity according to:

STD_PP

⁢

(

C

)

=

STD_PP

⁢

(

S

1

⁢

⁢

…

⁢

⁢

S

M

)

=

∑

(

PP

⁡

(

S

i

)

-

Mean_PP

⁢

(

C

)

)

2

M

-

1

where:

STD_PP(C) is a standard deviation of a category C.

19. A non-transitory computer readable storage medium comprising instructions that when executed configure a hardware processor to perform operations comprising:

receiving a request from a user device to add a publication to a publication corpus;

generating, based on the request, a publication semantic vector in a semantic vector space based on sequence sematic embedding of at least a portion of the publication;

comparing the publication semantic vector to each of a plurality of category vectors, each category vector resulting from a projection of a respective category of publications into the semantic vector space;

ranking a plurality of categories based on the comparison;

further ranking a portion of the ranked categories based on an expected perplexity of each of the categories in the portion and a perplexity of the at least the portion of the publication, the expected perplexity of each of the categories in the portion based on separate perplexities of each sentence of publications included in the respective category; and

causing display, on the user device, an output indicative of the further ranking.

20. The non-transitory computer readable storage medium of claim 19 , wherein the separate perplexities of each sentence are determined by:

PP

⁡

(

S

)

=

P

⁡

(

w

1

⁢

⁢

…

⁢

⁢

w

N

)

-

1

/

N

=

∏

i

=

1

N

⁢

⁢

1

P

⁡

(

w

i

|

w

1

⁢

⁢

…

⁢

⁢

w

i

-

1

)

N

where:

N is the number of words in the sentence, and

{w 1 , w 2 , . . . , w N } are individual words in the sentence,

and wherein an expected perplexity value of a category is based on a mean perplexity value determined according to:

Mean_PP

⁢

(

C

)

=

Mean_PP

⁢

(

S

1

⁢

⁢

…

⁢

⁢

S

M

)

=

∑

PP

⁡

(

S

i

)

M

where:

Mean_PP(C) is a mean perplexity of a category of the plurality of categories,

S 1 . . . S M are sentences included in the category, and

M is the number of sentences in the category.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2020
From: LIU, MINGKUAN
To: EBAY INC.
Reel/Frame 052176/0260 →
Continuity (3)
Continuation 15429564 · Feb 10, 2017
Provisional Application 62293922 · Feb 11, 2016
Related Publication 20200218750A1 · Jul 9, 2020
Cited By (1)
US 12,705,264