IP Library › Granted Patent US 10,223,354
Granted Patent B2
US 10,223,354 · App. 15/478,363 · Granted Mar 5, 2019

Unsupervised aspect extraction from raw data using word embeddings

Inventors: Ruidan He (Singapore, SG); Daniel Dahlmeier (Singapore, SG)
Assignee: SAP SE
G06F17/2785
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,223,354
App. No.
15/478,363
Granted
Mar 5, 2019
Kind
B2
Abstract

Methods, systems, and computer-readable storage media for receiving a vocabulary that includes text data that is provided as at least a portion of raw data, the raw data being provided in a computer-readable file, providing word embeddings based on the vocabulary, the word embeddings including word vectors for words included in the vocabulary, clustering word embeddings to provide a plurality of clusters, each cluster representing an aspect inferred from the vocabulary, determining a respective association score between each word in the vocabulary and a respective aspect, and providing a word ranking for each aspect based on the respective association scores.

Claims (34)

1. A computer-implemented method for unsupervised learning including aspect extraction from raw data to train a model, the method being executed by one or more processors and comprising:

receiving, by the one or more processors, a vocabulary, the vocabulary comprising text data that is provided as at least a portion of raw data, the raw data being provided in a computer-readable file;

providing, by the one or more processors, word embeddings based on the vocabulary, the word embeddings comprising word vectors for words included in the vocabulary;

clustering, by the one or more processors, word embeddings to provide a plurality of clusters, each cluster representing an aspect inferred from the vocabulary;

determining, by the one or more processors, an association strength matrix having dimension V×T, where V is a vocabulary size, and T is a number of aspects inferred from the vocabulary, each entry in the association strength matrix comprising a respective association score between each word in the vocabulary and a respective aspect based on a word vector and a centroid vector; and

providing, by the one or more processors, a word ranking for each aspect based on the respective association scores.

2. The method of claim 1 , wherein the word embeddings are provided based on one of a skip-gram model, and a continuous-bag-of-words (CBOW) model.

3. The method of claim 1 , further comprising incorporating domain-specific knowledge into the word embeddings using a graph-based learning objective that is used to refine word vectors relative to one another.

4. The method of claim 1 , wherein the word embeddings are clustered using k-gram clustering.

5. The method of claim 1 , wherein the vocabulary comprises fewer words than the raw data.

6. The method of claim 1 , wherein the raw data comprises review data.

7. A non-transitory computer-readable storage medium coupled to one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations for unsupervised learning including aspect extraction from raw data to train a model, the operations comprising:

receiving a vocabulary, the vocabulary comprising text data that is provided as at least a portion of raw data, the raw data being provided in a computer-readable file;

providing word embeddings based on the vocabulary, the word embeddings comprising word vectors for words included in the vocabulary;

clustering word embeddings to provide a plurality of clusters, each cluster representing an aspect inferred from the vocabulary;

determining an association strength matrix having dimension V×T, where V is a vocabulary size, and T is a number of aspects inferred from the vocabulary, each entry in the association strength matric comprising a respective association score between each word in the vocabulary and a respective aspect based on a word vector and a centroid vector; and

providing a word ranking for each aspect based on the respective association scores.

8. The computer-readable storage medium of claim 7 , wherein the word embeddings are provided based on one of a skip-gram model, and a continuous-bag-of-words (CBOW) model.

9. The computer-readable storage medium of claim 7 , wherein operations further comprise incorporating domain-specific knowledge into the word embeddings using a graph-based learning objective that is used to refine word vectors relative to one another.

10. The computer-readable storage medium of claim 7 , wherein the word embeddings are clustered using k-gram clustering.

11. The computer-readable storage medium of claim 7 , wherein the vocabulary comprises fewer words than the raw data.

12. The computer-readable storage medium of claim 7 , wherein the raw data comprises review data.

13. A system, comprising:

a computing device; and

a computer-readable storage device coupled to the computing device and having instructions stored thereon which, when executed by the computing device, cause the computing device to perform operations for unsupervised learning including aspect extraction from raw data to train a model, the operations comprising:

receiving a vocabulary, the vocabulary comprising text data that is provided as at least a portion of raw data, the raw data being provided in a computer-readable file;

providing word embeddings based on the vocabulary, the word embeddings comprising word vectors for words included in the vocabulary;

clustering word embeddings to provide a plurality of clusters, each cluster representing an aspect inferred from the vocabulary;

determining an association strength matrix having dimension V×T, where V is a vocabulary size, and T is a number of aspects inferred from the vocabulary, each entry in the association strength matric comprising a respective association score between each word in the vocabulary and a respective aspect based on a word vector and a centroid vector; and

providing a word ranking for each aspect based on the respective association scores.

14. The system of claim 13 , wherein the word embeddings are provided based on one of a skip-gram model, and a continuous-bag-of-words (CBOW) model.

15. The system of claim 13 , wherein operations further comprise incorporating domain-specific knowledge into the word embeddings using a graph-based learning objective that is used to refine word vectors relative to one another.

16. The system of claim 13 , wherein the word embeddings are clustered using k-gram clustering.

17. The system of claim 13 , wherein the vocabulary comprises fewer words than the raw data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 4, 2017
From: HE, RUIDAN; DAHLMEIER, DANIEL
To: SAP SE
Reel/Frame 042151/0864 →
Continuity (1)
Related Publication 20180285344A1 · Oct 4, 2018