IP Library Granted Patent US 12,518,098
Granted Patent B2
US 12,518,098 · App. 17/936,679 · Granted Jan 6, 2026

Fusion of word embeddings and word scores for text classification

Inventors: Ahmed Ataallah Ataallah Abobakr (Geelong, AU); Mark Edward Johnson (Sydney, AU); Thanh Long Duong (Seabrook, AU); Vladislav Blinov (Melbourne, AU); Yu-Heng Hong (Carlton, AU); Cong Duy Vu Hoang (Wantirna South, AU); Duy Vu (Melbourne, AU)
Assignee: Oracle International Corporation
G06F40/295G06F16/3329G06F16/35G06F40/205G06F40/263G06F40/30H04L51/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,518,098
App. No.
17/936,679
Granted
Jan 6, 2026
Kind
B2
Abstract

Techniques disclosed herein relate generally to text classification and include techniques for fusing word embeddings with word scores for text classification. In one particular aspect, a method for text classification is provided that includes obtaining an embedding vector for a textual unit, based on a plurality of word embedding vectors and a plurality of word scores. The plurality of word embedding vectors includes a corresponding word embedding vector for each of a plurality of words of the textual unit, and the plurality of word scores includes a corresponding word score for each of the plurality of words of the textual unit. The method also includes passing the embedding vector for the textual unit through at least one feed-forward layer to obtain a final layer output, and performing a classification on the final layer output.

Claims (51)

1 . A method of classifying a textual unit, the method comprising:

based on a plurality of word embedding vectors and a plurality of word scores, obtaining an embedding vector for the textual unit, wherein the embedding vector for the textual unit is based on a composite embedding vector and on information from a word score vector that includes the plurality of word scores, wherein the composite embedding vector is based on the plurality of word embedding vectors, and wherein obtaining the embedding vector for the textual unit comprises:

projecting the word score vector to obtain a projected vector having a dimensionality that is lower than a dimensionality of the word score vector; and

combining the composite embedding vector with the projected vector;

passing the embedding vector for the textual unit through at least one feed-forward layer to obtain a final layer output; and

performing a classification on the final layer output,

wherein, for each word among a plurality of words of the textual unit:

the plurality of word embedding vectors includes a corresponding word embedding vector, and

the plurality of word scores includes a corresponding word score.

2 . The method of claim 1 , wherein the at least one feed-forward layer comprises a multi-layer perceptron.

3 . The method of claim 1 , wherein performing the classification on the final layer output comprises applying a softmax function to the final layer output.

4 . The method of claim 1 , wherein the method comprises obtaining the plurality of word embedding vectors using a context-independent word embedding model.

5 . The method of claim 1 , wherein the method comprises obtaining the plurality of word embedding vectors using a trained FastText model.

6 . The method of claim 1 , wherein, for each word among the plurality of words, the corresponding word score is based on a term frequency of the word in the textual unit.

7 . The method of claim 1 , wherein, for each word among the plurality of words, the corresponding word score is based on a document frequency of the word.

8 . The method of claim 1 , wherein, for each word among the plurality of words, the corresponding word score is based on:

a term frequency of the word in the textual unit,

a document frequency of the word, and

one or more learned weighting parameters.

9 . The method of claim 1 , wherein, for each word among the plurality of words, the corresponding word score is based on a smooth inverse frequency of the word.

10 . The method of claim 1 , wherein obtaining the embedding vector for the textual unit comprises:

scaling each word embedding vector of the plurality of word embedding vectors with the corresponding word score of the plurality of word scores to obtain a plurality of scaled word embedding vectors, and

obtaining the embedding vector for the textual unit as a sum or average of the plurality of scaled word embedding vectors.

11 . The method of claim 1 , wherein the combining comprises at least one among concatenating, interpolating, gating, or attention.

12 . A system comprising:

one or more data processors; and

one or more non-transitory computer readable media storing instructions which, when executed by the one or more data processors, cause the one or more data processors to perform processing comprising:

based on a plurality of word embedding vectors and a plurality of word scores, obtaining an embedding vector for a textual unit, wherein the embedding vector for the textual unit is based on a composite embedding vector and on information from a word score vector that includes the plurality of word scores, wherein the composite embedding vector is based on the plurality of word embedding vectors, and wherein obtaining the embedding vector for the textual unit comprises:

projecting the word score vector to obtain a projected vector having a dimensionality that is lower than a dimensionality of the word score vector, and

combining the composite embedding vector with the projected vector;

passing the embedding vector for the textual unit through at least one feed-forward layer to obtain a final layer output; and

performing a classification on the final layer output,

wherein, for each word among a plurality of words of the textual unit:

the plurality of word embedding vectors includes a corresponding word embedding vector, and

the plurality of word scores includes a corresponding word score.

13 . The system of claim 12 , wherein the at least one feed-forward layer comprises a multi-layer perceptron.

14 . The system of claim 12 , wherein performing the classification on the final layer output comprises applying a softmax function to the final layer output.

15 . The system of claim 12 , wherein the processing comprises obtaining the plurality of word embedding vectors using a context-independent word embedding model.

16 . The system of claim 12 , wherein the processing comprises obtaining the plurality of word embedding vectors using a trained FastText model.

17 . The system of claim 12 , wherein, for each word among the plurality of words, the corresponding word score is based on a term frequency of the word in the textual unit.

18 . The system of claim 12 , wherein, for each word among the plurality of words, the corresponding word score is based on a document frequency of the word.

19 . A computer-program product tangibly embodied in one or more non-transitory machine-readable media, including instructions configured to cause one or more data processors to perform processing comprising:

based on a plurality of word embedding vectors and a plurality of word scores, obtaining an embedding vector for a textual unit, wherein the embedding vector for the textual unit is based on a composite embedding vector and on information from a word score vector that includes the plurality of word scores, wherein the composite embedding vector is based on a plurality of word embedding vectors, and wherein obtaining the embedding vector for the textual unit comprises:

projecting the word score vector to obtain a projected vector having a dimensionality that is lower than a dimensionality of the word score vector, and

combining the composite embedding vector with the projected vector;

passing the embedding vector for the textual unit through at least one feed-forward layer to obtain a final layer output; and

performing a classification on the final layer output,

wherein, for each word among a plurality of words of the textual unit:

the plurality of word embedding vectors includes a corresponding word embedding vector, and

the plurality of word scores includes a corresponding word score.

20 . The computer-program product of claim 19 , wherein the composite embedding vector is a sum or average of the plurality of word embedding vectors.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2022
From: ABOBAKR, AHMED ATAALLAH ATAALLAH; JOHNSON, MARK EDWARD; DUONG, THANH LONG; BLINOV, VLADISLAV; HONG, YU-HENG; HOANG, CONG DUY VU; VU, DUY
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 061478/0924 →
Continuity (2)
Provisional Application 63250274 · Sep 30, 2021
Related Publication 20230100508A1 · Mar 30, 2023
References Cited (23)
US 20160349274A1 · Williams · 2016 [cited by examiner]
US 20180165288A1 · Chang · 2018 [cited by examiner]
US 20210034707A1 · Podgorny · 2021 [cited by examiner]
US 20220310070A1 · Moritz · 2022 [cited by examiner]
US 20220366293A1 · Abhishek · 2022 [cited by examiner]
US 20230100508A1 · Abobakr · 2023 [cited by examiner]
CN 110390103A · 2019 [cited by examiner]
CN 111291788A · 2020 [cited by examiner]
CN 111460794A · 2020 [cited by examiner]
“A Guide to Deep Learning Layers: Four Deep Learning Layer Architectures Explained.”, Available Online at: https://adgefficiency.com/guide-deep-learning/, Nov. 16, 2020, 29 pages. [cited by applicant]
Arora et al., “A Simple but Tough-to-Beat Baseline for Sentence Embeddings”, International Conference on Learning Representations, Apr. 2017, pp. 1-16. [cited by applicant]
Bojanowski et al., “Enriching Word Vectors With Subword Information”, arXiv preprint arXiv:1607.04606v2, vol. 5, Jun. 19, 2017, 12 pages. [cited by applicant]
Cheng et al., “Wide & Deep Learning for Recommender Systems”, arXiv:1606.07792v1, 2016, 4 pages. [cited by applicant]
Cheng , “Wide Deep Learning: Better Together with TensorFlow”, Available Online at: https://ai.googleblog.com/2016/06/wide-deep-learning-better-together-with.html, Jun. 29, 2016, 5 pages. [cited by applicant]
Iyyer et al., “Deep Unordered Composition Rivals Syntactic Methods for Text Classification”, Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Confer… [cited by applicant]
Joulin et al., “Bag of Tricks for Efficient Text Classification”, arXiv:1607.01759v3, Jul. 2016, 5 pages. [cited by applicant]
Joulin et al., “FastText.zip: Compressing Text Classification Models”, Available Online at: https://arxiv.org/pdf/1612.03651v1.pdf, Dec. 12, 2016, pp. 1-13. [cited by applicant]
Minaee et al., “Deep Learning Based Text Classification: A Comprehensive Review”, arXiv:2004.03705v3, Association for Computing Machinery Computing Surveys, vol. 54, No. 3, Jan. 4, 2021, pp. 1-43. [cited by applicant]
Novotny et al., “Text Classification with Word Embedding Regularization and Soft Similarity Measure”, arXiv:2003.05019v1, 2020, 15 pages. [cited by applicant]
Wieting et al., “Towards Universal Paraphrastic Sentence Embeddings”, arXiv: 1511.08198v3, 2016, 19 pages. [cited by applicant]
Yao et al., “Graph Convolutional Networks for Text Classification”, arXiv: 1809.05679v3, 2018, 9 pages. [cited by applicant]
Zhang et al., “Every Document owns its Structure: Inductive Text Classification via Graph Neural Networks”, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 5-10, 2020, pp. 3… [cited by applicant]
Zhu et al., “Simple Spectral Graph Convolution”, International Conference on Learning Representations (ICLR), Sep. 28, 2020, 15 pages. [cited by applicant]