IP Library › Granted Patent US 12,333,238
Granted Patent B2
US 12,333,238 · App. 17/825,311 · Granted Jun 17, 2025

Embedding texts into high dimensional vectors in natural language processing

Inventors: Changchuan Yin (Hoffman Estates, IL); Shahzad Saeed (Plainfield, IL)
Assignee: AT&T Mobility II LLC
G06F40/126G06F40/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,238
App. No.
17/825,311
Granted
Jun 17, 2025
Kind
B2
Abstract

Concepts and technologies disclosed herein are directed to embedding texts into high dimensional vectors in natural language processing (“NLP”). According to one aspect, an NLP system can receive an input text that includes n number of words. The NLP system can encode the input text into a first matrix using a word embedding algorithm, such as Word2Vec algorithm. The NLP system can encode the input text into the Word2Vec by embedding each word in the n number of words of the input text into a k-dimensional Word2Vec vector using the Word2Vec algorithm. The NLP system also can decode the first matrix into a second matrix using a text embedding algorithm. In some embodiments, the second matrix is a congruence derivative matrix. The NLP system can then output the second matrix to a machine learning module that implements a machine learning technique such as short text classification.

Claims (41)

1. A method comprising:

receiving, by a natural language processing system comprising a processor, a first input text comprising n number of words;

receiving, by the natural language processing system, a second input text comprising m number of words, wherein the n number of words of the first input text is different than the m number of words of the second input text;

encoding, by the natural language processing system, the first input text into a first matrix using a word embedding algorithm, wherein the word embedding algorithm comprises a Word2Vec algorithm, wherein the first matrix of the first input text comprises a Word2Vec matrix of the first input text, and wherein encoding the first input text into the Word2Vec matrix of the first input text using the Word2Vec algorithm comprises embedding each word in the n number of words of the first input text into a k-dimensional Word2Vec vector using the Word2Vec algorithm resulting in the Word2Vec matrix of the first input text having a dimension of n×k;

encoding, by the natural language processing system, the second input text into a first matrix using the word embedding algorithm comprising the Word2Vec algorithm, wherein the first matrix of the second input text comprises a Word2Vec matrix of the second input text, wherein encoding the second input text into the Word2Vec matrix of the second input text using the Word2Vec algorithm comprises embedding each word in the m number of words of the second input text into a k-dimensional Word2Vec vector using the Word2Vec algorithm resulting in the Word2Vec matrix of the second input text having a dimension of m×k, and wherein the dimension of n×k of the Word2Vec matrix of the first input text is different from the dimension of m×k of the Word2Vec matrix of the second input text;

decoding, by the natural language processing system, the first matrix of the first input text into a second matrix of the first input text using a text embedding algorithm, wherein decoding the first matrix of the first input text into the second matrix of the first input text using the text embedding algorithm comprises decoding the Word2Vec matrix of the first input text into a first congruence derivative matrix using a congruence derivative vector representation, and wherein the first congruence derivative matrix has a first dimension;

decoding, by the natural language processing system, the first matrix of the second input text into a second matrix of the second input text using the text embedding algorithm, wherein decoding the first matrix of the second input text into the second matrix of the second input text using the text embedding algorithm comprises decoding the Word2Vec matrix of the second input text into a second congruence derivative matrix using the congruence derivative vector representation, wherein the second congruence derivative matrix has a second dimension, and wherein decoding the Word2Vec matrix of the first input text into the first congruence derivative matrix using the congruence derivative vector representation and decoding the Word2Vec matrix of the second input text into the second congruence derivative matrix using the congruence derivative vector representation results in the first dimension of the first congruence derivative matrix of the first input text being equal to the second dimension of the second congruence derivative matrix of the second input text; and

using, by a machine learning module executed by the natural language processing system, the first congruence derivative matrix and the second congruence derivative matrix as training data for a machine learning model, wherein the machine learning module requires matrices used for the training data to have uniform dimensions, and wherein using the first congruence derivative matrix and the second congruence derivative matrix as training data for a machine learning model comprises

creating, by the machine learning module of the natural language processing system, the machine learning model, and

passing, by the machine learning module of the natural language processing system, the first congruence derivative matrix and the second congruence derivative matrix through the machine learning model to train the machine learning model to determine mistakes in other input texts.

2. The method of claim 1 , wherein the first input text comprises multiple words in order, one or more sentences, one or more phrases, one or more paragraphs, or an entire document.

3. The method of claim 1 , wherein the second input text comprises multiple words in order, one or more sentences, one or more phrases, one or more paragraphs, or an entire document.

4. The method of claim 1 , wherein the machine learning module implements, based at least in part on the training data, short text classification.

5. A natural language processing system comprising

a processor; and

a memory comprising a machine learning module and instructions that, when executed by the processor, cause the processor to perform operations comprising

receiving a first input text comprising n number of words,

receiving a second input text comprising m number of words, wherein the n number of words of the first input text is different than the m number of words of the second input text,

encoding the first input text into a first matrix using a word embedding algorithm, wherein the word embedding algorithm comprises a Word2Vec algorithm, wherein the first matrix of the first input text comprises a Word2Vec matrix, and wherein encoding the first input text into the Word2Vec matrix of the first input text using the Word2Vec algorithm comprises embedding each word in the n number of words of the first input text into a k-dimensional Word2Vec vector using the Word2Vec algorithm resulting in the Word2Vec matrix of the first input text having a dimension of n×k,

encoding the second input text into a first matrix using the word embedding algorithm comprising the Word2Vec algorithm, wherein the first matrix of the second input text comprises a Word2Vec matrix of the second input text, wherein encoding the second input text into the Word2Vec matrix of the second input text using the Word2Vec algorithm comprises embedding each word in the m number of words of the second input text into a k-dimensional Word2Vec vector using the Word2Vec algorithm resulting in the Word2Vec matrix of the second input text having a dimension of m×k, and wherein the dimension of n×k of the Word2Vec matrix of the first input text is different from the dimension of m×k of the Word2Vec matrix of the second input text,

decoding the first matrix of the first input text into a second matrix of the first input text using a text embedding algorithm, wherein decoding the first matrix of the first input text into the second matrix of the first input text using the text embedding algorithm comprises decoding the Word2Vec matrix of the first input text into a first congruence derivative matrix using a congruence derivative vector representation, and wherein the first congruence derivative matrix has a first dimension

decoding the first matrix of the second input text into a second matrix of the second input text using the text embedding algorithm, wherein decoding the first matrix of the second input text into the second matrix of the second input text using the text embedding algorithm comprises decoding the Word2Vec matrix of the second input text into a second congruence derivative matrix using the congruence derivative vector representation, wherein the second congruence derivative matrix has a second dimension, and wherein decoding the Word2Vec matrix of the first input text into the first congruence derivative matrix using the congruence derivative vector representation and decoding the Word2Vec matrix of the second input text into the second congruence derivative matrix using the congruence derivative vector representation results in the first dimension of the first congruence derivative matrix of the first input text being equal to the second dimension of the second congruence derivative matrix of the second input text,

using, by the machine learning module, the first congruence derivative matrix and the second congruence derivative matrix as training data for a machine learning model, wherein the machine learning module requires matrices used for the training data to have uniform dimensions, and wherein using the first congruence derivative matrix and the second congruence derivative matrix as training data for a machine learning model comprises

creating, by the machine learning module, the machine learning model, and

passing, by the machine learning module, the first congruence derivative matrix and the second congruence derivative matrix through the machine learning model to train the machine learning model to determine mistakes in other input texts.

6. The natural language processing system of claim 5 , wherein the first input text comprises multiple words in order, one or more sentences, one or more phrases, one or more paragraphs, or an entire document.

7. The natural language processing system of claim 5 , wherein the second input text comprises multiple words in order, one or more sentences, one or more phrases, one or more paragraphs, or an entire document.

8. The natural language processing system of claim 5 , wherein the machine learning module implements, based at least in part on the training data, short text classification.

9. A computer-readable storage medium comprising a machine learning module and computer-executable instructions that, when executed by a processor, cause the processor to perform operations comprising:

receiving a first input text comprising n number of words;

receiving a second input text comprising m number of words, wherein the n number of words of the first input text is different than the m number of words of the second input text;

encoding the first input text into a first matrix using a word embedding algorithm, wherein the word embedding algorithm comprises a Word2Vec algorithm, wherein the first matrix of the first input text comprises a Word2Vec matrix of the first input text, and wherein encoding the first input text into the Word2Vec matrix of the first input text using the Word2Vec algorithm comprises embedding each word in the n number of words of the first input text into a k-dimensional Word2Vec vector using the Word2 Vec algorithm resulting in the Word2Vec matrix of the first input text having a dimension of n×k;

encoding the second input text into a first matrix using the word embedding algorithm comprising the Word2Vec algorithm, wherein the first matrix of the second input text comprises a Word2Vec matrix of the second input text, wherein encoding the second input text into the Word2Vec matrix of the second input text using the Word2Vec algorithm comprises embedding each word in the m number of words of the second input text into a k-dimensional Word2Vec vector using the Word2Vec algorithm resulting in the Word2Vec matrix of the second input text having a dimension of m×k, and wherein the dimension of n×k of the Word2Vec matrix of the first input text is different from the dimension m x k of the Word2Vec matrix of the second input text;

decoding the first matrix of the first input text into a second matrix of the first input text using a text embedding algorithm, wherein decoding the first matrix of the first input text into the second matrix of the first input text using the text embedding algorithm comprises decoding the Word2Vec matrix of the first input text into a first congruence derivative matrix using a congruence derivative vector representation, and wherein the first congruence derivative matrix has a first dimension;

decoding the first matrix of the second input text into a second matrix of the second input text using the text embedding algorithm, wherein decoding the first matrix of the second input text into the second matrix of the second input text using the text embedding algorithm comprises decoding the Word2Vec matrix of the second input text into a second congruence derivative matrix using the congruence derivative vector representation, wherein the second congruence derivative matrix has a second dimension, and wherein decoding the Word2Vec matrix of the first input text into the first congruence derivative matrix using the congruence derivative vector representation and decoding the Word2Vec matrix of the second input text into the second congruence derivative matrix using the congruence derivative vector representation results in the first dimension of the first congruence derivative matrix of the first input text being equal to the second dimension of the second congruence derivative matrix of the second input text; and

using, by the machine learning module, the first congruence derivative matrix and the second congruence derivative matrix as training data for a machine learning model, wherein the machine learning module requires matrices used for the training data to have uniform dimensions, and wherein using the first congruence derivative matrix and the second congruence derivative matrix as training data for a machine learning model comprises

creating, by the machine learning module, the machine learning model, and

passing, by the machine learning module, the first congruence derivative matrix and the second congruence derivative matrix through the machine learning model to train the machine learning model to determine mistakes in other input texts.

10. The computer-readable storage medium of claim 9 , wherein the first input text comprises multiple words in order, one or more sentences, one or more phrases, one or more paragraphs, or an entire document.

11. The computer-readable storage medium of claim 9 , wherein the second input text comprises multiple words in order, one or more sentences, one or more phrases, one or more paragraphs, or an entire document.

12. The computer-readable storage medium of claim 9 , wherein the machine learning module implements, based at least in part on the training data, short text classification.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 26, 2022
From: YIN, CHANGCHUAN; SAEED, SHAHZAD
To: AT&T MOBILITY II LLC
Reel/Frame 060026/0766 →
Continuity (1)
Related Publication 20240005082A1 · Jan 4, 2024
References Cited (48)
US 9037464B1 · Mikolov et al. · 2015 [cited by applicant]
US 11450124B1 · Zhang · 2022 [cited by examiner]
US 11625534B1 · Agarwal · 2023 [cited by examiner]
US 11847106B2 · Urdiales · 2023 [cited by examiner]
US 11972226B2 · Padfield · 2024 [cited by examiner]
US 20150066481A1 · Terrell · 2015 [cited by examiner]
US 20160321243A1 · Walia · 2016 [cited by examiner]
US 20180075368A1 · Brennan · 2018 [cited by examiner]
US 20180157644A1 · Mandt · 2018 [cited by examiner]
US 20180203851A1 · Wu · 2018 [cited by examiner]
US 20200126042A1 · Jawale · 2020 [cited by examiner]
US 20210192149A1 · Grail · 2021 [cited by examiner]
US 20210256069A1 · Grail · 2021 [cited by examiner]
US 20210319785A1 · Tandecki · 2021 [cited by examiner]
US 20210326675A1 · Oh · 2021 [cited by examiner]
US 20210374547A1 · Wang · 2021 [cited by examiner]
US 20220012763A1 · Sharma · 2022 [cited by examiner]
US 20220036003A1 · Bali · 2022 [cited by examiner]
US 20220076828A1 · Zhang · 2022 [cited by examiner]
US 20220100962A1 · Akhalwaya · 2022 [cited by examiner]
US 20220138424A1 · Gong · 2022 [cited by examiner]
US 20220180057A1 · Garcia Santa · 2022 [cited by examiner]
US 20220188513A1 · Kitamura · 2022 [cited by examiner]
US 20220198141A1 · Hudson · 2022 [cited by examiner]
US 20220237376A1 · Wang · 2022 [cited by examiner]
US 20220237567A1 · Tiwari · 2022 [cited by examiner]
US 20220247700A1 · Bhardwaj · 2022 [cited by examiner]
US 20220350825A1 · van de Nieuwegiessen · 2022 [cited by examiner]
US 20220358361A1 · Otsuka · 2022 [cited by examiner]
US 20220374426A1 · Thai · 2022 [cited by examiner]
US 20230012722A1 · Del Rosario · 2023 [cited by examiner]
US 20230034085A1 · Shalev · 2023 [cited by examiner]
US 20230074968A1 · Gookin · 2023 [cited by examiner]
US 20230195773A1 · Zhang · 2023 [cited by examiner]
US 20230214595A1 · Agarwal · 2023 [cited by examiner]
US 20230237272A1 · Teixeira de Abreu Pinho · 2023 [cited by examiner]
US 20230244868A1 · Yuan · 2023 [cited by examiner]
US 20230274102A1 · Marie · 2023 [cited by examiner]
US 20230274126A1 · Vijapur Gopinath Rao · 2023 [cited by examiner]
US 20230298627A1 · Marzorati · 2023 [cited by examiner]
US 20230334267A1 · Perez · 2023 [cited by examiner]
US 20230367969A1 · Chaturvedi · 2023 [cited by examiner]
US 20240111955A1 · Kirch · 2024 [cited by examiner]
Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” https://arxiv.org/pdf/1810.04805v2.pdf, May 24, 2019. [cited by applicant]
Mikolov et al., “Efficient Estimation of Word Representations in Vector Space,” https://arxiv.org/pdf/1301.3781.pdf, Sep. 7, 2013. [cited by applicant]
Mikolov et al., “Distributed Representations of Words and Phrases and their Compositionality,” Proceedings of the 26 [cited by applicant]
Pennington et al., “GloVe: Global Vectors for Word Representation,” Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, Oct. 25-29, 2014, pp. 1532-1543. [cited by applicant]
Yin et al., “Periodic power spectrum with applications in detection of latent periodicities in DNA sequences,” Journal of Mathematical Biology 73, Apr. 2016, pp. 1053-1079. [cited by applicant]