IP Library Granted Patent US 12,299,395
Granted Patent B2
US 12,299,395 · App. 18/148,374 · Granted May 13, 2025

Learning embedded representation of a correlation matrix to a network with machine learning

Inventors: Bhaskarjit Sarmah (Gurgaon, IN); Nayana Nair (Gurgaon, IN); Dhagash Mehta (Chester Springs, PA); Stefano Pasquali (New York, NY)
Assignee: BlackRock Finance, Inc.
G06F40/289G06N3/0495G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,395
App. No.
18/148,374
Granted
May 13, 2025
Kind
B2
Abstract

System, method, and a computer program product for generating embeddings are provided. A machine learning framework generates a fully connected network from a dataset associated with words. The words correspond to nodes in the fully connected network. The weights are associated with correlations between the nodes and correspond to the links in the fully connected network. The machine learning framework transforms the correlations corresponding to the links into distances. The machine learning framework generates a sparse network from the fully connected network based on the distances. From the sparse network, machine learning framework determines sentence structures by traversing the nodes. Using the sentence structures, the machine learning framework uses a neural network to generate embeddings in the embedded space.

Claims (48)

1. A method for generating embeddings, the method comprising:

generating a fully connected network associated with words, wherein a node in the fully connected network includes a word and a link between a pair of nodes is a correlation between a pair of words corresponding to the pair of nodes;

converting correlations associated with links in the fully connected network into distances;

converting, using a sparse algorithm, the fully connected network into a sparse network based on the distances;

traversing at least one node in the sparse network to generate sentence structures;

generating, using a neural network, the embeddings from the sentence structures in an embedded space; and

analyzing the embeddings to determine relationships between nodes.

2. The method of claim 1 , further comprising:

determining a first numerical value associated with the word at a first point in time;

determining a second numerical value associated with the word at a second point in time; and

determining a logarithmic return based on the first numerical value and the second numerical value.

3. The method of claim 1 , further comprising:

determining the correlation between the pair of words based on logarithmic returns associated with a first word and a second word in the pair of words.

4. The method of claim 1 , wherein the distances are based on logarithmic returns associated with the words in the fully connected network.

5. The method of claim 1 , wherein the fully connected network is represented by a correlation matrix that includes the words associated with the nodes for rows and columns and correlations for pairs of words as entries in the correlation matrix.

6. The method of claim 5 , further comprising:

converting the correlation matrix into a distance matrix, wherein the distance matrix includes distances as entries in the distance matrix for the pairs of words.

7. The method of claim 1 , wherein the sparse algorithm includes a minimum spanning tree algorithm that removes at least one link in the links from the fully connected network.

8. The method of claim 1 , wherein a node-to-vector algorithm traverses the sparse network to generate the sentence structures based on hyperparameters.

9. The method of claim 8 , further comprising:

tuning the hyperparameters, wherein the tunning the hyperparameters varies words in the sentence structures.

10. The method of claim 1 , wherein a word-to-vector algorithm and the neural network generates the embeddings associated with the words in the fully connected network.

11. The method of claim 1 , wherein the embeddings capture syntactic relationships among the words associated with the nodes in the fully connected network.

12. The method of claim 1 , wherein the words correspond to stocks and the correlations correspond to prices associated with the stocks.

13. A system for generating embeddings, the system comprising:

a memory configured to store a machine learning framework; and

a processor coupled to the memory and configured to cause the machine learning framework to perform operations, the operations comprising:

generating a fully connected network associated with words, wherein a node in the fully connected network includes a word and a link between a pair of nodes is a correlation between a pair of words corresponding to the pair of nodes;

converting correlations associated with links in the fully connected network into distances;

converting, using a sparse algorithm, the fully connected network into a sparse network based on the distances;

traversing at least one node in the sparse network to generate sentence structures; and

generating, using a neural network in the machine learning framework, the embeddings in an embedded space from the sentence structures.

14. The system of claim 13 , wherein the operations further comprise:

determining the correlation between the pair of words based on logarithmic returns associated with numerical features of a first word and a second word in the pair of words.

15. The system of claim 13 , wherein the fully connected network is represented by a correlation matrix that includes the words associated with nodes for rows and columns and correlations for pairs of words as entries in the correlation matrix.

16. The system of claim 15 , wherein the operations further comprise:

converting the correlation matrix into a distance matrix, wherein the distance matrix includes distances as entries in the distance matrix for the pairs of words.

17. The system of claim 13 , wherein the sparse algorithm includes a minimum spanning tree algorithm that removes a subset of links in the links from the fully connected network.

18. The system of claim 13 , wherein a node-to-vector algorithm traverses the sparse network to generate the sentence structures having predefined lengths.

19. The system of claim 13 , wherein the operations further comprise:

training the neural network to generate the embeddings from the sentence structures and corresponding target words.

20. A non-transitory computer readable medium having instructions stored thereon, that when executed by a processor cause the processor to perform operations for generating embeddings, the operations comprising:

generating a fully connected network associated with words, wherein each node in the fully connected network is a word and a link between a pair of nodes is a correlation of log returns of a pair of words corresponding to the pair of nodes;

generating a correlation matrix, wherein entries in the correlation matrix correspond to links in the fully connected network;

converting the correlation matrix into a distance matrix;

generating a sparse network from the distance matrix, wherein the sparse network has the same number of nodes and fewer links than the fully connected network;

generating, using a node-to-vector algorithm, sentence structures from the sparse network; and

generating using a word-to-vector algorithm, the embeddings for words from the sentence structures.

Assignments (2)
MERGER AND CHANGE OF NAME Recorded Oct 4, 2024
From: BLACKROCK, INC.; BANANA MERGER SUB, INC.
To: BLACKROCK FINANCE, INC.
Reel/Frame 069113/0616 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 25, 2023
From: SARMAH, BHASKARJIT; NAIR, NAYANA; MEHTA, DHAGASH; PASQUALI, STEFANO
To: BLACKROCK, INC.
Reel/Frame 062482/0472 →
Priority Claims (1)
IN 202211038745 · Jul 6, 2022 · national
Continuity (1)
Related Publication 20240012997A1 · Jan 11, 2024
References Cited (20)
CN 113032539A · 2021 [cited by examiner]
CN 109885683B · 2022 [cited by examiner]
WO WO2025029357A1 · 2025 [cited by examiner]
Long Short-term memory network for learning sentences similarity using deep contextual embeddings (Year: 2021). [cited by examiner]
Sentence embeddings and their relation with sentence structure (Year: 2024). [cited by examiner]
Word2vec node2vec graph2vec X2vec towards a theory of vector embeddings of structured data (Year: 2020). [cited by examiner]
Mantegna et al., “Hierarchical Structure in Financial Markets.” The European Physical Journal B, 1999, pp. 193-197. [cited by applicant]
Letizia et al., “Corporate Payments Networks and Credit Risk Rating.” EPJ Data Science, 2019, pp. 1-29. [cited by applicant]
Brandes, “On Variants of Shortest-Path Betweenness Centrality and their Generic Computation.” Social Networks, 2008, pp. 1-22. [cited by applicant]
Tang et al., “Complexities in Financial Network Topological Dynamics: Modeling of Emerging and Developed Stock Markets.” Complexity, 2018, pp. 1-31. [cited by applicant]
Huang et al., “A Financial Network Perspective of Financial Institutions' Systemic Risk Contributions.” Physica A, 2016, pp. 183-196. [cited by applicant]
Brede, “Networks-An Introduction.” Mark E.J. Newman, 2010, Oxford University Press, pp. 241-242. [cited by applicant]
Banerjee et al., “The Diffusion of Microfinance.” Science, 2013, vol. 341, 1236498, 9 pages. [cited by applicant]
Grover et al., “node2vec: Scalable Feature Learning for Networks.” Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2016, pp. 855-864. [cited by applicant]
Mikolov et al., “Distributed Representations of Words and Phrases and their Compositionality.” Advances in Neural Information Processing Systems, 2013, pp. 1-9. [cited by applicant]
Caselles-Dupre et al., “Word2vec Applied to Recommendation: Hyperparameters Matter.” Proceedings of the 12th ACM Conference on Recommender Systems, 2018, pp. 352-356. [cited by applicant]
Faruqui et al., “Problems With Evaluation of Word Embeddings Using Word Similarity Tasks.” arXiv preprint arXiv:1605.02276, 2016, 6 pages. [cited by applicant]
Dolphin et al., “Stock Embeddings: Learning Distributed Representations for Financial Assets.” arXiv preprint arXiv:2202.08968, 2022, 9 pages. [cited by applicant]
Bolukbasi et al., “Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings.” Advances in Neural Information Processing Systems, 2016, pp. 4349-4357. [cited by applicant]
Rosenberg et al., “V-Measure: A Conditional Entropy-Based External Cluster Evaluation Measure.” Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Lang… [cited by applicant]