IP Library Granted Patent US 12,451,251
Granted Patent B2
US 12,451,251 · App. 17/773,263 · Granted Oct 21, 2025

Method for determining a long-term survival prognosis of breast cancer patients, based on algorithms modelling biological networks

Inventors: Caterina Anna Maria La Porta (Milan, IT); Stefano Zapperi (Milan, IT); Francesc Font Clos (Milan, IT)
Assignee: COMPLEXDATA S.R.L.
G16H50/20G16B5/00G16B40/00G16H50/30G16H50/50
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,451,251
App. No.
17/773,263
Granted
Oct 21, 2025
Kind
B2
Abstract

A method is described for determining a survival prognosis of a patient suffering from a breast tumor, using processing carried out by electronic processing and/or calculation means. The method first comprises step (a) of defining a biological network representative of a particular biological process associated with the breast tumor. The biological network comprises a plurality of nodes, a set of directional relationships between these nodes and a set of genes associated with these nodes. The method also includes step (b) of accessing a data set related to the patient, comprising gene expressions in a biological sample of the tumor isolated from the patient; and step (c) of calculating a continuous expression value for the aforesaid nodes of the biological network. If the node is associated with only one gene and it is found that the gene is present in the biological sample, the continuous expression value of the node is calculated as the expression of the associated gene detected in the biological sample. If the node is associated with multiple genes, and it is found that at least one of the aforesaid genes is present in the biological sample, the continuous expression value of the node is calculated based on the expressions of the associated genes, present in the biological sample. If the node is not associated with any gene, or the associated gene is not found in the biological sample, the node is marked as a node not associated with a continuous expression value. The method then comprises the following steps, carried out by the electronic processing and/or calculation means: (d) binarizing the data set of continuous expression values calculated for each node of the biological network to which a continuous expression value is associated, based on a comparison of the continuous expression value with a respective threshold, to thus obtain a first binarized data set of the nodes, obtained based on the detections made; (e) calculating an aggressiveness score based on the aforesaid first binarized data set of the nodes; and finally (f) determining a survival prognosis result based on the aforesaid aggressiveness score calculated.

Claims (52)

1. A method for determining a survival prognosis of a patient suffering from a breast tumor, using processing carried out by electronic processing and/or calculation means, said method comprising the steps of:

(a) defining a biological network representative of a particular biological process associated with the breast tumor,

said biological network comprising a plurality of nodes, a set of directional relationships between said nodes and a set of genes associated with said nodes,

wherein each node represents a gene and/or a protein and/or a complex of multiple proteins and/or another molecule and/or an ion present in a cell or in contact therewith and/or a particular external condition to which the cell is subjected and/or states in which the cell can be found,

wherein each directional relationship is defined by a source node, a target node and an interaction type, and in which the interaction type comprises inhibition interaction, in an inhibition relationship, or excitation interaction, in an excitation relationship, or absence of interaction, said directional relationship being determined from the source node to the target node;

(b) accessing a patient-related data set, said data set comprising expressions of genes in a biological sample of said tumor isolated from the patient;

(c) calculating a continuous expression value for said nodes of the biological network, wherein said calculation step comprises:

if the node is associated with only one gene and if, based on said patient data set, it is found that the gene is present in the biological sample, calculating the continuous expression value of the node as the expression of the associated gene detected in the biological sample;

if the node is associated with multiple genes, and if, based on said patient data set, it is found that at least one of said genes is present in the biological sample, calculating the continuous expression value of the node based on the expressions of the associated genes, which are present in the biological sample;

if the node is not associated with any gene, or if, based on said patient data set, the associated gene is not found in the biological sample, marking the node as a node not associated with a continuous expression value;

(d) carrying out, by said electronic processing and/or calculation means, a binarization of the data set of continuous expression values calculated for each node of the biological network with which a continuous expression value is associated, based on a comparison of the continuous expression value to a respective threshold, in order to obtain a first binarized data set of the nodes, obtained on the basis of said calculating step (c);

(e) calculating, by said electronic processing and/or calculation means, an aggressiveness score, derivable from the first binarized data set of the nodes, based on said first binarized data set of the nodes; and

(f) determining a result of the survival prognosis based on said calculated aggressiveness score,

wherein step (e) of calculating an aggressiveness score comprises calculating the aggressiveness score as the energy of the patient's biological sample and/or the fraction of the out-of-balance nodes of the patient's biological sample and/or based on a projection of the data set on the main component thereof through a Principal Component Analysis methodology and/or based on a projection of the data set on a simulated binarized data set first principal component, through a Principal Component Analysis methodology,

wherein the step (e) of calculating an aggressiveness score further comprises:

simulating, through computational simulation, the biological network, through simulated samples which represent possible cellular states described by the biological network, in order to obtain a simulated binarized data set; and

defining the aggressiveness score based on a projection of the data set on the simulated binary data set first principal component, through a Principal Component Analysis methodology, and

wherein said step of simulating the biological network comprises:

numerically calculating the state of the biological network from a certain initial condition chosen randomly;

updating, by means of a simulation algorithm, the binarized value of each node, in a sequential manner, changing the state thereof so that there is an increase in the number of satisfied directional relationships among the directional relationships of which said node is the target; and

iterating said updating step, until a steady state of the biological network is achieved, in which each node is in a binary state which satisfies the majority of the directional relationships of which said node is the target, wherein said steady state corresponds to a possible state of the cell.

2. A method according to claim 1 , wherein step (f) of determining a survival prognosis result comprises calculating an estimated survival probability based on the calculated aggressiveness score.

3. A method according to claim 2 , wherein the step of calculating the estimated survival probability comprises calculating the estimated survival probability based on the aggressiveness score,

wherein a relationship between the estimated survival probability and the aggressiveness score is predetermined based on experimental data,

or wherein said relationship is obtained by processing with one or more trained predictive algorithms, wherein the one or more trained algorithms comprise machine learning algorithms, and wherein the training step is carried out on an available experimental data set, which are partitioned into a test set and a training set.

4. A method according to claim 2 , wherein the step of calculating an estimated survival probability comprises calculating the estimated survival probability as a function of time,

wherein the step of calculating an estimated probability further comprises:

classifying the patient into a high risk group or a low risk group, wherein the high and low risk groups are defined with respect to a lower quantile value threshold or a higher quantile value threshold of the survival probability distribution with respect to the aggressiveness score values;

assigning to the patient a probability prospect associated with the group to which the patient belongs, in which the patient was classified, said probability prospect associated with the group to which the patient belongs for each of the high-risk and low-risk groups being calculated upstream in a training step, based on experimental data.

5. A method according to claim 4 , further comprising the step of providing the patient with a personalized report, containing the calculated aggressiveness score value and/or the determined risk group to which the patient belongs.

6. A method according to claim 1 , adapted to determine a survival prognosis of a patient suffering from a specific molecular sub-type of breast tumor, wherein step (a) of defining comprises:

(a) defining, based on known medical/biological knowledge related to said molecular sub-type of breast tumor, a biological network representative of a particular biological process associated with the molecular sub-type of tumor.

7. A method according to claim 1 , wherein the step (c) of calculating the continuous expression value of the node, if the node is associated with multiple genes, comprises calculating the continuous expression value of the node as the maximum value among the expression values of the associated genes, or as the minimum value among the expression values of the associated genes, or as the average of the expression values of the associated genes.

8. A method according to claim 1 , further comprising a step (g) of defining, for each node with which a continuous expression value is associated, a respective threshold, and attributing to the node a first binary value, if the continuous expression value of the node is less than said respective threshold, and a second binary value, if the continuous expression value of the node is greater than said respective threshold.

9. A method according to claim 8 , wherein the first binary value is −1 and the second binary value is +1.

10. A method according to claim 8 , adapted to determine a survival prognosis of a patient suffering from a specific molecular sub-type of breast tumor, wherein step (a) of defining comprises:

(a) defining, based on known medical/biological knowledge related to said molecular sub-type of breast tumor, a biological network representative of a particular biological process associated with the molecular sub-type of tumor, and

wherein the step (g) of predefining the threshold comprises defining the threshold based on a preliminary processing carried out on a set of available experimental data related to a plurality of tumor samples of the particular molecular sub-type considered, from patients whose clinical history and survival time are known.

11. A method according to claim 10 , wherein the step of defining the threshold comprises defining the threshold based on machine learning algorithms using available experimental data sets of patients, whose clinical history is known,

and/or wherein the step of defining a threshold comprises defining the threshold based on a methodology based on Gaussian kernels or based on the average expression.

12. A method according to claim 10 , wherein the step of defining the threshold comprises defining the threshold based on a free energy binarization methodology,

wherein said free energy binarization methodology comprises calculating the threshold based on a minimization of the difference between the average energy of all the samples, belonging to said plurality of samples associated with the available experimental data set, and the sum of the entropy of all the nodes of the biological network multiplied by a correction parameter.

13. A method according to claim 1 , wherein the step (e) of calculating an aggressiveness score comprises calculating the aggressiveness score in a plurality of said methodologies, and the method further comprises:

carrying out a plurality of instances of the method according to claim 1 , using a respective aggressiveness score, among the aggressiveness scores calculated according to a respective methodology of the aforementioned calculation methodologies, on each of the cases of an experimental database of cases of patients affected by a specific molecular sub-type of breast tumor, the prognosis of which is known;

selecting, for the molecular sub-type of tumor considered, the calculation methodology which provides the most accurate prediction;

validating the selected calculation methodology on a further experimental data set of cases the prognosis of which is known, independent of said experimental database, to avoid “overfitting”.

14. A method according to claim 1 , wherein the step (e) of calculating an aggressiveness score comprises calculating the aggressiveness score as the energy of the patient's biological sample, defined as the number of relationships satisfied in the respective biological network minus the number of unsatisfied relationships in the respective biological network,

wherein a satisfied relationship is defined as an excitation relationship in which the source node and the target node take the same value, or as an inhibition relationship in which the source node and the target node take different values,

and wherein an unsatisfied relationship is defined as an excitation relationship in which the source node and the target node take different values, or an inhibition relationship in which the source node and the target node take the same value.

15. A method according to claim 1 , wherein the step (e) of calculating an aggressiveness score comprises calculating the aggressiveness score as the fraction of out-of-balance nodes of the patient's biological sample, wherein the fraction of out-of-balance nodes is defined as the number of out-of-balance nodes divided by the total number of nodes in the biological network.

16. A method according to claim 1 , wherein the step (e) of calculating an aggressiveness score comprises calculating the aggressiveness score based on a projection of the data set on the principal component thereof through a Principal Component Analysis methodology.

17. A method according to claim 1 , wherein said steps of numerically calculating, updating, and iterating are repeated several times, with different initial conditions, to represent a plurality of potential states in which the cell can be found.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2022
From: LA PORTA, CATERINA ANNA MARIA; ZAPPERI, STEFANO; FONT CLOS, FRANCESC
To: COMPLEXDATA S.R.L.
Reel/Frame 059831/0699 →
Priority Claims (1)
IT 102019000023946 · Dec 13, 2019 · national
Continuity (1)
Related Publication 20230145332A1 · May 11, 2023
References Cited (32)
US 11091809B2 · Harkin · 2021 [cited by examiner]
US 20150331992A1 · Arnon Jerby · 2015 [cited by examiner]
US 20160078167A1 · Rosner · 2016 [cited by examiner]
US 20160110496A1 · Taylor · 2016 [cited by examiner]
US 20180320237A1 · Snijders · 2018 [cited by examiner]
US 20190279769A1 · Conzen · 2019 [cited by examiner]
WO WO2008100352A2 · 2008 [cited by examiner]
WO WO2010104472A1 · 2010 [cited by examiner]
WO WO2013147330A1 · 2013 [cited by examiner]
WO 2019178217A1 · 2019 [cited by applicant]
WO WO2019233028A1 · 2019 [cited by examiner]
Al-Ejeh et al., “Meta-analysis of the global gene expression profile of triple-negative breast cancer identifies genes for the prognostication and treatment of aggressive breast cancer,” Oncogenesis (2014) 3, e100; doi:… [cited by examiner]
Guo et al., “Application of a co-expression network for the analysis of aggressive and non-aggressive breast cancer cell lines to predict the clinical outcome of patients, ” Molecular Medicine Reports 16: 7967-7978, 201… [cited by examiner]
Meng et al., “Biomarker discovery to improve prediction of breast cancer survival: using gene expression profiling, meta-analysis, and tissue validation,” Onco Targets and Therapy, , 6177-6185, DOI: 10.2147/OTT.S113855 … [cited by examiner]
Haibe-Kains et al., “Comparison of prognostic gene expression signatures for breast cancer,” BMC Genomics 2008, 9:394 doi: 10.1186/1471-2164-9-394 (Year: 2008). [cited by examiner]
Van de Vijver et al., “A Gene-Expression Signature as a Predictor of Survival in Breast Cancer,” N Engl J Med, vol. 347, No. 25.⋅ Dec. 19, 2002. (Year: 2002). [cited by examiner]
Riis et al., “Gene Expression Profile Analysis of T1 and T2 Breast Cancer Reveals Different Activation Pathways,” Hindawi Publishing Corporation, ISRN Oncology, vol. 2013, Article ID 924971, 12 pages, http://dx.doi.org/… [cited by examiner]
Tang et al., “Overexpression of ASPM, CDC20, and TTK Confer a Poorer Prognosis in Breast Cancer Identified by Gene Co-expression Network Analysis,” Front. Oncol. 9:310; doi: 10.3389/fonc.2019.00310. (Year: 2019). [cited by examiner]
Yang et al., “Network-Based Inference Framework for Identifying Cancer Genes from Gene Expression Data,” Hindawi Publishing Corporation, BioMed Research International, vol. 2013, Article ID 401649, 12 pages, http://dx.d… [cited by examiner]
Iuliano et al., “Combining Pathway Identification and Breast Cancer Survival Prediction via Screening-Network Methods,” Front. Genet. 9:206; doi: 10.3389/fgene.2018.00206. (Year: 2018). [cited by examiner]
Pepke et al., “Comprehensive discovery of subsample gene expression components by information explanation: therapeutic implications in cancer,” Pepke and Ver Steeg BMC Medical Genomics (2017) 10:12; DOI 10.1186/s12920-0… [cited by examiner]
Zhang et al., “Identification of novel prognostic indicators for triple-negative breast cancer patients through integrative analysis of cancer genomics data and protein interactome data,” Oncotarget, vol. 7, No. 44 (Yea… [cited by examiner]
International Search Report and Written Opinion for International Patent Application No. PCT/IB2020/061014, mailed Feb. 25, 2021, 18 pages. [cited by applicant]
H Bonsang-Kitzis et al: “Biological network-driven gene selection identifies a stromal imune module as a key determinant of triple-negative breast carcinoma prognosis”, Oncoimmunology, vol. 5, No. 1, Jan. 2, 2016 (Jan. … [cited by applicant]
Brian D. Lehmann et al: “Refinement of Triple-Negative Breast Cancer Molecular Subtypes: Implications for Neoadjuvant Chemotherapy Selection”, PLOS ONE, vol. 11, No. 6, Jun. 16, 2016 (Jun. 16, 2016), p. e0157368, XP0553… [cited by applicant]
Ozturk Kivilcim et al: “The Emerging Potential for Network Analysis to Inform Precision Cancer Medicine”, Journal of Molecular Biology, Academic Press, United Kingdom, vol. 430, No. 18, Jun. 15, 2018 (Jun. 15, 2018), pp… [cited by applicant]
Van't Veer LJ, Dai H, van de Vijver MJ, He YD, Hart AA, et al. (2002) “Gene expression profiling predicts clinical outcome of breast cancer”. Nature 415: 530-536. [cited by applicant]
Paik S, Shak S, Tang G, Kim C, Baker J, et al. (2004) “A multigene assay to predict recurrence of tamoxifen-treated, hode-negative breast cancer”. N Engl J Med 351: 2817-2826. [cited by applicant]
Lehmann B.D., Bauer J.A., Chen X., Sanders M.E., Chakravarthy A.B., Shyr Y. and Pietenpol J.A. 2011: “Identification of Human Triple-Negative Breast Cancer Subtypes and Preclinical Models for Selection of Targeted Thera… [cited by applicant]
Alter O., Brown P.O., and Botstein D. 2000: “Singular Value Decomposition for Genome-Wide Expression Data Processing and Modeling.” Proceedings of the National Academy of Sciences of the United States of America 97 (18)… [cited by applicant]
Berry M., Dumais S., O'Brien G. 1995: “Using Linear Algebra for Intelligent Information Retrieval.” SIAM Review 37 (4): 573-95. [cited by applicant]
Font-Clos F., Zapperi S., La Porta C.A.M. 2018: “Topography of Epithelial-Mesenchymal Plasticity.” Proceedings of the National Academy of Sciences of the United States of America 115 (23): 5902-7. [cited by applicant]