NEOANTIGEN IDENTIFICATION FOR T-CELL THERAPY
A method for identifying T-cells that are antigen-specific for at least one neoantigen that is likely to be presented on surfaces of tumor cells of a subject. Peptide sequences of tumor neoantigens are obtained by sequencing the tumor cells of the subject. The peptide sequences are input into a machine-learned presentation model to generate presentation likelihoods for the tumor neoantigens, each presentation likelihood representing the likelihood that a neoantigen is presented by an MHC allele on the surfaces of the tumor cells of the subject. A subset of the neoantigens is selected based on the presentation likelihoods. T-cells that are antigen-specific for at least one of the neoantigens in the subset are identified. These T-cells can be expanded for use in T-cell therapy. TCRs of these identified T-cells can also be sequenced and cloned into new T-cells for use in T-cell therapy.
1 - 39 . (canceled)
40 . A method comprising:
infusing T-cells into a subject, wherein the T-cells were produced by:
obtaining data representing peptide sequences of each of a set of neoantigens;
encoding the peptide sequences of each of the neoantigens into a corresponding numerical vector, each numerical vector including information regarding a plurality of amino acids that make up the peptide sequence and a set of positions of the amino acids in the peptide sequence;
inputting the numerical vectors, using a computer processor, into a machine-learned presentation model to generate a set of presentation likelihoods for the set of neoantigens, each presentation likelihood in the set representing the likelihood that a corresponding neoantigen is presented by one or more MHC alleles on the surface of the tumor cells of the subject, the machine-learned presentation model comprising:
a plurality of parameters identified at least based on a training data set comprising:
for each sample in a plurality of samples, a label obtained by mass spectrometry measuring presence of peptides bound to at least one MHC allele in a set of MHC alleles identified as present in the sample; and
for each of the samples, training peptide sequences encoded as numerical vectors including information regarding a plurality of amino acids that make up the peptides and a set of positions of the amino acids in the peptides;
a function representing a relation between the numerical vector received as input and the presentation likelihood generated as output based on the numerical vector and the parameters;
selecting a subset of the set of neoantigens based on the set of presentation likelihoods to generate a set of selected neoantigens;
identifying one or more T-cells that are antigen-specific for at least one of the neoantigens in the subset; and
returning the one or more identified T-cells.
41 . The method of claim 40 , wherein inputting the numerical vector into the machine-learned presentation model comprises:
applying the machine-learned presentation model to the peptide sequence of the neoantigen to generate a dependency score for each of the one or more MHC alleles indicating whether the MHC allele will present the neoantigen based on the particular amino acids at the particular positions of the peptide sequence.
42 . The method of claim 41 , wherein inputting the numerical vector into the machine-learned presentation model further comprises:
transforming the dependency scores to generate a corresponding per-allele likelihood for each MHC allele indicating a likelihood that the corresponding MHC allele will present the corresponding neoantigen; and
combining the per-allele likelihoods to generate the presentation likelihood of the neoantigen.
43 . The method of claim 40 , wherein the set of presentation likelihoods are further identified by at least one or more allele noninteracting features, and further comprising:
applying the machine-learned presentation model to the allele noninteracting features to generate a dependency score for the allele noninteracting features indicating whether the peptide sequence of the corresponding neoantigen will be presented based on the allele noninteracting features.
44 . The method of claim 40 , wherein the one or more MHC alleles include two or more different MHC alleles.
45 . The method of claim 40 , wherein the peptide sequences comprise peptide sequences having lengths other than 9 amino acids.
46 . The method of claim 40 , wherein encoding the peptide sequence comprises encoding the peptide sequence using a one-hot encoding scheme.
47 . The method of claim 40 , wherein the plurality of samples comprise at least one of:
(a) one or more cell lines engineered to express a single MHC allele;
(b) one or more cell lines engineered to express a plurality of MHC alleles;
(c) one or more human cell lines obtained or derived from a plurality of patients;
(d) fresh or frozen tumor samples obtained from a plurality of patients; and
(e) fresh or frozen tissue samples obtained from a plurality of patients.
48 . The method of claim 40 , wherein the training data set further comprises at least one of:
(a) data associated with peptide-MHC binding affinity measurements for at least one of the peptides; and
(b) data associated with peptide-MHC binding stability measurements for at least one of the peptides.
49 . The method of claim 40 , wherein the set of presentation likelihoods are further identified by at least expression levels of the one or more MHC alleles in the subject, as measured by RNA-seq or mass spectrometry.
50 . The method of claim 40 , wherein selecting the set of selected neoantigens comprises:
selecting neoantigens that have an increased likelihood of being presented on the tumor cell surface relative to unselected neoantigens based on the machine-learned presentation model;
selecting neoantigens that have an increased likelihood of being capable of inducing a tumor-specific immune response in the subject relative to unselected neoantigens based on the machine-learned presentation model;
selecting neoantigens that have an increased likelihood of being capable of being presented to naive T-cells by professional antigen presenting cells (APCs) relative to unselected
neoantigens based on the presentation model, optionally wherein the APC is a dendritic cell (DC); or
selecting neoantigens that have a decreased likelihood of being subject to inhibition via central or peripheral tolerance relative to unselected neoantigens based on the machine-learned presentation model.
51 . The method of claim 40 , wherein the one or more tumor cells are selected from the group consisting of: lung cancer, melanoma, breast cancer, ovarian cancer, prostate cancer, kidney cancer, gastric cancer, colon cancer, testicular cancer, head and neck cancer, pancreatic cancer, brain cancer, B-cell lymphoma, acute myelogenous leukemia, chronic myelogenous leukemia, chronic lymphocytic leukemia, and T-cell lymphocytic leukemia, non-small cell lung cancer, and small cell lung cancer.
52 . The method of claim 40 , wherein the machine-learned presentation model is a neural network model.
53 . The method of claim 52 , wherein the neural network model includes a plurality of network models for the MHC alleles, each network model assigned to a corresponding MHC allele of the MHC alleles and including a series of nodes arranged in one or more layers.
54 . The method of claim 52 , wherein the neural network model is trained by updating the parameters of the neural network model, and wherein the parameters of at least two network models are jointly updated for at least one training iteration.
55 . The method of claim 40 , wherein identifying the one or more T-cells comprises co-culturing the one or more T-cells with one or more of the neoantigens in the subset under conditions that expand the one or more T-cells.