IP Library Granted Patent US 12,632,623
Granted Patent B2
US 12,632,623 · App. 18/814,374 · Granted May 19, 2026

Modeling physical systems with large language machine-learned models

Inventors: Paul Maragakis (New York, NY); Andreas Kraemer (Brooklyn, NY); James P. Roney (Cambridge, MA); Peter Skopp (Westport, CT)
Assignee: D. E. Shaw Research, LLC
G06F30/27G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,623
App. No.
18/814,374
Granted
May 19, 2026
Kind
B2
Abstract

A predictive system may access a set of physical structures corresponding to a physical system. Each physical structure representative of a configuration. The predictive system may encode the accessed physical structures to produce a set of encoded physical structures by encoding, for each accessed physical structure, a position of each constituent unit of the physical system within the accessed physical structure. The predictive system may train a machine-learned model using the encoded physical structures. The predictive system may retrain the machine-learned model by iteratively: accessing a set of two or more candidate physical structures, determining a first energy difference among the set of candidate physical structures, obtaining a second energy difference between a set of physical structures corresponding to the set of candidate physical structures using a method to calculate reference energy values.

Claims (58)

1 . A method comprising:

accessing, by a predictive system, a set of physical structures corresponding to a physical system, each physical structure representative of a configuration of the physical system;

encoding, by the predictive system, the accessed physical structures to produce a set of encoded physical structures by encoding, for each accessed physical structure, a position of each constituent unit of the physical system within the accessed physical structure;

training, by the predictive system, a machine-learned transformer model, the machine-learned transformer model comprising one or more attention blocks configured to model interactions of constituent units by generating activation outputs as relative energy values that correspond to a physical structure distribution; and

retraining, by the predictive system, the machine-learned transformer model by iteratively:

accessing a set of candidate physical structures for the physical system in a training sample;

converting, through forward propagation of the machine-learned transformer model, the candidate physical structures into a plurality of token sequences and corresponding activation outputs, wherein each token sequence represents a candidate physical structure of the physical system,

determining a first energy difference between two candidate physical structures, the first energy difference being an aggregation of the corresponding activation outputs from the attention blocks;

comparing the first energy difference to a second energy difference between a set of physical structures corresponding to the set of candidate physical structures, wherein the second energy difference is derived from a ground truth distribution in the training sample using an energy function;

backpropagating the second energy difference to the transformer model to adjust one or more parameters in the attention blocks of the machine-learned transformer model; and

retraining the machine-learned transformer model using the set of candidate physical structures, the first energy difference, and the second energy difference; and

receiving a sequence of a target molecule; and

using the one or more attention blocks in the trained machine-learned transformer model to determine energy values over a set of physical structures of the target molecule.

2 . The method of claim 1 , wherein encoding a particular accessed physical structure of the accessed physical structures comprises:

converting one or more structural expressions of the physical system into a sequence string representation of the physical system;

tokenizing the sequence string representation of the physical system to produce a tokenized structural expression;

tokenizing, for each constituent unit in the physical system, coordinates for the constituent unit within the particular accessed physical structure to produce tokenized coordinates for the constituent unit; and

combining the tokenized structural expression and the tokenized coordinates for each constituent unit to produce a tokenized physical structure corresponding to the particular accessed physical structure.

3 . The method of claim 2 , wherein tokenizing coordinates for the constituent unit within the particular accessed physical structure comprises:

pixelating a rendered sphere to produce a set of pixels each corresponding to a location on the surface of the rendered sphere;

tokenizing, for a first constituent unit in the physical system, coordinates at a center of the sphere; and

tokenizing, for each additional constituent unit in the physical system, coordinates corresponding to a pixel selected from the set of pixels based on a location of the additional constituent unit relative to the center of the sphere.

4 . The method of claim 2 , wherein tokenizing coordinates for the constituent unit within the particular accessed physical structure comprises using a Cartesian coordinate system, an xyz coordinate system, an octree coordinate system, a polar coordinate system, a cylindrical coordinate system, or a barycentric coordinate system to generate the coordinates.

5 . The method of claim 2 , wherein the sequence string representation of the physical system includes an ordered set of constituent units, and wherein the coordinates for a constituent unit comprise coordinates relative to a center of a coordinate sphere.

6 . The method of claim 1 , wherein the machine-learned transformer model is iteratively retrained until one or more retraining criteria is satisfied.

7 . The method of claim 6 , wherein the retraining criteria is satisfied when performance measurement of a holdout set starts to decrease.

8 . The method of claim 1 , wherein retraining the machine-learned transformer model comprises modifying weights of one or more layers of the machine-learned transformer model to minimize a difference between the first energy difference and the second energy difference over subsequent iterations.

9 . The method of claim 1 , wherein retraining the machine-learned transformer model comprises modifying weights of one or more layers of the machine-learned transformer model by:

generating a set of token sequences representing two or more physical structures;

store activation outputs corresponding to the tokens in the token sequences;

determining an aggregated energy state for each token sequence; and

determining backpropagated gradients based on comparing the aggregated energy states for the token sequences.

10 . The method of claim 1 , wherein the first energy difference is determined based at least in part on a Boltzmann probability ratio between the set of candidate physical structures.

11 . The method of claim 1 , wherein the machine-learned transformer model is configured to:

determine, for a candidate physical structure, a force associated with each constituent unit in the physical system based on a relative energy associated with a plurality of neighboring positions for the constituent unit.

12 . The method of claim 11 , wherein the machine-learned transformer model is retrained further based on a difference in forces associated with the candidate physical structures.

13 . The method of claim 11 , wherein the force associated with each constituent unit in the machine-learned transformer model is based on a collective gradient of the relative energy associated with the plurality of neighboring positions for the constituent unit, wherein the collective gradient is determined based on a vector to a neighboring position from which a finite difference gradient is calculated.

14 . The method of claim 11 , wherein a determined force associated with a constituent unit in the physical system is based on forces associated with one or more preceding constituent units in a sequence string representation of the physical system.

15 . The method of claim 11 , wherein the relative energies associated with each of the plurality of neighboring positions for the constituent unit are determined using a softmax output of the machine-learned transformer model.

16 . The method of claim 11 , wherein a determined force associated with a constituent unit in the physical system is based on, for each of the plurality of neighboring positions for the constituent unit, the relative energy associated with the neighboring position and a distance between the constituent unit and the neighboring position.

17 . The method of claim 11 , wherein a determined force generated by the machine-learned transformer model associated with a constituent unit in the physical system is based additionally on reference forces applied from a solvent or solution on the constituent unit in the physical system.

18 . The method of claim 11 , wherein a determined force generated by the machine-learned transformer model associated with a constituent unit in the physical system is based additionally on inherent forces in a physical system comprising the physical system.

19 . The method of claim 1 , wherein a constituent unit in an encoded physical structure is encoded with multiple tokens, each token representing decreasing significance of the constituent unit's coordinates.

20 . The method of claim 1 , further comprising fine tuning, by the predictive system, the machine-learned transformer model by using experimentally determined structures of the physical system, optionally wherein the experimentally determined structures comprises a crystallographic structure of the physical system.

21 . The method of claim 1 , further comprising:

generating, by the predictive system, a probability distribution of structures for a physical system or a portion thereof.

22 . The method of claim 21 , further comprising:

comparing a target probability distribution and the generated probability distribution to predict a binding affinity between a first physical system or a portion thereof and a second physical system or a portion thereof.

23 . The method of claim 1 , wherein obtaining the second energy difference comprises performing a quantum mechanical energy calculation.

24 . The method of claim 1 , wherein encoding an accessed physical structure comprises using a roto-translational transformation to encode spatial data of one or more constituent units in the physical system.

25 . The method of claim 1 , wherein the machine-learned transformer model is a pretrained machine-learned language model and retraining the machine-learned transformer model comprises fine tuning the machine-learned transformer model using training samples of physical systems.

26 . The method of claim 1 , further comprising:

computing importance weights based on the first and second energy differences, determining confidence levels, or selecting configurations based on the first and second energy differences.

27 . The method of claim 1 , wherein encoding the accessed physical structures comprises generating a series of tokens, and generating the series of tokens comprises:

generating, for a first constituent unit, a first set of tokens; and

generating, for a second constituent unit, a second set of tokens based on the first set of tokens.

28 . The method of claim 1 , further comprising:

generating, using the machine-learned transformer model, a molecular sequence that is predicted to be binding to the target molecule.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2024
From: MARAGAKIS, PAUL; KRAEMER, ANDREAS; RONEY, JAMES P.; SKOPP, PETER
To: D. E. SHAW RESEARCH, LLC
Reel/Frame 069635/0980 →
Continuity (1)
Related Publication 20260057148A1 · Feb 26, 2026
References Cited (45)
US 7684975B2 · Aoki et al. · 2010 [cited by applicant]
US 8463759B2 · Johnson · 2013 [cited by applicant]
US 9537504B1 · Guilford et al. · 2017 [cited by applicant]
US 11615246B2 · Reisswig · 2023 [cited by applicant]
US 20180082171A1 · Merity et al. · 2018 [cited by applicant]
US 20180373705A1 · Kwon · 2018 [cited by applicant]
US 20220398445A1 · Polleri et al. · 2022 [cited by applicant]
US 20230238085A1 · Jang et al. · 2023 [cited by applicant]
US 20240184982A1 · Mathewson et al. · 2024 [cited by applicant]
US 20250175193A1 · Najjar et al. · 2025 [cited by applicant]
Fu, Cong et al. “Fragment and Geometry Aware Tokenization of Molecules for Structure-Based Drug Design Using Language Models.” arXiv.org (2024): n. pag. Print. (Year: 2024). [cited by examiner]
Chennakesavalu, Shriram et al. “Energy Rank Alignment: Using Preference Optimization to Search Chemical Space at Scale.” arXiv.org (2024): n. pag. Print. (Year: 2024). [cited by examiner]
Hedelius, Bryce E, Damon Tingey, and Dennis Della Corte. “TrIPTransformer Interatomic Potential Predicts Realistic Energy Surface Using Physical Bias.” Journal of chemical theory and computation 20.1 (2024): 199-211. We… [cited by examiner]
Seunghoon Yi, Youngwoo Cho et al. Towards Physically Reliable Molecular Representation Learning, Proceedings of the Thirty-Ninth Conference on Uncertainty in Artificial Intelligence, PMLR 216:2433-2443, 2023. (Year: 202… [cited by examiner]
Frank, J. Thorben et al. “A Euclidean Transformer for Fast and Stable Machine Learned Force Fields.” Nature communications 15.1 (2024): n. pag. Web. (Aug. 6, 2024). (Year: 2024). [cited by examiner]
Antunes, L.M et al., “Crystal Structure Generation with Autoregressive Large Language Modeling,” arXiv:2307.04340v2, Jul. 18, 2023, pp. 1-19. [cited by applicant]
Billera, L. et al., “The Continuous Language of Protein Structure,” bioRxiv, May 11, 2024, pp. 1-13. [cited by applicant]
Bran, A.M. et al., “Transformers and Large Language Models for Chemistry and Drug Discovery,” arXiv:2310.06083v1, Oct. 9, 2023, pp. 1-22. [cited by applicant]
Choudhary, K., “AtomGPT: Atomistic Generative Pre-trained Transformer for Forward and Inverse Materials Design,” arXiv:2405.03680v1, May 6, 2024, pp. 1-20. [cited by applicant]
Ciccotti, G. et al., “Blue Moon Sampling, Vectorial Reaction Coordinates, and Unbiased Constrained Dynamics,” ChemPhysChem, 6, Sep. 2005, pp. 1809-1814. [cited by applicant]
Felardos, L et al., “Designing losses for data-free training of normalizing flows on boltzmann distributions,” arXiv:2301.05475v1, Jan. 13, 2023, pp. 1-31. [cited by applicant]
Flam-Shepherd, D. et al., “Atom-by-atom protein generation and beyond with language models,” arXiv:2308.09482v1, Aug. 16, 2023, pp. 1-18. [cited by applicant]
Flam-Shepherd, D. et al., “Language models can generate molecules, materials, and protein binding sites directly in three dimensions as XYZ, CIF, and PDB files,” arXiv:2305.05708v1, May 9, 2023, pp. 1-14. [cited by applicant]
Flam-Shepherd, D. et al., “Language models can learn complex molecular distributions,” Nature Communications, 13, Jun. 2022, pp. 1-10. [cited by applicant]
Fu, C. et al., “Fragment and geometry aware tokenization of molecules for structure-based drug design using language models,” arXiv:2408.09730, Aug. 19, 2024, pp. 1-22. [cited by applicant]
Gaujac, B. et al., “Learning the Language of Protein Structure,” arXiv:2405.15840v1, May 24, 2024, pp. 1-20. [cited by applicant]
Gebauer, N. et al., “Symmetry-adapted generation of 3d point sets for the targeted discovery of molecules,” 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Dec. 2019, pp. 1-13. [cited by applicant]
Gruver, N. et al., “Fine-Tuned Language Models Generate Stable Inorganic Materials as Text,” arXiv:2402.04379v1, Feb. 6, 2024, pp. 1-20. [cited by applicant]
Hayes, T. et al., “Simulating 500 million years of evolution with a language model,” bioRxiv preprint, Jul. 2024, pp. 1-68. [cited by applicant]
Hu, E.J. et al., “Amortizing Intractable Inference in Large Language Models,” arXiv:2310.04363v2, Mar. 13, 2024, pp. 1-31. [cited by applicant]
Köhler, J. et al., “Flow-matching: Efficient coarse-graining of molecular dynamics without forces,” Journal of Chemical Theory and Computation, Jan. 2023, pp. 942-952. [cited by applicant]
Krämer, A. et al., “Statistically Optimal Force Aggregation for Coarse-Graining Molecular Dynamics,” Journal of Physical Chemistry Letters, 14, Apr. 20, 2023, pp. 3970-3979. [cited by applicant]
Li, X. et al., “Geometry Informed Tokenization of Molecules for Language Model Generation,” arXiv:2408.10120, Aug. 19, 2024, pp. 1-34. [cited by applicant]
Lin, Z. et al., “Evolutionary-scale prediction of atomic-level protein structure with a language model,” bioRxiv preprint, Oct. 31, 2022, pp. 1-28. [cited by applicant]
Livine, M. et al., “nach0: multimodal natural and chemical languages foundation model,” Chemical Science, vol. 15, May 8, 2024, pp. 8380-8389. [cited by applicant]
Noé, F. et al., “Boltzmann generators: Sampling equilibrium states of many-body systems with deep learning,” Science, vol. 365, Sep. 2019, pp. 1-18. [cited by applicant]
Noid, W.G et al., “The multiscale coarse-graining method. I. A rigorous bridge between atomistic and coarse-grained models,” Journal of Chemical Physics, 128:244114, Jun. 2008, pp. 1-11. [cited by applicant]
Richter, L. et al., “VarGrad: A Low-Variance Gradient Estimator for Variational Inference,” Advances in Neural Information Processing Systems, Dec. 2020, pp. 1-12. [cited by applicant]
Touvron, H. et al., “Llama 2: Open foundation and fine-tuned chat models,” arXiv:2307.09288v2, Jul. 19, 2023, pp. 1-77. [cited by applicant]
Vaswani, A. et al., “Attention Is All You Need,” 31st Conference on Neural Information Processing Systems (NIPS 2017), Dec. 2017, pp. 1-11. [cited by applicant]
Wang, J. et al., “Machine Learning of Coarse-Grained Molecular Dynamics Force Fields,” ACS Central Science, 5, Apr. 15, 2019, pp. 755-767. [cited by applicant]
Zholus, A. et al., “BindGPT: A Scalable Framework for 3D Molecular Design via Language Modeling and Reinforcement Learning,” arXiv:2406.03686v1, Jun. 2024, pp. 1-23. [cited by applicant]
Patent Cooperation Treaty, International Search Report, PCT Application No. PCT/US2025/043199, Dec. 9, 2025, 123 pages. [cited by applicant]
Qiao et al. “OrbNet: Deep Learning for Quantum Chemistry Using Symmetry-Adapted Atomic-Orbital Features”, Jan. 18, 2022, 11 pages. [cited by applicant]
Wang et al “Token-Mal 1.0: Tokenized drug design with large language model”, Jul. 10, 2024, 59 pages. [cited by applicant]