IP Library Granted Patent US 12,541,488
Granted Patent B2
US 12,541,488 · App. 19/094,850 · Granted Feb 3, 2026

System and method for generating thoughts with large language models using codewords

Inventors: Brian Galvin (Silverdale, WA); Alan McCord (Forney, TX)
Assignee: ATOMBEAM TECHNOLOGIES INC.
G06F16/212G06F40/284G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,541,488
App. No.
19/094,850
Granted
Feb 3, 2026
Kind
B2
Abstract

This invention presents an optimized approach for training and operating Large Language Models (LLMs) using codewords. By converting traditional token-based LLMs to codeword-based systems, the method achieves significant efficiency gains. The process involves tokenizing training data and assigning codewords to tokens. LLMs are then trained and operated using these compact codewords instead of conventional tokens. During operation, prompts are converted to codewords, processed by the LLM, and the outputs are converted back to text. This approach reduces the overall cost of training and operating LLMs by approximately, offering a more efficient solution for large-scale language processing tasks.

Claims (43)

1 . A computer system comprising:

a hardware memory, wherein the computer system is configured to execute software instructions stored on nontransitory machine-readable storage media that:

tokenize a set of training data into a plurality of training tokens;

create a codeword dictionary by assigning unique codewords to each of the plurality of training tokens;

convert all training tokens into a plurality of training codewords using the codeword dictionary;

train a large language model using the plurality of training codewords;

receive a text prompt from a user;

tokenize the prompt into a plurality of prompt tokens;

convert the plurality of prompt tokens into a plurality of prompt codewords using the codeword dictionary;

process the plurality of prompt codewords through a large language model to generate a plurality of thought codewords representing intermediate reasoning steps;

associate each thought codeword with a corresponding portion of the prompt;

encode the thought codewords with metadata; and

store the thought codewords and their associated metadata.

2 . The system of claim 1 , wherein the text prompt is received, tokenized, and converted from tokens to codewords and from codewords back to tokens on an edge device.

3 . The system of claim 2 , wherein the codeword dictionary is a local codeword dictionary lookup on the edge device.

4 . The system of claim 1 , wherein the large language model uses a transformer architecture.

5 . The system of claim 1 , wherein the large language model uses a latent transformer architecture.

6 . The computer system of claim 1 , further configured to:

process the sequence of prompt codewords through the large language model to generate a codeword response; and

convert the codeword response into a text response.

7 . The computer system of claim 1 , further configured to:

generate a response based on both the prompt codewords and the generated thought codewords.

8 . A computer-implemented method comprising the steps of:

tokenizing a set of training data into a plurality of training tokens;

creating a codeword dictionary by assigning unique codewords to each of the plurality of training tokens;

converting all training tokens into a plurality of training codewords using the codeword dictionary;

training a large language model using the plurality of training codewords;

receiving a text prompt from a user;

tokenizing the prompt into a plurality of tokens;

converting the plurality of tokens into a plurality of prompt codewords using the codeword dictionary;

processing the plurality of prompt codewords through a large language model to generate a plurality of thought codewords representing intermediate reasoning steps;

associating each thought codeword with a corresponding portion of the prompt;

encoding the thought codewords with metadata; and

storing the thought codewords and their associated metadata.

9 . The method of claim 8 , wherein the text prompt is received, tokenized, and converted from tokens to codewords and from codewords back to tokens on an edge device.

10 . The method of claim 9 , wherein the codeword dictionary is a local codeword dictionary lookup on the edge device.

11 . The method of claim 8 , wherein the large language model uses a transformer based architecture.

12 . The method of claim 8 , wherein the large language model uses a variational autoencoder based architecture.

13 . The method of claim 8 , further comprising the steps of:

processing the sequence of prompt codewords through the large language model to generate a codeword response; and

convert the codeword response into a text response.

14 . The method of claim 8 , further comprising the steps of:

generating a response based on both the prompt codewords and the generated thought codewords.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 7, 2025
From: GALVIN, BRIAN; MCCORD, ALAN
To: ATOMBEAM TECHNOLOGIES INC.
Reel/Frame 071621/0470 →
Continuity (4)
Continuation In Part 18791465 · Aug 1, 2024
Continuation In Part 18736498 · Jun 6, 2024
Provisional Application 63651359 · May 23, 2024
Related Publication 20250363078A1 · Nov 27, 2025
References Cited (46)
US 11893057B2 · Donaldson et al. · 2024 [cited by applicant]
US 11893981B1 · Clark et al. · 2024 [cited by applicant]
US 11934792B1 · Adato et al. · 2024 [cited by applicant]
US 11954102B1 · Palaniappan et al. · 2024 [cited by applicant]
US 11977854B2 · Tunstall-Pedoe et al. · 2024 [cited by applicant]
US 11989507B2 · Tunstall-Pedoe et al. · 2024 [cited by applicant]
US 11989527B2 · Tunstall-Pedoe et al. · 2024 [cited by applicant]
US 12001462B1 · Madisetti · 2024 [cited by examiner]
US 12008621B1 · Atef · 2024 [cited by applicant]
US 12038918B1 · Zhang et al. · 2024 [cited by applicant]
US 12039264B2 · Prasad · 2024 [cited by examiner]
US 12067362B2 · Tunstall-Pedoe et al. · 2024 [cited by applicant]
US 12073180B2 · Tunstall-Pedoe et al. · 2024 [cited by applicant]
US 20230129094A1 · Lauritzen et al. · 2023 [cited by applicant]
US 20230135179A1 · Mielke et al. · 2023 [cited by applicant]
US 20230237053A1 · Dangoor et al. · 2023 [cited by applicant]
US 20230259714A1 · Lange · 2023 [cited by applicant]
US 20230262234A1 · Amini et al. · 2023 [cited by applicant]
US 20230350936A1 · Alayrac et al. · 2023 [cited by applicant]
US 20230376365A1 · Linquist et al. · 2023 [cited by applicant]
US 20230394188A1 · Zhang et al. · 2023 [cited by applicant]
US 20230421373A1 · Ryan et al. · 2023 [cited by applicant]
US 20240038226A1 · Nouri et al. · 2024 [cited by applicant]
US 20240070270A1 · Mace et al. · 2024 [cited by applicant]
US 20240070394A1 · Peng et al. · 2024 [cited by applicant]
US 20240073056A1 · Ishchenko et al. · 2024 [cited by applicant]
US 20240185137A1 · Atlan et al. · 2024 [cited by applicant]
US 20240220712A1 · Zass · 2024 [cited by applicant]
US 20240242030A1 · Lao et al. · 2024 [cited by applicant]
US 20240256582A1 · Jain et al. · 2024 [cited by applicant]
US 20240273282A1 · Muralidharan et al. · 2024 [cited by applicant]
US 20240273286A1 · Lu et al. · 2024 [cited by applicant]
US 20240273291A1 · Smith et al. · 2024 [cited by applicant]
US 20240273306A1 · Somaiya et al. · 2024 [cited by applicant]
US 20240289395A1 · Zhou · 2024 [cited by examiner]
US 20240289558A1 · Muraoka · 2024 [cited by examiner]
US 20240290331A1 · Lu et al. · 2024 [cited by applicant]
US 20240296425A1 · Rosenkranz et al. · 2024 [cited by applicant]
US 20240320268A1 · Mallipeddi · 2024 [cited by applicant]
US 20250021761A1 · Santhanam · 2025 [cited by examiner]
US 20250131261A1 · Harang · 2025 [cited by examiner]
US 20250147737A1 · Maturana · 2025 [cited by examiner]
US 20250165444A1 · Rahimov · 2025 [cited by examiner]
CN 117371453A · 2024 [cited by applicant]
CN 119312910A · 2025 [cited by examiner]
WO WO2025073037A1 · 2025 [cited by examiner]