IP Library › Granted Patent US 12,664,440
Granted Patent B2
US 12,664,440 · App. 18/914,257 · Granted Jun 23, 2026

System and method for a large codeword model for deep learning

Inventor: Brian Galvin (Silverdale, WA)
Assignee: ATOMBEAM TECHNOLOGIES INC.
G06N3/096G06F18/23G06N3/044G06N3/0455
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,664,440
App. No.
18/914,257
Filed
Oct 13, 2024
Granted
Jun 23, 2026
Kind
B2
Art Unit
2148
USPC
706/25
Abstract

Modality agnostic Large Codeword Model (“LCM”) is an advanced deep learning architecture that processes discrete, compressed data representations called codewords across multiple modalities. Unlike traditional models using raw tokens and dense embeddings, LCMs efficiently handle diverse input types including text, images, audio, and video. The system employs a modality agnostic encoder, unified codebook, and multimodal machine learning core to capture inherent data structures and patterns. This approach enables more generalizable and interpretable feature learning, facilitating transfer learning across domains. The LCM's scalable and flexible architecture includes components for modality-specific processing, cross-modal attention, and joint representation learning. With its computational efficiency and versatility, the Modality Agnostic LCM offers significant potential for various AI applications, including natural language processing, computer vision, and multimodal reasoning.

Claims (43)

1 . A system for a generic compound large codeword model for natively multimodal deep learning, comprising one or more computers with executable instructions that, when executed, cause the system to:

receive a plurality of inputs of different modalities;

convert the plurality of inputs to a unified representation using a modality-agnostic encoder prior to tokenization, wherein the unified representation is a fixed-size tensor in which different sections encode information from different modalities in a common semantic space;

tokenize the unified representation into a plurality of sourceblocks;

assign the plurality of sourceblocks a plurality of codewords, where each sourceblock is mapped to a particular codeword through a unified codebook that is prefix-free and entropy-coded using Huffman or arithmetic coding, wherein the unified codebook is stored in a codebook library;

cluster the plurality of codewords based on semantic similarity and learn a single embedding vector for each codeword cluster rather than learning embeddings for individual codewords, thereby reducing the number of embedding vectors processed by the multimodal machine learning core;

process the plurality of codewords through a multimodal machine learning core comprising an embedding layer and a cross-modal attention mechanism configured to operate on the cluster-level embedding vectors in the shared semantic space;

generate a codeword response to the plurality of inputs using the multimodal machine learning core;

translate the codeword response into a translated response which matches one or more modalities of the inputs; and

train the multimodal machine learning core simultaneously on the multiple modalities using a joint loss function that applies uncertainty weighting to automatically balance training contributions from the different modalities within the multimodal machine learning core.

2 . The system of claim 1 , wherein the multimodal machine learning core has a transformer-based machine learning architecture.

3 . The system of claim 1 , wherein the multimodal machine learning core has a variational autoencoder-based machine learning architecture.

4 . The system of claim 1 , wherein the multimodal machine learning core has a recurrent neural network-based machine learning architecture.

5 . The system of claim 1 , further comprising a plurality of unified codebooks and a plurality of multimodal machine learning cores, wherein each unified codebook and multimodal machine learning core is configured to process inputs of different modalities.

6 . The system of claim 5 , further comprising a codeword translator which translates codewords between any plurality of modalities.

7 . The system of claim 1 , wherein the embedding layer is a unified embedding layer that maps codewords from all modalities using a single set of learned embedding parameters.

8 . The system of claim 1 , wherein the plurality of inputs comprises at least two of: text data, image data, audio data, and video data.

9 . The system of claim 1 , wherein the unified codebook is capable of mapping sourceblocks from multiple modalities to codewords within a shared codeword space.

10 . The system of claim 1 , wherein the modality-agnostic encoder comprises modality-specific processing channels that feed into a fusion layer, the fusion layer combining outputs from the modality specific processing channels into the unified representation.

11 . The system of claim 1 , wherein the system is configured to perform transfer learning between different modalities using the unified codeword representation.

12 . The system of claim 1 , wherein the system is configured to generate outputs in multiple modalities based on a single codeword response.

13 . A method for a generic compound large codeword model for natively multimodal deep learning, comprising the steps of:

receiving a plurality of inputs of different modalities;

converting the plurality of inputs to a unified representation using a modality-agnostic encoder prior to tokenization, wherein the unified representation is a fixed-size tensor in which different sections encode information from different modalities in a common semantic space;

tokenizing the unified representation into a plurality of sourceblocks;

assigning the plurality of sourceblocks a plurality of codewords, where each sourceblock is mapped to a particular codeword through a unified codebook that is prefix-free and entropy-coded using Huffman or arithmetic coding, wherein the unified codebook is stored in a codebook library;

clustering the plurality of codewords based on semantic similarity and learn a single embedding vector for each codeword cluster rather than learning embeddings for individual codewords, thereby reducing the number of embedding vectors processed by the multimodal machine learning core;

processing the plurality of codewords through a multimodal machine learning core comprising an embedding layer and a cross-modal attention mechanism;

generating a codeword response to the plurality of inputs using the multimodal machine learning core;

translating the codeword response into a translated response which matches one or more modalities of the inputs; and

training the multimodal machine learning core simultaneously on the multiple modalities using a joint loss function that applies uncertainty weighting to automatically balance training contributions from the different modalities within the multimodal machine learning core.

14 . The method of claim 13 , wherein the multimodal machine learning core has a transformer-based machine learning architecture.

15 . The method of claim 13 , wherein the multimodal machine learning core has a variational autoencoder-based machine learning architecture.

16 . The method of claim 13 , wherein the multimodal machine learning core has a recurrent neural network-based machine learning architecture.

17 . The method of claim 13 , further comprising using a plurality of unified codebooks and a plurality of multimodal machine learning cores, wherein each unified codebook and multimodal machine learning core is configured to process inputs of different modalities.

18 . The method of claim 17 , further comprising translating codewords between any plurality of modalities using a codeword translator.

19 . The method of claim 13 , wherein the embedding layer is a unified embedding layer that maps codewords from all modalities using a single set of learned embedding parameters.

20 . The method of claim 13 , wherein the plurality of inputs comprises at least two of:

text data, image data, audio data, and video data.

21 . The method of claim 13 , wherein the multimodal machine learning core comprises a cross-modal attention mechanism capable of attending to information across different modalities.

22 . The method of claim 13 , further comprising using modality-specific processing channels that feed into a fusion layer, the fusion layer combining outputs from the modality specific processing channels into the unified representation.

23 . The method of claim 13 , further comprising performing transfer learning between different modalities using the unified codeword representation.

24 . The method of claim 13 , further comprising generating outputs in multiple modalities based on a single codeword response.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2024
From: GALVIN, BRIAN
To: ATOMBEAM TECHNOLOGIES INC.
Reel/Frame 069266/0204 →
Continuity (3)
Continuation In Part 18736498 · Jun 6, 2024
Provisional Application 63651359 · May 23, 2024
Related Publication 20250363385A1 · Nov 27, 2025
References Cited (51)
US 4780718A · Hudson et al. · 1988 [cited by applicant]
US 5708436A · Loiz et al. · 1998 [cited by applicant]
US 7411540B1 · Lopez et al. · 2008 [cited by applicant]
US 7629922B2 · Winstead et al. · 2009 [cited by applicant]
US 7876257B2 · Vetro et al. · 2011 [cited by applicant]
US 9524392B2 · Naehrig et al. · 2016 [cited by applicant]
US 10789427B2 · Shazeer · 2020 [cited by examiner]
US 11451242B2 · Choi et al. · 2022 [cited by applicant]
US 11656353B2 · Li et al. · 2023 [cited by applicant]
US 20040017307A1 · Cirillo et al. · 2004 [cited by applicant]
US 20040160353A1 · Cirillo et al. · 2004 [cited by applicant]
US 20080231504A1 · Sartor et al. · 2008 [cited by applicant]
US 20110012778A1 · Nguyen et al. · 2011 [cited by applicant]
US 20150054678A1 · Wakayama · 2015 [cited by applicant]
US 20170048537A1 · Boufounos et al. · 2017 [cited by applicant]
US 20180196609A1 · Niesen · 2018 [cited by applicant]
US 20190140658A1 · Cooper · 2019 [cited by examiner]
US 20200258296A1 · Pennings et al. · 2020 [cited by applicant]
US 20200395955A1 · Choi et al. · 2020 [cited by applicant]
US 20210390269A1 · Rezagholizadeh · 2021 [cited by examiner]
US 20220156631A1 · Kanso et al. · 2022 [cited by applicant]
US 20220404490A1 · Evans et al. · 2022 [cited by applicant]
US 20230131694A1 · Saber et al. · 2023 [cited by applicant]
US 20230169623A1 · Chen et al. · 2023 [cited by applicant]
US 20230184927A1 · Chen et al. · 2023 [cited by applicant]
US 20240152695A1 · Shukla · 2024 [cited by examiner]
US 20240185037A1 · Park et al. · 2024 [cited by applicant]
US 20240195438A1 · Isik et al. · 2024 [cited by applicant]
US 20250103886A1 · Norouzzadeh · 2025 [cited by examiner]
EP 3364212A1 · 2018 [cited by applicant]
GB 2620921A · 2024 [cited by applicant]
WO 2020104416A1 · 2020 [cited by applicant]
Bordes, Patrick. “Deep Multimodal Learning for Joint Textual and Visual Reasoning” (Year: 2020). [cited by examiner]
Vaswani, Ashish, et al. “Attention is all you need.” (Year: 2017). [cited by examiner]
Caglayan, et al., “Multimodal attention for neural machine translation.” (Year: 2016). [cited by examiner]
Sulubacak, et al., “Multimodal machine translation through visuals and speech.” (Year: 2020). [cited by examiner]
Wang, “T-CVAE: Transformer-based conditioned variational autoencoder for story completion.” (Year: 2019). [cited by examiner]
Duan, et al., “Multi-modal alignment using representation codebook.” (Year: 2022). [cited by examiner]
Yu, et al., “Multi-scale multi-modal dictionary BERT for effective text-image retrieval in multimedia advertising.” (Year: 2022). [cited by examiner]
Dabre et al., “A survey of multilingual neural machine translation.” (Year: 2020). [cited by examiner]
Sennrich, et al., “Linguistic input features improve neural machine translation.” (Year: 2016). [cited by examiner]
Du, et al., “Pinyin as subword unit for chinese-sourced neural machine translation.” (Year: 2017). [cited by examiner]
Bordes et al., “Deep Multimodal Learning for Joint Textual and Visual Reasoning”. (Year: 2020). [cited by examiner]
Khan et al., “Coding textual inputs boosts the accuracy of neural networks” (Year: 2020). [cited by examiner]
Balaneshin-Kordan, “Deep Neural Architecture for Multi-Modal Retrieval based on Joint Embedding Space for Text and Images”, WSDM '18, p. 28-36. [cited by applicant]
Khan, “Coding Textual Inputs Boosts the Accuracy of Neural Networks”, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, 2020, p. 1350-1360. [cited by applicant]
Messina, “Towards Efficient Cross-Modal Visual Textual Retrieval using Transformer-Encoder Deep Features”, IEEE Xplore, 2024. [cited by applicant]
Seo, “How Does a Transformer Learn Compression? An Attention Study on Huffman and LZ4”, IEEE Access, 2023. [cited by applicant]
Vaswani, “Attention Is All You Need”, 31st Conference on Neural Information Processing Systems, 2017. [cited by applicant]
Wang, “T-CVAE: Transformer-Based Conditioned Variational Autoencoder for Story Completion”, JJCAI-19, p. 5233-5239. [cited by applicant]
Wieting et al, “A Bilingual Generative Transformer for Semantic Sentence Embedding”, 2020. [cited by applicant]