IP Library Granted Patent US 12,261,631
Granted Patent B2
US 12,261,631 · App. 18/909,976 · Granted Mar 25, 2025

Deep learning using large codeword model with homomorphically compressed data

Inventor: Brian Galvin (Silverdale, WA)
Assignee: ATOMBEAM TECHNOLOGIES INC
H03M7/3059G06N20/00H03M7/6005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,261,631
App. No.
18/909,976
Filed
Oct 9, 2024
Granted
Mar 25, 2025
Kind
B2
Art Unit
2845
USPC
707/693
Abstract

A system and method for deep learning using a large codeword model with homomorphically compressed and dyadically encrypted data is disclosed. The system preprocesses input data, applies homomorphic-dyadic compression and encryption, tokenizes the compressed data into sourceblocks, and assigns codewords using a codebook. These codewords are processed through a machine learning core, which can be either a conventional transformer-based architecture or a latent transformer core utilizing a variational autoencoder. The system enables secure operations on encrypted data, preserving privacy while allowing complex computations. The processed output is decrypted, decompressed, and translated to match the input modality. A neural upsampler may further enhance the output. The machine learning core is continuously trained using the processed data and additional training data, improving performance over time.

Claims (72)

1. A system for deep learning using a large codeword model with homomorphically compressed dyadically encrypted data, comprising:

a computing device comprising at least a memory and a processor;

a plurality of programming instructions stored in the memory and operable on the processor, wherein the plurality of programming instructions, when operating on the processor, cause the computing device to:

receive a plurality of inputs;

preprocess the inputs to generate a plurality of input data sets;

compress and encrypt the input data sets by:

analyzing the input data sets to determine their properties;

creating transformation matrices based on the properties of the input data;

transforming the input data into modified distributions;

generating main data streams of transformed data and secondary data streams of transformation information; and

compressing the main data streams;

tokenize the compressed main data streams into a plurality of sourceblocks;

assign the plurality of sourceblocks a plurality of codewords, where each sourceblock is mapped to a particular codeword through a codebook;

process the plurality of codewords through a machine learning core to generate a codeword response;

translate the codeword response into a translated response which matches the modality of the inputs;

decompress and decrypt the translated response; and

train the machine learning core using the decompressed and decrypted response and a plurality of training data.

2. The system of claim 1 , wherein the machine learning core is a conventional transformer-based architecture comprising:

an embedding layer;

a positional encoding layer; and

a series of transformer layers.

3. The system of claim 2 , further comprising a syntactic splitting component that splits the codewords into smaller units before processing through the conventional transformer-based architecture.

4. The system of claim 1 , wherein the machine learning core is a latent transformer core comprising:

a variational autoencoder with an encoder and a decoder; and

a transformer that processes latent space vectors, wherein the transformer does not include an embedding layer and a positional encoding layer.

5. The system of claim 4 , wherein processing the plurality of codewords through the machine learning core comprises:

generating a plurality of latent space vectors by processing the plurality of codewords through the variational autoencoder's encoder;

learning relationships between the plurality of latent space vectors by processing them through the transformer;

using the learned relationships to generate a plurality of output latent space vectors; and

generating the codeword response by passing the output latent space vectors through the variational autoencoder's decoder.

6. The system of claim 5 , further comprising a syntactic splitting component that splits the latent space vectors into smaller units before processing through the transformer.

7. The system of claim 1 , wherein compressing and encrypting the input data sets further comprises:

combining the compressed main data streams and the secondary data streams into output streams; and

implementing security measures to protect the output streams.

8. The system of claim 7 , wherein the security measures comprise providing cryptographically secure random numbers for use in data transformation and implementing protections against side-channel attacks.

9. The system of claim 1 , wherein transforming the input data into modified distributions comprises transforming the input data into dyadic distributions.

10. The system of claim 1 , further comprising a neural upsampler that processes the codeword response to generate a reconstructed output containing more information than the translated response.

11. A method for deep learning using a large codeword model with homomorphically compressed dyadically encrypted data, comprising the steps of:

receiving a plurality of inputs;

preprocessing the inputs to generate a plurality of input data sets;

compressing and encrypting the input data sets by:

analyzing the input data sets to determine their properties;

creating transformation matrices based on the properties of the input data;

transforming the input data into modified distributions;

generating main data streams of transformed data and secondary data streams of transformation information; and

compressing the main data streams;

tokenizing the compressed main data streams into a plurality of sourceblocks;

assigning the plurality of sourceblocks a plurality of codewords, where each sourceblock is mapped to a particular codeword through a codebook;

processing the plurality of codewords through a machine learning core to generate a codeword response;

translating the codeword response into a translated response which matches the modality of the inputs;

decompressing and decrypting the translated response; and

training the machine learning core using the decompressed and decrypted response and a plurality of training data.

12. The method of claim 11 , wherein the machine learning core is a conventional transformer-based architecture comprising:

an embedding layer;

a positional encoding layer; and

a series of transformer layers.

13. The method of claim 12 , further comprising splitting the codewords into smaller units before processing through the conventional transformer-based architecture.

14. The method of claim 11 , wherein the machine learning core is a latent transformer core comprising:

a variational autoencoder with an encoder and a decoder; and

a transformer that processes latent space vectors, wherein the transformer does not include an embedding layer and a positional encoding layer.

15. The method of claim 14 , wherein processing the plurality of codewords through the machine learning core comprises:

generating a plurality of latent space vectors by processing the plurality of codewords through the variational autoencoder's encoder;

learning relationships between the plurality of latent space vectors by processing them through the transformer;

using the learned relationships to generate a plurality of output latent space vectors; and

generating the codeword response by passing the output latent space vectors through the variational autoencoder's decoder.

16. The method of claim 15 , further comprising splitting the latent space vectors into smaller units before processing through the transformer.

17. The method of claim 11 , wherein compressing and encrypting the input data sets further comprises:

combining the compressed main data streams and the secondary data streams into output streams; and

implementing security measures to protect the output streams.

18. The method of claim 17 , wherein the security measures comprise providing cryptographically secure random numbers for use in data transformation and implementing protections against side-channel attacks.

19. The method of claim 11 , wherein transforming the input data into modified distributions comprises transforming the input data into dyadic distributions.

20. The method of claim 11 , further comprising processing the codeword response through a neural upsampler to generate a reconstructed output containing more information than the translated response.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2024
From: GALVIN, BRIAN
To: ATOMBEAM TECHNOLOGIES INC.
Reel/Frame 069266/0193 →
Continuity (36)
Continuation In Part 18770652 · Jul 12, 2024
Continuation In Part 18755653 · Jun 26, 2024
Continuation In Part 18737906 · Jun 7, 2024
Continuation In Part 18736498 · Jun 6, 2024
Continuation In Part 18657683 · May 7, 2024
Continuation In Part 18648340 · Apr 27, 2024
Continuation In Part 18427716 · Jan 30, 2024
Continuation In Part 18410980 · Jan 11, 2024
Continuation In Part 18537728 · Dec 12, 2023
Continuation In Part 18503135 · Nov 6, 2023
Continuation 18305305 · Apr 21, 2023
Continuation In Part 18190044 · Mar 24, 2023
Continuation In Part 17875201 · Jul 27, 2022
Continuation In Part 17727913 · Apr 25, 2022
Continuation 17514913 · Oct 29, 2021
Continuation 17458747 · Aug 27, 2021
Continuation 17404699 · Aug 17, 2021
Continuation In Part 17404699 · Aug 17, 2021
Continuation In Part 17234007 · Apr 19, 2021
Continuation In Part 17180439 · Feb 19, 2021
Continuation In Part 16923039 · Jul 7, 2020
Continuation In Part 16923039 · Jul 7, 2020
Continuation In Part 16716098 · Dec 16, 2019
Continuation 16455655 · Jun 27, 2019
Continuation In Part 16455655 · Jun 27, 2019
Continuation In Part 16200466 · Nov 26, 2018
Continuation In Part 15975741 · May 9, 2018
Provisional Application 63651359 · May 23, 2024
Provisional Application 63485518 · Feb 16, 2023
Provisional Application 63388411 · Jul 12, 2022
Provisional Application 63232041 · Aug 11, 2021
Provisional Application 63140111 · Jan 21, 2021
Provisional Application 63027166 · May 19, 2020
Provisional Application 62926723 · Oct 28, 2019
Provisional Application 62578824 · Oct 30, 2017
Related Publication 20250047295A1 · Feb 6, 2025
References Cited (39)
US 4780718A · Hudson et al. · 1988 [cited by applicant]
US 5708436A · Loiz et al. · 1998 [cited by applicant]
US 7411540B1 · Lopez et al. · 2008 [cited by applicant]
US 7629922B2 · Winstead et al. · 2009 [cited by applicant]
US 7876257B2 · Vetro et al. · 2011 [cited by applicant]
US 9524392B2 · Naehrig et al. · 2016 [cited by applicant]
US 11086843B2 · Swaminathan · 2021 [cited by examiner]
US 11100420B2 · Dirac · 2021 [cited by examiner]
US 11216742B2 · Bhattacharyya · 2022 [cited by examiner]
US 11451242B2 · Choi et al. · 2022 [cited by applicant]
US 11557323B1 · Stewart · 2023 [cited by examiner]
US 11656353B2 · Li et al. · 2023 [cited by applicant]
US 20040017307A1 · Cirillo et al. · 2004 [cited by applicant]
US 20040160353A1 · Cirillo et al. · 2004 [cited by applicant]
US 20080231504A1 · Sartor et al. · 2008 [cited by applicant]
US 20100117875A1 · Oslick · 2010 [cited by examiner]
US 20110012778A1 · Nguyen et al. · 2011 [cited by applicant]
US 20150054678A1 · Wakayama · 2015 [cited by applicant]
US 20170048537A1 · Boufounos et al. · 2017 [cited by applicant]
US 20180196609A1 · Niesen · 2018 [cited by applicant]
US 20200258296A1 · Pennings et al. · 2020 [cited by applicant]
US 20200395955A1 · Choi et al. · 2020 [cited by applicant]
US 20220156631A1 · Kanso et al. · 2022 [cited by applicant]
US 20220404490A1 · Evans et al. · 2022 [cited by applicant]
US 20230131694A1 · Saber et al. · 2023 [cited by applicant]
US 20230169623A1 · Chen et al. · 2023 [cited by applicant]
US 20230184927A1 · Chen et al. · 2023 [cited by applicant]
US 20240185037A1 · Park et al. · 2024 [cited by applicant]
US 20240195438A1 · Isik et al. · 2024 [cited by applicant]
EP 3364212A1 · 2018 [cited by applicant]
GB 2620921A · 2024 [cited by applicant]
WO 2020104416A1 · 2020 [cited by applicant]
Balaneshin-Kordan, Saeid et al., Deep ueral Architecture for Multi-Modal Retrieval based on Joint Embedding Space for Text and Images, Association for Computing Machinery, Feb. 5-9, 2018, pp. 1-9, Marina Del Rey, CA, US… [cited by applicant]
Kahn, Abdul Rafae et al., Coding Textual Inputs Boosts the Accuracy of Neural Networks, 2020 Conference on Empirical Methods in Natural Language Processing, Nov. 16-20, 2020, pp. 1350-1360. [cited by applicant]
Messina, Nicola et al., Towards Efficient Cross-Modal Visual Textual Retrieval using Transformer-Encoder Deep Features, 2021 International Conference on Content-Based Multimedia Indexing, 2021, pp. 1-6, United States. [cited by applicant]
Seo, Beomseok et al., How Does A Transformer Learn Compression? An Attention Study on Huffman and LZ4, Department of Electronic and Electrical Engineering, Dec. 12, 2023, pp. 1-10, vol. 11, Seoul, South Korea. [cited by applicant]
Vaswani, Ashish et al., Attention is All You Need, 31st Conference on Neural Information Processing Systems, 2017, pp. 1-11, Long Beach, CA, USA. [cited by applicant]
Wang, Tianming et al., T-CVAE: Transformer-Based Conditioned Variational Autoencoder for Story Completion, Proceedings of the Twenty-Eigth Joint Conference on Artificial Intelligence, pp. 5233-5239. [cited by applicant]
Wieting, John et al., A Bilingual Generative Transformer for Semantic Sentence Embedding, Nov. 19, 2020, pp. 1-14. [cited by applicant]
Cited By (1)
US 12,505,580