IP Library › Granted Patent US 12,634,606
Granted Patent B2
US 12,634,606 · App. 18/561,985 · Granted May 19, 2026

In-network optical inference

Inventors: Manya Ghobadi (Cambridge, MA); Zhizhen Zhong (Cambridge, MA); Weiyang Wang (Cambridge, MA); Liane Sarah Beland Bernstein (Cambridge, MA); Alexander Sludds (Cambridge, MA); Ryan Hamerly (Cambridge, MA); Dirk Robert Englund (Brookline, MA)
Assignees: Massachusetts Institute of Technology; NTT Research, Incorporated
H04Q11/0005G06N5/04H04B10/524H04Q2011/0039H04Q2011/0041
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,634,606
App. No.
18/561,985
Granted
May 19, 2026
Kind
B2
Abstract

In-network Optical Inference (IOI) provides low-latency machine learning inference by leveraging programmable switches and optical matrix multiplication. IOI uses a transceiver module, called a Neuro Transceiver, with an optical processor to perform linear operations, such as matrix multiplication, in the optical domain. IOI's transceiver modules can be plugged into programmable packet switches, which are programmed to perform non-linear activations in the electronic domain and to respond to inference queries. Processing inference queries at the programmable packet switches inside the network, without sending them to cloud or edge inference servers, significantly reduces end-to-end inference latency experienced by users.

Claims (68)

1 . A method of inference processing, the method comprising:

transmitting a first data packet containing an input vector to a layer of neurons in an artificial neural network from a client to a programmable packet switch comprising a programmable packet processing chip and an optical computing unit;

removing source and/or destination metadata from the first data packet;

after removing the source and/or destination metadata from the first data packet, concatenating, by the programmable packet processing chip, a weight vector corresponding to the layer of neurons with the first data packet;

transmitting the first data packet and the weight vector from the programmable packet processing chip to the optical computing unit;

computing a product of the input vector and the weight vector with the optical computing unit;

transmitting a second data packet containing the product of the input vector and the weight vector from the optical computing unit to the programmable packet processing chip; and

applying a nonlinear activation function to the product with the programmable packet processing chip.

2 . The method of claim 1 , further comprising:

storing the destination metadata in a memory while computing the product of the input vector and the weight vector and applying the nonlinear activation function;

adding the destination metadata to an output packet containing an output of the artificial neural network; and

transmitting the output packet from the programmable packet switch to a destination specified in the destination metadata.

3 . The method of claim 1 , wherein applying the nonlinear activation function comprises retrieving information from a match-action table programmed into a memory of the programmable packet processing chip.

4 . The method of claim 1 , wherein the layer of neurons is a first layer of neurons of the artificial neural network, the weight vector is a first weight vector, and further comprising:

concatenating, by the programmable packet processing chip, a second weight vector corresponding to a second layer of neurons in the artificial neural network to the second data packet; and

transmitting the second data packet and the second weight vector from the programmable packet processing chip to the optical computing unit.

5 . The method of claim 1 , further comprising, before transmitting the second data packet:

digitizing the product of the input vector and the weight vector; and

storing the product in a memory.

6 . A method of inference processing, the method comprising:

transmitting a first data packet containing an input vector to a layer of neurons in an artificial neural network from a client to a programmable packet switch comprising a programmable packet processing chip and an optical computing unit;

concatenating, by the programmable packet processing chip, a weight vector corresponding to the layer of neurons with the first data packet;

transmitting the first data packet and the weight vector from the programmable packet processing chip to the optical computing unit;

computing a product of the input vector and the weight vector with the optical computing unit;

transmitting a second data packet containing the product of the input vector and the weight vector from the optical computing unit to the programmable packet processing chip; and

applying a nonlinear activation function to the product with the programmable packet processing chip,

wherein computing the product of the input vector and the weight vector comprises:

modulating, with a first modulator, an optical pulse with a waveform proportional to the input vector;

modulating, with a second modulator in optical communication with the first modulator, the optical pulse with a waveform proportional to the weight vector; and

detecting the optical pulse with a photodetector in optical communication with the second modulator.

7 . The method of claim 6 , further comprising:

delaying the optical pulse between the first modulator and the second modulator by a delay equal to a duration of the waveform proportional to the input vector.

8 . A programmable packet switch comprising:

a programmable packet processing chip to concatenate a first data packet representing an input vector to a first layer of neurons of an artificial neural network with a first weight vector corresponding to the first layer of neurons; and

an optical processor, operably coupled to the programmable packet processing chip, to receive the first data packet and the first weight vector from the programmable packet processing chip, to compute a product of the input vector and the first weight vector, and to transmit the product of the input vector and the first weight vector to the programmable packet processing chip as a second data packet,

wherein the programmable packet processing chip is configured to remove source and/or destination metadata from the first data packet before transmitting the first data packet and the first weight vector to the optical processor.

9 . The programmable packet switch of claim 8 , wherein the programmable packet switch further comprises:

a memory to store the destination metadata while the optical processor computes the product of the input vector and the first weight vector,

wherein the programmable packet processing chip is configured to add the destination metadata stored in the memory to an output packet containing an output of the artificial neural network.

10 . The programmable packet switch of claim 8 , wherein the programmable packet processing chip is configured to perform a nonlinear activation on the product of the input vector and the first weight vector, thereby producing an output of the first layer of neurons.

11 . The programmable packet switch of claim 10 , wherein the programmable packet processing chip comprises a memory to store a match-action table representing the nonlinear activation.

12 . The programmable packet switch of claim 10 , wherein the programmable packet processing chip is further configured to concatenate the output of the first layer of neurons with a second weight vector corresponding to a second layer of neurons of the artificial neural network and to transmit the output of the first layer of neurons and the second weight vector to the optical processor.

13 . A programmable packet switch of claim 8 comprising:

a programmable packet processing chip to concatenate a first data packet representing an input vector to a first layer of neurons of an artificial neural network with a first weight vector corresponding to the first layer of neurons; and

an optical processor, operably coupled to the programmable packet processing chip, to receive the first data packet and the first weight vector from the programmable packet processing chip, to compute a product of the input vector and the first weight vector, and to transmit the product of the input vector and the first weight vector to the programmable packet processing chip as a second data packet,

wherein the optical processor comprises:

a first modulator to modulate an optical beam with a first analog waveform representing the input vector;

a second modulator, in optical communication with the first modulator, to modulate the optical beam with a second analog waveform representing the first weight vector; and

a photodetector, in optical communication with the second modulator, to detect the optical beam.

14 . The programmable packet switch of claim 13 , further comprising:

at least one digital-to-analog converter, operably coupled to the first modulator and to the second modulator, to convert the input vector into the first analog waveform and to convert the first weight vector into the second analog waveform.

15 . The programmable packet switch of claim 13 , further comprising:

an analog-to-digital converter, operably coupled to the photodetector, to convert a photocurrent generated by the photodetector in response to detecting the optical beam into a digital signal representing the product of the input vector and the first weight vector.

16 . The programmable packet switch of claim 13 , wherein the optical processor further comprises:

an optical delay line, coupling the first modulator to the second modulator, to delay the optical beam by a delay equal to a duration of the first analog waveform.

17 . A method of inference processing, the method comprising:

receiving a packet with a header comprising source/destination metadata and a payload comprising an input to a deep neural network (DNN);

removing the source/destination metadata from the header;

adding a weight vector corresponding to a first layer of the DNN to the header;

transmitting the packet to an optical processor comprising a first modulator, a second modulator in series with the first modulator, and a photodetector;

converting the input to the DNN into a first analog waveform;

converting the weight vector into a second analog waveform;

modulating, by the first modulator, an amplitude of an optical beam with the first analog waveform;

modulating, by the second modulator, the amplitude of the optical beam with the second analog waveform;

transducing, by the photodetector, the optical beam into an electrical signal with an amplitude representing a product of the input to the DNN and the weight vector; and

performing a nonlinear activation on the electrical signal to produce an output of the first layer of the DNN.

18 . The method of claim 17 , further comprising:

transmitting a packet containing the output of the first layer of the DNN and a weight vector corresponding to a second layer of the DNN to the header to the optical processor.

Assignments (3)
CONFIRMATORY LICENSE Recorded May 14, 2024
From: MASSACHUSETTS INSTITUTE OF TECHNOLOGY
To: US DEPARTMENT OF ENERGY
Reel/Frame 067404/0115 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 29, 2023
From: HAMERLY, RYAN
To: MASSACHUSETTS INSTITUTE OF TECHNOLOGY; NTT RESEARCH, INCORPORATED
Reel/Frame 065700/0262 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 29, 2023
From: GHOBADI, MANYA; ZHONG, ZHIZHEN; WANG, WEIYANG; BERNSTEIN, LIANE SARAH BELAND; SLUDDS, ALEXANDER; ENGLUND, DIRK ROBERT
To: MASSACHUSETTS INSTITUTE OF TECHNOLOGY
Reel/Frame 065700/0302 →
Continuity (2)
Provisional Application 63191120 · May 20, 2021
Related Publication 20240244354A1 · Jul 18, 2024
References Cited (33)
US 11604978B2 · Hamerly et al. · 2023 [cited by applicant]
US 12020150B2 · Guo · 2024 [cited by examiner]
US 12033065B2 · Kenney · 2024 [cited by examiner]
US 20190102672A1 · Bifulco · 2019 [cited by examiner]
US 20190317802A1 · Bachmutsky et al. · 2019 [cited by applicant]
US 20200250534A1 · Shen et al. · 2020 [cited by applicant]
US 20210064958A1 · Lin · 2021 [cited by examiner]
US 20210357737A1 · Hamerly et al. · 2021 [cited by applicant]
US 20220180175A1 · Guo · 2022 [cited by applicant]
US 20230274156A1 · Hamerly · 2023 [cited by examiner]
Agrawal et al., “Intel Tofino2—A 12.9 Tbps P4-Programmable Ethernet Switch.” 2020 IEEE Hot Chips 32 Symposium (HCS). IEEE Computer Society, 2020, 32 pages. [cited by applicant]
Bansal et al., “Chauffeurnet: Learning to drive by imitating the best and synthesizing the worst.” arXiv preprint arXiv:1812.03079 (2018), 20 pages. [cited by applicant]
Cheng et al., “Wide & deep learning for recommender systems.” Proceedings of the 1st workshop on deep learning for recommender systems. 2016, 4 pages. [cited by applicant]
Gebara et al., “Challenging the Stateless Quo of Programmable Switches.” Proceedings of the 19th ACM Workshop on Hot Topics in Networks. 2020, 7 pages. [cited by applicant]
Hamerly et al., “Large-scale optical neural networks based on photoelectric multiplication.” Physical Review X 9.2 (2019): 021032, 12 pages. [cited by applicant]
Han et al., “A survey of label-noise representation learning: Past, present and future.” arXiv preprint arXiv:2011.04406 (2020), 24 pages. [cited by applicant]
Han et al., “Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding.” arXiv preprint arXiv:1510.00149 (2015), 14 pages. [cited by applicant]
Han et al., “Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding.” international conference on learning representations (2015), 14 pages. [cited by applicant]
Iandola et al., “SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and< 0.5 MB model size.” arXiv preprint arXiv:1602.07360 (2016), 13 pages. [cited by applicant]
Intel Newsroom, Intel Demonstrates Industry-First Co-Packaged Optics Ethernet Switch. Mar. 5, 2020. Accessed at https://newsroom.intel.com/news/intel-demonstrates-industry-first-co-packaged-optics-ethernet-switch/#gs.2p… [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2022/030254 mailed Nov. 16, 2022, 11 pages. [cited by applicant]
Junos Trio Programmable Silicon Optimized for the Universal Edge. Juniper Networks Oct. 2009. 12 pages. [cited by applicant]
Karimi et al., “Deep learning with noisy labels: Exploring techniques and remedies in medical image analysis.” Medical Image Analysis 65 (2020): 101759, 49 pages. [cited by applicant]
Kim, “Powered by ai: Oculus insight” Facebook AI, 2019. Accessed at https://ai.facebook.com/blog/powered-byai-oculus-insight/, 2019, 5 pages. [cited by applicant]
Kim, The scalable neural architecture behind alexa's ability to select skills. Amazon Science Jun. 7, 2018. Accessed at https://www.amazon.science/blog/the-scalable-neural-architecture-behind-alexas-ability-to-select-sk… [cited by applicant]
Naumov et al., “Deep learning recommendation model for personalization and recommendation systems.” arXiv preprint arXiv:1906.00091 (2019), 10 pages. [cited by applicant]
Nvidia, Cloud, and Data Center. “Nvidia V100 Tensor Core GPU.” Accessed at https://www.nvidia.com/en-us/data-center/a100/ on May 30, 2021. 9 pages. [cited by applicant]
Pandit et al., Modeling the impact of cpu properties to optimize and predict packet-processing performance. Intel Case Study 2018. Accessed at https://www.intel.com/content/dam/www/public/us/en/documents/case-studies/at… [cited by applicant]
Siri Team. Hey siri: An on-device dnn-powered voice trigger for apple's personal assistant. Apple Machine Learning, Oct. 2017. Accessed at https://machinelearning.apple.com/research/hey-siri. 11 pages. [cited by applicant]
Wang et al., “Monolithic lithium niobate photonic circuits for Kerr frequency comb generation and modulation.” Nature communications 10.1 (2019): 1-6. [cited by applicant]
Wetzstein et al., “Inference in artificial intelligence with deep optics and photonics.” Nature 588.7836 (2020): 39-47. [cited by applicant]
Wu et al., “Machine learning at facebook: Understanding inference at the edge.” 2019 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 2019, pp. 331-344. [cited by applicant]
Xiong et al., “Do switches dream of machine learning? Toward in-network classification.” Proceedings of the 18th ACM workshop on hot topics in networks. 2019, pp. 25-33. [cited by applicant]