IP Library Granted Patent US 12,328,649
Granted Patent B2
US 12,328,649 · App. 17/286,330 · Granted Jun 10, 2025

Optimization of edge computing distributed neural processor for wearable devices

Inventor: Jie Gu (Evanston, IL)
Assignee: Northwestern University
H04W4/38A61B5/486A61F2/72G06F18/214G06F18/251G06N3/047G06N3/063G06N3/08G11C11/54
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,328,649
App. No.
17/286,330
Granted
Jun 10, 2025
Kind
B2
Abstract

Systems and/or methods may include an edge-computing distributed neural processor to effectively reduce the data traffic and physical wiring congestion. A local and global networking architecture may reduce traffic among multi-chips in edge computing. A mixed-signal feature extraction approach with assistance of neural network distortion recovery is also described to reduce the silicon area. High precision in signal features classification with a low bit processing circuitry may be achieved by compensating with a recursive stochastic rounding routine, and provide on-chip learning to re-classify the sensor signals.

Claims (42)

1. A neural processor for edge computing in a distributed neural network, comprising an integrated chip that comprises:

mixed-signal processing circuitry at an input that extracts features from multi-channels of incoming analog signals received from a sensor, wherein the features that are extracted comprise statistical values in a digital format from the incoming analog signals;

on-chip memory banks that store weighted ranks to correspond to the features that are extracted for the sensor;

a local neural network layer formed by processing circuitry comprising a plurality of neuron nodes that process the features that are extracted based on their weighted ranks; and

a global network layer formed by the processing circuitry comprising global neuron nodes at an output that process and classify the features that are extracted of the sensor and communicate with at least one other neural processor within the distributed neural network,

wherein the mixed-signal processing circuitry comprises an on-chip multi-channels voltage controlled oscillator-based front end,

wherein each channel of the voltage controlled oscillator-based front end further includes at least a voltage controlled oscillator clocked by a same on-chip clock generator, a plurality of comparators and counters and a single-differential converter.

2. The neural processor according to claim 1 , wherein the mixed-signal processing circuitry is devoid of an analog to digital converter in the neural processor, wherein the mixed-signal processing circuitry extracts the features from the incoming analog signals as time-domain features of: a mean, a variance, a slope absolute value, a histograms and a zero crossings.

3. The neural processor according to claim 2 , wherein the mean is generated by the voltage-controlled oscillator and an output of a counter which calculates averages counts of the voltage-controlled oscillator within an overlapped time window.

4. The neural processor according to claim 2 , wherein the variance is generated by the voltage-controlled oscillator and another reference voltage-controlled oscillator in conjunction with a bidirectional counter that accumulate a distance from the mean over a time window.

5. The neural processor according to claim 2 , wherein the slope absolute value is generated by a bidirectional counter which compares a difference in voltage between two-timing windows.

6. The neural processor according to claim 1 , wherein the local neural network layer comprises an input layer having It neuron nodes and a hidden layer having at least a first local layer of Ni neuron nodes, wherein the input layer having It neuron nodes receives the features that are extracted from the mixed-signal processing circuitry, and each of the Ni neuron nodes is configurable to receive processed signals from one or more of It neuron nodes, wherein Ni<It.

7. The neural processor according to claim 6 , wherein the hidden layer comprises a second local layer of Nk neuron nodes, wherein each of the Nk neuron nodes is configurable to receive processed signals from one or more of the Ni neuron nodes, wherein Nk<Ni.

8. The neural processor according to claim 7 , wherein the global network layer at the output is connected to a global clock line and to a global data line, wherein the global clock line sends or receives a global clock signal used for inter-chip communication, and the global data line is configured to communicate by sending computed sensor data from the neuron nodes of the one or both of the first and the second local layers of the neural processor to another neural processor, or receive computed sensor data from another neural processor.

9. The neural processor according to claim 1 , wherein the weighted ranks correspond to a number of bits for each neuron.

10. The neural processor according to claim 9 , wherein the weighted rank for the sensor is updated in machine learning for reclassifying the features that are extracted of the sensor.

11. The neural processor according to claim 9 , wherein the neural processor is programmed to execute a stochastic rounding process or stochastic batching processing techniques to improve a precision due to a reduction in a plurality of bit numbers to total neuron nodes in a hidden layer.

12. The neural processor according to claim 9 , wherein an eight-bit on-chip learning is enabled by a stochastic rounding process implemented through an on-chip random number generator using linear feedback shift register.

13. The neural processor according to claim 9 , wherein an eight-bit on-chip learning is enabled by pre-loading globally trained weights, where accuracy is improved through sequentially sending batch training data into the neuron nodes in a hidden layer, and a random number generator based on linear feedback shift register is used to randomize training sequence for each batch during the on-chip learning.

14. The neural processor according to claim 13 , wherein the on-chip memory banks are overwritten by the pre-loaded globally trained weights during on-chip learning.

15. The neural processor according to claim 9 , wherein the on-chip memory banks are crossbar connected with the plurality of neuron nodes in a hidden layer and global processing layer to form recurrent neural network to allow bidirectional signal propagations to support learning operations.

16. A neural processor for edge computing in a distributed neural network, comprising an integrated chip that comprises:

mixed-signal processing circuitry at an input that extracts features from multi-channels of incoming analog signals received from a sensor, wherein the features that are extracted comprise statistical values in a digital format from the incoming analog signals;

on-chip memory banks that store weighted ranks to correspond to the features that are extracted for the sensor;

a local neural network layer formed by processing circuitry comprising a plurality of neuron nodes that process the features that are extracted based on their weighted ranks; and

a global network layer formed by the processing circuitry comprising global neuron nodes at an output that process and classify the features that are extracted of the sensor and communicate with at least one other neural processor within the distributed neural network,

wherein a distributed neural network architecture is formed by a plurality of neural processors that each extracts and processes features from a respective sensor local to a respective neural processor, wherein one of the plurality of the neural processors comprising a master neural processor which is responsible for starting communication and providing a global clock signal to synchronize remaining neural processors within the distributed neural network.

17. The neural processor according to claim 16 , wherein one of the neural processors in the distributed neural network architecture sequentially sends its hidden layer neuron data output to a global data line and all remaining neural processors in the distributed neural network architecture read the data output from the one neural processor from the global data line.

18. A neural processor for edge computing in a distributed neural network, comprising an integrated chip that comprises:

mixed-signal processing circuitry at an input that extracts features from multi-channels of incoming analog signals received from a sensor, wherein the features that are extracted comprise statistical values in a digital format from the incoming analog signals;

on-chip memory banks that store weighted ranks to correspond to the features that are extracted for the sensor;

a local neural network layer formed by processing circuitry comprising a plurality of neuron nodes that process the features that are extracted based on their weighted ranks; and

a global network layer formed by the processing circuitry comprising global neuron nodes at an output that process and classify the features that are extracted of the sensor and communicate with at least one other neural processor within the distributed neural network,

wherein the local neural network layer comprises an input layer having It neuron nodes and a hidden layer having at least a first local layer of Ni neuron nodes, wherein the input layer having It neuron nodes receives the features that are extracted from the mixed-signal processing circuitry, and each of the Ni neuron nodes is configurable to receive processed signals from one or more of It neuron nodes, wherein Ni<It,

wherein the hidden layer comprises a second local layer of Nk neuron nodes, wherein each of the Nk neuron nodes is configurable to receive processed signals from one or more of the Ni neuron nodes, wherein Nk<Ni,

wherein the neuron nodes at the hidden layer are configurable for regrouping and reconnecting through crossbar connections into different topologies to achieve trade off and optimization during on-chip learning.

19. A method of processing signals from a biomedical device, comprising:

attaching a biomedical device to a human body part, wherein the biomedical device comprises a neural processor coupled to at least one sensor which sends multi-channel analog signals of detected physiological activities to the neural processor, and wherein the biomedical device is one of a plurality of biomedical devices that form a distributed neural network;

directly extracting, by mixed-signal processing circuitry of the neural processor, features from multi-channel analog signals received from the at least one sensor, wherein the features that are extracted are statistical values in a digital format from analog signals, wherein the mixed-signal processing circuitry consists of an on-chip multi-channels voltage controlled oscillator-based frontend, wherein each channel of the voltage controlled oscillator-based front end further includes at least a voltage controlled oscillator clocked by a same on-chip clock generator, a plurality of comparators and counters and a single-differential converter;

executing program code stored in on-chip memory banks to configure the neural processor to process the features that are extracted, wherein the features that are extracted are processed according to weighted ranks corresponding to the features that are extracted and the weighted ranks are locally stored in the on-chip memory banks;

the processing of the features that are extracted comprising processing by a local neural network layer and a global network layer of the neural processor, wherein the local neural network layer is formed by processing circuitry comprising a plurality of neuron nodes that process the features that are extracted according to their weighted ranks, and the global network layer is formed by processing circuitry comprising global neuron nodes at an output to process and classify the features that are extracted of the at least one sensor; and

communicating through a global data line of the neural processor, the features that are extracted and classified with at least one other biomedical device within the distributed neural network.

Assignments (2)
CONFIRMATORY LICENSE Recorded Feb 13, 2025
From: NORTHWESTERN UNIVERSITY
To: NATIONAL SCIENCE FOUNDATION
Reel/Frame 070205/0424 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 19, 2021
From: GU, JIE
To: NORTHWESTERN UNIVERSITY
Reel/Frame 055968/0183 →
Continuity (2)
Provisional Application 62748075 · Oct 19, 2018
Related Publication 20210383201A1 · Dec 9, 2021
References Cited (32)
US 7048697B1 · Mitsuru · 2006 [cited by examiner]
US 20030236760A1 · Nugent · 2003 [cited by applicant]
US 20110300851A1 · Krishnaswamy et al. · 2011 [cited by applicant]
US 20140031952A1 · Harshbarger · 2014 [cited by examiner]
US 20170220923A1 · Bae · 2017 [cited by examiner]
US 20180330238A1 · Luciw · 2018 [cited by examiner]
US 20190197549A1 · Sharma · 2019 [cited by examiner]
Hwanjo Yu, Jinoh Oh, and Wook-Shin Han. 2009. Efficient feature weighting methods for ranking. In Proceedings of the 18th ACM conference on Information and knowledge management (CIKM '09). Association for Computing Mach… [cited by examiner]
M. Seiffert, F. Holstein, R. Schlosser and J. Schiller, “Next Generation Cooperative Wearables: Generalized Activity Assessment Computed Fully Distributed Within a Wireless Body Area Network,” in IEEE Access, vol. 5, pp… [cited by examiner]
M. Magno, M. Pritz, P. Mayer and L. Benini, “DeepEmote: Towards multi-layer neural networks in a low power wearable multi-sensors bracelet,” 2017 7th IEEE International Workshop on Advances in Sensors and Interfaces (IW… [cited by examiner]
Dubey et al., “Fog computing in medical internet-of-things: architecture, implementation, and applications.”, In: Handbook of Large-Scale Distributed Computing in Smart Healthcare dated Jun. 24, 2017, Retrieved on Dec. … [cited by applicant]
Yu et al., “Efficient feature weighting methods for ranking.”, In: Proceedings of the 18th ACM conference on Information and knowledge management dated Nov. 6, 2009, Retrieved on Dec. 13, 2019 from <http://citeseerx.ist… [cited by applicant]
International Search Report and Written Opinion dated Jan. 9, 2020 for PCT Application No. PCT/US2019/057255, 9 pages. [cited by applicant]
D. Farina, et al., “The extraction of neural information from the surface EMG for the control of upper-limb prostheses: emerging avenues and challenges,” IEEE Transactions on Neural Systems and Rehabilitation Engineerin… [cited by applicant]
N. Helleputte, et al., “A 345 μW multi-sensor biomedical SoC with bioimpedance, 3-channel ECG, motion artifact reduction, and integrated Dsp”, IEEE Journal of Solid-State Circuits, vol. 50, No. 1, pp. 230-244, Jan. 2015. [cited by applicant]
A. Young, et al., “Analysis of using EMG and mechanical sensors to enhance intent recognition in powered lower limb prostheses,” Journal of Neural Engineering, vol. 11, No. 5, Sep. 2014. [cited by applicant]
S. Wurth, et al., “A real-time comparison between direct control, sequential pattern recognition control and simultaneous pattern recognition control using a fitts' law style assessment procedure”, Journal of NeuroEngin… [cited by applicant]
A. Adewuyi, et al., “An analysis of intrinsic and extrinsic hand muscle EMG for improved pattern recognition control”, IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol. 24, No. 4, pp. 485-494, 201… [cited by applicant]
N. Krausz, et al., “Depth sensing for improved control of lower limb prostheses”, IEEE Transactions on Biomedical Engineering, vol. 62, No. 11, pp. 2576-2587, 2015. [cited by applicant]
M. Atzori, et al., “Electromyography data for non-invasive naturally-controlled robotic hand prostheses,” in Scientific Data, 1:140053, Dec. 2014. [cited by applicant]
N. Krausz, L. Hargrove, “Recognition of ascending stairs from 2D images for control of powered lower limb prostheses”, IEEE Inter. Conf. in Medicine and Biology Society (EMBC), 2015. [cited by applicant]
A. Jamthe, et al., “Harnessing big data for wireless body area network applications”, International Conf. on Computational Intelligence and Communication Networks (INFOCOM), 2015. [cited by applicant]
G. Almashaqbeh, et al., “A cloud-based interference-aware remote health monitoring system for non-hospitalized patients”, Symposium on Selected Areas in Communications, 2014. [cited by applicant]
H. Dubey, et al., “Fog Computing in Medical Internet-of-Things: Architecture, Implementation, and Applications”, arXiv: 1706.08012, 2017. [cited by applicant]
W. Shi, et al., “Edge computing: vision and challenges”, IEEE Internet of Things Journal, vol. 3, No. 5, pp. 637-646, Oct. 2016. [cited by applicant]
M. Satyanarayanan, “The emergence of edge computing”, Computer, vol. 50, No. 1, pp. 30-39, Jan. 2017. [cited by applicant]
B. Calhoun, et al., “Body sensor networks: a holistic approach from silicon to users”, Proceedings of the IEEE, vol. 100, No. 1, pp. 91-106, Jan. 2012. [cited by applicant]
K. AL-Tamimi, et al., “Preweighted Linearized VCO Analog-to Digital Converter,” in IEEE Transactions on Very Large Scale Integration (VLSI) Systems, vol. 435, pp. 1983-1987, Jun. 2017. [cited by applicant]
N. Desai, et al., “A scalable, 2.9 mW, 1 Mb/s e-textiles body area network transceiver with remotely-powered nodes and bi-directional data communication”, IEEE Journal of Solid-State Circuits, vol. 49, No. 9, pp. 1995-2… [cited by applicant]
J. Yoo, et al., “An 8-channel scalable EEG acquisition SoC with fully integrated patient-specific seizure classification and recording processor,” ISSCC, pp. 292-294, Feb. 2014. [cited by applicant]
S. Yin, et al., “A 1.06 μW smart ECG processor in 65 nm CMOS for realtime biometric authentication and personal cardiac monitoring,” Symposium on VLSI Circuits, Jun. 2017. [cited by applicant]
S. Benatti, et al., “A sub-10mW real-time implementation for EMG hand gesture recognition based on a multi-core biomedical SoC,” IWASI, pp. 139-144, Jul. 2017. [cited by applicant]