IP Library Granted Patent US 12,670,404
Granted Patent B2
US 12,670,404 · App. 17/469,573 · Granted Jun 30, 2026

Method and system for training a neural network model using knowledge distillation

Inventors: Peyman Passban (Montreal, CA); Yimeng Wu (Montreal, CA); Mehdi Rezagholizadeh (Montreal, CA)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06N3/088G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,404
App. No.
17/469,573
Filed
Sep 8, 2021
Granted
Jun 30, 2026
Kind
B2
Art Unit
2129
USPC
706/25
Abstract

An agnostic combinatorial knowledge distillation (CKD) method for transferring trained knowledge of neural model from a complex model (teacher) to a less complex model (student) is described. In addition to training the student to generate a final output that approximates both the teacher's final output and a ground truth of a training input, the method further maximizes knowledge transfer by training hidden layers of the student to generate outputs that approximate a representation of a subset of teacher hidden layers are mapped to each of the student hidden layers for a given training input.

Claims (834)

1 . A method of knowledge distillation from a teacher neural model having a plurality of teacher hidden layers to a student neural model having a plurality of student hidden layers, the method comprising:

training the teacher neural model, wherein the teacher neural model is configured to receive a training input and generate a teacher output for the training input; and

training the student neural model on a plurality of training inputs, wherein the student neural model is configured to receive inputs and generate a corresponding student output, comprising:

processing each training input using the teacher neural model to generate the teacher output for the training input, each of the plurality of teacher hidden layers generating a teacher hidden layer output to obtain a plurality of teacher hidden layer outputs;

mapping a subset of the plurality of teacher hidden layers to each of the plurality of student hidden layers;

calculating, for each of the plurality of student hidden layers, a representation of the teacher hidden layer outputs of the subset of the plurality of teacher hidden layers mapped to each of the plurality of student hidden layers; and

training the student to, for each of the training inputs, generate a student output that approximates the teacher output for the training input, wherein each of the plurality of student hidden layers, for each of the training inputs, is trained to generate a student hidden layer output that approximates the representation of the subset of the plurality of teacher hidden layers mapped to the each of the plurality of student hidden layers,

wherein the mapping comprises, for each of the student hidden layers, assigning an attention weight (ϵ ij ) to the teacher hidden layer output of each of the subset of the plurality of teacher hidden layers,

wherein each attention weight (ϵ ij ) is computed by:

ϵ

i

j

=

e

φ

i

j

k

=

1

k

=

|

H

T

|

e

φ

ik

where

φ

i

j

=

Φ

(

h

i

s

,

h

j

T

)

is an energy score between an output generated by an i th hidden layer of the student

(

h

i

s

)

 and an output generated by a j th hidden layer of the teacher

(

h

j

T

)

,

 the energy score being indicative of a similarity between the two generated outputs,

φ

i

k

=

Φ

(

h

i

s

,

h

k

T

)

is an energy score between the output generated by the i th hidden layer of the student

(

h

i

s

)

 and an output generated by the k th hidden layer of the teacher

(

h

k

T

)

,

 the energy score being indicative of a similarity between the two generated outputs,

Φ(h S , h T ) is an energy function of an output generated by a hidden layer of the student and an output generated by a hidden layer of the teacher,

h

k

T

where k∈{1, . . . , |H T |} is a k th layer of the teacher that belongs to a set of all hidden layers of the teacher (H T ), and

|H T | is a size of the set H T representing a total number of hidden layers of the teacher.

2 . The method of claim 1 , further comprising training the student to, for each of the training inputs, generate the student output to approximate a ground truth of the training input.

3 . The method of claim 1 , wherein training the student to generate the student output that approximates the teacher output for the training input further comprises:

computing a knowledge distillation (KD) loss between the student output and the teacher output;

computing a standard loss between the student output and the ground truth;

computing a combinatorial KD (CKD) loss between each of the plurality of student hidden layers and the subset of the teacher hidden layers mapped to each of the plurality of student layer;

calculating a total loss as a weighted average of the KD loss, the standard loss and the CKD loss; and

adjusting parameters of the student to minimize the total loss.

4 . The method of claim 3 , wherein the CKD loss is computed by:

L

CKD

=

h

i

s

H

s

MSE

(

h

i

s

,

f

i

T

)

where

L CKD is the CKD loss,

MSE( ) is a mean-square error function,

h

i

s

is the output generated by the i th hidden layer of the student,

f

i

T

is an output generated by the subset of hidden teacher layers which are associated with or assigned to the i th hidden layer of the student by a mapping function M which is computed by

f

i

T

=

F

(

H

T

(

i

)

)

,

where

H

T

(

i

)

=

{

j

M

(

i

)

}

,

F( ) is a fusion function that fuses/aggregates a subset of hidden teacher layers which are assigned to a particular hidden layer of the student model provided by a first

(

h

1

T

)

 and third

(

h

3

T

)

 hidden layers of the teacher,

an output of F( ) being mapped to a second hidden layer of the student

(

h

2

s

)

,

H S is a set of all hidden layers of the student

H T is the set of all hidden layers of the teacher,

H T (i) is a subset of hidden layers of the teacher selected mapped to the i th hidden layer of the student, and

M( ) is a mapper function that takes an index referencing to the hidden layer of the student and returns a set of indices for the teacher.

5 . The method of claim 4 , wherein the fusion function F( ) includes a concatenation operation followed by a linear projection layer.

6 . The method of claim 5 , wherein the fusion function F( ) is defined by:

F

(

h

1

T

,

h

3

T

)

=

m

u

l

(

W

,

[

h

1

T

;

h

3

T

]

)

+

b

where

“;” is a concatenation operator,

mul( ) is the matrix multiplication operation, and

W and b are learnable parameters.

7 . The method of claim 4 , wherein the mapping function M( ) defines a combination policy for combining the hidden layers of the teacher.

8 . The method of claim 7 , wherein the combination policy is any one of overlap combination, regular combination, skip combination, and cross combination.

9 . The method of claim 1 , wherein the mapping further comprises, defining a combination policy for mapping the teacher hidden layers to each of the plurality of student hidden layers.

10 . The method of claim 9 , wherein the combination policy is any one of overlap combination, regular combination, skip combination, and cross combination.

11 . The method of claim 3 , wherein the CKD loss is computed by:

L

C

K

D

*

(

H

S

,

H

T

)

=

h

i

s

H

s

MSE

(

h

i

s

,

f

i

*

T

)

where

L CKD* is the CKD loss,

MSE( ) is a mean-square error function,

h

i

s

is the output generated by the i th hidden layer of the student,

f

i

*

T

is a combined attention-based representation of the set of all hidden layers (H T ) of the teacher for an i th hidden layer of the teacher, and

H S is a set of all hidden layers of the student.

12 . The method of claim 11 , wherein the student output and the teacher output have identical dimensions

(

"\[LeftBracketingBar]"

h

i

s

"\[RightBracketingBar]"

=

"\[LeftBracketingBar]"

h

i

T

"\[RightBracketingBar]"

)

,

and

f

i

*

T

is computed by:

f

i

*

T

=

h

j

T

H

*

T

ϵ

i

j

h

j

T

where

ϵ ij is the attention weight, and the attention weight indicates how much the j th hidden layer of the teacher

(

h

j

T

)

 contributes to the knowledge distillation process of the i th hidden layer of the student

(

h

j

s

)

,

h

j

T

is a j th hidden layer of the teacher, and

H* T is the set of all hidden layers from the teacher that is assigned to the i th hidden layer of the student.

13 . The method of claim 11 , wherein the student output and the teacher output have different dimensions

(

"\[LeftBracketingBar]"

h

i

s

"\[RightBracketingBar]"

"\[LeftBracketingBar]"

h

j

T

"\[RightBracketingBar]"

)

,

and

f

i

*

T

is computed by:

f

i

*

T

=

h

j

T

H

T

ϵ

i

j

(

W

i

h

j

T

)

where

ϵ ij is the attention weight, and the attention weight indicates how much the j th hidden layer of the teacher

(

h

j

T

)

 contributes to the knowledge distillation process of the i th hidden layer of the student

(

h

i

s

)

,

h

j

T

is a j th hidden layer of the teacher,

H T is the set of all hidden layers of the teacher that is assigned to the i th hidden layer of the student, and

W

i

R

|

h

t

s

|

×

|

h

j

T

|

and is a weight value for the i th hidden layer of the teacher.

14 . The method of claim 13 , wherein a sum of all attention weights ϵ ij is 1.

15 . The method of claim 1 , wherein the energy function

Φ

(

h

i

s

,

h

j

T

)

is computed as the dot product of the output of the i th hidden layer of the student

(

h

i

s

)

and the output generated by the j th hidden layer of the teacher

(

h

j

T

)

by:

Φ

(

h

i

s

,

h

j

T

)

<

h

i

s

,

h

j

T

>

.

16 . The method of claim 15 , wherein the energy function

Φ

(

h

i

s

,

h

j

T

)

is computed as the dot product of the output of the i th hidden layer of the student

(

h

i

s

)

and a weighted value of the output generated by the j th hidden layer of the teacher

(

W

i

h

j

T

)

by:

Φ

(

h

i

s

,

h

j

T

)

<

h

i

s

,

W

i

h

j

T

>

.

17 . A method of knowledge distillation from a plurality of teacher neural models each having a plurality of teacher hidden layers to a student neural model having a plurality of student hidden layers, the method comprising:

inferring the plurality of teacher neural models, wherein each of the plurality of the teacher neural models is configured to receive an input and generate a teacher output; and

training the student neural model on a plurality of training inputs, wherein the student neural model is configured to receive inputs and generate a student output, comprising:

processing each training input using the plurality of teacher neural models to generate a plurality of teacher outputs for the training input, each of the plurality of teacher hidden layers of each of the plurality of teacher neural models generating a teacher hidden layer output;

mapping a subset of the plurality of teacher hidden layers of the plurality of teacher neural models to each of the plurality of student hidden layers;

calculating a representation of the teacher hidden layer outputs of the subset of the plurality of teacher hidden layers mapped to each of the plurality of student hidden layers; and

training the student neural model to generate a student output that approximates the plurality of teacher outputs for the training input, wherein each of the plurality of student hidden layers, for each of the training inputs, is trained to generate a student hidden layer output that approximates the representation of the subset of the plurality of teacher hidden layers mapped to the each of the plurality of student hidden layers,

wherein the mapping comprises, for each of the student hidden layers, assigning an attention weight (ϵ pq ) to the teacher hidden layer outputs of each of the subset of the plurality of teacher hidden layers,

wherein each attention weight (ϵ pq ) is computed by:

pq

=

e

φ

pq

i

=

1

K

e

φ

iq

where

φ

pq

=

Φ

(

T

p

(

x

i

q

)

,

S

(

x

i

q

)

)

;

x

i

q

D

q

is an energy score between an inferred output of a p th teacher T p , represented by

T

p

(

x

i

q

)

,

 and an inferred output of a student, represented by

S

(

x

i

q

)

,

 the energy score being indicative of a similarity between the respective inferred outputs,

φ iq is an energy score between an inferred output of an i th teacher and the inferred output of the student, where is K is a total number of teachers,

Φ

(

T

p

(

x

i

q

)

,

S

(

x

i

q

)

)

is an energy function of the inferred output of T p and the inferred output of S,

training data samples

{

(

x

i

q

)

}

i

=

1

N

q

of training dataset D q are sent to T p and S to obtain the respective inferred outputs, where D q is the training dataset for the K number of teachers, and

N q is a number of training data samples in the training dataset D q for the q th teacher.

18 . The method of claim 17 , wherein training the student neural model to generate a student output that approximates the plurality of teacher outputs for the training input further comprises computing a weighted knowledge distillation (KD) loss for each teacher neural model, where the weighted KD loss is computed by:

L

K

D

*

q

=

p

=

1

p

=

K

x

i

q

D

q

ϵ

p

q

L

K

D

(

T

p

(

x

i

q

)

,

S

(

x

i

q

)

)

where

L

K

D

*

q

is the weighted KD loss,

T

p

(

x

i

q

)

is the inferred output of the p th teacher T p ,

S

(

x

i

q

)

is the inferred output of the student S,

L KD is a loss function, and

∈ pq is the weight value of the teacher hidden layers of each p th teacher T p of the plurality of teacher neural models.

19 . The method of claim 17 , wherein

T

p

(

x

i

q

)

and

S

(

x

i

q

)

are vectors and the

Φ

(

T

p

(

x

i

q

)

,

S

(

x

i

q

)

)

function is a dot product of the

T

p

(

x

i

q

)

vector and the

S

(

x

i

q

)

vector.

20 . The method of claim 17 , wherein

T

p

(

x

i

q

)

and

S

(

x

i

q

)

are vectors and the

Φ

(

T

p

(

x

i

q

)

,

S

(

x

i

q

)

)

function is a neural network suitable for measuring a similarity of the

T

p

(

x

i

q

)

vector and the

S

(

x

i

q

)

vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 17, 2021
From: PASSBAN, PEYMAN; WU, YIMENG; REZAGHOLIZADEH, MEHDI
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 058547/0992 →
Continuity (2)
Provisional Application 63076335 · Sep 9, 2020
Related Publication 20220076136A1 · Mar 10, 2022
References Cited (52)
US 20190205748A1 · Fukuda et al. · 2019 [cited by applicant]
CN 106548190A · 2017 [cited by applicant]
CN 109165738A · 2019 [cited by applicant]
CN 111105008A · 2020 [cited by applicant]
CN 111242297A · 2020 [cited by applicant]
CN 111611377A · 2020 [cited by applicant]
Sun, Siqi, et al. “Patient knowledge distillation for bert model compression.” arXiv preprint arXiv:1908.09355 (2019) (Year: 2019). [cited by examiner]
MIT OpenCourseWare (https://ocw.mit.edu/courses/18-01sc-single-variable-calculus-fall-2010/14798b1a310b0f2917e2d1cecd7494bf_MIT18_01SCF10_Ses61a.pdf, 2010) (Year: 2010). [cited by examiner]
Yue (Yue, Kaiyu, Jiangfan Deng, and Feng Zhou. “Matching Guided Distillation.” arXiv preprint arXiv:2008.09958 (2020)). (Year: 2020). [cited by examiner]
Freitag (Freitag, Markus, Yaser Al-Onaizan, and Baskaran Sankaran. “Ensemble distillation for neural machine translation.” arXiv preprint arXiv:1702.01802 (2017) (Year: 2017). [cited by examiner]
T. Moriya, H. Sato, T. Tanaka, T. Ashihara, R. Masumura and Y. Shinohara, “Distilling Attention Weights for CTC-Based ASR Systems,” ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processi… [cited by examiner]
Aguilar, G.; Ling, Y.; Zhang, Y.; Yao, B.; Fan, X.; and Guo, C., Knowledge Distillation from Internal Representations. In AAAI, pp. 7350-7357, 2020. [cited by applicant]
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, Neural machine translation by jointly learning to align and translate, in 3rd International Conference on Learning Representations, ICLR, arXiv preprint arXiv:1409.047… [cited by applicant]
Bar-Haim, R.; Dagan, I.; Dolan, B.; Ferro, L.; Giampiccolo, D.; Magnini, B.; and Szpektor, I., The second pascal recognising textual entailment challenge. In Proceedings of the second PASCAL challenges workshop on recog… [cited by applicant]
Bentivogli, L.; Clark, P.; Dagan, I.; and Giampiccolo, D., The Fifth PASCAL Recognizing Textual Entailment Challenge. In TAC., 2009. [cited by applicant]
Buc Cristian Buciluá, Rich Caruana, and Alexandru Niculescu-Mizil, Model compression, In Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining, pp. 535-541, 2006. [cited by applicant]
Cer, D.; Diab, M.; Agirre, E.; Lopez-Gazpio, I.; and Specia, L., Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation. arXiv preprint arXiv:1708.00055, 2017. [cited by applicant]
Dagan, I.; Glickman, O.; and Magnini, B., The PAS-CAL recognising textual entailment challenge. In Machine Learning Challenges Workshop, 177-190. Springer, 2005. [cited by applicant]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, In Proceedings of the 2019 Conference f the North American Chapter of t… [cited by applicant]
Dolan, W. B.; and Brockett, C., Automatically constructing a corpus of sentential paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP2005), 2005. [cited by applicant]
Markus Freitag, Yaser Al-Onaizan, and Baskaran Sankaran, Ensemble distillation for neural machine translation, ArXiv preprint arXiv:1702.01802, 2017. [cited by applicant]
Furlanello, T.; Lipton, Z. C.; Tschannen, M.; Itti, L.; and Anandkumar, A., Born Again Neural Networks, 2018. [cited by applicant]
Giampiccolo, D.; Magnini, B.; Dagan, I.; and Dolan, B., The third pascal recognizing textual entailment challenge. In Proceedings of the ACL-PASCAL workshop on textual entailment and paraphrasing, 1-9. Association for C… [cited by applicant]
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean, Distilling the knowledge in a neural network, arXiv preprint arXiv:1503.02531, 2015. [cited by applicant]
Iyer, S.; Dandekar, N.; and Csernai, K., First Quora Dataset Release: Question Pairs. URL https://data.quora.com/First-Quora-Dataset-Release-Question-Pairs, 2017. [cited by applicant]
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu, Tinybert: Distilling bert for natural language understanding. arXiv preprint arXiv:1909.10351, 2019. [cited by applicant]
Yoon Kim and Alexander M Rush, Sequence-Level Knowledge Distillation, In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 1317-1327, 2016. [cited by applicant]
Liu, X.; He, P.; Chen, W.; and Gao, J., Improving MultiTask Deep Neural Networks via Knowledge Distillation for Natural Language Understanding. ArXiv abs/1904.09482, 2019. [cited by applicant]
Peters, M. E.; Neumann, M.; Iyyer, M.; Gardner, M.; Clark, C.; Lee, K.; and Zettlemoyer, L., Deep contextualized word representations. In Proc. of NAACL, 2018. [cited by applicant]
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P., Squad: 100,000+ questions for machine comprehension of text. arXiv preprint arXiv: 1606.05250, 2016. [cited by applicant]
Sanh, V.; Debut, L.; Chaumond, J.; and Wolf, T., DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. ArXiv abs/1910.01108, 2019. [cited by applicant]
Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C. D.; Ng, A.; and Potts, C., Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical met… [cited by applicant]
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu, Patient Knowledge Distillation for BERT Model Compression. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th Internation… [cited by applicant]
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou, Mobilebert: a compact task-agnostic bert for resource-limited devices. arXiv preprint arXiv:2004.02984, 2020. [cited by applicant]
Xu Tan, Yi Ren, Di He, Tao Qin, and Tie-Yan Liu, Multilingual Neural Machine Translation with Knowledge Distillation. In International Conference on Learning Representations. URL https://openreview.net/forum?id=S1gUsoR9… [cited by applicant]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, Attention is all you need. In Advances in neural information processing systems, pp. 5998-6008… [cited by applicant]
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R., GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. ArXiv abs/1804.07461, 2018. [cited by applicant]
Wang, W.; Wei, F.; Dong, L.; Bao, H.; Yang, N.; and Zhou, M., Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers. arXiv preprint arXiv:2002.10957, 2020. [cited by applicant]
Warstadt, A.; Singh, A.; and Bowman, S. R., Neural Network Acceptability Judgments. arXiv preprint arXiv:1805.12471, 2018. [cited by applicant]
Hao-Ran Wei, Shujian Huang, Ran Wang, Xinyu Dai, and Jiajun Chen, Online Distilling from Checkpoints for Neural Machine Translation. In Proceedings of the 2019 Conference of the North American Chapter of the Association… [cited by applicant]
Williams, A.; Nangia, N.; and Bowman, S., A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Comput… [cited by applicant]
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning, What does BERTlook at? An analysis of BERT's attention, In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Netwo… [cited by applicant]
Sepp Hochreiter and Jürgen Schmidhuber, Long short-term memory. Neural computation, 9(8): pp. 1735-1780, 1997. [cited by applicant]
Marcin Junczys-Dowmunt, Roman Grundkiewicz, Tomasz Dwojak, Hieu Hoang, Kenneth Heafield, Tom Neckermann, Frank Seide, Ulrich Germann, Alham Fikri Aji, Nikolay Bogoychev, Andre´ F. T. Martins, and Alexandra Birch, Marian… [cited by applicant]
Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senel-Iart, and Alexander M. Rush, Opennmt: Open-source toolkit for neural machine translation. In Proc. ACL, 2017. [cited by applicant]
Taku Kudo and John Richardson, Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Languag… [cited by applicant]
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli, Fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of NAACL-HLT 2019: Demonstrations, 2… [cited by applicant]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu, Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting on association for computational linguistics, pp. 31… [cited by applicant]
Ashish Vaswani, Samy Bengio, Eugene Brevdo, Fran¬cois Chollet, Aidan N. Gomez, Stephan Gouws, Llion Jones, Łukasz Kaiser, Nal Kalchbrenner, Niki Parmar, Ryan Sepassi, Noam Shazeer, and Jakob Uszkoreit, Tensor2tensor for… [cited by applicant]
Clark, K., Luong, M. T., Khandelwal, U., Manning, C. D., & Le, Q. V., Bam! born-again multi-task networks for natural language understanding. arXiv preprint arXiv:1907.04829, 2019. [cited by applicant]
Matt Post, A Call for Clarity in Reporting BLEU Scores, Proceedings of the Third Conference on Machine Translation (WMT18), 6 pages, 2018. [cited by applicant]
Svante Wold, Kim Esbensen and Paul Geladi, Principal component analysis, Chemometrics and Intelligent Laboratory Systems, vol. 2, Issues 1-3, pp. 37-52, 1987. [cited by applicant]