IP Library Granted Patent US 11,681,901
Granted Patent B2
US 11,681,901 · App. 16/879,934 · Granted Jun 20, 2023

Quantifying the predictive uncertainty of neural networks via residual estimation with I/O kernel

Inventors: Xin Qiu (San Francisco, CA); Risto Miikkulainen (Stanford, CA); Elliot Meyerson (San Francisco, CA)
Assignee: Cognizant Technology Solutions U.S. Corporation
G06N3/047G06F17/16G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,681,901
App. No.
16/879,934
Granted
Jun 20, 2023
Kind
B2
Abstract

A residual estimation with an I/O kernel (“RIO”) framework provides estimates of predictive uncertainty of neural networks, and reduces their point-prediction errors. The process captures neural network (“NN”) behavior by estimating their residuals with an I/O kernel using a modified Gaussian process (“GP”). RIO is applicable to real-world problems, and, by using a sparse GP approximation, scales well to large datasets. RIO can be applied directly to any pretrained NNs without modifications to model architecture or training pipeline.

Claims (20)

1. A process for estimating residuals of a neural network (NN) model to determine uncertainty in the NN model's predictions of a value of at least one of a physical, chemical or electrical variable, the process comprising:

training, by a processor, a NN model to make one or more predictions of the value of the at least one of a physical, chemical or electrical variable using a training data ( , y) input set, wherein the training data input set includes (x i ,y i )) n i=1 wherein y i are the expected outcomes by the NN model given input x i;

storing, by the processor, an output data set from the NN model, including the one or more predictions resulting from operation on the training data input set, wherein the output data set includes (x i , ŷ i )} n i=1 , wherein ŷ i are the predicted outcomes by the NN model given input x i ; and

training, by the processor, a Gaussian process (GP) to estimate residuals of the NN model when applied to raw input data x* using the training data input set (x i , y i )) n i=1 and the output data set (x i , ŷ i )) n i=1,

wherein the training of the Gaussian process (GP) includes:

calculating, by the processor, residuals r={r i =y i −ŷ i } n i=1 wherein r denotes the vector of all residuals and ŷ denotes the vector of all NN model predictions;

calculating, by the processor, an n×n covariance matrix at all pairs of training points based on a composite kernel K c (( , ŷ), ( , ŷ)), where each entry is given by k c ((x i , ŷ i ), (x j , ŷ j ))=k in (x i , x j )+k out (ŷ i , ŷ j ), for i,j=1,2, . . . , n; and

optimizing, by a gradient-based optimizer, GP hyperparameters σ 2 in , l in , σ 2 out , l out , and σ 2 n by maximizing log marginal likelihood

log p ( r| ,ŷ )=−½ r T ( K c (( , ŷ ), ( , ŷ ))+σ 2 n I ) −1 r −½log| K c (( , ŷ ), ( ,ŷ ))+σ 2 n I|−n /2log 2π.

2. The process according to claim 1 , wherein the NN model is a fully connected feed-forward network.

3. The process according to claim 1 , further comprising:

applying, by the processor, the trained Gaussian process (GP) to predictions ŷ * of a neural network (NN) model applied to raw input data x * , the applying including:

calculating, by the processor, residual mean;

calculating, by the processor, residual variance; and

returning, by the processor, distribution of calibrated prediction ŷ′ * .

4. The process according to claim 3 , further comprising

calculating, by the processor, residual mean in accordance with {circumflex over ( r )} * =k T * (K c (( , ŷ), ( , ŷ))+σ 2 n I) −1 r;

calculating, by the processor, residual variance in accordance with var({circumflex over (r)} * )=k c ((x * , ŷ * ), (x * , ŷ * ))−k T * (K c (( , ŷ), ( , ŷ))+σ 2 n I) −1 k * ; and

returning, by the processor, a distribution of calibrated prediction ŷ′ 2 in accordance with ŷ′ * ˜ (ŷ * +{circumflex over ( r )} * , var({circumflex over (r)} * )).

5. The process according to claim 3 , wherein the NN model is a fully connected feed-forward network.

Assignments (5)
CORRECTIVE ASSIGNMENT TO CORRECT THE APPLICATION NUMBER 16878934 PREVIOUSLY RECORDED AT REEL: 052725 FRAME: 0439. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT . Recorded May 4, 2022
From: QIU, XIN
To: COGNIZANT TECHNOLOGY SOLUTIONS U.S. CORPORATION
Reel/Frame 059853/0805 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 5, 2022
From: QIU, XIN
To: COGNIZANT TECHNOLOGY SOLUTIONS U.S. CORPORATION
Reel/Frame 058559/0302 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2020
From: QIU, XIN
To: COGNIZANT TECHNOLOGY SOLUTIONS U.S. CORPORATION
Reel/Frame 052725/0439 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2020
From: MIIKKULAINEN, RISTO
To: COGNIZANT TECHNOLOGY SOLUTIONS U.S. CORPORATION
Reel/Frame 052725/0483 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2020
From: MEYERSON, ELLIOT
To: COGNIZANT TECHNOLOGY SOLUTIONS U.S. CORPORATION
Reel/Frame 052725/0761 →
Continuity (2)
Provisional Application 62851782 · May 23, 2019
Related Publication 20200372327A1 · Nov 26, 2020