IP Library › Granted Patent US 11,914,672
Granted Patent B2
US 11,914,672 · App. 17/488,796 · Granted Feb 27, 2024

Method of neural architecture search using continuous action reinforcement learning

Inventors: Mohammad Salameh (Edmonton, CA); Keith George Mills (Leduc, CA); Di Niu (Edmonton, CA)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06F18/214G06F18/217G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,914,672
App. No.
17/488,796
Granted
Feb 27, 2024
Kind
B2
Abstract

A method and system for generating neural architectures to perform a particular task. An actor neural network, as part of a continuous action reinforcement learning (RL) agent, generates a randomized continuous actions parameters to encourage exploration of a search space to generate candidate architectures without bias. The continuous action parameters are discretized and applied to a search space to generate candidate architectures, the performance of which for performing the particular task is evaluated. Corresponding reward and state are determined based on the performance. A critic neural network, as part of the continuous action RL agent, learns a mapping of the continuous action to a reward using modified Deep Deterministic Policy Gradient (DDPG) with quantile loss function by sampling a list of top performing architectures. The actor neural network is updated with the learned mapping.

Claims (148)

1. A method for neural architectural search (NAS) for performing a task, the method comprising:

(i) generating, by an actor neural network having actor parameters in accordance with current values of the actor parameters, a set of continuous neural network architecture parameters comprising score distributions over possible values for configuring a plurality of architecture cells of a trained search space;

(ii) discretizing the set of continuous architecture parameters into a set of discrete neural network architecture parameters;

(iii) generating a candidate architecture by configuring the trained search space using the discrete neural network architecture parameters, which specify a subset of the plurality of architecture cells to be active;

(iv) evaluating a performance of the candidate architecture at performing the task;

(v) determining a reward and a state for the discrete neural network architecture parameters based on the performance;

(vi) storing an experience tuple comprising the continuous neural network architecture parameters, the reward, and the state in a buffer storage;

(vii) learning a mapping, by a critic neural network, between network architectures and performance; and

(viii) updating the actor neural network with the learned mapping from the critic neural network.

2. The method of claim 1 , wherein the generating the set of continuous neural network architecture parameters comprises incorporating a randomized noise value into the set of continuous neural network architecture parameters.

3. The method of claim 1 , wherein the search space is a weight-sharing supernet, and is trained by:

in each training session in a plurality of training sessions:

generating, from a set of training data, a batch of training data comprising a plurality of training data samples; and

for each training data sample:

generating a set of continuous neural network architecture parameters comprising score distributions over possible values for configuring the plurality of architecture cells;

discretizing the set of continuous architecture parameters into a set of discrete neural network architecture parameters;

selecting a candidate architecture by assigning the discrete neural network architecture parameters to the supernet;

evaluating a performance of the selected candidate architecture at performing the task with a performance metric;

determining a loss value as a function of the difference between the performance metric and validation data; and

updating a subset of the weight values of the supernet to minimize the loss value.

4. The method of claim 3 , wherein the updating further comprises only updating the weight values of the supernet that are associated with the candidate architecture.

5. The method of claim 1 , further comprising:

storing, based on the performance of the candidate architecture, a list of top performing candidate architectures into an architecture history storage.

6. The method of claim 5 , wherein the storing comprises:

comparing the performance of the candidate architecture with a performance of a worst stored architecture;

if the performance of the candidate architecture is better than the performance of the worst stored architecture, replacing the worst stored architecture with the candidate architecture; and

sorting the list of top performing architecture based on performance.

7. The method of claim 1 , wherein the discretizing uses a many-to-one mapping algorithm.

8. The method of claim 1 , wherein the learning comprises:

sampling a batch from the buffer storage; and

for each experience tuple in the batch, performing operations comprising:

predicting a reward of the candidate architecture based on a current mapping;

determining a check loss using quantile regression as a function of the predicted reward and the reward from each experience tuple; and

updating the current mapping to minimize the check loss.

9. The method of claim 8 , wherein the check loss is determined using the following equation:

ℒ

c

⁢

r

⁢

i

⁢

t

⁢

i

⁢

c

=

1

❘

"\[LeftBracketingBar]"

B

R

❘

"\[RightBracketingBar]"

⁢

∑

i

∈

B

R

u

i

(

τ

-

1

⁢

(

u

i

<

0

)

)

,

where critic is the check loss, B R is the batch of training data, τ is a decimal value τ∈[0,1] that corresponds to a desired quantile level of the reward from each experience tuple, and u i is a difference between the predicted reward and the reward from each experience tuple.

10. The method of claim 9 , wherein the parameter u i is determined using the following equation:

u i =r i −Q ( a i ),

where r i is a mapped reward for the i th action Q(a i ) is the predicted reward value for a i , and u i is the difference between the mapped reward for the i th action a i and the predicted reward value for a i .

11. The method of claim 9 , wherein the parameter T is used to cause the critic to learn a mapping from a desired performance quantile of candidate architectures.

12. The method of claim 9 , wherein the task is image classification and reward value r t may be determined in accordance with the following equation:

r t =100 Acc(α t d ) ,

where Acc(α t d ) is the accuracy value of the candidate architecture selected based on the discrete architecture parameters α t d in its decimal form.

13. The method of claim 1 , wherein the updating comprises:

determining a loss value using the following equation:

ℒ

a

⁢

c

⁢

t

⁢

o

⁢

r

=

1

❘

"\[LeftBracketingBar]"

B

R

❘

"\[RightBracketingBar]"

⁢

∑

i

∈

B

R

Q

⁡

(

μ

⁡

(

s

i

)

)

,

where actor is the loss value of the actor neural network, B R is the batch of training data, and Q(μ(s i )) is a predicted reward by critic neural network of each output μ(s i ) of the actor neural network for a state corresponding to one of the experience tuples of the batch training data B R .

14. The method of claim 2 , wherein the randomized noise value is incorporated into the set of continuous neural network architecture parameters in accordance with a probability value associated with each continuous neural network architecture parameter.

15. The method of claim 14 , further comprising:

initializing the probability value indicative of high probability; and

annealing the probability value to a minimum value over a plurality of cycles.

16. The method of claim 15 , wherein the annealing further comprises applying a cosine annealing schedule.

17. The method of claim 1 , wherein the operations (i) to (viii) are repeatedly performed.

18. The method of claim 17 , wherein each experience tuple is comprised of the state, action, and reward (s t , a t , r t ) for each step t, wherein the state s t defines a set of channel-wise average of discrete neural network architecture parameters, the action a t is a set of continuous neural architecture parameters, and the reward r t defines the reward.

19. A computing device, comprising:

one or more processors configured to:

generate, by an actor neural network having actor parameters in accordance with current values of the actor parameters, a set of continuous neural network architecture parameters comprising score distributions over possible values for configuring a plurality of architecture cells of a trained search space;

discretize the set of continuous architecture parameters into a set of discrete neural network architecture parameters;

generate a candidate architecture by configuring the trained search space using the discrete neural network architecture parameters, which specify a subset of the plurality of architecture cells to be active;

evaluate a performance of the candidate architecture at performing a task;

determine a reward and a state for the discrete neural network architecture parameters based on the performance;

store an experience tuple comprising the continuous neural network architecture parameters, the reward, and the state in a buffer storage;

learn a mapping, by a critic neural network, between network architectures and performance; and

update the actor neural network with the learned mapping from the critic neural network.

20. A non-transitory machine-readable storage medium having tangibly stored thereon executable instructions for execution by a processor of a computing device that, in response to execution by the processor, cause the computing device to:

generate, by an actor neural network having actor parameters in accordance with current values of the actor parameters, a set of continuous neural network architecture parameters comprising score distributions over possible values for configuring a plurality of architecture cells of a trained search space;

discretize the set of continuous architecture parameters into a set of discrete neural network architecture parameters;

generate a candidate architecture by configuring the trained search space using the discrete neural network architecture parameters, which specify a subset of the plurality of architecture cells to be active;

evaluate a performance of the candidate architecture at performing a task;

determine a reward and a state for the discrete neural network architecture parameters based on the performance;

store an experience tuple comprising the continuous neural network architecture parameters, the reward, and the state in a buffer storage;

learn a mapping, by a critic neural network, between network architectures and performance; and

update the actor neural network with the learned mapping from the critic neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2021
From: SALAMEH, MOHAMMAD; MILLS, KEITH GEORGE; NIU, DI
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 057994/0931 →
Continuity (1)
Related Publication 20230096654A1 · Mar 30, 2023