Management apparatus, lithography apparatus. management method, and article manufacturing method
A management apparatus includes a learning device. The learning device is configured to, in a case where a reward obtained from a control result of a controlled object by a controller configured to control the controlled object using a neural network, for which a parameter value is decided by reinforcement learning, does not satisfy a predetermined criterion, redecide the parameter value by reinforcement learning.
1 . A management apparatus comprising:
a processor; and
a memory including instructions, which when executed by the processor, cause the management apparatus to perform operations comprising:
performing a processing sequence of executing processing for a processing target object, wherein a stage is controlled, in accordance with a driving command and an output of a sensor that detects a state of the stage, using a controller comprising a compensator which includes a neural network, for which a parameter value has been decided by reinforcement learning; and
obtaining, after the processing sequence completes, a reward from a control result of the stage in the processing sequence, wherein, in a case where the reward does not satisfy a predetermined criterion, performing the reinforcement learning to redecide the parameter value.
2 . The management apparatus according to claim 1 , wherein
the stage includes a holder configured to hold the processing target object,
and in the processing sequence of executing processing for the processing target object, the holder is controlled so as to move the holder, and
in a case where a reward obtained from a control result of the holder in the processing sequence does not satisfy the predetermined criterion, the parameter value is redecided by reinforcement learning.
3 . The management apparatus according to claim 2 , wherein
the processing sequence includes a plurality of sub-sequences,
the predetermined criterion includes a plurality of criteria each corresponding to each of the plurality of sub-sequences, and
in a case where a reward obtained from a control result of the holder in each of the plurality of sub-sequences does not satisfy a corresponding criterion among the plurality of criteria, the parameter value is redecided by reinforcement learning.
4 . The management apparatus according to claim 3 , wherein
the processing sequence is a sequence for transferring a pattern of an Original to a substrate as the processing target object, and
the plurality of sub-sequences include a conveyance sequence in which the substrate is conveyed, a measurement sequence in which an alignment error between the substrate and the Original is measured, and an exposure sequence in which the pattern of the Original is projected onto the substrate and the substrate is exposed.
5 . The management apparatus according to claim 4 , wherein
among the plurality of criteria, a criterion corresponding to the conveyance sequence is related to a time required for a control error of the holder to converge to a predetermined value or less.
6 . The management apparatus according to claim 4 , wherein
among the plurality of criteria, a criterion corresponding to the measurement sequence is related to a control error of the holder during measurement of the alignment error between the substrate and the Original.
7 . The management apparatus according to claim 4 , wherein
among the plurality of criteria, a criterion corresponding to the exposure sequence is related to a synchronous error between the substrate and the Original during exposure of the substrate.
8 . The management apparatus according to claim 2 , wherein
the parameter value is redecided by reinforcement learning after the processing sequence ends.
9 . The management apparatus according to claim 1 , wherein
The stage includes a holder configured to hold the processing target object,
and in a period in which the processing sequence of executing processing for the processing target object is not executed, the holder is controlled so as to move the holder, and
in a case where a reward obtained from a control result of the holder in the period does not satisfy the predetermined criterion, the parameter value is redecided by reinforcement learning.
10 . The management apparatus according to claim 1 , wherein
a position of the stage is controlled.
11 . The management apparatus according to claim 1 , wherein the compensator is configured to generate a first command value in accordance with the driving command and the output of the sensor, and
wherein the controller further comprises a second compensator configured to generate a second command value based on the driving command and the output of the sensor, and an adder configured to generate a command value based on the first command value and the second command value, and
the command value is supplied to a driver configured to drive the stage.
12 . The management apparatus according to claim 1 , wherein the reward is obtained from the control result of the stage in the processing sequence after the processing target object is unloaded.
13 . The management apparatus according to claim 1 , wherein during the processing sequence, the reinforcement learning, including the decision of the parameter value, is not performed when the reward satisfies the predetermined criterion.
14 . A lithography apparatus for performing processing of transferring a pattern of an Original to a substrate, the apparatus comprising:
performing a processing sequence of executing processing for the substrate, wherein the processing includes a stage;
a controller comprising a compensator which includes a neural network for which a parameter value has been decided by reinforcement learning, wherein the stage in the processing sequence is controlled by the controller in accordance with a driving command and an output of a sensor that detects a state of the stage; and
obtaining, after the processing sequence completes, a reward from a control result of the stage in the processing sequence, wherein, in a case where the reward obtained does not satisfy a predetermined criterion, performing the reinforcement learning to redecide the parameter value.
15 . The lithography apparatus according to claim 14 , wherein
the stage includes a holder configured to hold the substrate,
in the processing sequence of executing the processing, the holder is controlled so as to move the holder, and
in a case where a reward obtained from a control result of the holder in the processing sequence does not satisfy the predetermined criterion, the parameter value is redecided by reinforcement learning.
16 . The lithography apparatus according to claim 15 , wherein
the processing sequence includes a plurality of sub-sequences,
the predetermined criterion includes a plurality of criteria each corresponding to each of the plurality of sub-sequences, and
in a case where a reward obtained from a control result of the holder in each of the plurality of sub-sequences does not satisfy a corresponding criterion among the plurality of criteria, the parameter value is redecided by reinforcement learning.
17 . The lithography apparatus according to claim 16 , wherein
the plurality of sub-sequences include a conveyance sequence in which the substrate is conveyed, a measurement sequence in which an alignment error between the substrate and the Original is measured, and an exposure sequence in which the pattern of the Original is projected onto the substrate and the substrate is exposed.
18 . The lithography apparatus according to claim 17 , wherein
among the plurality of criteria, a criterion corresponding to the conveyance sequence is related to a time required for a control error of the holder to converge to a predetermined value or less.
19 . The lithography apparatus according to claim 17 , wherein
among the plurality of criteria, a criterion corresponding to the measurement sequence is related to a control error of the holder during measurement of the alignment error between the substrate and the Original.
20 . The lithography apparatus according to claim 17 , wherein
among the plurality of criteria, a criterion corresponding to the exposure sequence is related to a synchronous error between the substrate and the Original during exposure of the substrate.
21 . A management method comprising:
a causing step of causing a controller to perform a processing sequence of executing processing for a processing target object, wherein a stage is controlled by the controller in the processing sequence in accordance with a driving command and an output of a sensor that detects a state of the stage, and the controller comprises a compensator which includes a neural network, for which a parameter value has been decided by reinforcement learning;
an acquiring step of acquiring a control result of the stage in the processing sequence from the controller, after the processing sequence completes; and
a learning step of, in a case where a reward obtained from the control result does not satisfy a predetermined criterion, redeciding the parameter value by reinforcement learning.
22 . An article manufacturing method comprising:
a transfer step of transferring a pattern of an Original to a substrate using a lithography apparatus defined in claim 14 ; and
a processing step of processing the substrate having undergone the transfer step,
wherein an article is obtained from the substrate having undergone the processing step.