IP Library Granted Patent US 12675082
Granted Patent B2
US 12675082 · App. 18/033,302 · Granted Jul 7, 2026

Training device, control system, training method, and recording medium

Inventors: Shumpei Kubosawa (Tokyo, JP); Takashi Onishi (Tokyo, JP)
Assignee: NEC CORPORATION
G05B13/0265
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12675082
App. No.
18/033,302
Granted
Jul 7, 2026
Kind
B2
Abstract

A training device updates, among a plurality of models, a model for a region that includes a given sample based on the sample. The plurality of models are provided for each region obtained by dividing a state space that includes a sample indicates a state about a control object and the plurality of models represent an evaluation of an action of the control object in response to control over the control object. The training device evaluates an action of the control object in a given state, based on a model for a region that includes a sample indicating the state.

Claims (39)

1 . A training device comprising:

a memory configured to store instructions; and

a processor configured to execute the instructions to:

update, among a plurality of models, one model for one region based on one sample included in the one region,

wherein a plurality of regions including the one region are obtained by dividing a state space, wherein each of the plurality of regions includes a sample indicating a state about a control object, wherein each of the plurality of models is provided for each of the plurality of regions, and wherein each of the plurality of models represents an evaluation of an action of the control object in response to control over the control object;

evaluate one action of the control object in one state indicated by the one sample based on the one model;

update the one region based on the one sample;

update division of the plurality of regions based on a difference between a value calculated by applying the one sample to the model and a value obtained based on the one sample; and

control the control object based on the evaluation of the action of the control object in response to the control over the control object,

wherein the control object is a vehicle, the state space corresponds to vehicle travel control, and the plurality of regions include a first state in which a road surface is dry and a second state in which the road surface is wet.

2 . The training device according to claim 1 , wherein the processor is configured to execute the instructions to determine whether or not the one sample is included in the one region among the plurality of regions in the state space.

3 . The training device according to claim 1 , wherein the processor is configured to execute the instructions to update the region based on a Mahalanobis distance between the one sample in the state space and a representative point for determining the region.

4 . The training device according to claim 1 , wherein the processor is configured to execute the instructions to determine a value of an error counter variable which is provided for each representative point for determining the region and indicates a degree of difference between each of the representative points and the one sample, based on a distance between the one sample in the state space and the representative point and a difference between a value calculated by applying the one sample to the model and the value obtained based on the one sample, and determines an installation position of a new representative point in the state space based on the value of the error counter variable.

5 . The training device according to claim 1 , wherein the processor is configured to execute the instructions to divide the region when an index value based on a distance between the one sample in the state space and a representative point for determining the region satisfies a predetermined condition.

6 . The training device according to claim 1 , wherein the model associated with at least one of the regions is a linear model.

7 . The training device according to claim 1 , wherein the model associated with at least one of the regions is a nonlinear model.

8 . A control system comprising:

the training device according to claim 1 ; and

the control object.

9 . A training device comprising:

a memory configured to store instructions; and

a processor configured to execute the instructions to:

determine a region that includes a given sample when there is a model that is provided in each of a plurality of regions obtained by dividing a state space that includes a sample indicating a state about a control object and the model represents an evaluation of an action of the control object in response to control over the control object;

evaluate the action based on the given sample and the model for the determined region;

update the region based on the sample;

update the division of the plurality of regions based on a difference between a value calculated by applying the given sample to the model and a value obtained based on the given sample: and

control the control object based on the evaluation of the action of the control object in response to the control over the control object,

wherein the control object is a vehicle, the state space corresponds to vehicle travel control, and the plurality of regions include a first state in which a road surface is dry and a second state in which the road surface is wet.

10 . A control system comprising:

the training device according to claim 9 ; and

the control object.

11 . A training method comprising:

updating, among a plurality of models, one model for one region based on one sample included in the one region,

wherein a plurality of regions including the one region are obtained by dividing a state space, wherein each of the plurality of regions includes a sample indicating a state about a control object, wherein each of the plurality of models is provided for each of the plurality of regions, and wherein each of the plurality of models represents an evaluation of an action of the control object in response to control over the control object;

evaluating one action of the control object in one state indicated by the one sample based on the one model;

updating the one region based on the one sample;

updating division of the plurality of regions based on a difference between a value calculated by applying the one sample to the model and a value obtained based on the one sample; and

controlling the control object based on the evaluation of the action of the control object in response to the control over the control object,

wherein the control object is a vehicle, the state space corresponds to vehicle travel control, and the plurality of regions include a first state in which a road surface is dry and a second state in which the road surface is wet.