Information processing apparatus, information processing method, mobile object control device, and mobile object control method
An information processing apparatus capable of controlling a mobile object on the basis of an instruction by an utterance of a user identifies which scene a use scene of a target user is among a plurality of use scenes in a case where the mobile object is used, acquires utterance information of the target user, and selects a different machine learning model according to the identified use scene of the target user. The information processing apparatus estimates an intent of an utterance of the target user by using the selected machine learning model.
1. An information processing apparatus capable of controlling a self-driving vehicle on the basis of an instruction by an utterance of a user, the information processing apparatus comprising:
one or more processors; and
a memory storing instructions which, when the instructions are executed by the one or more processors, cause the information processing apparatus to:
identify a use scene of a target user among a plurality of use scenes, wherein each of the use scenes is a state of a user comprising a state before boarding, a state during boarding, or a state after alighting a self-driving vehicle;
acquire utterance information of the target user;
select a different machine learning model from a plurality of machine learning models according to the identified use scene of the target user, wherein each one of the machine learning models relates to a corresponding one of the state before boarding, the state during boarding, or the state after alighting the self-driving vehicle; and
estimate an intent of an utterance of the target user by using the selected machine learning model.
2. The information processing apparatus according to claim 1 , wherein
each one of the machine learning models has a different intent class to be estimated for each of the use scenes with which the respective machine learning models are associated.
3. The information processing apparatus according to claim 2 , wherein
the instructions cause the information processing apparatus to estimate the intent of the target user by using one of the machine learning models that outputs the likelihood for predetermined intent classes among all intent classes associated with the plurality of use scenes.
4. The information processing apparatus according to claim 1 , wherein
the instructions cause the information processing apparatus to estimate the intent of the utterance of the target user in consideration of calculation using an initial state probability distribution set to an intent class as a prior distribution to output of the selected machine learning model.
5. The information processing apparatus according to claim 4 , wherein
the initial state probability distribution set as the prior distribution is separately determined for each of the use scenes.
6. The information processing apparatus according to claim 1 , wherein
the instructions cause the information processing apparatus to estimate the intent of the utterance of the target user in consideration of calculation using a state transition probability distribution between intent classes to output of the selected machine learning model.
7. The information processing apparatus according to claim 6 , wherein
the state transition probability distribution is separately determined for each one of the use scenes.
8. The information processing apparatus according to claim 1 , wherein
in a case where an intent of an utterance at time t is estimated, the instructions cause the information processing apparatus to estimate the intent of the utterance of the target user in consideration of an estimation result estimated for an utterance immediately before the utterance at the time t to output of the selected machine learning model.
9. The information processing apparatus according to claim 1 , wherein
each of the machine learning models is learned using learning data different for each corresponding ones of the use scenes, and the learning data includes a label indicating the corresponding one of the use scenes.
10. The information processing apparatus according to claim 1 , wherein
the instructions cause the information processing apparatus to identify one of the use scenes of the target user on a basis of information from a vehicle associated with the target user.
11. An information processing method in an information processing apparatus capable of controlling a self-driving vehicle on the basis of an instruction by an utterance of a user, the information processing method comprising:
identifying a use scene of a target user among a plurality of use scenes, wherein each of the use scenes is a state of a user comprising a state before boarding, a state during boarding, or a state after alighting a self-driving vehicle;
acquiring utterance information of the target user;
selecting a different machine learning model from a plurality of machine learning models according to the identified use scene of the target user, wherein each one of the machine learning models relates to a corresponding one of the state before boarding, the state during boarding, or the state after alighting the self-driving vehicle; and
estimating an intent of an utterance of the target user by using the selected machine learning model.
12. A control device of a self-driving vehicle that is controllable on the basis of an instruction by an utterance of a user, the control device comprising:
one or more processors; and
a memory storing instructions which, when the instructions are executed by the one or more processors, cause the control device to:
identify a use scene of a target user among a plurality of use scenes, wherein each of the use scenes is a state of a user comprising a state before boarding, a state during boarding, or a state after alighting a self-driving vehicle;
acquire utterance information of the target user;
select a different machine learning model from a plurality of machine learning models according to the identified use scene of the target user, wherein each one of the machine learning models relates to a corresponding one of the state before boarding, the state during boarding, or the state after alighting the self-driving vehicle; and
estimate an intent of an utterance of the target user by using the selected machine learning model.
13. A method for controlling a self-driving vehicle that is controllable on the basis of an instruction by an utterance of a user, the method comprising:
identifying a use scene of a target user among a plurality of use scenes, wherein each of the use scenes is a state of a user comprising a state before boarding, a state during boarding, or a state after alighting a self-driving vehicle;
acquiring utterance information of the target user;
selecting a different machine learning model from a plurality of machine learning models according to the identified use scene of the target user, wherein each one of the machine learning models relates to a corresponding one of the state before boarding, the state during boarding, or the state after alighting the self-driving vehicle; and
estimating an intent of an utterance of the target user by using the selected machine learning model.