Learning machine training based on plan types
A system for training a learning machine accesses a training database of reference metadata that describes reference plans that include reference first-type plans and reference second-type plans. Such plans may be travel plans or other plans. The system trains the learning machine to distinguish candidate first-type plans from candidate second-type plans. The training of the learning machine is based on a set of decision trees generated from randomly selected subsets of the reference metadata, and the randomly selected subsets each describe a corresponding randomly selected portion of the reference plans. The system then modifies the trained learning machine based on asymmetrical penalties for incorrectly distinguishing candidate first-type plans from candidate second-type plans. The system then provides the modified learning machine for run-time use in classifying plans.
1 . A computer-implemented method comprising:
accessing, by one or more processors, a training database of reference metadata descriptive of reference travel plans that include reference first-type travel plans and reference second-type travel plans, the reference metadata indicating sizes of destination airports, the indicated sizes respectively corresponding to each reference travel plan among the reference travel plans;
training, by the one or more processors, a learning machine to distinguish candidate first-type travel plans from candidate second-type travel plans, the learning machine being trained based on the indicated sizes of the destination airports and subsets of the reference metadata that is descriptive of the reference travel plans, the subsets each describing a corresponding portion of the reference travel plans that include the reference first-type travel plans and the reference second-type travel plans; and
modifying, by the one or more processors, the trained learning machine based on asymmetrical penalties for incorrectly distinguishing candidate first-type travel plans from candidate second-type travel plans, the asymmetrical penalties including unequal first and second penalties, the first penalty penalizing incorrectly classifying a candidate first-type travel plan being greater than the second penalty penalizing incorrectly classifying a candidate second-type travel plan.
2 . The computer-implemented method of claim 1 , further comprising:
providing, by the one or more processors, the modified learning machine trained to distinguish candidate first-type travel plans from candidate second-type travel plans based on the asymmetrical penalties for incorrectly distinguishing candidate first-type travel plans from candidate second-type travel plans.
3 . The computer-implemented method of claim 1 , wherein the candidate first-type travel plans comprise non-reimbursable non-business travel plans, and the candidate second-type travel plans comprise reimbursable business travel plans.
4 . The computer-implemented method of claim 1 , wherein the learning machine is trained based on decision trees generated from randomly selected subsets of the reference metadata.
5 . The computer-implemented method of claim 1 , wherein:
the reference metadata of the reference travel plans indicates source entities that each reserved a corresponding reference travel plan among the reference travel plans for a corresponding user; and
the training of the learning machine to distinguish candidate first-type travel plans from candidate second-type travel plans is based on the indicated source entities that each reserved a corresponding reference travel plan for a corresponding user.
6 . The computer-implemented method of claim 1 , wherein:
the reference metadata of the reference travel plans indicates ratios of layovers to destination cities, the indicated ratios respectively corresponding to each reference travel plan among the reference travel plans; and
the training of the learning machine to distinguish candidate first-type travel plans from candidate second-type travel plans is based on the indicated ratios of layovers to destination cities.
7 . The computer-implemented method of claim 1 , wherein:
the reference metadata of the reference travel plans indicates total counts of destination cities, the indicated total counts respectively corresponding to each reference travel plan among the reference travel plans; and
the training of the learning machine to distinguish candidate first-type travel plans from candidate second-type travel plans is based on the indicated total counts of destination cities.
8 . The computer-implemented method of claim 1 , wherein:
the reference metadata of the reference travel plans includes indications of whether conventions occurred in destination cities respectively corresponding to each reference travel plan among the reference travel plans; and
the training of the learning machine to distinguish candidate first-type travel plans from candidate second-type travel plans is based on the indications of whether conventions occurred in the destination cities.
9 . A system comprising:
one or more processors; and
a memory storing instructions that, when executed by at least one processor among the one or more processors, cause the system to perform operations comprising:
accessing a training database of reference metadata descriptive of reference travel plans that include reference first-type travel plans and reference second-type travel plans, the reference metadata indicating sizes of destination airports, the indicated sizes respectively corresponding to each reference travel plan among the reference travel plans;
training a learning machine to distinguish candidate first-type travel plans from candidate second-type travel plans, the learning machine being trained based on the indicated sizes of the destination airports and subsets of the reference metadata that is descriptive of the reference travel plans, the subsets each describing a corresponding portion of the reference travel plans that include the reference first-type travel plans and the reference second-type travel plans; and
modifying the trained learning machine based on asymmetrical penalties for incorrectly distinguishing candidate first-type travel plans from candidate second-type travel plans, the asymmetrical penalties including unequal first and second penalties, the first penalty penalizing incorrectly classifying a candidate first-type travel plan being greater than the second penalty penalizing incorrectly classifying a candidate second-type travel plan.
10 . The system of claim 9 , wherein the operations further comprise:
providing the modified learning machine trained to distinguish candidate first-type travel plans from candidate second-type travel plans based on the asymmetrical penalties for incorrectly distinguishing candidate first-type travel plans from candidate second-type travel plans.
11 . The system of claim 9 , wherein the candidate first-type travel plans comprise non-reimbursable non-business travel plans, and the candidate second-type travel plans comprise reimbursable business travel plans.
12 . The system of claim 9 , wherein the learning machine is trained based on decision trees generated from randomly selected subsets of the reference metadata.
13 . The system of claim 9 , wherein:
the reference metadata of the reference travel plans indicates source entities that each reserved a corresponding reference travel plan among the reference travel plans for a corresponding user; and
the training of the learning machine to distinguish candidate first-type travel plans from candidate second-type travel plans is based on the indicated source entities that each reserved a corresponding reference travel plan for a corresponding user.
14 . The system of claim 9 , wherein:
the reference metadata of the reference travel plans indicates ratios of layovers to destination cities, the indicated ratios respectively corresponding to each reference travel plan among the reference travel plans; and
the training of the learning machine to distinguish candidate first-type travel plans from candidate second-type travel plans is based on the indicated ratios of layovers to destination cities.
15 . The system of claim 9 , wherein:
the reference metadata of the reference travel plans indicates total counts of destination cities, the indicated total counts respectively corresponding to each reference travel plan among the reference travel plans; and
the training of the learning machine to distinguish candidate first-type travel plans from candidate second-type travel plans is based on the indicated total counts of destination cities.
16 . The system of claim 9 , wherein:
the reference metadata of the reference travel plans includes indications of whether conventions occurred in destination cities respectively corresponding to each reference travel plan among the reference travel plans; and
the training of the learning machine to distinguish candidate first-type travel plans from candidate second-type travel plans is based on the indications of whether conventions occurred in the destination cities.
17 . A non-transitory machine-readable storage medium comprising instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising:
accessing a training database of reference metadata descriptive of reference travel plans that include reference first-type travel plans and reference second-type travel plans, the reference metadata indicating sizes of destination airports, the indicated sizes respectively corresponding to each reference travel plan among the reference travel plans;
training a learning machine to distinguish candidate first-type travel plans from candidate second-type travel plans, the learning machine being trained based on the indicated sizes of the destination airports and subsets of the reference metadata that is descriptive of the reference travel plans, the subsets each describing a corresponding portion of the reference travel plans that include the reference first-type travel plans and the reference second-type travel plans; and
modifying the trained learning machine based on asymmetrical penalties for incorrectly distinguishing candidate first-type travel plans from candidate second-type travel plans, the asymmetrical penalties including unequal first and second penalties, the first penalty penalizing incorrectly classifying a candidate first-type travel plan being greater than the second penalty penalizing incorrectly classifying a candidate second-type travel plan.
18 . The non-transitory machine-readable storage medium of claim 17 , wherein the operations further comprise:
providing the modified learning machine trained to distinguish candidate first-type travel plans from candidate second-type travel plans based on the asymmetrical penalties for incorrectly distinguishing candidate first-type travel plans from candidate second-type travel plans.