METHODS AND SYSTEMS FOR COMBINING VEHICLE DATA
Methods and systems are provided for automatically comparing, combining and fusing vehicle data. First data is obtained pertaining to a first plurality of vehicles. Second data is obtained pertaining to a second plurality of vehicles. The first data and the second data are compared and combined based on syntactic similarity between respective data elements of the first data and the second data collected during different stages of vehicle life cycle development.
1 . A method comprising:
obtaining first data comprising data elements pertaining to a first plurality of vehicles;
obtaining second data comprising data elements pertaining to a second plurality of vehicles; and
combining the first data and the second data, via a processor, based on syntactic similarity between respective data elements of the first data and the second data.
2 . The method of claim 1 , wherein the first data and the second data are obtained from different sources.
3 . The method of claim 1 , wherein:
the first data comprises design failure mode and effects analysis (DFMEA) data that is generated using vehicle warranty claims; and
the second data comprises vehicle field data.
4 . The method of claim 1 , wherein the step of combining the first data and the second data comprises:
calculating, via the processor, a measure of syntactic similarity pertaining to respective data elements of the first data and the second data; and
determining, via the processor, that the respective data elements of the first data and the second data are related to one another based on the calculated measure of the syntactic similarity.
5 . The method of claim 4 , wherein the step of calculating the measure of the syntactic similarity comprises calculating, via the processor, the measure of syntactic similarity between terms associated with vehicle symptoms derived from the respective data elements of the first data and the second data.
6 . The method of claim 4 , wherein:
the step of calculating the measure of the syntactic similarity comprises calculating, via the processor, a Jaccard Distance between terms derived from the respective data elements of the first data and the second data; and
the step of determining that the respective data elements are related comprises determining, via the processor, that the respective data elements of the first data and the second data are related if the Jaccard Distance exceeds a predetermined threshold.
7 . The method of claim 6 , wherein the step of determining that the respective data elements are related comprises:
determining, via the processor, that the respective data elements of the first data and the second data are synonymous if the Jaccard Distance exceeds the predetermined threshold.
8 . The method of claim 6 , wherein:
the respective data elements of the first data and the second data comprise strings representing vehicle parts, vehicle systems, and vehicle actions; and
the step of calculating the Jaccard Distance comprises calculating, via the processor, the Jaccard Distance between the respective strings of the respective data elements of the first data and the second data.
9 . A program product comprising:
a program configured to at least facilitate:
obtaining first data comprising data elements pertaining to a first plurality of vehicles;
obtaining second data comprising data elements pertaining to a second plurality of vehicles; and
combining the first data and the second data based on syntactic similarity between respective data elements of the first data and the second data; and
a non-transitory, computer readable storage medium storing the program.
10 . The program product of claim 9 , wherein
the first data comprises design failure mode and effects analysis (DFMEA) data that is generated using vehicle warranty claims; and
the second data comprises vehicle field data.
11 . The program product of claim 9 , wherein the program is further configured to at least facilitate:
calculating a measure of syntactic similarity between respective data elements of the first data and the second data; and
determining that the respective data elements of the first data and the second data are related to one another based on the calculated measure of the syntactic similarity.
12 . The program product of claim 11 , wherein the program is further configured to at least facilitate:
calculating a Jaccard Distance between respective data elements of the first data and the second data; and
determining that the respective data elements of the first data and the second data are related if the Jaccard Distance exceeds a predetermined threshold.
13 . The program product of claim 12 , wherein the program is further configured to at least facilitate determining that the respective data elements of the first data and the second data are synonymous if the Jaccard Distance exceeds the predetermined threshold.
14 . The program product of claim 12 wherein:
the respective data elements of the first data and the second data comprise strings representing vehicle parts, vehicle systems, and vehicle actions; and
the program is further configured to at least facilitate calculating the Jaccard Distance between the respective strings of the respective data elements of the first data and the second data.
15 . A system comprising:
a memory storing:
first data comprising data elements pertaining to a first plurality of vehicles;
second data comprising data elements pertaining to a second plurality of vehicles; and
a processor coupled to the memory and configured to combine the first data and the second data based on syntactic similarity between respective data elements of the first data and the second data.
16 . The system of claim 15 , wherein
the first data comprises design failure mode and effects analysis (DFMEA) data that is generated using vehicle warranty claims; and
the second data comprises vehicle field data.
17 . The system of claim 15 , wherein the processor is further configured to:
calculate a measure of syntactic similarity between respective data elements of the first data and the second data; and
determine that the respective data elements of the first data and the second data are related to one another based on the calculated measure of the syntactic similarity.
18 . The system of claim 17 , wherein the processor is further configured to:
calculate a Jaccard Distance between respective data elements of the first data and the second data; and
determine that the respective data elements of the first data and the second data are related if the Jaccard Distance exceeds a predetermined threshold.
19 . The system of claim 18 , wherein the processor is further configured to determine that the respective data elements of the first data and the second data are synonymous if the Jaccard Distance exceeds the predetermined threshold.
20 . The system of claim 18 , wherein:
the respective data elements of the first data and the second data comprise strings representing vehicle parts, vehicle systems, and vehicle actions; and
the processor is further configured to calculate the Jaccard Distance between the respective strings of the respective data elements of the first data and the second data.