Similarity search of industrial components models
A computer implemented method for improving a similarity search of an industrial component model including obtaining a set of industrial component models, each having associated attributes and a similarity embedding, receiving a similarity request using a given industrial component model as an input, the output of said similarity request being a first subset of industrial component models selected from the set of industrial component models based on the comparison between similarity embeddings and the similarity embedding of the input industrial component model, receiving a second subset of industrial component models from said first subset of industrial component models based on an interchangeability criteria of the input industrial component model with any industrial component model of said second subset of industrial component models, associating a similarity attribute to the input industrial component model, and computing a new set of similarity embeddings.
1 . A computer implemented method for improving a similarity search of an industrial component model, comprising:
obtaining a set of industrial component models, each defining a respective industrial component and each having associated attributes and a similarity embedding that is an embedding of at least a portion of said associated attributes, each respective industrial component being at least one among a mechanical part and an electronic part;
receiving a similarity request using a first industrial component model as an input, an output of said similarity request being a first subset of industrial component models selected from the set of industrial component models based on a comparison between similarity embeddings of the set of industrial component models and a first similarity embedding of the first industrial component model by applying similarity measurement, wherein the first subset of industrial component models includes industrial component models of the set of industrial component models with measured similarity above a given threshold and/or a fixed number of industrial component models of the set of industrial component models with the highest measured similarities;
receiving a second subset of industrial component models selected by a user from said first subset of industrial component models based on an interchangeability criterion of the first industrial component model with any industrial component model of said second subset of industrial component models, the interchangeability criterion being satisfied if the industrial components respectively defined by the industrial component models of the second subset of industrial component models are interchangeable from a design standpoint with the industrial component defined by the first industrial component model;
generating a unique hash code distinct from the first similarity embedding and based on said second subset of industrial component models;
adding said unique hash code as a similarity data to the associated attributes of each industrial component model of said second subset of industrial component models, and
updating the similarity embeddings of the industrial component models of the set of industrial component models based on said similarity data.
2 . The computer implemented method for improving the similarity search of the industrial component model according to claim 1 , wherein the similarity embeddings are embedded by vectorization of the industrial component models attributes, and by embedding of the resulting vectorized data.
3 . The computer implemented method for improving the similarity search of the industrial component model according to claim 2 , wherein the embedding is performed by a context sensitive autoencoder comprising an encoder and a decoder which are both neural networks
wherein
the input of the context sensitive autoencoder is said vectorized data and constitutes the input of the encoder, and the output of the encoder constitutes the similarity embeddings,
the input of the decoder is the similarity embeddings, and the output of the decoder is a term frequency-inverse document frequency of said vectorized data, and
said encoder and decoder are tuned such that the term frequency-inverse document frequency of said vectorized data best approximates said vectorized data.
4 . The computer implemented method for improving the similarity search of the industrial component model according to claim 2 , wherein the embedding is performed by performing a principal component analysis on a concatenation of an L1-normalization of the attributes with a chosen weight multiplied by the L1-normalization of the similarity data.
5 . The computer implemented method for improving the similarity search of the industrial component model according to claim 2 , wherein the vectorization is performed by a doc2vec vectorization of text attributes, and by a vectorization of the similarity data which includes adding a column for each unique hash code, and, for each industrial component model, filling this column with 1 if the industrial component model is associated with this unique hash code, and 0 otherwise.
6 . The computer implemented method for improving the similarity search of the industrial component model according to claim 2 , wherein the vectorization is performed by applying a Bidirectional Encoder Representations from Transformers technique to the industrial component models.
7 . A non-transitory computer readable medium having stored thereon a computer program having instructions for improving a similarity search of an industrial component model that when executed by a computer causes the computer to implement a method comprising:
obtaining a set of industrial component models, each defining a respective industrial component and each having associated attributes and a similarity embedding that is an embedding of at least a portion of said associated attributes, each respective industrial component being at least one among a mechanical part and an electronic part;
receiving a similarity request using a first industrial component model as an input, an output of said similarity request being a first subset of industrial component models selected from the set of industrial component models based on a comparison between similarity embeddings of the set of industrial component models and a first similarity embedding of the first industrial component model by applying similarity measurement, wherein the first subset of industrial component models includes industrial component models of the set of industrial component models with measured similarity above a given threshold and/or a fixed number of industrial component models of the set of industrial component models with the highest measured similarities;
receiving a second subset of industrial component models selected by a user from said first subset of industrial component models based on an interchangeability criterion of the first industrial component model with any industrial component model of said second subset of industrial component models, the interchangeability criterion being satisfied if the industrial components respectively defined by the industrial component models of the second subset of industrial component models are interchangeable from a design standpoint with the industrial component defined by the first industrial component model;
generating a unique hash code distinct from the first similarity embedding and based on said second subset of industrial component models;
adding said unique hash code as a similarity data to the associated attributes of each industrial component model of said second subset of industrial component models; and
updating the similarity embeddings of the industrial component models of the set of industrial component models based on said similarity data.
8 . The non-transitory computer readable medium according to claim 7 , wherein the similarity embeddings are embedded by vectorization of the industrial component models attributes, and by embedding of the resulting vectorized data.
9 . The non-transitory computer readable medium according to claim 8 , wherein the embedding is performed by a context sensitive autoencoder comprising an encoder and a decoder which are both neural networks
wherein
the input of the context sensitive autoencoder is said vectorized data and constitutes the input of the encoder, and the output of the encoder constitutes the similarity embeddings,
the input of the decoder is the similarity embeddings, and the output of the decoder is a term frequency-inverse document frequency of said vectorized data, and
said encoder and decoder are tuned such that the term frequency-inverse document frequency of said vectorized data best approximates said vectorized data.
10 . The non-transitory computer readable medium according to claim 8 , wherein the embedding is performed by performing a principal component analysis on a concatenation of an L1-normalization of the attributes with a chosen weight multiplied by the L1-normalization of the similarity data.
11 . The non-transitory computer readable medium according to claim 8 , wherein the vectorization is performed by a doc2vec vectorization of text attributes, and by a vectorization of the similarity data which includes adding a column for each unique hash code, and, for each industrial component model, filling this column with 1 if the industrial component model is associated with this unique hash code, and 0 otherwise.
12 . The non-transitory computer readable medium according to claim 8 , wherein the vectorization is performed by applying a Bidirectional Encoder Representations from Transformers technique to the industrial component models.
13 . A computer system comprising:
a processor coupled to a memory, the memory having recorded thereon instructions for improving a similarity search of an industrial component model that when executed by the processor causes the processor to be configured to:
obtain a set of industrial component models, each defining a respective industrial component and each having associated attributes and a similarity embedding that is an embedding of at least a portion of said associated attributes, each respective industrial component being at least one among a mechanical part and an electronic part;
receive a similarity request using a first industrial component model as an input, an output of said similarity request being a first subset of industrial component models selected from the set of industrial component models based on a comparison between similarity embeddings of the set of industrial component models and a first similarity embedding of the first industrial component model by applying similarity measurement, wherein the first subset of industrial component models includes industrial component models of the set of industrial component models with measured similarity above a given threshold and/or a fixed number of industrial component models of the set of industrial component models with the highest measured similarities;
receive a second subset of industrial component models selected by a user from said first subset of industrial component models based on an interchangeability criterion of the first industrial component model with any industrial component model of said second subset of industrial component models, the interchangeability criterion being satisfied if the industrial components respectively defined by the industrial component models of the second subset of industrial component models are interchangeable from a design standpoint with the industrial component defined by the first industrial component model;
generate a unique hash code distinct from the first similarity embedding and based on said second subset of industrial component models;
add said unique hash code as a similarity data to the associated attributes of each industrial component model of said second subset of industrial component models, and
update the similarity embeddings of the industrial component models of the set of industrial component models based on said similarity data.
14 . The computer system according to claim 13 , wherein the similarity embeddings are embedded by vectorization of the industrial component models attributes, and by embedding of the resulting vectorized data.
15 . The computer system according to claim 14 , wherein the embedding is performed by a context sensitive autoencoder comprising an encoder and a decoder which are both neural networks
wherein
the input of the context sensitive autoencoder is said vectorized data and constitutes the input of the encoder, and the output of the encoder constitutes the similarity embeddings,
the input of the decoder is the similarity embeddings, and the output of the decoder is a term frequency-inverse document frequency of said vectorized data, and
said encoder and decoder are tuned such that the term frequency-inverse document frequency of said vectorized data best approximates said vectorized data.
16 . The computer system according to claim 14 , wherein the embedding is performed by performing a principal component analysis on a concatenation of an L1-normalization of the attributes with a chosen weight multiplied by the L1-normalization of the similarity data.
17 . The computer system according to claim 14 , wherein the vectorization is performed by a doc2vec vectorization of text attributes, and by a vectorization of the similarity data which includes adding a column for each unique hash code, and, for each industrial component model, filling this column with 1 if the industrial component model is associated with this unique hash code, and 0 otherwise.
18 . The computer system according to claim 14 , wherein the vectorization is performed by applying a Bidirectional Encoder Representations from Transformers technique to the industrial component models.