IP Library Granted Patent US 11,256,755
Granted Patent B2
US 11,256,755 · App. 16/807,543 · Granted Feb 22, 2022

Tag mapping process and pluggable framework for generating algorithm ensemble

Inventors: Ian Moore (Burnaby, CA); Massoud Seifi (Burnaby, CA); Alex Clark (Burnaby, CA)
Assignee: General Electric Company
G06F16/90335G06F16/907
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,256,755
App. No.
16/807,543
Granted
Feb 22, 2022
Kind
B2
Abstract

The example embodiments are directed to a system and method for tag mapping. In one example, the method includes receiving a request to perform tag mapping for a target tag of a master data set, the target tag representing a target component of an asset, querying a customer data for a plurality of candidate tag records based on the target tag, tokenizing the plurality of candidate tag records included in the customer data set, reducing an amount of the tokenized tag records in the customer data set based on the target tag and each tokenized candidate tag record, performing tag mapping with the reduced amount of tokenized tag records to identify at least one candidate tag that is a possible match to the target tag, and outputting information concerning the identified at least one matching candidate tag.

Claims (30)

1. A computing system comprising:

a processor configured to

tokenize a plurality of candidate tag records;

reduce the tokenized tag records in the candidate data set based on a target tag and the plurality of tokenized candidate tag records; and

execute a machine learning algorithm which performs automated tag mapping with the reduced tokenized tag records to identify at least one candidate tag that is a match to the target tag.

2. The computing system of claim 1 , wherein each candidate tag record comprises a descriptive identifier of an industrial asset or a part of the industrial asset.

3. The computing system of claim 1 , wherein the processor is configured to break-up a descriptive identifier within a candidate data tag record into a group of tokens.

4. The computing system of claim 1 , wherein the processor is configured to remove one or more of punctuation and digits from a descriptive identifier within a candidate data tag record.

5. The computing system of claim 1 , wherein the processor is configured to generate a unique identifier for a tokenized candidate tag record, and store the unique identifier within the tokenized candidate tag record.

6. The computing system of claim 1 , wherein the processor is configured to reduce a number of the tokenized tag records based on semantic relationships between descriptive identifiers of industrial assets included in the tokenized tag records.

7. The computing system of claim 1 , wherein the processor is configured to execute a first machine learning algorithm on tokenized descriptive identifiers within the candidate tag records to reduce the plurality of tokenized tag records to a subset of tokenized candidate tag records.

8. The computing system of claim 7 , wherein the processor is further configured to execute a second machine learning algorithm on the tokenized descriptive identifiers within the subset of tokenized tag records to reduce the subset of tokenized tag records to a predetermined number of best matching tag records.

9. A method comprising:

tokenizing, via a processing device, a plurality of candidate tag records;

reducing, via the processing device, the tokenized tag records in the candidate data set based on a target tag and the plurality of tokenized candidate tag records; and

executing a machine learning algorithm, via the processing device, which performs automated tag mapping with the reduced tokenized tag records to identify at least one candidate tag that is a match to the target tag.

10. The method of claim 9 , wherein each candidate tag record comprises a descriptive identifier of an industrial asset or a part of the industrial asset.

11. The method of claim 9 , wherein the tokenizing comprises breaking-up a descriptive identifier within a candidate data tag record into a group of tokens.

12. The method of claim 9 , wherein the tokenizing comprises removing one or more of punctuation and digits from a descriptive identifier within a candidate data tag record.

13. The method of claim 9 , wherein the tokenizing comprises generating a unique identifier for a tokenized candidate tag record, and storing the unique identifier within the tokenized candidate tag record.

14. The method of claim 9 , wherein the reducing comprises reducing a number of tokenized tag records based on semantic relationships between descriptive identifiers of industrial assets included in the tokenized tag records.

15. The method of claim 9 , wherein the reducing comprises executing a first machine learning algorithm on tokenized descriptive identifiers within the candidate tag records to reduce the plurality of tokenized tag records to a subset of tokenized candidate tag records.

16. The method of claim 15 , wherein the reducing further comprises executing a second machine learning algorithm on the tokenized descriptive identifiers within the subset of tokenized tag records to reduce the subset of tokenized tag records to a predetermined number of best matching tag records.

17. A non-transitory computer-readable storage medium comprising instructions which when executed by a processor cause a computer to perform a method comprising:

tokenizing, via a processing device, a plurality of candidate tag records;

reducing, via the processing device, the tokenized tag records in the candidate data set based on a target tag and the plurality of tokenized candidate tag records; and

executing a machine learning algorithm, via the processing device, which performs automated tag mapping with the reduced tokenized tag records to identify at least one candidate tag that is a match to the target tag.

18. The non-transitory computer-readable storage medium of claim 17 , wherein each candidate tag record comprises a descriptive identifier of an industrial asset or a part of the industrial asset.

19. The non-transitory computer-readable storage medium of claim 17 , wherein the tokenizing comprises breaking-up a descriptive identifier within a candidate data tag record into a group of tokens.

20. The non-transitory computer-readable storage medium of claim 17 , wherein the tokenizing comprises removing one or more of punctuation and digits from a descriptive identifier within a candidate data tag record.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 17, 2023
From: GENERAL ELECTRIC COMPANY
To: GE DIGITAL HOLDINGS LLC
Reel/Frame 065612/0085 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2020
From: MOORE, IAN; SEIFI, MASSOUD; CLARK, ALEX
To: GENERAL ELECTRIC COMPANY
Reel/Frame 051993/0938 →
Continuity (2)
Continuation 15635828 · Jun 28, 2017
Related Publication 20200201916A1 · Jun 25, 2020