IP Library › Granted Patent US 12,045,570
Granted Patent B2
US 12,045,570 · App. 17/650,825 · Granted Jul 23, 2024

Multi-class text classifier

Inventor: Thomas Allen Wentworth (New York, NY)
Assignee: S&P Global Inc.
G06F40/284G06F40/205G06N5/022G06Q10/0838
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,045,570
App. No.
17/650,825
Granted
Jul 23, 2024
Kind
B2
Abstract

A method, system, and product for multi-class text classification. The method includes receiving an uncoded shipping container description comprising an uncoded text; and assigning a predicted HS code segment to the uncoded container description. Assigning includes cleaning the uncoded text of the uncoded container description by removing 1-2 letter words, numbers and symbols from the text; tokenizing the cleaned uncoded text to define a plurality of unencoded tokens by parsing the cleaned uncoded text into single words and bigrams; summing the scores for each of a plurality of HS code segments associated with each of the plurality of unencoded tokens across all of the plurality of unencoded tokens; and determining the predicted HS code segment based on the highest summation. The method can include receiving an uncoded shipping container description comprising an uncoded text; and assigning a predicted HS code segment to the uncoded container description.

Claims (112)

1. A computer-implemented method for text classification, comprising:

receiving a training set of shipping container descriptions comprising a plurality of shipping container descriptions each of the plurality of shipping container descriptions comprising a Harmonized System (HS) code and a text;

recording the HS code of each of the plurality of shipping container descriptions to define a plurality of HS code labels;

cleaning the text of each of the plurality of shipping container descriptions by removing 1-2 letter words, numbers and symbols from the text of each of the plurality of shipping container descriptions;

tokenizing the cleaned text of each of the plurality of shipping container descriptions to define a plurality of tokens by parsing the cleaned text into single words and bigrams;

accumulating coincidences between each of the plurality of HS code labels and each of the plurality of tokens, with regard to each of the plurality of shipping container descriptions; and

scoring each of the accumulated coincidences to define a score representing correlation between each of the plurality of HS code labels and each of the plurality of tokens based on a total number of occurrences of each of the plurality of tokens and a number of the accumulated coincidences between each of the plurality of HS code labels and each of the plurality of tokens.

2. The computer-implemented method of claim 1 , wherein

the scoring defines the score according to: token i label j score=min(cap,totalCount i /(totalCount i −token i label j Count+c)),

wherein token i label j score is the score, cap is a limit, totalCount i is the total number of occurrences of each of the plurality of tokens, token i label j Count is the number of the accumulated coincidences between each of the plurality of HS code labels and each of the plurality of tokens, i is a first integer, j is a second integer, and c is a constant, where min is the minimum function that returns the smallest of its inputs.

3. The computer-implemented method of claim 2 , wherein cap is 10 min(totalCount i /2, 10/2) .

4. The computer-implemented method of claim 3 , wherein c is 0.1.

5. The computer-implemented method of claim 1 , wherein

the scoring defines the score according to: token i label j score=min(10 min(totalCount i /2, 10/2) ,

totalCount i n /(totalCount i n −token i label j Count n +0.1)),

wherein token i label j score is the score, i is a first integer, j is a second integer, n is a third integer, totalCount i is the total number of occurrences of each of the plurality of tokens, and token i label j Count is the number of the accumulated coincidences between each of the plurality of HS code labels and each of the plurality of tokens, where min is the minimum function that returns the smallest of its inputs.

6. The computer-implemented method of claim 1 , wherein

the scoring defines the score according to: token i label j score=min(10 min(totalCount i /2, 10/2) ,

(totalCount i /(totalCount i −token i label j Count+0.1) n )),

wherein token i label j score is the score, i is a first integer, j is a second integer, n is a third integer, totalCount i is the total number of occurrences of each of the plurality of tokens, and token i label j Count is the number of the accumulated coincidences between each of the plurality of HS code labels and each of the plurality of tokens, where min is the minimum function that returns the smallest of its inputs.

7. The computer-implemented method of claim 1 , wherein

the scoring defines the score according to: token i label j score=token i label j Count/total Count j ,

wherein token i label j score is the score, i is a first integer, j is a second integer, totalCount i is the total number of occurrences of each of the plurality of tokens, totalCount i is the total number of occurrences of each of the plurality of HS code labels, and token i label j Count is the number of the accumulated coincidences between each of the plurality of HS code labels and each of the plurality of tokens.

8. The computer-implemented method of claim 1 , further comprising automatic rolling retraining.

9. The computer-implemented method of claim 1 , further comprising:

receiving an uncoded shipping container description comprising an uncoded text; and

assigning a predicted HS code segment to the uncoded shipping container description comprising:

cleaning the uncoded text of the uncoded shipping container description by removing 1-2 letter words, numbers and symbols from the text;

tokenizing the cleaned uncoded text to define a plurality of unencoded tokens by parsing the cleaned uncoded text into single words and bigrams;

summing the scores for each of a plurality of HS code segments associated with each of the plurality of unencoded tokens across all of the plurality of unencoded tokens; and

determining the predicted HS code segment based on the highest summation.

10. The computer-implemented method of claim 9 , further comprising repeating receiving and assigning using a further HS code segment.

11. The computer-implemented method of claim 9 , wherein the HS code segment is a complete HS code.

12. The computer-implemented method of claim 9 , further comprising searching uncoded descriptions using HS codes.

13. The computer-implemented method of claim 9 , further comprising automatically physically moving a shipping container associated with the uncoded shipping container description.

14. A computer system, comprising:

a hardware processor; and

a text classifier, in communication with the hardware processor, wherein the text classifier is configured:

to receive a training set of shipping container descriptions comprising a plurality of shipping container descriptions each of the plurality of shipping container descriptions comprising a Harmonized System (HS) code and a text;

to record the HS code of each of the plurality of shipping container descriptions to define a plurality of HS code labels;

to clean the text of each of the plurality of shipping container descriptions by removing 1-2 letter words, numbers and symbols from the text of each of the plurality of shipping container descriptions;

to tokenize the cleaned text of each of the plurality of shipping container descriptions to define a plurality of tokens by parsing the cleaned text into single words and bigrams;

to accumulate coincidences between each of the plurality of HS code labels and each of the plurality of tokens, with regard to each of the plurality of shipping container descriptions; and

to score each of the accumulated coincidences to define a score representing correlation between each of the plurality of HS code labels and each of the plurality of tokens based on a total number of occurrences of each of the plurality of tokens and a number of the accumulated coincidences between each of the plurality of HS code labels and each of the plurality of tokens.

15. The computer system of claim 14 , wherein

the scoring defines the score according to: token i label j score=min(cap,totalCount i /(totalCount i −token i label j Count+c)),

wherein token i label j score is the score, cap is a limit, totalCount i is the total number of occurrences of each of the plurality of tokens, token i label j Count is the number of the accumulated coincidences between each of the plurality of HS code labels and each of the plurality of tokens, i is a first integer, j is a second integer, and c is a constant, where min is the minimum function that returns the smallest of its inputs.

16. The computer system of claim 15 , wherein cap is 10 min(totalCount i /2, 10/2) .

17. The computer system of claim 16 , wherein c is 0.1.

18. The computer system of claim 14 , wherein

the scoring defines the score according to: token i label j score=min(10 min(totalCount i /2, 10/2) ,

totalCount i n /(totalCount i n −token i label j Count n +0.1)),

wherein token i label j score is the score, i is a first integer, j is a second integer, n is a third integer, totalCount i is the total number of occurrences of each of the plurality of tokens, and token i label j Count is the number of the accumulated coincidences between each of the plurality of HS code labels and each of the plurality of tokens, where min is the minimum function that returns the smallest of its inputs.

19. The computer system of claim 14 , wherein

the scoring defines the score according to: token i label j score=min(10 min(totalCount i /2, 10/2) ,

(totalCount i /(totalCount i −token i label j Count+0.1) n )),

wherein token i label j score is the score, i is a first integer, j is a second integer, n is a third integer, totalCount i is the total number of occurrences of each of the plurality of tokens, and token i label j Count is the number of the accumulated coincidences between each of the plurality of HS code labels and each of the plurality of tokens, where min is the minimum function that returns the smallest of its inputs.

20. The computer system of claim 14 , wherein

the scoring defines the score according to: token i label j score=token i label j Count/totalCount j ,

wherein token i label j score is the score, i is a first integer, j is a second integer, totalCount i is the total number of occurrences of each of the plurality of tokens, totalCount j is the total number of occurrences of each of the plurality of HS code labels, and token i label j Count is the number of the accumulated coincidences between each of the plurality of HS code labels and each of the plurality of tokens.

21. The computer system of claim 14 , wherein the text classifier is configured to retrain on an automatic rolling basis.

22. The computer system of claim 14 , wherein the text classifier is configured:

to receive an uncoded shipping container description comprising an uncoded text; and

to assign a predicted HS code segment to the uncoded shipping container description comprising:

to clean the uncoded text of the uncoded shipping container description by removing 1-2 letter words, numbers and symbols from the text;

to tokenize the cleaned uncoded text to define a plurality of unencoded tokens by parsing the cleaned uncoded text into single words and bigrams;

to sum the scores for each of a plurality of HS code segments associated with each of the plurality of unencoded tokens across all of the plurality of unencoded tokens; and

to determine the predicted HS code segment based on the highest summation.

23. The computer system of claim 22 , wherein the text classifier is configured:

to repeat receiving and assigning using a further HS code segment.

24. The computer system of claim 22 , wherein the HS code segment is a complete HS code.

25. The computer system of claim 22 , wherein the text classifier is configured:

to search uncoded descriptions using HS codes.

26. The computer system of claim 22 , wherein the text classifier is configured:

to automatically physically move a shipping container associated with the uncoded shipping container description.

27. A computer program product comprising:

a computer readable storage media; and

program code, stored on the computer readable storage media, for classifying text, the program code comprising:

code for receiving a training set of shipping container descriptions comprising a plurality of shipping container descriptions each of the plurality of shipping container descriptions comprising a Harmonized System (HS) code and a text;

code for recording the HS code of each of the plurality of shipping container descriptions to define a plurality of HS code labels;

code for cleaning the text of each of the plurality of shipping container descriptions by removing 1-2 letter words, numbers and symbols from the text of each of the plurality of shipping container descriptions;

code for tokenizing the cleaned text of each of the plurality of shipping container descriptions to define a plurality of tokens by parsing the cleaned text into single words and bigrams;

code for accumulating coincidences between each of the plurality of HS code labels and each of the plurality of tokens, with regard to each of the plurality of shipping container descriptions; and

code for scoring each of the accumulated coincidences to define a score representing correlation between each of the plurality of HS code labels and each of the plurality of tokens based on a total number of occurrences of each of the plurality of tokens and a number of the accumulated coincidences between each of the plurality of HS code labels and each of the plurality of tokens.

28. The computer program product of claim 27 , wherein

the scoring defines the score according to: token i label j score=min(cap,totalCount i /(totalCount i −token i label j Count+c)),

wherein token i label j score is the score, cap is a limit, totalCount i is the total number of occurrences of each of the plurality of tokens, token i label j Count is the number of the accumulated coincidences between each of the plurality of HS code labels and each of the plurality of tokens, i is a first integer, j is a second integer, and c is a constant, where min is the minimum function that returns the smallest of its inputs.

29. The computer program product of claim 28 , wherein cap is 10 min(totalCount i /2, 10/2) .

30. The computer program product of claim 29 , wherein c is 0.1.

31. The computer program product of claim 27 , wherein

the scoring defines the score according to: token i label j score=min(10 min(totalCount i /2, 10/2) ,

totalCount i n /(totalCount i n −token i label j Count n +0.1)),

wherein token i label j score is the score, i is a first integer, j is a second integer, n is a third integer, totalCount i is the total number of occurrences of each of the plurality of tokens, and token i label j Count is the number of the accumulated coincidences between each of the plurality of HS code labels and each of the plurality of tokens, where min is the minimum function that returns the smallest of its inputs.

32. The computer program product of claim 27 , wherein

the scoring defines the score according to: token i label j score=min(10 min(totalCount i /2, 10/2) ,

(total Count i /(totalCount i −token i label j Count+0.1) n )),

wherein token i label j score is the score, i is a first integer, j is a second integer, n is a third integer, totalCount i is the total number of occurrences of each of the plurality of tokens, and token i label j Count is the number of the accumulated coincidences between each of the plurality of HS code labels and each of the plurality of tokens, where min is the minimum function that returns the smallest of its inputs.

33. The computer program product of claim 27 , wherein

the scoring defines the score according to: token i label j score=token i label j Count/totalCount j ,

wherein token i label j score is the score, i is a first integer, j is a second integer, totalCount i is the total number of occurrences of each of the plurality of tokens, totalCount j is the total number of occurrences of each of the plurality of HS code labels, and token i label j Count is the number of the accumulated coincidences between each of the plurality of HS code labels and each of the plurality of tokens.

34. The computer program product of claim 27 , the program code further comprising code for automatic rolling retraining.

35. The computer program product of claim 27 , the program code further comprising:

code for receiving an uncoded shipping container description comprising an uncoded text; and

code for assigning a predicted HS code segment to the uncoded shipping container description comprising:

cleaning the uncoded text of the uncoded shipping container description by removing 1-2 letter words, numbers and symbols from the text;

tokenizing the cleaned uncoded text to define a plurality of unencoded tokens by parsing the cleaned uncoded text into single words and bigrams;

summing the scores for each of a plurality of HS code segments associated with each of the plurality of unencoded tokens across all of the plurality of unencoded tokens; and

determining the predicted HS code segment based on the highest summation.

36. The computer program product of claim 35 , the program code further comprising code to repeat receiving and assigning using a further HS code segment.

37. The computer program product of claim 35 , wherein the HS code segment is a complete HS code.

38. The computer program product of claim 35 , the program code further comprising code to search uncoded descriptions using HS codes.

39. The computer program product of claim 35 , the program code further comprising code to automatically physically move a shipping container associated with the uncoded shipping container description.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 22, 2023
From: WENTWORTH, THOMAS ALLEN
To: S&P GLOBAL INC.
Reel/Frame 064660/0664 →
Continuity (1)
Related Publication 20230259706A1 · Aug 17, 2023