IP Library Granted Patent US 11,645,447
Granted Patent B2
US 11,645,447 · App. 16/846,756 · Granted May 9, 2023

Encoding textual information for text analysis

Inventors: Halid Yerebakan (Malvern, PA); Yoshihisa Shinagawa (Downingtown, PA)
Assignee: Siemens Healthcare GmbH
G06F40/126G06F40/205G06F40/284G06F40/58
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,645,447
App. No.
16/846,756
Granted
May 9, 2023
Kind
B2
Abstract

A computer-implemented method of encoding a word for use in a method of text analysis comprises receiving input text to be analysed, the input text comprising a first word which is not represented in a vocabulary set stored on a storage. The vocabulary set comprises a plurality of words and an associated word embedding vector for each word in the set. The method comprises identifying the first word as a word which is not represented in the vocabulary set and determining one or more sub-words within the first word with which to encode the first word. Each of the one or more sub-words corresponds with a word represented in the vocabulary set and having an embedding vector in the vocabulary set. The method comprises determining an encoding for the first word based on the one or more sub-words.

Claims (40)

1. A computer-implemented method of encoding a word for use in a method of text analysis, the method comprising:

receiving input text to be analysed, the input text comprising a first word which is not represented in a vocabulary set stored on a storage, the vocabulary set comprising a plurality of words and an associated word embedding vector for each word in the vocabulary set;

identifying, the first word as a word which is not represented in the vocabulary set;

selecting one of different segmentations of the first word with which to encode the first word, wherein the different segmentations are different from each other, wherein each of the different segmentations comprises one or more sub-words, wherein a combination of all the one or more sub-words in the segmentation forms the first word, each of the one or more sub-words corresponding with a word represented in the vocabulary set and having an embedding vector in the vocabulary set; and

determining an encoding for the first word based on the one or more sub-words of the selected segmentation.

2. The method according to claim 1 , wherein selecting one of the different segmentations of the first word to obtain the one or more sub-words comprises:

selecting the segmentation of the first word based on pre-determined criteria relating to the different segmentations.

3. The method according to claim 2 , wherein a pre-determined criterion based upon which the segmentation of the first word is selected is a frequency of the sub-words produced by each of the different segmentations.

4. The method according to claim 3 , wherein selecting one of the different segmentations of the first word comprises

determining a scoring value for each of the different segmentations based on a number of sub-words and the frequency associated with the sub-words produced by each segmentation and

selecting the segmentation based on a comparison of the scoring values.

5. The method according to claim 4 wherein the scoring values are determined based on a weighted sum of a first term relating to the number of sub-words produced by each segmentation and a second term relating to the frequency of the sub-words produced by each segmentation.

6. The method according to claim 1 , wherein selecting one of the different segmentations of the first word comprises selecting a segmentation comprising a lowest number of sub-words.

7. The method according to claim 1 , wherein the vocabulary set comprises individual characters of a natural language of the input text and has embedding vectors associated with the individual characters, and wherein the first word may be segmented to provide one or more sub-words comprising a plurality of characters and one or more sub-words comprising an individual character.

8. The method according to claim 1 , wherein the vocabulary set comprises fewer than 500,000 words.

9. The method according to claim 8 wherein the vocabulary set is a Fasttext 100k vocabulary set or a Glove vocabulary set.

10. A system for applying a computer implemented text analysis process to natural language text to obtain a text analysis result, comprising:

a non-transitory memory device for storing computer readable program code;

a storage device including a vocabulary set stored on the storage device, the vocabulary set comprising a plurality of words and an associated word embedding vector for each word in the vocabulary set; and

a processor device in communication with the non-transitory memory device and the storage device, the processor device being operative with the computer readable program code to perform steps including:

receiving input text to be analysed, the input text comprising a first word which is not represented in the vocabulary set;

identifying the first word as a word which is not represented in the vocabulary set;

selecting one of different segmentations of the first word with which to encode the first word, wherein the different segmentations are different from each other, wherein each of the different segmentations comprises one or more sub-words, wherein a combination of all the one or more sub-words in the segmentation forms the first word, each of the one or more sub-words corresponding with a word represented in the vocabulary set and having an embedding vector in the vocabulary set;

determining an encoding for the first word based on the one or more sub-words;

identifying and encoding each word in the natural language text which does not correspond with a word in the vocabulary set to obtain an encoding for each word;

determining for each encoding an embedding vector; and

determining, based on the embedding vectors, a text analysis result.

11. The system according to claim 10 , wherein the text analysis result is a text classification model and the step of determining based on the embedding vectors further comprises obtaining a classification of the natural language text.

12. The system according to claim 10 , wherein the text analysis result is a translation model and the step of determining based on the embedding vectors further comprises obtaining a translation of the natural language text from a first natural language into a second natural language.

13. The system according to claim 10 , wherein selecting one of the different segmentations of the first word comprises:

selecting the segmentation of the first word based on pre-determined criteria relating to the different segmentations.

14. The system according to claim 13 , wherein a pre-determined criterion based upon which one of the different segmentations of the first word is selected relates to a frequency associated with the one or more sub-words produced by each of the different segmentations.

15. The system according to claim 10 , wherein selecting one of the different segmentations of the first word comprises selecting a segmentation comprising a lowest number of sub-words.

16. The system according to claim 10 , wherein selecting one of the different segmentations of the first word comprises:

determining a scoring value for each of the different segmentations based on number of sub-words and frequency associated with the sub-words produced by each segmentation; and

selecting the segmentation based on a comparison of the scoring values.

17. The system according to claim 16 wherein the scoring values are determined based on a weighted sum of a first term relating to the number of sub-words produced by each segmentation and a second term relating to the frequency of the sub-words produced by each segmentation.

18. The system according to claim 10 , wherein the vocabulary set comprises individual characters of a natural language of the input text and has embedding vectors associated with the individual characters, and wherein the first word may be segmented to provide one or more sub-words comprising a plurality of characters and one or more sub-words comprising an individual character.

19. The system according to claim 10 , wherein the vocabulary set comprises fewer than 500,000 words.

20. The system according to claim 10 wherein the vocabulary set is a Fasttext 100k vocabulary set or a Glove vocabulary set.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2023
From: SIEMENS HEALTHCARE GMBH
To: SIEMENS HEALTHINEERS AG
Reel/Frame 066267/0346 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2020
From: SIEMENS MEDICAL SOLUTIONS USA, INC.
To: SIEMENS HEALTHCARE GMBH
Reel/Frame 054139/0517 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 18, 2020
From: YEREBAKAN, HALID; SHINAGAWA, YOSHIHISA
To: SIEMENS MEDICAL SOLUTIONS USA, INC.
Reel/Frame 053811/0571 →
Priority Claims (1)
WO 19170335 · Apr 18, 2019 · international
Continuity (1)
Related Publication 20200334410A1 · Oct 22, 2020
Cited By (2)
US 12,547,834 US 12,675,634