IP Library Granted Patent US 10,740,561
Granted Patent B1
US 10,740,561 · App. 16/670,533 · Granted Aug 11, 2020

Identifying entities in electronic medical records

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,740,561
App. No.
16/670,533
Granted
Aug 11, 2020
Kind
B1
Abstract

Disclosed herein are methods, systems, and apparatus, including computer programs encoded on computer storage media, for entity prediction. One of the methods includes performing word segmentation on text to be predicted to obtain a plurality of words. For each word of the plurality of words, a determination is made whether the word has a pre-trained word vector. In response to determining that the word has a pre-trained word vector, the pre-trained word vector for the word is obtained. In response to determining that the word does not have a pre-trained word vector, a word vector for the word is determined based on a pre-trained stroke vector. The word vector and the pre-trained stroke vector are trained based on a text sample and a word vector model. An entity associated with the text is predicted by inputting word vectors of the plurality of words into an entity prediction model.

Claims (49)

1. A computer-implemented method for entity prediction, the method comprising:

segmenting a text written in a stroke-based writing system to obtain a plurality of words in the text;

for each word of the plurality of words obtained from the text written in a stroke-based writing system,

determining whether the word is associated with a pre-trained word vector,

if the word is associated with a pre-trained word vector, obtaining the pre-trained word vector for the word,

if the word is not associated with a pre-trained word vector,

disassembling the word into constituent strokes,

forming a plurality of n-strokes based on the constituent strokes,

for each n-stroke of the plurality of n-strokes, obtaining a corresponding n-stroke vector from an n-stroke table output by a pre-trained model, and

based on the plurality of corresponding n-stroke vectors, determining a word vector for the word; and

predicting an entity associated with the text by inputting word vectors of the plurality of words into an entity prediction model.

2. The computer-implemented method of claim 1 , wherein the pre-trained model comprises a cw2vec model.

3. The computer-implemented method of claim 1 , wherein the entity prediction model comprises a bi-directional long short term memory, conditional random field (BiLSTM-CRF) model.

4. The computer-implemented method of claim 1 , wherein the text comprises a plurality of Chinese characters, and wherein each n-stroke comprises one or more Chinese language strokes.

5. The computer-implemented method of claim 1 , wherein the pre-trained model is trained on a medical record text sample, the text comprises a medical record text, and the entity comprises a medically-related entity.

6. The computer-implemented method of claim 5 , wherein the medically-related entity comprises one or more of an anatomical term, a medical condition, a medical procedure, a medical staff name, a provider name, a diagnose, and a medication name.

7. A non-transitory, computer-readable storage medium storing one or more instructions executable by a computer system to perform operations comprising:

segmenting a text written in a stroke-based writing system to obtain a plurality of words in the text;

for each word of the plurality of words obtained from the text written in a stroke-based writing system,

determining whether the word is associated with a pre-trained word vector,

if the word is associated with a pre-trained word vector, obtaining the pre-trained word vector for the word,

if the word is not associated with a pre-trained word vector,

disassembling the word into constituent strokes,

forming a plurality of n-strokes based on the constituent strokes,

for each n-stroke of the plurality of n-strokes, obtaining a corresponding n-stroke vector from an n-stroke table output by a pre-trained model, and

based on the plurality of corresponding n-stroke vectors, determining a word vector for the word; and

predicting an entity associated with the text to be predicted by inputting word vectors of the plurality of words into an entity prediction model.

8. The non-transitory, computer-readable storage medium of claim 7 , wherein the pre-trained model comprises a cw2vec model.

9. The non-transitory, computer-readable storage medium of claim 7 , wherein the entity prediction model comprises a bi-directional long short term memory, conditional random field (BiLSTM-CRF) model.

10. The non-transitory, computer-readable storage medium of claim 7 , wherein the text comprises a plurality of Chinese characters, and wherein each n-stroke comprises one or more Chinese language strokes.

11. The non-transitory, computer-readable storage medium of claim 7 , wherein the pre-trained model is trained on a medical record text sample, the text comprises a medical record text, and the entity comprises a medically-related entity.

12. The non-transitory, computer-readable storage medium of claim 11 , wherein the medically-related entity comprises one or more of an anatomical term, a medical condition, a medical procedure, a medical staff name, a provider name, a diagnose, and a medication name.

13. A computer-implemented system, comprising:

one or more computers; and

one or more computer memory devices interoperably coupled with the one or more computers and having tangible, non-transitory, machine-readable media storing one or more instructions that, when executed by the one or more computers, perform one or more operations comprising:

segmenting a text written in a stroke-based writing system to obtain a plurality of words in the text;

for each word of the plurality of words obtained from the text written in a stroke-based writing system,

determining whether the word is associated with a pre-trained word vector, if the word is associated with a pre-trained word vector, obtaining the pre-trained word vector for the word,

if the word is not associated with a pre-trained word vector,

disassembling the word into constituent strokes,

forming a plurality of n-strokes based on the constituent strokes,

for each n-stroke of the plurality of n-strokes, obtaining a corresponding n-stroke vector from an n-stroke table output by a pre-trained model, and

based on the plurality of corresponding n-stroke vectors, determining a word vector for the word; and

predicting an entity associated with the text by inputting word vectors of the plurality of words into an entity prediction model.

14. The system of claim 13 , wherein the pre-trained model comprises a cw2vec model.

15. The system of claim 13 , wherein the entity prediction model comprises a bi-directional long short term memory, conditional random field (BiLSTM-CRF) model.

16. The system of claim 13 , wherein the text comprises a plurality of Chinese characters, and wherein each n-stroke comprises one or more Chinese language strokes.

17. The system of claim 13 , wherein the pre-trained model is trained on a medical record text sample, the text comprises a medical record text, and the entity comprises a medically-related entity.

18. The system of claim 17 , wherein the medically-related entity comprises one or more of an anatomical term, a medical condition, a medical procedure, a medical staff name, a provider name, a diagnose, and a medication name.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 10, 2020
From: ADVANTAGEOUS NEW TECHNOLOGIES CO., LTD.
To: ADVANCED NEW TECHNOLOGIES CO., LTD.
Reel/Frame 053754/0625 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 31, 2020
From: ALIBABA GROUP HOLDING LIMITED
To: ADVANTAGEOUS NEW TECHNOLOGIES CO., LTD.
Reel/Frame 053743/0464 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 14, 2020
From: CAO, SHAOSHENG; ZHOU, JUN
To: ALIBABA GROUP HOLDING LIMITED
Reel/Frame 051503/0015 →