IP Library Granted Patent US 11,461,668
Granted Patent B1
US 11,461,668 · App. 16/565,268 · Granted Oct 4, 2022

Recognizing entities based on word embeddings

Inventor: Gokhuldass Mohandas (Palo Alto, CA)
Assignee: Ciitizen, LLC
G06N5/02G06F40/295G06N7/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,461,668
App. No.
16/565,268
Granted
Oct 4, 2022
Kind
B1
Abstract

Some embodiments provide a non-transitory machine-readable medium that stores a program. The program receives a set of words. The program further retrieves an entry from a knowledge base comprising a plurality of entries. Each entry includes a text description of a concept. The program also determines an embedding for the entry based on the text description of the concept. The program further iteratively determines an embedding for a word in the set of words, increasing a size of a window of words in the set of words, and calculating a confidence score for the entry with respect to the word based on the embedding for the entry and the embedding for the word until a successive calculated confidence score decreases below a previous calculated confidence score. The program also determines that a window of words in the set of words having a previous size represents an entity.

Claims (49)

1. A non-transitory machine-readable medium storing a program executable by at least one processing unit of a device, the program comprising sets of instructions for:

receiving a set of words;

retrieving an entry from a knowledge base comprising a plurality of entries, each entry comprising a text description of a concept;

determining an embedding for the entry based on the text description of the concept;

iteratively determining an embedding for a word in the set of words, increasing a size of a window of words in the set of words, and calculating a confidence score for the entry with respect to the word based on the embedding for the entry and the embedding for words in the window of words until a successive calculated confidence score decreases below a previous calculated confidence score; and

determining that a window of words in the set of words having a previous size represents an entity.

2. The non-transitory machine-readable medium of claim 1 , wherein determining the embedding for the entry based on the text description of the concept comprises determining an embedding for each word in a set of words in the description of the concept, wherein the program further comprises a set of instructions for generating an embedding for the entry based on the determined embeddings for each word in the set of words.

3. The non-transitory machine-readable medium of claim 2 , wherein generating the embedding for the entry comprises:

calculating an average of the determined embeddings for each word in the set of words; and

using the average as the embedding for the entry.

4. The non-transitory machine-readable medium of claim 1 , wherein the program further comprises a set of instructions for, before iteratively determining an embedding for a word in the set of words, increasing a size of a window of words in the set of words, and calculating a confidence score for the entry with respect to the word based on the embedding for the entry and the embedding for the words in the window of words, removing words from the set of words based on a list of stop words.

5. The non-transitory machine-readable medium of claim 1 , wherein the previous calculated confidence score is calculated for the embedding for the window of words in the set of words having the previous size.

6. The non-transitory machine-readable medium of claim 1 , wherein the program further comprises sets of instructions for:

setting the size of the window of words to a default size; and

resetting the size of the window of words to the default size when a particular calculated confidence score for a particular word is less than a defined threshold score.

7. The non-transitory machine-readable medium of claim 1 , wherein the knowledge base is a medical terminology knowledge base, wherein each entry in the knowledge base further comprises a unique identifier associated with the concept described by the text description.

8. A method comprising:

receiving a set of words;

retrieving an entry from a knowledge base comprising a plurality of entries, each entry comprising a text description of a concept;

determining an embedding for the entry based on the text description of the concept;

iteratively determining an embedding for a word in the set of words, increasing a size of a window of words in the set of words, and calculating a confidence score for the entry with respect to the word based on the embedding for the entry and the embedding for words in the window of words until a successive calculated confidence score decreases below a previous calculated confidence score; and

determining that a window of words in the set of words having a previous size represents an entity.

9. The method of claim 8 , wherein determining the embedding for the entry based on the text description of the concept comprises determining an embedding for each word in a set of words in the description of the concept, wherein the method further comprises generating an embedding for the entry based on the determined embeddings for each word in the set of words.

10. The method of claim 9 , wherein generating the embedding for the entry comprises:

calculating an average of the determined embeddings for each word in the set of words; and

using the average as the embedding for the entry.

11. The method of claim 8 further comprising, before iteratively determining an embedding for a word in the set of words, increasing a size of a window of words in the set of words, and calculating a confidence score for the entry with respect to the word based on the embedding for the entry and the embedding for the words in the window of words, removing words from the set of words based on a list of stop words.

12. The method of claim 8 , wherein the previous calculated confidence score is calculated for the embedding for the window of words in the set of words having the previous size.

13. The method of claim 8 further comprising:

setting the size of the window of words to a default size; and

resetting the size of the window of words to the default size when a particular calculated confidence score for a particular word is less than a defined threshold score.

14. The method of claim 8 , wherein the knowledge base is a medical terminology knowledge base, wherein each entry in the knowledge base further comprises a unique identifier associated with the concept described by the text description.

15. A system comprising:

a set of processing units; and

a non-transitory machine-readable medium storing instructions that when executed by at least one processing unit in the set of processing units cause the at least one processing unit to:

receive a set of words;

retrieve an entry from a knowledge base comprising a plurality of entries, each entry comprising a text description of a concept;

determine an embedding for the entry based on the text description of the concept;

iteratively determine an embedding for a word in the set of words, increasing a size of a window of words in the set of words, and calculating a confidence score for the entry with respect to the word based on the embedding for the entry and the embedding for words in the window of words until a successive calculated confidence score decreases below a previous calculated confidence score; and

determine that a window of words in the set of words having a previous size represents an entity.

16. The system of claim 15 , wherein determining the embedding for the entry based on the text description of the concept comprises determining an embedding for each word in a set of words in the description of the concept, wherein the instructions further cause the at least one processing unit to generate an embedding for the entry based on the determined embeddings for each word in the set of words.

17. The system of claim 16 , wherein generating the embedding for the entry comprises:

calculating an average of the determined embeddings for each word in the set of words; and

using the average as the embedding for the entry.

18. The system of claim 15 , wherein the instructions further cause the at least one processing unit to, before iteratively determining an embedding for a word in the set of words, increasing a size of a window of words in the set of words, and calculating a confidence score for the entry with respect to the word based on the embedding for the entry and the embedding for the words in the window of words, remove words from the set of words based on a list of stop words.

19. The system of claim 15 , wherein the previous calculated confidence score is calculated for the embedding for the window of words in the set of words having the previous size.

20. The system of claim 15 , wherein the instructions further cause the at least one processing unit to:

set the size of the window of words to a default size; and

reset the size of the window of words to the default size when a particular calculated confidence score for a particular word is less than a defined threshold score.

Assignments (7)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2024
From: INVITAE CORPORATION; CIITIZEN, LLC
To: CITIZEN HEALTH, INC.
Reel/Frame 066087/0060 →
RELEASE OF SECURITY INTEREST AT R/F 63787/0148 Recorded Dec 14, 2023
From: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION
To: CIITIZEN, LLC
Reel/Frame 066017/0791 →
SECURITY INTEREST Recorded Mar 7, 2023
From: CIITIZEN, LLC
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 062907/0924 →
RELEASE OF SECURITY INTEREST Recorded Mar 2, 2023
From: PERCEPTIVE CREDIT HOLDINGS III, LP
To: CIITIZEN, LLC
Reel/Frame 062861/0976 →
MERGER AND CHANGE OF NAME Recorded Oct 22, 2021
From: CIITIZEN CORPORATION; CAYMAN MERGER SUB B LLC
To: CIITIZEN, LLC
Reel/Frame 057881/0810 →
SECURITY INTEREST Recorded Oct 22, 2021
From: CIITIZEN, LLC
To: PERCEPTIVE CREDIT HOLDINGS III, LP
Reel/Frame 057877/0241 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2019
From: MOHANDAS, GOKHULDASS
To: CIITIZEN CORP.
Reel/Frame 050880/0169 →
Cited By (3)
US 12,205,026 US 12,505,281 US 12,579,182