IP Library Granted Patent US 10,628,529
Granted Patent B2
US 10,628,529 · App. 16/283,707 · Granted Apr 21, 2020

Device and method for natural language processing

Inventors: Vitalii Zhelezniak (London, GB); Alexsandar Savkov (London, GB); Francesco Moramarco (London, GB); Jack Flann (London, GB); Nils Hammerla (London, GB)
Assignee: Babylon Partners Limited
G06F17/2785
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,628,529
App. No.
16/283,707
Granted
Apr 21, 2020
Kind
B2
Abstract

Methods for determining whether two sets of words are similar are provided. In one aspect, a method includes receiving a first set of words and a second set of words, whichare subsets of a vocabulary, and each of the first and second sets of words include word embeddings corresponding to each word. The method also includes determining a word membership function for each word in the vocabulary. Determining the word membership includes determining a set of similarity values, each representing the similarity between the word and a respective word in the vocabulary. The method also includes determining a membership function for the first and second sets of words based on the determined word membership functions, and determining a set-based coefficient for the similarity between the first and second sets of words based on the membership function. Systems and devices are also provided.

Claims (60)

1. A computer-implemented method for determining semantic similarity between a first set of words and a second set of words, the method comprising:

receiving the first set of words and the second set of words, wherein the first and second sets of words are subsets of a vocabulary and each of the first and second sets of words comprise word embeddings corresponding to each word, the word embeddings comprising hidden parameters;

determining a word membership function for each word in the first and second sets of words, wherein determining a word membership function for a word comprises determining a set of similarity values, each similarity value representing the semantic similarity between the word and a respective word in the vocabulary;

determining a membership function for the first set of words and a membership function for the second set of words based on the determined word membership functions, wherein determining a membership function for the first and second sets of words comprises, for each of the first and second set of words, determining a set of similiarity values, each similiarity value representing the semantic similiarity between the respective set of words and a respective word in the vocabulary; and

determining a set-based coefficient for the semantic similarity between the first and second sets of words based on the membership function for the first set of words and the membership function for the second set of words.

2. The method of claim 1 wherein determining a membership function for the first set of words and a membership function for the second set of words comprises, for each set of words, determining a fuzzy union between the word membership functions for the respective set of words.

3. The method of claim 2 wherein determining the fuzzy union between the word membership functions for the respective set of words comprises determining the triangular conorm between the word membership functions for the respective set of words.

4. The method of claim 3 wherein the triangular conorm is the maximum triangular conorm and wherein determining the fuzzy union between the word membership functions for the respective set of words comprises determining, for each word in the vocabulary, the maximum semantic similarity value taken from the semantic similarity values for the word relative to each word in the set of words.

5. The method of claim 1 wherein determining the set-based coefficient comprises determining the intersection between the first set of words and the second set of words.

6. The method of claim 1 wherein the set-based coefficient comprises one of a Jaccard similarity coefficient, a cosine similarity coefficient, a Sørensen-Dice similarity index and an overlap coefficient.

7. The method of claim 1 wherein determining the set-based coefficient comprises:

for each word in the vocabulary, determining a maximum semantic similarity value from the determined semantic similarity values for the first set of words relative to the respective word in the vocabulary;

for each word in the vocabulary, determining a maximum semantic similarity value from the determined semantic similarity values for the second set of words relative to the respective word in the vocabulary;

for each word in the vocabulary, determining a highest semantic similarity value taken from the maximum semantic similarity values for the word;

for each word in the vocabulary, determining a lowest semantic similarity taken from the maximum semantic similarity values for the word;

determining an intersection between the first and second sets of words by determining a sum of each of the lowest semantic similarity values;

determining a union between the first and second sets of words by determining a sum of each of the highest semantic similarity values; and

determining the set-based coefficient by dividing the intersection by the union.

8. The method of claim 1 wherein the semantic similarity values are determined based on a dot product membership function or a cosine membership function.

9. The method of claim 8 wherein the dot product membership function μwi between two word embeddings Wi and Wj is one of:

μ w i ( w j )= W i ·W j

or

μ w i ( w j )=α i α j W i ·W j

wherein αi and αj are weights corresponding to Wi and Wj respectively.

10. The method of claim 8 wherein the cosine membership function μwi between two word embeddings Wi and Wj is:

μ

w

i

(

w

j

)

=

cos

(

W

i

,

W

j

)

+

1

2

11. The method of claim 1 wherein the vocabulary consists of the first set of words and the second set of words.

12. A system for determining semantic similarity between a first set of words and a second set of words, the system comprising a processor configured to:

receive the first set of words and the second set of words, wherein the first and second sets of words are subsets of a vocabulary and each of the first and second sets of words comprise word embeddings corresponding to each word, the word embeddings comprising hidden parameters;

determine a word membership function for each word in the first and second sets of words, wherein determining a word membership function for a word comprises determining a set of similarity values, each similarity value representing the semantic similarity between the word and a respective word in the vocabulary;

determine a membership function for the first set of words and a membership function for the second set of words based on the determined word membership functions, wherein determining a membership function for the first and second sets of words comprises, for each of the first and second sets of words, determining a set of similiarity values, each similiarity value representing the semantic similiarity between the respective set of words and a respective word in vocabulary; and

determine a set-based coefficient for the semantic similarity between the first and second sets of words based on the membership function for the first set of words and the membership function for the second set of words.

13. A non-transient computer readable medium comprising instructions that, when executed by a computer, cause the computer to implement the method of claim 1 .

14. A computer implemented method for retrieving content in response to receiving a natural language query, the method comprising:

receiving a natural language query submitted by a user using a user interface;

generating an embedded sentence from said query, the embedded sentence comprising hidden parameters;

determining a semantic similarity between the embedded sentence derived from the received natural language query and embedded sentences from queries saved in a database;

determining a set-based coefficient for the similarity between a first and a second set of words based on a membership function for the first set of words and a membership function for the second set of words, wherein determining a membership function for the first and second sets of words comprises, for each of the first and second sets of words, determining a set of similarity values, each similarity value representing the semantic similarity between the respective set of words and a respective word in the vocabulary;

retrieving a response for an embedded sentence determined to be semantically similar to one of the saved queries; and

providing the response to the user via the user interface.

Assignments (4)
CHANGE OF NAME Recorded Aug 13, 2025
From: EMED POPULATION HEALTH, LLC
To: EMED POPULATION HEALTH, INC.
Reel/Frame 072434/0946 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2025
From: EMED HEALTHCARE UK, LIMITED
To: EMED POPULATION HEALTH, LLC
Reel/Frame 071207/0882 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2023
From: BABYLON PARTNERS LIMITED
To: EMED HEALTHCARE UK, LIMITED
Reel/Frame 065597/0640 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 1, 2019
From: ZHELEZNIAK, VITALII; SAVKOV, ALEKSANDAR; MORAMARCO, FRANCESCO; FLANN, JACK; HAMMERLA, NILS
To: BABYLON PARTNERS LIMITED
Reel/Frame 048759/0823 →
Priority Claims (1)
GB 1808056.4 · May 17, 2018 · national
Continuity (1)
Related Publication 20190354588A1 · Nov 21, 2019