IP Library › Granted Patent US 11,853,702
Granted Patent B2
US 11,853,702 · App. 17/161,778 · Granted Dec 26, 2023

Self-supervised semantic shift detection and alignment

Inventors: Pin-Yu Chen (White Plains, NY); Maurício Gruppi (Troy, NY); Sibel Adali (Slingerlands, NY)
Assignees: International Business Machines Corporation; RENSSELAER POLYTECHNIC INSTITUTE
G06F40/30G06F17/16G06N3/04G06N5/01G06N20/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,853,702
App. No.
17/161,778
Granted
Dec 26, 2023
Kind
B2
Abstract

Generate, for each of the words of a common vocabulary of first and second text corpora, a first word embedding vector in the first text corpus and a second word embedding vector in the second text corpus. Generate, for each word in a random sample of non-landmark words, an artificially shifted word embedding vector by modifying the first word embedding vector for that word. Train a machine learning classifier to predict whether an artificial shift has been injected for a given word, based on the artificially shifted word embedding vector and the second word embedding vector for the given word. Predict semantic shifts for at least a plurality of the words of the common vocabulary by providing the first word embedding vectors and the second word embedding vectors for at least the plurality of the words of the common vocabulary as input to the trained machine learning classifier.

Claims (49)

1. A method comprising:

obtaining first and second text corpora;

identifying a common vocabulary of the two text corpora;

identifying a plurality of landmark words and a plurality of non-landmark words in the common vocabulary;

generating, for each of the words of the common vocabulary, a first word embedding vector in the first text corpus and a second word embedding vector in the second text corpus;

generating, for each word in a random sample of the non-landmark words, an artificially shifted word embedding vector by modifying the first word embedding vector for that word;

training a machine learning classifier to predict whether an artificial shift has been injected for a given word, based on the artificially shifted word embedding vector and the second word embedding vector for the given word; and

predicting semantic shifts for at least a plurality of the words of the common vocabulary by providing the first word embedding vectors and the second word embedding vectors for at least the plurality of the words of the common vocabulary as input to the trained machine learning classifier.

2. The method of claim 1 , wherein modifying the first word embedding vector comprises adding to the first word embedding vector a multiple of the second word embedding vector, wherein the multiple is less than 1.

3. The method of claim 1 , wherein generating the first word embedding vector and the second word embedding vector for a given word comprises counting co-occurrences of other words within a predetermined number of words from each occurrence of the given word in the respective corpus.

4. The method of claim 1 , wherein the machine learning classifier is a neural network.

5. The method of claim 1 , wherein the machine learning classifier is a decision tree.

6. The method of claim 1 , wherein the machine learning classifier is a support vector machine.

7. The method of claim 1 , further comprising: producing a list of words for which semantic shifts are predicted, and reporting the list of words to a user.

8. The method of claim 1 , further comprising, during computerized natural language processing of a body of text, improving the natural language processing by, for at least one word for which one of said semantic shifts is predicted, replacing the at least one word with a synonym for the unshifted meaning.

9. The method of claim 1 , further comprising:

updating the plurality of landmark words to include words for which a semantic shift is not predicted; and

aligning the word embedding vectors of the first and second text corpora based on the plurality of landmark words.

10. A computer program product comprising one or more computer readable storage media that embody computer executable instructions, which when executed by a computer cause the computer to perform a method comprising:

obtaining first and second text corpora;

identifying a common vocabulary of the two text corpora;

identifying a plurality of landmark words and a plurality of non-landmark words in the common vocabulary;

generating, for each of the words of the common vocabulary, a first word embedding vector in the first text corpus and a second word embedding vector in the second text corpus;

generating, for each word in a random sample of the non-landmark words, an artificially shifted word embedding vector by modifying the first word embedding vector for that word;

training a machine learning classifier to predict whether an artificial shift has been injected for a given word, based on the artificially shifted word embedding vector and the second word embedding vector for the given word; and

predicting semantic shifts for at least a plurality of the words of the common vocabulary by providing the first word embedding vectors and the second word embedding vectors for at least the plurality of the words of the common vocabulary as input to the trained machine learning classifier.

11. The computer-readable medium of claim 10 , wherein modifying the first word embedding vector comprises adding to the first word embedding vector a multiple of the second word embedding vector, wherein the multiple is less than 1.

12. The computer-readable medium of claim 10 , wherein generating the first word embedding vector and the second word embedding vector for a given word comprises counting co-occurrences of other words within a predetermined number of words from each occurrence of the given word in the respective corpus.

13. The computer-readable medium of claim 10 , wherein the machine learning classifier is a neural network.

14. The computer-readable medium of claim 10 , wherein the machine learning classifier is a decision tree.

15. The computer-readable medium of claim 10 , wherein the machine learning classifier is a support vector machine.

16. The computer-readable medium of claim 10 , wherein the method further comprises: producing a list of words for which semantic shifts are predicted, and reporting the list of words to a user.

17. The computer-readable medium of claim 10 , wherein the method further comprises: for at least one word for which a semantic shift is predicted, replacing the at least one word with a synonym for the unshifted meaning.

18. The computer-readable medium of claim 10 , wherein the method further comprises:

updating the plurality of landmark words to include words for which a semantic shift is not predicted; and

aligning the word embedding vectors of the first and second text corpora based on the plurality of landmark words.

19. An apparatus comprising:

a memory embodying computer-executable instructions; and

at least one processor, coupled to the memory, and operative by the computer-executable instructions to facilitate a method comprising:

obtaining first and second text corpora;

identifying a common vocabulary of the two text corpora;

identifying a plurality of landmark words and a plurality of non-landmark words in the common vocabulary;

generating, for each of the words of the common vocabulary, a first word embedding vector in the first text corpus and a second word embedding vector in the second text corpus;

generating, for each word in a random sample of the non-landmark words, an artificially shifted word embedding vector by modifying the first word embedding vector for that word;

training a machine learning classifier to predict whether an artificial shift has been injected for a given word, based on the artificially shifted word embedding vector and the second word embedding vector for the given word; and

predicting semantic shifts for at least a plurality of the words of the common vocabulary by providing the first word embedding vectors and the second word embedding vectors for at least the plurality of the words of the common vocabulary as input to the trained machine learning classifier.

20. The apparatus of claim 19 , wherein the method further comprises:

updating the plurality of landmark words to include words for which a semantic shift is not predicted; and

aligning the word embedding vectors of the first and second text corpora based on the plurality of landmark words.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 2, 2021
From: ADALI, SIBEL; GRUPPI, MAURICIO
To: RENSSELAER POLYTECHNIC INSTITUTE
Reel/Frame 055108/0758 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2021
From: CHEN, PIN-YU
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 055072/0742 →
Continuity (1)
Related Publication 20220245348A1 · Aug 4, 2022