IP Library Granted Patent US 11,893,347
Granted Patent B2
US 11,893,347 · App. 17/335,823 · Granted Feb 6, 2024

Contrastive meta-learning for zero-shot learning

Inventors: Tassilo Klein (Berlin, DE); Moin Nabi (Berlin, DE)
Assignee: SAP SE
G06F40/284G06F40/30G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,893,347
App. No.
17/335,823
Granted
Feb 6, 2024
Kind
B2
Abstract

Disclosed herein are system, method, and computer program product embodiments for utilizing non-RAM memory to implement machine learning configured with a meta-learning training set (small dataset), to create a common-sense predictive language model, thus boosting the performance for downstream tasks. An embodiment operates by receiving a base sentence and perturbation sentences as an input and tokenizing the input to generate a sequence of tokens. Tokens of the semantic perturbation sentences are embedded with tokens of the base sentence as contextually similar tokens pairs to generate training data and classified to capture relationships of the base sentence and the perturbation sentences to generate a classification, which is used to train a language model.

Claims (51)

1. A computer implemented method for natural language processing, the method comprising:

receiving, by a tokenization module, a base sentence and one or more sentences comprising a semantic perturbation of the base sentence as an input, wherein the semantic perturbation of the base sentence comprises one or more linguistic deviations of the base sentence from a first version;

tokenizing, by the tokenization module, the input to generate a sequence of tokens;

embedding, by a machine learning engine, tokens of the semantic perturbation with tokens of the base sentence as tokens pairs to generate training data;

classifying, by a classifier, the semantic perturbation of the token pairs to capture relationships of the base sentence and the one or more sentences to generate a classification; and

training, by the machine learning engine, a language model based at least in part on the training data and the classification; and

wherein at least one of the receiving, tokenizing, determining, embedding and training are performed by one or more computers.

2. The method of claim 1 , the tokenizing the input comprising:

splitting the base sentence and the one or more sentences into smaller units.

3. The method of claim 2 , wherein the smaller units include any of:

individual words, terms, numbers or punctuation marks.

4. The method of claim 1 , the embedding comprising pairing contextually similar tokens.

5. The method of claim 4 , the classifying further comprising:

recognizing contextually similar tokens.

6. The method of claim 1 , the classifying further comprising:

mapping common-sense concepts associated with a specific language structure to capture the relationships between concepts of the base sentence and the one or more sentences.

7. The method of claim 1 , the training further comprising:

limiting a distance of relative positions of token pairs in the sequence of tokens.

8. A system, comprising:

a memory; and

at least one processor coupled to the memory and configured to:

receive a base sentence and one or more sentences comprising a semantic perturbation of the base sentence as an input, wherein the semantic perturbation of the base sentence comprises one or more linguistic deviations of the base sentence from a first version;

tokenize the input to generate a sequence of tokens;

embed tokens of the semantic perturbation with tokens of the base sentence as tokens pairs to generate training data;

classify the semantic perturbation of the token pairs to capture relationships of the base sentence and the one or more sentences to generate a classification; and

train a language model based, at least in part, on the training data and the classification.

9. The system of claim 8 , wherein to tokenize the input, the at least one processor is configured to:

split the base sentence and the one or more sentences into smaller units.

10. The system of claim 9 , wherein the smaller units include any of:

individual words, terms, numbers or punctuation marks.

11. The system of claim 8 , wherein to embed the one or more semantic perturbations, the at least one processor is configured to pair contextually similar tokens.

12. The system of claim 11 , wherein to classify the semantic perturbation of the token pairs, the at least one processor is configured to:

recognize contextually similar tokens.

13. The system of claim 8 , wherein to execute to train the language model, the at least one processor is configured to:

map common-sense concepts associated with a specific language structure to capture the relationships between concepts and sentences.

14. The system of claim 8 , wherein to train the language model, the at least one processor is configured to:

limit a distance of relative positions of token pairs in the sequence of tokens.

15. A non-transitory computer-readable device having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising:

receiving a base sentence and one or more sentences comprising a semantic perturbation of the base sentence as an input, wherein the semantic perturbation of the base sentence comprises one or more linguistic deviations of the base sentence from a first version:

tokenizing the input to generate a sequence of tokens;

embedding tokens of the semantic perturbation with tokens of the base sentence as tokens pairs to generate training data;

classifying the semantic perturbation of the token pairs to capture relationships of the base sentence and the one or more sentences to generate a classification; and

training a language model based at least in part on the training data and the classification.

16. The non-transitory computer-readable device of claim 15 , the tokenizing the input comprising:

splitting the base sentence and one or more sentences into smaller units.

17. The non-transitory computer-readable device of claim 16 , wherein the smaller units include any of: individual words, terms, numbers or punctuation marks.

18. The non-transitory computer-readable device of claim 15 , wherein the embedding comprises pairing contextually similar tokens.

19. The non-transitory computer-readable device of claim 15 , wherein the classifying further comprises:

recognizing contextually similar tokens.

20. The non-transitory computer-readable device of claim 15 , the raining further comprising:

mapping common-sense concepts associated with a specific language structure to capture the relationships between concepts of the base sentence and the one or more sentences.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2021
From: KLEIN, TASSILO; NABI, MOIN
To: SAP SE
Reel/Frame 056426/0570 →
Continuity (1)
Related Publication 20220382979A1 · Dec 1, 2022
Cited By (1)
US 12,436,745