IP Library Granted Patent US 11,928,432
Granted Patent B2
US 11,928,432 · App. 17/319,189 · Granted Mar 12, 2024

Multi-modal pre-training model acquisition method, electronic device and storage medium

Inventors: Fei Yu (Beijing, CN); Jiji Tang (Beijing, CN); Weichong Yin (Beijing, CN); Yu Sun (Beijing, CN); Hao Tian (Beijing, CN); Hua Wu (Beijing, CN); Haifeng Wang (Beijing, CN)
Assignee: BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO., LTD.
G06F40/284G06F40/30G06N5/04G06N20/00G06V10/811G06V20/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,928,432
App. No.
17/319,189
Granted
Mar 12, 2024
Kind
B2
Abstract

A multi-modal pre-training model acquisition method, an electronic device and a storage medium, which relate to the fields of deep learning and natural language processing, are disclosed. The method may include: determining, for each image-text pair as training data, to-be-processed fine-grained semantic word in the text; masking the to-be-processed fine-grained semantic words; and training the multi-modal pre-training model using the training data with the fine-grained semantic words masked.

Claims (40)

1. A multi-modal pre-training model acquisition method, comprising:

determining, for each image-text pair as training data, to-be-processed fine-grained semantic words in the text;

masking the to-be-processed fine-grained semantic words; and

training the multi-modal pre-training model using the training data with the fine-grained semantic words masked,

wherein determining the to-be-processed fine-grained semantic words in the text comprises:

acquiring a scene graph corresponding to the text, wherein the scene graph comprises: entity nodes, attribute tuples and relationship triples, each attribute tuple is composed of one entity node and one attribute node, and each relationship triple is composed of two entity nodes and one relationship node;

selecting a predetermined number of entity nodes, attribute tuples and relationship triples from the scene graph, and taking entity words in the text corresponding to the selected entity nodes, attribute words in the text corresponding to attribute nodes in the selected attribute tuples, and relationship words in the text corresponding to relationship nodes in the selected relationship triples, as the to-be-processed fine-grained semantic words.

2. The method according to claim 1 , wherein the selecting the predetermined number of entity nodes, attribute tuples and relationship triples from the scene graph comprises:

determining a number of nodes to be selected according to a total number of nodes comprised in the scene graph as the predetermined number; and randomly selecting the predetermined number of entity nodes, attribute tuples and relationship triples from the scene graph.

3. The method according to claim 1 , wherein training tasks of the multi-modal pre-training model comprise: entity prediction, attribute prediction and relationship prediction; and

wherein the multi-modal pre-training model predicts masked words in the text according to a context of the text and corresponding image content.

4. The method according to claim 1 , further comprising:

after completion of the training of the multi-modal pre-training model, fine-tuning, for any downstream task, the multi-modal pre-training model according to the training data corresponding to the downstream task.

5. An electronic device, comprising:

at least one processor; and

a memory in communication connection with the at least one processor; wherein

the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to carry out a multi-modal pre-training model acquisition method, which comprises:

determining, for each image-text pair as training data, to-be-processed fine-grained semantic words in the text;

masking the to-be-processed fine-grained semantic words; and

training the multi-modal pre-training model using the training data with the fine-grained semantic words masked,

wherein determining the to-be-processed fine-grained semantic words in the text comprises:

acquiring a scene graph corresponding to the text, wherein the scene graph comprises: entity nodes, attribute tuples and relationship triples, each attribute tuple is composed of one entity node and one attribute node, and each relationship triple is composed of two entity nodes and one relationship node;

selecting a predetermined number of entity nodes, attribute tuples and relationship triples from the scene graph, and taking entity words in the text corresponding to the selected entity nodes, attribute words in the text corresponding to attribute nodes in the selected attribute tuples, and relationship words in the text corresponding to relationship nodes in the selected relationship triples, as the to-be-processed fine-grained semantic words.

6. The electronic device according to claim 5 , wherein the selecting the predetermined number of entity nodes, attribute tuples and relationship triples from the scene graph comprises:

determining a number of nodes to be selected according to a total number of nodes comprised in the scene graph as the predetermined number; and randomly selecting the predetermined number of entity nodes, attribute tuples and relationship triples from the scene graph.

7. The electronic device according to claim 5 , wherein training tasks of the multi-modal pre-training model comprise: entity prediction, attribute prediction and relationship prediction; and

wherein the multi-modal pre-training model predicts masked words in the text according to a context of the text and corresponding image content.

8. The electronic device according to claim 5 , wherein the method further comprises:

after completion of the training of the multi-modal pre-training model, fine-tuning, for any downstream task, the multi-modal pre-training model according to the training data corresponding to the downstream task.

9. A non-transitory computer-readable storage medium comprising instructions, which, when executed by a computer, cause the computer to carry out a multi-modal pre-training model acquisition method, which comprises:

determining, for each image-text pair as training data, to-be-processed fine-grained semantic words in the text;

masking the to-be-processed fine-grained semantic words; and

training the multi-modal pre-training model using the training data with the fine-grained semantic words masked,

wherein determining the to-be-processed fine-grained semantic words in the text comprises:

acquiring a scene graph corresponding to the text, wherein the scene graph comprises: entity nodes, attribute tuples and relationship triples, each attribute tuple is composed of one entity node and one attribute node, and each relationship triple is composed of two entity nodes and one relationship node;

selecting a predetermined number of entity nodes, attribute tuples and relationship triples from the scene graph, and taking entity words in the text corresponding to the selected entity nodes, attribute words in the text corresponding to attribute nodes in the selected attribute tuples, and relationship words in the text corresponding to relationship nodes in the selected relationship triples, as the to-be-processed fine-grained semantic words.

10. The non-transitory computer-readable storage medium according to claim 9 , wherein the selecting the predetermined number of entity nodes, attribute tuples and relationship triples from the scene graph comprises:

determining a number of nodes to be selected according to a total number of nodes comprised in the scene graph as the predetermined number; and randomly selecting the predetermined number of entity nodes, attribute tuples and relationship triples from the scene graph.

11. The non-transitory computer-readable storage medium according to claim 9 , wherein training tasks of the multi-modal pre-training model comprise: entity prediction, attribute prediction and relationship prediction; and

wherein the multi-modal pre-training model predicts masked words in the text according to a context of the text and corresponding image content.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2021
From: YU, FEI; TANG, JIJI; YIN, WEICHONG; SUN, YU; TIAN, HAO; WU, HUA; WANG, HAIFENG
To: BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO., LTD.
Reel/Frame 056227/0671 →
Priority Claims (1)
CN 202010676107.3 · Jul 14, 2020 · national
Continuity (1)
Related Publication 20220019744A1 · Jan 20, 2022