IP Library › Granted Patent US 11,775,766
Granted Patent B2
US 11,775,766 · App. 17/249,718 · Granted Oct 3, 2023

Method and apparatus for improving model based on pre-trained semantic model

Inventors: Xuyi Chen (Beijing, CN); Shiwei Huang (Beijing, CN)
Assignee: BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO., LTD.
G06F40/30G06N3/045G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,775,766
App. No.
17/249,718
Granted
Oct 3, 2023
Kind
B2
Abstract

Embodiments of a method and an apparatus for improving a model based on a pre-trained semantic model are provided. The method may include: based on the pre-trained semantic model, obtaining an initial improved model, where semantic result information of an input vector is determined in the initial improved model based on a hash search method; and based on a model distillation method, training the initial improved model to obtain an improved model. Some embodiments can obtain the semantic result information of the input vector by performing the hash search method on the input vector, replace the original complex iterative calculation process of a semantic model, and obtain the improved model with few model parameters and high compression ratio.

Claims (32)

1. A method for improving a model based on a pre-trained semantic model, the method comprising:

obtaining, based on the pre-trained semantic model, an initial improved model, wherein semantic result information of an input vector is determined in the initial improved model based on a hash search method; and

training, based on a model distillation method, the initial improved model to obtain an improved model, wherein determining the semantic result information of the input vector in the initial improved model based on the hash search method comprises:

obtaining, based on a fully connected layer, a transformation vector of the input vector in the initial improved model; and

determining a target position corresponding to the transformation vector in a hash storage module storing semantic information according to the hash search method, and using semantic information stored in the target position as the semantic result information of the input vector.

2. The method according to claim 1 , wherein the pre-trained semantic model is obtained based on the model distillation method.

3. The method according to claim 1 , wherein the training, based on a model distillation method, the initial improved model to obtain an improved model, comprises:

using the semantic model as a teacher network model to train the initial improved model used as a student network model to obtain the improved model.

4. The method according to claim 1 , wherein the method further comprises:

inputting a to-be-recognized vector into the improved model to obtain semantic result information of the to-be-recognized vector.

5. An electronic device, comprising:

at least one processor; and

a memory communicating with the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

obtaining, based on a pre-trained semantic model, an initial improved model, wherein semantic result information of an input vector is determined in the initial improved model based on a hash search method; and

training, based on a model distillation method, the initial improved model to obtain an improved model, wherein determining the semantic result information of the input vector in the initial improved model based on the hash search method comprises:

obtaining, based on a fully connected layer, a transformation vector of the input vector in the initial improved model; and

determining a target position corresponding to the transformation vector in a hash storage module storing semantic information according to the hash search method, and using semantic information stored in the target position as the semantic result information of the input vector.

6. The electronic device according to claim 5 , wherein the pre-trained semantic model is obtained based on the model distillation method.

7. The electronic device according to claim 5 , wherein the training, based on a model distillation method, the initial improved model to obtain an improved model, comprises:

using the semantic model as a teacher network model to train the initial improved model used as a student network model to obtain the improved model.

8. The electronic device according to claim 5 , wherein the operations further comprise:

inputting a to-be-recognized vector into the improved model to obtain semantic result information of the to-be-recognized vector.

9. A non-transitory computer readable storage medium storing computer instructions, wherein the computer instructions, when executed by a computer, cause the computer to perform operations, the operations comprising:

obtaining, based on a pre-trained semantic model, an initial improved model, wherein semantic result information of an input vector is determined in the initial improved model based on a hash search method; and

training, based on a model distillation method, the initial improved model to obtain an improved model, wherein determining the semantic result information of the input vector in the initial improved model based on the hash search method comprises:

obtaining, based on a fully connected layer, a transformation vector of the input vector in the initial improved model; and

determining a target position corresponding to the transformation vector in a hash storage module storing semantic information according to the hash search method, and using semantic information stored in the target position as the semantic result information of the input vector.

10. The storage medium according to claim 9 , wherein the pre-trained semantic model is obtained based on the model distillation method.

11. The storage medium according to claim 9 , wherein the training, based on a model distillation method, the initial improved model to obtain an improved model, comprises:

using the semantic model as a teacher network model to train the initial improved model used as a student network model to obtain the improved model.

12. The storage medium according to claim 9 , wherein the operations further comprise:

inputting a to-be-recognized vector into the improved model to obtain semantic result information of the to-be-recognized vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2021
From: CHEN, XUYI; HUANG, SHIWEI
To: BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO., LTD.
Reel/Frame 056800/0192 →
Priority Claims (1)
CN 202010555885.7 · Jun 17, 2020 · national
Continuity (1)
Related Publication 20210397794A1 · Dec 23, 2021
Cited By (1)
US 12,645,471