IP Library › Granted Patent US 11,501,078
Granted Patent B2
US 11,501,078 · App. 16/693,689 · Granted Nov 15, 2022

Method and device for performing reinforcement learning on natural language processing model and storage medium

Inventor: Zhuang Qian (Beijing, CN)
Assignee: Beijing Xiaomi Intelligent Technology Co., Ltd.
G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,501,078
App. No.
16/693,689
Filed
Nov 25, 2019
Granted
Nov 15, 2022
Kind
B2
Examiner
HE, JIALONG
Art Unit
2659
USPC
706/20
Abstract

A method for natural language, includes: determining a slot tagging result output by a Bi-directional Long Short-Term Memory-Conditional Random Field algorithm (BiLSTM-CRF) model after slot tagging on conversation data input by a user; determining reward information based on the slot tagging result and a reward of the user for the slot tagging result; and performing reinforcement learning on the BiLSTM-CRF model according to the reward information.

Claims (45)

1. A method for natural language processing, executed by a chatbot in a man-machine conversation device comprising a central control device, the method comprising:

determining a slot tagging result output by a Bi-directional Long Short-Term Memory-Conditional Random Field algorithm (BiLSTM-CRF) model, wherein the BiLSTM-CRF model performs slot tagging on conversation data input by a user and outputs the slot tagging result;

outputting, by the chatbot, the slot tagging result output by the BiLSTM-CRF model to the central control device;

acquiring a target slot tagging result determined by the central control device from a received slot tagging result set for the conversation data, wherein the slot tagging result set comprises the slot tagging result output by the BiLSTM-CRF model and one or more slot tagging results output by one or more other chatbots, and the target slot tagging result is output as a reply result of the man-machine conversation device for the user;

determining reward information by:

responsive to inconsistency of the target slot tagging result and the slot tagging result output by the BiLSTM-CRF model, determining inconsistency reward information to be negative reward information; or

responsive to consistency of the target slot tagging result and the slot tagging result output by the BiLSTM-CRF model, determining the reward information according to a reward operation of the user for the reply result; and

performing reinforcement learning on the BiLSTM-CRF model according to the reward information.

2. The method of claim 1 , wherein determining the reward information according to the reward operation of the user for the reply result comprises one of:

responsive to that a positive reward rate of the user is more than or equal to a preset threshold value, determining the reward information to be positive reward information; or

responsive to that the positive reward rate is less than the preset threshold value, determining the reward information to be the negative reward information,

wherein the positive reward rate is determined according to the reward operation of the user for the reply result within a time period.

3. The method of claim 1 , wherein performing model reinforcement learning according to the reward information comprises:

providing a CRF layer in the BiLSTM-CRF model with the reward information, to perform model reinforcement training according to the reward information.

4. A man-machine conversation device that comprises a chatbot and a central control device, the man-machine conversation device comprising:

a processor; and

a memory configured to store instructions executable by the processor,

wherein the processor is configured to:

determine a slot tagging result output by a Bi-directional Long Short-Term Memory-Conditional Random Field algorithm (BiLSTM-CRF) model, wherein the BiLSTM-CRF model performs slot tagging on conversation data input by a user and outputs the slot tagging result;

output the slot tagging result output by the BiLSTM-CRF model to the central control device;

acquire a target slot tagging result determined by the central control device in a received slot tagging result set for the conversation data, wherein the slot tagging result set comprises the slot tagging result output by the BiLSTM-CRF model and one or more slot tagging results output by one or more other chatbots, and the target slot tagging result is output as a reply result of the man-machine conversation device for the user;

determine reward information by:

responsive to inconsistency of the target slot tagging result and the slot tagging result output by the BiLSTM-CRF model, determining inconsistency reward information to be negative reward information; or

responsive to consistency of the target slot tagging result and the slot tagging result output by the BiLSTM-CRF model, determining the reward information according to a reward operation of the user for the reply result; and

perform reinforcement learning on the BiLSTM-CRF model according to the reward information.

5. The man-machine conversation device of claim 4 , wherein the processor is configured to:

responsive to that a positive reward rate of the user is more than or equal to a preset threshold value, determine the reward information to be positive reward information; and

responsive to that the positive reward rate is less than the preset threshold value, determine the reward information to be the negative reward information,

wherein the positive reward rate is determined according to the reward operation of the user for the reply result within a time period.

6. The man-machine conversation device of claim 4 , wherein the processor is further configured to:

provide a CRF layer in the BiLSTM-CRF model with the reward information, for the CRF layer to perform model reinforcement training according to the reward information.

7. A non-transitory computer-readable storage medium having stored therein instructions that, when executed by a processor of a man-machine conversation device comprising a chatbot and a central control device, cause the man-machine conversation device to perform a method comprising:

determining a slot tagging result output by a Bi-directional Long Short-Term Memory-Conditional Random Field algorithm (BiLSTM-CRF) model, wherein the BiLSTM-CRF model performs slot tagging on conversation data input by a user and outputs the slot tagging result;

outputting, by the chatbot, the slot tagging result output by the BiLSTM-CRF model to the central control device;

acquiring a target slot tagging result determined by the central control device from a received slot tagging result set for the conversation data, wherein the slot tagging result set comprises the slot tagging result output by the BiLSTM-CRF model and one or more slot tagging results output by one or more other chatbots, and the target slot tagging result is output as a reply result of the man-machine conversation device for the user;

determining reward information by:

responsive to inconsistency of the target slot tagging result and the slot tagging result output by the BiLSTM-CRF model, determining inconsistency reward information to be negative reward information; or

responsive to consistency of the target slot tagging result and the slot tagging result output by the BiLSTM-CRF model, determining the reward information according to a reward operation of the user for the reply result; and

performing reinforcement learning on the BiLSTM-CRF model according to the reward information.

8. The non-transitory computer-readable storage medium of claim 7 , wherein determining the reward information according to the reward operation of the user for the reply result comprises one of:

responsive to that a positive reward rate of the user is more than or equal to a preset threshold value, determining the reward information to be positive reward information; or

responsive to that the positive reward rate is less than the preset threshold value, determining the reward information to be the negative reward information,

wherein the positive reward rate is determined according to the reward operation of the user for the reply result within a time period.

9. The non-transitory computer-readable storage medium of claim 7 , wherein performing model reinforcement learning according to the reward information comprises:

providing a CRF layer in the BiLSTM-CRF model with the reward information, to perform model reinforcement training according to the reward information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 25, 2019
From: QIAN, ZHUANG
To: BEIJING XIAOMI INTELLIGENT TECHNOLOGY CO., LTD.
Reel/Frame 051102/0358 →
Priority Claims (1)
CN 201910687763.0 · Jul 29, 2019 · national
Continuity (1)
Related Publication 20210034966A1 · Feb 4, 2021