IP Library Granted Patent US 11,514,271
Granted Patent B2
US 11,514,271 · App. 16/720,374 · Granted Nov 29, 2022

System and method for automatically adjusting strategies

Inventor: Jinjian Zhai (Union City, CA)
Assignee: Beijing DiDi Infinity Technology and Development Co., Ltd.
G06K9/6264G06F16/9035G06F40/40G06K9/6224G06K9/6267G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,514,271
App. No.
16/720,374
Granted
Nov 29, 2022
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for automatically adjusting strategies. One of the methods includes: determining one or more characteristics of a plurality of complaints, wherein each of the complaints corresponds to an order; classifying the plurality of complaints into a plurality of categories based on the one or more characteristics by using a trained classifier; selecting a category from the plurality of categories based on a number of complaints in the selected category; from a group of strategies each associated with one or more conditions and one or more actions, identifying a candidate strategy causing the complaints of the selected category, wherein the one or more actions are executed in response to the one or more conditions being satisfied; and optimizing the candidate strategy using a reinforcement learning model at least based on a plurality of historical orders.

Claims (50)

1. A computer-implemented method for automatically adjusting strategies, comprising:

determining one or more characteristics of a plurality of complaints, wherein each of the complaints corresponds to an order;

classifying the plurality of complaints into a plurality of categories based on the one or more characteristics by using a trained classifier;

selecting a category from the plurality of categories based on a number of complaints in the selected category;

from a group of strategies each associated with one or more conditions and one or more actions, identifying a candidate strategy causing the complaints of the selected category, wherein the one or more actions are executed in response to the one or more conditions being satisfied, and the identifying comprises:

in response to the complaints complaining about a false positive error, selecting the candidate strategy based on a number of orders that have applied the candidate strategy, and

in response to the complaints complaining about a false negative error, selecting the candidate strategy based on a number of orders that have skipped the candidate strategy; and

optimizing the candidate strategy using a reinforcement learning model at least by changing the one or more conditions of the candidate strategy based on a plurality of historical orders.

2. The method of claim 1 , wherein the determining one or more characteristics of a plurality of complaints comprises, for each of the complaints, using Natural Language Processing (NLP) to:

extract one or more first features from a content of the each complaint;

extract one or more second features from the order corresponding to the each complaint; and

extract one or more third features from a user profile associated with the order corresponding to the each complaint, wherein the one or more characteristics comprise the first, second, and third features.

3. The method of claim 1 , wherein the classifier is trained as using semi-supervised machine learning based on a first plurality of historical complaints with corresponding categorical labels and a second plurality of historical complaints that are not labeled.

4. The method of claim 1 , wherein the classifier comprises an unsupervised machine learning model trained to group the complaints based on vector representations of the one or more characteristics of the plurality of complaints.

5. The method of claim 1 , wherein the selecting a category from the categories based on a number of complaints in the selected category comprises:

selecting the category if the number of complaints in the category is greater than a threshold.

6. The method of claim 1 , wherein the selecting a category from the plurality of categories based on a number of complaints in the selected category comprises:

selecting the category if an increase of the number of complaints in the selected category during a period of time is greater than a threshold.

7. The method of claim 1 , wherein:

the one or more conditions are based on one or more of the following parameters: time of the order, pickup location, and destination.

8. The method of claim 1 , wherein the optimizing a candidate strategy using a reinforcement learning model comprises:

building one or more search graphs using Monte Carlo Graph Search (MCGS) algorithm based on a plurality of historical orders.

9. The method of claim 1 , wherein the optimizing a candidate strategy comprises:

determining a false positive rate and a false negative rate by examining the optimized candidate strategy against the plurality of historical orders, wherein each of the plurality of historical orders is labeled with whether an action associated with the candidate strategy should have been performed.

10. The method of claim 1 , wherein the optimizing a candidate strategy at least lowers a false positive rate.

11. A non-transitory computer-readable storage medium for automatically adjusting strategies, configured with instructions executable by one or more processors to cause the one or more processors to perform operations comprising:

determining one or more characteristics of a plurality of complaints, wherein each of the complaints corresponds to an order;

classifying the plurality of complaints into a plurality of categories based on the one or more characteristics by using a trained classifier;

selecting a category from the plurality of categories based on a number of complaints in the selected category;

from a group of strategies each associated with one or more conditions and one or more actions, identifying a candidate strategy causing the complaints of the selected category, wherein the one or more actions are executed in response to the one or more conditions being satisfied; and

optimizing the candidate strategy using a reinforcement learning model at least by changing the one or more conditions of the candidate strategy based on a plurality of historical orders, wherein the optimizing comprises:

determining a false positive rate and a false negative rate by examining the optimized candidate strategy against the plurality of historical orders, wherein each of the plurality of historical orders is labeled with whether an action associated with the optimized candidate strategy should have been performed.

12. The non-transitory computer readable storage medium of claim 11 , wherein the determining one or more characteristics of a plurality of complaints comprises, for each of the complaints, using Natural Language Processing (NLP) to:

extract one or more first features from a content of the each complaint;

extract one or more second features from the order corresponding to the each complaint; and

extract one or more third features from a user profile associated with the order corresponding to the each complaint, wherein the one or more characteristics comprise the first, second, and third features.

13. The non-transitory computer readable storage medium of claim 11 , wherein the classifier is trained as using semi-supervised machine learning based on a first plurality of historical complaints with corresponding categorical labels and a second plurality of historical complaints that are not labeled.

14. The non-transitory computer readable storage medium of claim 11 , wherein the classifier comprises an unsupervised machine learning model trained to group the complaints based on vector representations of the one or more characteristics of the plurality of complaints.

15. The non-transitory computer readable storage medium of claim 11 , wherein the selecting a category from the categories based on a number of complaints in the selected category comprises:

selecting the category if the number of complaints in the category is greater than a threshold.

16. The non-transitory computer readable storage medium of claim 11 , wherein the optimizing the candidate strategy using a reinforcement learning model comprises:

building one or more search graphs using Monte Carlo Graph Search (MCGS) algorithm based on a plurality of historical orders.

17. The non-transitory computer readable storage medium of claim 11 , wherein the optimizing a candidate strategy at least lowers a false positive rate.

18. A method for automatically adjusting strategies, the method comprising:

determining one or more characteristics of a plurality of feedbacks, wherein each of the feedbacks corresponds to an order;

classifying the plurality of feedbacks into a plurality of categories based on the one or more characteristics using a classifier;

selecting a category from the plurality of categories based on a number of feedbacks in the selected category;

from a group of strategies each associated with one or more conditions and one or more actions, identifying a candidate strategy resulting in the feedbacks of the selected category, wherein the one or more actions are executed in response to the one or more conditions being satisfied; and

in response to the feedbacks of the selected category comprising complaints, optimizing the candidate strategy using a reinforcement learning model at least by changing the one or more conditions of the candidate strategy based on a plurality of historical orders, wherein the optimizing comprises:

determining a false positive rate and a false negative rate by examining the optimized candidate strategy against the plurality of historical orders, wherein each of the plurality of historical orders is labeled with whether an action associated with the optimized candidate strategy should have been performed.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 9, 2020
From: DIDI (HK) SCIENCE AND TECHNOLOGY LIMITED
To: BEIJING DIDI INFINITY TECHNOLOGY AND DEVELOPMENT CO., LTD.
Reel/Frame 053180/0456 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 29, 2020
From: DIDI RESEARCH AMERICA, LLC
To: DIDI (HK) SCIENCE AND TECHNOLOGY LIMITED
Reel/Frame 053081/0934 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2019
From: ZHAI, JINJIAN
To: DIDI RESEARCH AMERICA, LLC
Reel/Frame 051330/0843 →
Continuity (1)
Related Publication 20210192292A1 · Jun 24, 2021
Cited By (1)
US 12,429,836