IP Library › Granted Patent US 11,551,236
Granted Patent B2
US 11,551,236 · App. 16/911,183 · Granted Jan 10, 2023

Method and system for user protection in ride-hailing platforms based on semi-supervised learning

Inventor: Jinjian Zhai (Union City, CA)
Assignee: Beijing DiDi Infinity Technology and Development Co., Ltd.
G06Q30/0185G06F21/44G06N20/00G06Q50/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,551,236
App. No.
16/911,183
Granted
Jan 10, 2023
Kind
B2
Abstract

Methods, systems, and apparatus for detecting malicious activities in a ride-hailing platforms are described. An exemplary method comprises: identifying a set of trips from historical data to form training data; training a classifier based on a plurality of features of the set of trips in the training data to identify whether a given trip is malicious or benign; deploying the classifier to classify new trips in the ride-hailing platform for a first period of time to obtain a plurality of malicious trip candidates; storing the plurality of malicious trip candidates in a staging database for a second period of time for data cleansing based on supplementary data collected during the second period of time; fetching, from the staging database, a set of malicious trip candidates that have been stored in the staging database longer than the second period of time; and re-training the classifier.

Claims (69)

1. A computer-implemented method for detecting malicious activities in a ride-hailing platform, the method comprising:

identifying a set of trips from historical data to form training data, wherein the set of trips comprise one or more malicious trips and one or more benign trips;

training a classifier based on a plurality of features of the set of trips in the training data to identify whether a given trip is malicious or benign, wherein the plurality of features comprise identification numbers of devices associated with the set of trips, and latitude/longitude information of the set of trips in the training data;

deploying the classifier to classify new trips in the ride-hailing platform for a first period of time to obtain a plurality of malicious trip candidates;

storing the plurality of malicious trip candidates in a staging database for a second period of time for data cleansing based on supplementary data collected during the second period of time for the plurality of malicious trip candidates, wherein a malicious trip candidate is removed from the staging database when the supplementary data indicate that the malicious trip candidate is false-positively classified;

fetching, from the staging database, a set of malicious trip candidates that have been stored in the staging database longer than the second period of time; and

re-training the classifier based on the plurality of features of trips in the set of malicious trip candidates.

2. The method of claim 1 , wherein the classifier classifies the new trips as malicious or benign with corresponding confidence scores,

the storing the plurality of malicious trip candidates into a staging database for a second period of time for data cleansing comprises assigning a Time-To-Live (TTL) to each of the plurality of malicious trip candidates based on the confidence score of the each malicious trip candidate, and

the fetching a set of malicious trip candidates that have been stored in the staging database longer than the second period of time from the staging database comprises fetching a malicious trip that has been stored in the staging database longer than the assigned TTL.

3. The method of claim 1 , further comprising:

deploying the classifier to classify new trips to obtain a plurality of benign trip candidates with corresponding confidence scores; and

sampling, from the plurality of benign trip candidates, a set of benign trip candidates with highest classification confidence scores; and

updating the training data by adding the set of benign trip candidates.

4. The method of claim 1 , wherein the supplementary data is extracted from a complaint proving the malicious trip candidate is false-positively identified.

5. The method of claim 1 , wherein the identifying a set of trips from historical data comprises:

for a historical trip, obtaining latitude information and longitude information across a plurality of points in time;

obtaining an instantaneous velocity of the historical trip based on a first order derivative of the latitude information and longitude information;

obtaining an instantaneous acceleration of the historical trip based on a second order derivative of the latitude information and longitude information; and

determining the historical trip as a malicious trip when the instantaneous velocity or the instantaneous acceleration is greater than a corresponding threshold.

6. The method of claim 1 , wherein the plurality of features further comprise:

a list of applications installed on each of the devices, and access permissions to system services granted to the list of applications.

7. The method of claim 1 , wherein the first period time is a day, and the second period of time is two weeks.

8. The method of claim 1 , wherein the identifier numbers of the devices comprise International Mobile Equipment Identity (IMEI) numbers, and the deploying the classifier to classify new trips comprises:

extracting IMEI numbers of a plurality of devices associated of a new trip at real-time;

determining whether one of the plurality of devices has an IMEI number that matches an item in a blacklist; and

sending a warning message to each of the plurality of devices other than the one device.

9. The method of claim 1 , wherein the plurality of features further comprise MD5-lists of versions of one or more applications installed on the devices.

10. The method of claim 1 , further comprising:

obtaining an android application package (APK) of one of the applications installed on the devices;

determining one or more Uniform Resource Locators (URLs), Internet Protocol (IP) addresses, or domain names that the APK designed to visit; and

determining whether a trip associated with the device is malicious based on the one or more URLs, IP addresses, or domain names.

11. A system comprising one or more processors and one or more non-transitory computer-readable memories coupled to the one or more processors, the one or more non-transitory computer-readable memories storing instructions that, when executed by the one or more processors, cause the system to perform operations comprising:

identifying a set of trips from historical data to form training data, wherein the set of trips comprise one or more malicious trips and one or more benign trips;

training a classifier based on a plurality of features of the set of trips in the training data to identify whether a given trip is malicious or benign, wherein the plurality of features comprise identification numbers of devices associated with the set of trips, and latitude/longitude information of the set of trips in the training data;

deploying the classifier to classify new trips in a ride-hailing platform for a first period of time to obtain a plurality of malicious trip candidates;

storing the plurality of malicious trip candidates in a staging database for a second period of time for data cleansing based on supplementary data collected during the second period of time for the plurality of malicious trip candidates, wherein a malicious trip candidate is removed from the staging database when the supplementary data indicate that the malicious trip candidate is false-positively classified;

fetching, from the staging database, a set of malicious trip candidates that have been stored in the staging database longer than the second period of time; and

re-training the classifier based on the plurality of features of trips in the set of malicious trip candidates.

12. The system of claim 11 , wherein the classifier classifies the new trips as malicious or benign with corresponding confidence scores,

the storing the plurality of malicious trip candidates into a staging database for a second period of time for data cleansing comprises assigning a Time-To-Live (TTL) to each of the plurality of malicious trip candidates based on the confidence score of the each malicious trip candidate, and

the fetching a set of malicious trip candidates that have been stored in the staging database longer than the second period of time from the staging database comprises fetching a malicious trip that has been stored in the staging database longer than the assigned TTL.

13. The system of claim 11 , wherein the operations further comprise:

deploying the classifier to classify new trips to obtain a plurality of benign trip candidates with corresponding confidence scores; and

sampling, from the plurality of benign trip candidates, a set of benign trip candidates with highest classification confidence scores; and

updating the training data by adding the set of benign trip candidates.

14. The system of claim 11 , wherein the identifying a set of trips from historical data comprises:

for a historical trip, obtaining latitude information and longitude information across a plurality of points in time;

obtaining an instantaneous velocity of the historical trip based on a first order derivative of the latitude information and longitude information;

obtaining an instantaneous acceleration of the historical trip based on a second order derivative of the latitude information and longitude information; and

determining the historical trip as a malicious trip when the instantaneous velocity or the instantaneous acceleration is greater than a corresponding threshold.

15. The system of claim 11 , wherein the plurality of features further comprise:

a list of applications installed on each of the devices, and access permissions to system services granted to the list of applications.

16. The system of claim 11 , wherein the plurality of features further comprise MD5-lists of versions of one or more applications installed on the devices.

17. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

identifying a set of trips from historical data to form training data, wherein the set of trips comprise one or more malicious trips and one or more benign trips;

training a classifier based on a plurality of features of the set of trips in the training data to identify whether a given trip is malicious or benign, wherein the plurality of features comprise identification numbers of devices associated with the set of trips, and latitude/longitude information of the set of trips in the training data;

deploying the classifier to classify new trips in a ride-hailing platform for a first period of time to obtain a plurality of malicious trip candidates;

storing the plurality of malicious trip candidates in a staging database for a second period of time for data cleansing based on supplementary data collected during the second period of time for the plurality of malicious trip candidates, wherein a malicious trip candidate is removed from the staging database when the supplementary data indicate that the malicious trip candidate is false-positively classified;

fetching, from the staging database, a set of malicious trip candidates that have been stored in the staging database longer than the second period of time; and

re-training the classifier based on the plurality of features of trips in the set of malicious trip candidates.

18. The storage medium of claim 17 , wherein the classifier classifies the new trips as malicious or benign with corresponding confidence scores,

the storing the plurality of malicious trip candidates into a staging database for a second period of time for data cleansing comprises assigning a Time-To-Live (TTL) to each of the plurality of malicious trip candidates based on the confidence score of the each malicious trip candidate, and

the fetching a set of malicious trip candidates that have been stored in the staging database longer than the second period of time from the staging database comprises fetching a malicious trip that has been stored in the staging database longer than the assigned TTL.

19. The storage medium of claim 17 , wherein the operations further comprise:

deploying the classifier to classify new trips to obtain a plurality of benign trip candidates with corresponding confidence scores; and

sampling, from the plurality of benign trip candidates, a set of benign trip candidates with highest classification confidence scores; and

updating the training data by adding the set of benign trip candidates.

20. The storage medium of claim 17 , wherein the plurality of features further comprise: a list of applications installed on each of the devices, and access permissions to system services granted to the list of applications.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 24, 2020
From: ZHAI, JINJIAN
To: BEIJING DIDI INFINITY TECHNOLOGY AND DEVELOPMENT CO., LTD.
Reel/Frame 053031/0163 →
Continuity (1)
Related Publication 20210406916A1 · Dec 30, 2021