IP Library › Granted Patent US 10,599,954
Granted Patent B2
US 10,599,954 · App. 15/961,721 · Granted Mar 24, 2020

Method and apparatus of discovering bad case based on artificial intelligence, device and storage medium

Inventor: Xiaoxiong Sun (Haidian District Beijing, CN)
Assignee: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
G06K9/6264G06F15/76G06F40/295G06K9/6265G06K9/6885G06F40/205G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,599,954
App. No.
15/961,721
Filed
Apr 24, 2018
Granted
Mar 24, 2020
Kind
B2
Art Unit
2659
USPC
704/9
Abstract

The present disclosure provides a method and apparatus of discovering a bad case based on artificial intelligence, a device and a storage medium, wherein the method comprises: performing named entity recognition for a to-be-recognized query, and respectively obtaining a confidence level of each character in the query; respectively obtaining a probability value of each character of forming a word with a neighboring character in the query; determining whether there is a bad case according to the confidence level and the probability value. The solution of the present disclosure may be applied to save man power costs, and improve the processing efficiency and enhance a discovery rate of bad cases.

Claims (123)

1. A method of discovering a bad case based on artificial intelligence, wherein the method comprises:

performing named entity recognition for a to-be-recognized query, and respectively obtaining a confidence level of each character in the query;

respectively obtaining a probability value of each character of forming a word with a neighboring character in the query; and

determining whether there is a bad case according to the confidence level and the probability value,

wherein before the performing named entity recognition for a to-be-recognized query, the method further comprises:

training to obtain a probability value evaluating model:

the respectively obtaining a probability value of each character of forming a word with a neighboring character in the query comprises:

according to the probability value evaluating model, respectively determining the probability value of each character of forming a word with a neighboring character in the query,

and wherein the according to the probability value evaluating model, respectively determining the probability value of each character of forming a word with a neighboring character in the query comprises:

considering each character in the query as a candidate character, and respectively performing the following processing for each candidate character:

determining a character which is spaced apart from a candidate character by less than or equal to M characters in the query as a neighboring character of the candidate character, M being a natural number; and

segmenting the query to obtain a segment which comprises the candidate word and at least one neighboring character;

regarding each segment, determining a similar word similar to the segment and a similar probability value of each similar word according to the probability value evaluating model;

selecting a similar probability value with a maximum value as a probability value of the candidate character forming a word with the neighboring character.

2. The method according to claim 1 , wherein

the probability value evaluating model comprises a word embedding model.

3. The method according to claim 1 , wherein

the segmenting the query to obtain a segment comprises:

for each neighboring word, determining a location of the neighboring character;

if the neighboring character is located before the candidate character, segmenting the query to obtain a segment starting from the neighboring character and ending at the candidate character;

if the neighboring character is located behind the candidate character, segmenting the query to obtain a segment starting from the candidate character and ending at the neighboring character.

4. The method according to claim 1 , wherein

the determining whether there is a bad case according to the confidence level and the probability value comprises:

considering each character in the query as a candidate character, and respectively performing the following processing for each candidate character:

calculating a difference between the confidence level of the candidate character and a preset first threshold to obtain a first difference, and determining a first parameter according to the first difference;

calculating a difference between a preset second threshold and the probability value corresponding to the candidate character to obtain a second difference, and determining a second parameter according to the second difference;

determining a third parameter according to the first parameter and the second parameter;

summating the third parameter corresponding to each candidate character in the query, comparing a sum with a preset third threshold, and determining a bad case if the sum is larger than the third threshold.

5. The method according to claim 4 , wherein

the determining a first parameter according to the first difference comprises:

if the first difference is larger than 0, setting a value of the first parameter as 1;

if the first difference is equal to 0, setting the value of the first parameter as 0;

if the first difference is smaller than 0, setting the value of the first parameter as −1;

the determining a second parameter according to the second difference comprises:

if the second difference is larger than 0, setting the value of the second parameter as 1;

if the second difference is equal to 0, setting the value of the second parameter as 0;

if the second difference is smaller than 0, setting the value of the second parameter as −1.

6. The method according to claim 5 , wherein

the determining the third parameter according to the first parameter and the second parameter comprises:

if both the first parameter and second parameter are smaller than 0, setting a value of the third parameter as 1, otherwise as 0;

a value of the third threshold is 1.

7. A computer device, comprising a memory, a processor and a computer program which is stored on the memory and runs on the processor, wherein the processor, upon executing the program, implements the following operation:

performing named entity recognition for a to-be-recognized query, and respectively obtaining a confidence level of each character in the query;

respectively obtaining a probability value of each character of forming a word with a neighboring character in the query; and

determining whether there is a bad case according to the confidence level and the probability value,

wherein before the performing named entity recognition for a to-be-recognized query, the method further comprises:

training to obtain a probability value evaluating model;

the respectively obtaining a probability value of each character of forming a word with a neighboring character in the query comprises:

according to the probability value evaluating model, respectively determining the probability value of each character of forming a word with a neighboring character in the query,

and wherein the according to the probability value evaluating model, respectively determining the probability value of each character of forming a word with a neighboring character in the query comprises:

considering each character in the query as a candidate character, and respectively performing the following processing for each candidate character:

determining a character which is spaced apart from a candidate character by less than or equal to M characters in the query as a neighboring character of the candidate character, M being a natural number;

segmenting the query to obtain a segment which comprises the candidate word and at least one neighboring character;

regarding each segment, determining a similar word similar to the segment and a similar probability value of each similar word according to the probability value evaluating model; and

selecting a similar probability value with a maximum value as a probability value of the candidate character forming a word with the neighboring character.

8. The computer device according to claim 7 , wherein

the probability value evaluating model comprises a word embedding model.

9. The computer device according to claim 7 , wherein

the segmenting the query to obtain a segment comprises:

for each neighboring word, determining a location of the neighboring character;

if the neighboring character is located before the candidate character, segmenting the query to obtain a segment starting from the neighboring character and ending at the candidate character;

if the neighboring character is located behind the candidate character, segmenting the query to obtain a segment starting from the candidate character and ending at the neighboring character.

10. The computer device according to claim 7 , wherein

the determining whether there is a bad case according to the confidence level and the probability value comprises:

considering each character in the query as a candidate character, and respectively performing the following processing for each candidate character:

calculating a difference between the confidence level of the candidate character and a preset first threshold to obtain a first difference, and determining a first parameter according to the first difference;

calculating a difference between a preset second threshold and the probability value corresponding to the candidate character to obtain a second difference, and determining a second parameter according to the second difference;

determining a third parameter according to the first parameter and the second parameter;

summating the third parameter corresponding to each candidate character in the query, comparing a sum with a preset third threshold, and determining a bad case if the sum is larger than the third threshold.

11. The computer device according to claim 10 , wherein

the determining a first parameter according to the first difference comprises:

if the first difference is larger than 0, setting a value of the first parameter as 1;

if the first difference is equal to 0, setting the value of the first parameter as 0;

if the first difference is smaller than 0, setting the value of the first parameter as −1;

the determining a second parameter according to the second difference comprises:

if the second difference is larger than 0, setting the value of the second parameter as 1;

if the second difference is equal to 0, setting the value of the second parameter as 0;

if the second difference is smaller than 0, setting the value of the second parameter as −1.

12. The computer device according to claim 11 , wherein

the determining the third parameter according to the first parameter and the second parameter comprises:

if both the first parameter and second parameter are smaller than 0, setting a value of the third parameter as 1, otherwise as 0;

a value of the third threshold is 1.

13. A non-transitory computer-readable storage medium on which a computer program is stored, wherein the program, when executed by a processor, implements the following operation:

performing named entity recognition for a to-be-recognized query, and respectively obtaining a confidence level of each character in the query;

respectively obtaining a probability value of each character of forming a word with a neighboring character in the query; and

determining whether there is a bad case according to the confidence level and the probability value,

wherein before the performing named entity recognition for a to-be-recognized query, the method further comprises:

training to obtain a probability value evaluating model;

the respectively obtaining a probability value of each character of forming a word with a neighboring character in the query comprises:

according to the probability value evaluating model, respectively determining the probability value of each character of forming a word with a neighboring character in the query,

and wherein the according to the probability value evaluating model, respectively determining the probability value of each character of forming a word with a neighboring character in the query comprises:

considering each character in the query as a candidate character, and respectively performing the following processing for each candidate character:

determining a character which is spaced apart from a candidate character by less than or equal to M characters in the query as a neighboring character of the candidate character, M being a natural number;

segmenting the query to obtain a segment which comprises the candidate word and at least one neighboring character;

regarding each segment, determining a similar word similar to the segment and a similar probability value of each similar word according to the probability value evaluating model; and

selecting a similar probability value with a maximum value as a probability value of the candidate character forming a word with the neighboring character.

14. The non-transitory computer-readable storage medium according to claim 13 , wherein

the probability value evaluating model comprises a word embedding model.

15. The non-transitory computer-readable storage medium according to claim 13 , wherein

the segmenting the query to obtain a segment comprises:

for each neighboring word, determining a location of the neighboring character;

if the neighboring character is located before the candidate character, segmenting the query to obtain a segment starting from the neighboring character and ending at the candidate character;

if the neighboring character is located behind the candidate character, segmenting the query to obtain a segment starting from the candidate character and ending at the neighboring character.

16. The non-transitory computer-readable storage medium according to claim 13 , wherein

the determining whether there is a bad case according to the confidence level and the probability value comprises:

considering each character in the query as a candidate character, and respectively performing the following processing for each candidate character:

calculating a difference between the confidence level of the candidate character and a preset first threshold to obtain a first difference, and determining a first parameter according to the first difference;

calculating a difference between a preset second threshold and the probability value corresponding to the candidate character to obtain a second difference, and determining a second parameter according to the second difference;

determining a third parameter according to the first parameter and the second parameter;

summating the third parameter corresponding to each candidate character in the query, comparing a sum with a preset third threshold, and determining a bad case if the sum is larger than the third threshold.

17. The non-transitory computer-readable storage medium according to claim 16 , wherein

the determining a first parameter according to the first difference comprises:

if the first difference is larger than 0, setting a value of the first parameter as 1;

if the first difference is equal to 0, setting the value of the first parameter as 0;

if the first difference is smaller than 0, setting the value of the first parameter as −1;

the determining a second parameter according to the second difference comprises:

if the second difference is larger than 0, setting the value of the second parameter as 1;

if the second difference is equal to 0, setting the value of the second parameter as 0;

if the second difference is smaller than 0, setting the value of the second parameter as −1.

18. The non-transitory computer-readable storage medium according to claim 17 , wherein

the determining the third parameter according to the first parameter and the second parameter comprises:

if both the first parameter and second parameter are smaller than 0, setting a value of the third parameter as 1, otherwise as 0;

a value of the third threshold is 1.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 25, 2018
From: SUN, XIAOXIONG
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 045632/0813 →
Priority Claims (1)
CN 2017 1 0311895 · May 5, 2017 · national
Continuity (1)
Related Publication 20180322370A1 · Nov 8, 2018
Cited By (1)
US 12,639,550