IP Library Granted Patent US 11,217,226
Granted Patent B2
US 11,217,226 · App. 16/667,022 · Granted Jan 4, 2022

System to detect and reduce understanding bias in intelligent virtual assistants

Inventor: Ian Beaver (Spokane, WA)
Assignee: VERINT AMERICAS INC.
G10L15/063G06F17/16G06F40/20G10L15/18G10L2015/0636
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,217,226
App. No.
16/667,022
Granted
Jan 4, 2022
Kind
B2
Abstract

Disclosed is a system and method for detecting and addressing bias in training data prior to building language models based on the training data. Accordingly system and method, detect bias in training data for Intelligent Virtual Assistant (IVA) understanding and highlight any found. Suggestions for reducing or eliminating them may be provided This detection may be done for each model within the Natural Language Understanding (NLU) component. For example, the language model, as well as any sentiment or other metadata models used by the NLU, can introduce understanding bias. For each model deployed, training data is automatically analyzed for bias and corrections suggested.

Claims (32)

1. A computer product comprising computer executable code embodied in a non-transitory computer readable medium that, when executing on one or more computing devices performs a method of automatically detecting bias in multi-class training data for training a language model, the method comprising:

digitally processing training data to identify if the training data comprises co-occurrence bias by:

building a co-occurrence matrix of all class labels in a given classification task in the training data,

normalizing entries of class labels in the co-occurrence matrix, and

identifying normalized entry values in the co-occurrence matrix above a predetermined threshold;

and

wherein, if the training data comprises co-occurrence bias, adjusting the training data to compensate for the bias identified by adding examples of class label combinations for entry values below the predetermined threshold until the normalized entries of all class values are above the predetermined threshold.

2. The computer program product of claim 1 , wherein the digitally processing the training data comprises scanning the training data with a bias scoring system.

3. The computer program product of claim 1 , further comprising flagging for further review by a human reviewer co-occurrence pairs having a number of co-occurrences above the predetermined threshold.

4. The computer program product of claim 1 , further comprising digitally processing training data to identify if the training data comprises a class population bias by comparing distribution of a given one of the class labels to a representation threshold value.

5. The computer program product of claim 4 , wherein the determination includes one of deeming the co-occurrence or class population bias being artificial requiring repair and deeming the co-occurrence or class population bias being accurate allowing disregarding.

6. The computer program product of claim 4 , wherein comprising adjusting the training data comprises adding more samples of an underpopulated label if class population bias is identified.

7. The computer program product of claim 4 , wherein adjusting the training data comprises deleting samples of an overpopulated label if class population bias is identified.

8. The computer program product of claim 4 , wherein upon identifying class population bias, adjusting the training data comprises adding samples of an underrepresented class labels to the training data.

9. The computer program product of claim 1 , further comprising referring detected bias to a human reviewer for further determination of bias.

10. The computer program product of claim 1 , further comprising adjusting the training data to compensate for the bias identified by deleting examples of class label combinations for entry values above the predetermined threshold until the normalized entries of all class values are below the predetermined threshold.

11. A method of automatically detecting bias in training data for training a language model, comprising:

digitally processing training data to identify if the training data comprises co-occurrence bias by:

building a co-occurrence matrix of all class labels in a given classification task in the training data,

normalizing entries of class labels in the co-occurrence matrix, and

identifying normalized entry values in the co-occurrence matrix above a predetermined threshold;

and

wherein, if the training data comprises co-occurrence bias, adjusting the training data to compensate for the bias identified by adding examples of class labels for entry values below the predetermined threshold until the normalized entries of all class values are above the predetermined threshold.

12. The method of claim 11 , wherein the digitally processing the training data comprises scanning the training data with a bias scoring system.

13. The method of claim 11 , further comprising flagging for further review by a human reviewer co-occurrence pairs having a number of co-occurrences above the predetermined threshold.

14. The method of claim 11 , further comprising digitally processing training data to identify if the training data comprises a class population bias by comparing distribution of a given one of the class labels to a representation threshold value.

15. The method of claim 14 , wherein the determination includes one of deeming the co-occurrence or class population bias being artificial requiring repair and deeming the co-occurrence or class population bias being accurate allowing disregarding.

16. The method of claim 14 , wherein comprising adjusting the training data comprises adding more samples of an underpopulated label if class population bias is identified.

17. The method of claim 14 , wherein adjusting the training data comprises deleting samples of an overpopulated label if class population bias is identified.

18. The method of claim 14 , wherein upon identifying class population bias, adjusting the training data comprises adding samples of an underrepresented class labels to the training data.

19. The method of claim 11 , further comprising referring detected bias to a human reviewer for further determination of bias.

20. The method of claim 11 , further comprising adjusting the training data to compensate for the bias identified by deleting examples of class label combinations for entry values above the predetermined threshold until the normalized entries of all class values are below the predetermined threshold.

Assignments (2)
SECURITY INTEREST Recorded Dec 23, 2025
From: VERINT AMERICAS INC.
To: ALTER DOMUS (US) LLC, AS COLLATERAL AGENT
Reel/Frame 074034/0292 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 30, 2020
From: BEAVER, IAN
To: VERINT AMERICAS INC.
Reel/Frame 051664/0681 →
Continuity (2)
Provisional Application 62752668 · Oct 30, 2018
Related Publication 20200143794A1 · May 7, 2020
Cited By (2)
US 12,430,590 US 12,494,293