IP Library › Granted Patent US 11,562,137
Granted Patent B2
US 11,562,137 · App. 16/847,895 · Granted Jan 24, 2023

System to correct model drift for natural language understanding

Inventors: Utkarsh Raj (Charlotte, NC); Maharaj Mukherjee (Poughkeepsie, NY)
Assignee: Bank of America Corporation
G06F40/216G06F40/242G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,562,137
App. No.
16/847,895
Filed
Apr 14, 2020
Granted
Jan 24, 2023
Kind
B2
Art Unit
2655
USPC
704/9
Abstract

A system retrains a natural language understanding (NLU) model by regularly analyzing electronic documents including web publications such as online newspapers, blogs, social media posts, etc. to understand how word and phrase usage is evolving. Generally, the system determines the frequency of words and phrases in the electronic documents and updates an NLU dictionary depending on whether certain words or phrases are being used more frequently or less frequently. This dictionary is then used to retrain the NLU model, which is then applied to predict the meaning of text or speech communicated by a people group. By analyzing electronic documents such as web publications, the system is able to stay up-to-date on the vocabulary of the people group and make correct predictions as the vocabulary changes (e.g., due to natural disaster). In this manner, the safety of the people is improved.

Claims (71)

1. An apparatus comprising:

a memory configured to store a hash table, the hash table comprising a plurality of indices and a plurality of frequencies, each frequency of the plurality of frequencies is assigned to an index of the plurality of indices; and

a hardware processor communicatively coupled to the memory, the hardware processor configured to:

receive an electronic document;

determine a number of occurrences of a first one or more words in the electronic document;

update a stored frequency of the first one or more words based on the determined number of occurrences by:

applying a hash function to the first one or more words to produce an index for the first one or more words, the index for the first one or more words is in the plurality of indices; and

updating a frequency of the plurality of frequencies assigned to the index for the first one or more words based on the number of occurrences of the first one or more words;

compare the frequency to a first threshold;

in response to determining that the frequency exceeds the first threshold, add the first one or more words to a dictionary for a natural language understanding process;

retrain the natural language understanding process with the dictionary after the first one or more words is added to the dictionary; and

use the dictionary having the one or more added words and the retrained natural language understanding process to determine a meaning of a portion of the electronic document;

wherein the processor is further configured to:

determine a number of occurrences of a second one or more words in the electronic document;

update a stored frequency of the second one or more words based on the determined number of occurrences of the second one or more words;

compare the frequency of the second one or more words to a second threshold; and

in response to determining that the frequency of the second one or more words is below the second threshold, remove the second one or more words from the dictionary.

2. The apparatus of claim 1 , wherein the second threshold is below the first threshold.

3. The apparatus of claim 1 , wherein the processor is further configured to:

determine a topic of the first one or more words; and

assign the first one or more words to the topic in the dictionary.

4. The apparatus of claim 1 , wherein the processor is further configured to:

determine a number of occurrences of a phrase in the electronic document;

update a stored frequency of the phrase based on the determined number of occurrences of the phrase;

compare the frequency of the phrase to the first threshold; and

in response to determining that the frequency of the phrase exceeds the first threshold, add the phrase to the dictionary.

5. A method comprising:

receiving, by a hardware processor communicatively coupled to a memory, an electronic document, wherein the memory stores a hash table, the hash table comprising a plurality of indices and a plurality of frequencies, each frequency of the plurality of frequencies is assigned to an index of the plurality of indices;

determining, by the processor, a number of occurrences of a first one or more words in the electronic document;

updating, by the processor, a stored frequency of the first one or more words based on the determined number of occurrences by:

applying a hash function to the first one or more words to produce an index for the first one or more words, the index for the first one or more words is in the plurality of indices; and

updating a frequency of the plurality of frequencies assigned to the index for the first one or more words based on the number of occurrences of the first one or more words;

comparing, by the processor, the frequency to a first threshold;

in response to determining that the frequency exceeds the first threshold, adding, by the processor, the first one or more words to a dictionary for a natural language understanding process;

retraining the natural language understanding process with the dictionary after the first one or more words is added to the dictionary; and

using, by the processor, the dictionary having the one or more added words and the retrained natural language understanding process to determine a meaning of a portion of the electronic document;

determining, by the processor, a number of occurrences of a second one or more words in the electronic document;

updating, by the processor, a stored frequency of the second one or more words based on the determined number of occurrences of the second one or more words;

comparing, by the processor, the frequency of the second one or more words to a second threshold; and

in response to determining that the frequency of the second one or more words is below the second threshold, removing, by the processor, the second one or more words from the dictionary.

6. The method of claim 5 , wherein the second threshold is below the first threshold.

7. The method of claim 5 , further comprising:

determining, by the processor, a topic of the first one or more words; and

assigning, by the processor, the first one or more words to the topic in the dictionary.

8. The method of claim 5 , further comprising:

determining, by the processor, a number of occurrences of a phrase in the electronic document;

updating, by the processor, a stored frequency of the phrase based on the determined number of occurrences of the phrase;

comparing, by the processor, the frequency of the phrase to the first threshold; and

in response to determining that the frequency of the phrase exceeds the first threshold, adding, by the processor, the phrase to the dictionary.

9. A system comprising:

a device; and

a model drift correction tool comprising a hardware processor communicatively coupled to a memory, the memory configured to store a hash table, the hash table comprising a plurality of indices and a plurality of frequencies, each frequency of the plurality of frequencies is assigned to an index of the plurality of indices;

the hardware processor configured to:

receive an electronic document;

determine a number of occurrences of a first one or more words in the electronic document;

update a stored frequency of the first one or more words based on the determined number of occurrences by:

applying a hash function to the first one or more words to produce an index for the first one or more words, the index for the first one or more words is in the plurality of indices; and

updating a frequency of the plurality of frequencies assigned to the index for the first one or more words based on the number of occurrences of the first one or more words;

compare the frequency to a first threshold;

in response to determining that the frequency exceeds the first threshold, add the first one or more words to a dictionary for a natural language understanding process;

retrain the natural language understanding process with the dictionary after the first one or more words is added to the dictionary; and

use the dictionary having the one or more added words and the retrained natural language understanding process to determine a meaning of a message provided by the device;

wherein the processor is further configured to:

determine a number of occurrences of a second one or more words in the electronic document;

update a stored frequency of the second one or more words based on the determined number of occurrences of the second one or more words;

compare the frequency of the second one or more words to a second threshold; and

in response to determining that the frequency of the second one or more words is below the second threshold, remove the second one or more words from the dictionary.

10. The system of claim 9 , wherein the second threshold is below the first threshold.

11. The system of claim 9 , wherein the processor is further configured to:

determine a topic of the first one or more words; and

assign the first one or more words to the topic in the dictionary.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 14, 2020
From: RAJ, UTKARSH; MUKHERJEE, MAHARAJ
To: BANK OF AMERICA CORPORATION
Reel/Frame 052389/0562 →
Continuity (1)
Related Publication 20210319174A1 · Oct 14, 2021