IP Library › Granted Patent US 11,928,433
Granted Patent B2
US 11,928,433 · App. 17/985,425 · Granted Mar 12, 2024

Systems and methods for term prevalence-volume based relevance

Inventor: Richard Kerr (Farmers Branch, TX)
G06F40/284G06F16/148G06F16/2237
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,928,433
App. No.
17/985,425
Granted
Mar 12, 2024
Kind
B2
Abstract

Techniques for prevalence-volume based relevance are provided. Corresponding systems and methods may include ingesting a corpus of documents; receiving a search operator; segmenting the corpus of documents into (i) a first set of documents that matches the search operator, and (ii) a second set of documents that do not match the search operator; extracting a first and second token list of tokens; calculating a prevalence-volume value for tokens included in the first and second token lists; generating a prevalence-volume ratio (PVR) matrix that associates tokens included in the first and/or second token lists with a PVR value, wherein the PVR value for a particular token is a ratio between the prevalence-volume value of the particular token for the first set of documents and the prevalence-volume value of the particular token for the second set of documents; and associating the search operator with the generated PVR matrix.

Claims (66)

1. A computer-implemented method comprising:

ingesting, via one or more processors, a corpus of documents;

receiving, via the one or more processors, a plurality of search operators associated with respective job categories;

for search operators in the plurality of search operators:

segmenting, via the one or more processors, the corpus of documents into (i) a first set of documents that matches the search operator, and (ii) a second set of documents that do not match the search operator,

extracting, via the one or more processors, (i) a first token list of tokens included in the first set of documents and (ii) a second token list of tokens included in the second set of documents,

calculating, via the one or more processors, a prevalence-volume value for tokens included in the first and second token lists,

based on the calculated prevalence-volume values, generating, via the one or more processors, a prevalence-volume ratio (PVR) matrix wherein the PVR matrix associates tokens included in the first and/or second token lists with a PVR value, wherein the PVR value for a particular token is a ratio between the calculated prevalence-volume value of the particular token for the first set of documents and the calculated prevalence-volume value of the particular token for the second set of documents, and

associating, via the one or more processors, the search operator with the generated PVR matrix;

obtaining, via the one or more processors, a user-uploaded document;

generating, via the one or more processors, a PVR score for the user-uploaded document using the PVR matrices for the plurality of search operators; and

presenting, via the one or more processors, a ranked list of job categories based upon the generated PVR scores.

2. The computer-implemented method of claim 1 , further comprising:

transmitting, to a client device, an indication of tokens included in the PVR matrix.

3. The computer-implemented method of claim 1 , wherein generating the PVR score for the user-uploaded document for a search operator in the plurality of search operators comprises:

extracting, via the one or more processors, a document token list for the user uploaded-document;

for tokens included in the document token list, obtaining, via the one or more processors, the corresponding PVR value from the PVR matrix for the search operator; and

generating, via the one or more processors, PVR score for the user-uploaded document by summing the obtained PVR values for the tokens included in the document token list.

4. The computer-implemented method of claim 3 , further comprising:

determining, via the one or more processors, whether the PVR score for the user-uploaded document exceeds a threshold PVR score when applying the PVR matrix for the search operator; and

presenting, via the one or more processors, an indication of whether the user-uploaded document exceeds the threshold PVR score for the search operator.

5. The computer-implemented method of claim 1 , wherein the user-uploaded document is a resume.

6. The computer-implemented method of claim 1 , wherein presenting the ranked list of job categories comprises:

presenting, via the one or more processors, a normalized indication of the PVR score for the job categories included in the ranked list.

7. The computer-implemented method of claim 1 , wherein generating the PVR matrix comprises:

excluding, via the one or more processors, tokens that have a PVR value below a threshold PVR value from inclusion in the PVR matrix.

8. The computer-implemented method of claim 1 , wherein receiving the plurality of search operators comprises:

receiving, from a client device, the plurality of search operators.

9. The computer-implemented method of claim 1 , wherein ingesting the corpus of documents comprises:

scraping, via the one or more processors, a website or a third party database to generate the documents in the corpus of documents.

10. The computer-implemented method of claim 1 , wherein extracting the token list comprises:

excluding, via the one or more processors, words that meet one or more exclusion criteria.

11. A system comprising:

one or more processors; and

one or more non-transitory memories coupled to the one or more processors and storing processor-executable instructions thereon that, when executed by the one or more processors, cause the system to:

ingest a corpus of documents;

receive a plurality of search operators associated with respective job categories;

for search operators in the plurality of search operators:

segment the corpus of documents into (i) a first set of documents that matches the search operator, and (ii) a second set of documents that do not match the search operator,

extract (i) a first token list of tokens included in the first set of documents and (ii) a second token list of tokens included in the second set of documents,

calculate a prevalence-volume value for tokens included in the first and second token lists,

based on the calculated prevalence-volume values, generate a prevalence-volume ratio (PVR) matrix wherein the PVR matrix associates tokens included in the first and/or second token lists with a PVR value, wherein the PVR value for a particular token is a ratio between the calculated prevalence-volume value of the particular token for the first set of documents and the calculated prevalence-volume value of the particular token for the second set of documents, and

associate the search operator with the generated PVR matrix;\

obtain a user-uploaded document;

generate a PVR score for the user-uploaded document using the PVR matrices for the plurality of search operators; and

present a ranked list of job categories based upon the generated PVR scores.

12. The system of claim 11 , wherein the instructions, when executed, cause the system to:

transmit, to a client device, an indication of tokens included in the PVR matrix.

13. The system of claim 11 , wherein to generate the PVR score for the user-uploaded document for a search operator in the plurality of search operators, the instructions, when executed, cause the system to:

extract a document token list for the user uploaded-document;

for tokens included in the document token list, obtain the corresponding PVR value from the PVR matrix for the search operator; and

generate PVR score for the user-uploaded document by summing the obtained PVR values for the tokens included in the document token list.

14. The system of claim 13 , wherein the instructions, when executed, cause the system to:

determine whether the PVR score for the user-uploaded document exceeds a threshold PVR score when applying the PVR matrix for the search operator; and

present an indication of whether the user-uploaded document exceeds the threshold PVR score for the search operator.

15. The system of claim 11 , wherein the user-uploaded document is a resume.

16. The system of claim 11 , wherein to present the ranked list of job categories, the instructions, when executed, cause the system to:

present a normalized indication of the PVR score for the job categories included in the ranked list.

17. The system of claim 11 , wherein to generate the PVR matrix, the instructions, when executed, cause the system to:

exclude tokens that have a PVR value below a threshold PVR value from inclusion in the PVR matrix.

18. The system of claim 11 , wherein to receive the plurality of search operators, the instructions, when executed, cause the system to:

receive, from a client device, the plurality of search operators.

19. The system of claim 11 , wherein to ingest the corpus of documents, the instructions, when executed, cause the system to:

scrape a website or a third party database to generate the documents in the corpus of documents.

20. The system of claim 11 , wherein to extract the token list, the instructions, when executed, cause the system to:

exclude words that meet one or more exclusion criteria.

Continuity (2)
Continuation 16775455 · Jan 29, 2020
Related Publication 20230073243A1 · Mar 9, 2023