IP Library › Granted Patent US 11,748,410
Granted Patent B2
US 11,748,410 · App. 17/499,710 · Granted Sep 5, 2023

System and method for pre-indexing filtering and correction of documents in search systems

Inventors: Bruce Edward Kiefer (Denver, CO); Gregory John Berka (Centennial, CO)
Assignee: OPEN TEXT HOLDINGS, INC.
G06F16/901G06F16/93G06F16/144G06F16/2465G06F16/316G06F40/10G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,748,410
App. No.
17/499,710
Granted
Sep 5, 2023
Kind
B2
Abstract

Embodiments as disclosed herein provide a search system with an pre-indexing filter that provides both a sophisticated and contextually tailored approach to filtering documents and a corrector that is adapted to alter a document that has been designated to be filtered out from the indexing process and determine if the altered document should be indexed. The alteration of the document may be tied to the attributes, rules or thresholds used to initially filter the document from the indexing process. The filtering criteria can thus be tailored to a specific context such that both the initial filtering and the alteration process may be better suited for application in that context.

Claims (42)

1. A search system, comprising:

a processor;

a data store, having an index of a corpus stored thereon, wherein the corpus comprises a set of documents; and

a non-transitory computer readable medium, having instructions executable on the processor for:

receiving a first set of tokens for a document from a text extractor;

providing the first set of tokens to a detector for determining if each of the received first set of tokens is of an associated type;

determining an associated detector score based on the determination if each of the received first set of tokens is of the associated type of token, the detector score associated with the associated type;

determining a filter score based on the detector score produced by the detector, wherein the filter score is based on a scoring rule associated with the detector or the associated type;

determining whether the document should be indexed based on the application of a verdict rule, wherein the verdict rule includes an expression for evaluating the filter score associated with the detector; and

when it is determined that the document should be indexed, providing the document to an indexer adapted to index the first set of tokens for the document.

2. The search system of claim 1 , wherein when it is determined that the document should not be indexed, removing at least some of the first set of tokens or a second set of tokens determined by the detector from the first set of tokens to determine a third set of tokens, and indexing or discarding the document based on a suitability criteria.

3. The search system of claim 2 , wherein indexing the document comprises indexing the first set of tokens or third set of tokens.

4. The search system of claim 2 , wherein the first set of tokens or the second set of tokens is removed based on a correction specification specifying an attribute or type of token to remove.

5. The search system of claim 2 , wherein an altered document comprising the third set of tokens is created.

6. The search system of claim 1 , wherein the associated type comprises an attribute or type that is associated with a token or the document.

7. The search system of claim 1 , wherein the suitability criteria is document size or a number of the third tokens.

8. A non-transitory computer readable medium, comprising instructions for:

receiving a first set of tokens for a document from a text extractor, wherein the document is one of a corpus of documents associated with an index;

providing the first set of tokens to a detector for determining if each of the received first set of tokens is of an associated type;

determining an associated detector score based on the determination if each of the received first set of tokens is of the associated type of token, the detector score associated with the associated type;

determining a filter score based on the detector score produced by the detector, wherein the filter score is based on a scoring rule associated with the detector or the associated type;

determining whether the document should be indexed based on the application of a verdict rule, wherein the verdict rule includes an expression for evaluating the filter score associated with the detector; and

when it is determined that the document should be indexed, providing the document to an indexer adapted to index the first set of tokens for the document.

9. The non-transitory computer readable medium of claim 8 , wherein when it is determined that the document should not be indexed, removing at least some of the first set of tokens or a second set of tokens determined by the detector from the first set of tokens to determine a third set of tokens, and indexing or discarding the document based on a suitability criteria.

10. The non-transitory computer readable medium of claim 9 , wherein indexing the document comprises indexing the first set of tokens or third set of tokens.

11. The non-transitory computer readable medium of claim 9 , wherein the first set of tokens or the second set of tokens is removed based on a correction specification specifying an attribute or type of token to remove.

12. The non-transitory computer readable medium of claim 9 , wherein an altered document comprising the third set of tokens is created.

13. The non-transitory computer readable medium of claim 8 , wherein the associated type comprises an attribute or type that is associated with a token or the document.

14. The non-transitory computer readable medium of claim 8 , wherein the suitability criteria is document size or a number of the third tokens.

15. A method, comprising:

receiving a first set of tokens for a document from a text extractor, wherein the document is one of a corpus of documents associated with an index;

providing the first set of tokens to a detector for determining if each of the received first set of tokens is of an associated type;

determining an associated detector score based on the determination if each of the received first set of tokens is of the associated type of token, the detector score associated with the associated type;

determining a filter score based on the detector score produced by the detector, wherein the filter score is based on a scoring rule associated with the detector or the associated type;

determining whether the document should be indexed based on the application of a verdict rule, wherein the verdict rule includes an expression for evaluating the filter score associated with the detector; and

when it is determined that the document should be indexed, providing the document to an indexer adapted to index the first set of tokens for the document.

16. The method of claim 15 , wherein when it is determined that the document should not be indexed, removing at least some of the first set of tokens or a second set of tokens determined by the detector from the first set of tokens to determine a third set of tokens, and indexing or discarding the document based on a suitability criteria.

17. The method of claim 16 , wherein indexing the document comprises indexing the first set of tokens or third set of tokens.

18. The method of claim 16 , wherein the first set of tokens or the second set of tokens is removed based on a correction specification specifying an attribute or type of token to remove.

19. The method of claim 16 , wherein an altered document comprising the third set of tokens is created.

20. The method of claim 15 , wherein the associated type comprises an attribute or type that is associated with a token or the document.

21. The method of claim 15 , wherein the suitability criteria is document size or a number of the third tokens.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2021
From: KIEFER, BRUCE EDWARD; BERKA, GREGORY JOHN
To: OPEN TEXT HOLDINGS, INC.
Reel/Frame 057920/0031 →
Continuity (2)
Continuation 16582882 · Sep 25, 2019
Related Publication 20220067094A1 · Mar 3, 2022
Cited By (2)
US 12,307,136 US 12,488,195