IP Library Granted Patent US 12,254,032
Granted Patent B2
US 12,254,032 · App. 17/962,157 · Granted Mar 18, 2025

System and method for hybrid multilingual search indexing

Inventor: Geoffrey Michael Obbard (Waterloo, CA)
Assignee: OPEN TEXT CORPORATION
G06F16/3337G06F16/31G06F40/263
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,032
App. No.
17/962,157
Granted
Mar 18, 2025
Kind
B2
Abstract

System and method for the indexing and searching of multilingual documents are disclosed.

Claims (53)

1. A system, comprising:

a processor; and

a computer readable medium storing instructions translatable by the processor to implement a multilingual search engine, comprising instructions for:

receiving a multilingual document;

determining a set of fragments of a content of the multilingual document, wherein each of the set of fragments comprises a portion of the content of the multilingual document;

determining an index language associated with each fragment of the determined set of fragments of the multilingual document;

based on the index language, indexing each fragment in a multilingual object index in association with the multilingual document, wherein the multilingual object index includes tokens from the multilingual document in multiple languages determined according to their respective index language;

receiving a search query; and

performing the search query by:

determining a set of search fragments of the search query, wherein each of the set of search fragments comprises a portion of the search query;

determining a search language associated with each search fragment of the determined set of search fragments; and

based on the search language associated with each search fragment, indexing each search fragment in a multilingual search index in association with the search query; and

performing a search of the multilingual object index according to the search query based on the multilingual search index.

2. The system of claim 1 , wherein indexing each fragment is performed using a language specific model for the index language.

3. The system of claim 2 , wherein the indexing of each search fragment is preformed using a language specific model for the search language.

4. The system of claim 3 , wherein the language specific model for the index language is different that then language specific model for the search language.

5. The system of claim 1 , wherein each fragment is indexed in a single field.

6. The system of claim 1 , wherein the multilingual search index comprises tokens of the search fragments according to their respective associated search language.

7. The system of claim 6 , wherein the multilingual search index is temporarily stored while performing the search.

8. A method, comprising:

receiving a multilingual document;

determining a set of fragments of a content of the multilingual document, wherein each of the set of fragments comprises a portion of the content of the multilingual document;

determining an index language associated with each fragment of the determined set of fragments of the multilingual document;

based on the index language, indexing each fragment in a multilingual object index in association with the multilingual document, wherein the multilingual object index includes tokens from the multilingual document in multiple languages determined according to their respective index language;

receiving a search query; and

performing the search query by:

determining a set of search fragments of the search query, wherein each of the set of search fragments comprises a portion of the search query;

determining a search language associated with each search fragment of the determined set of search fragments; and

based on the search language associated with each search fragment, indexing each search fragment in a multilingual search index in association with the search query; and

performing a search of the multilingual object index according to the search query based on the multilingual search index.

9. The method of claim 8 , wherein indexing each fragment is performed using a language specific model for the index language.

10. The method of claim 9 , wherein the indexing of each search fragment is preformed using a language specific model for the search language.

11. The method of claim 10 , wherein the language specific model for the index language is different than language specific model for the search language.

12. The method of claim 8 , wherein each fragment is indexed in a single field.

13. The method of claim 8 , wherein the multilingual search index comprises tokens of the search fragments according to their respective associated search language.

14. The method of claim 13 , wherein the multilingual search index is temporarily stored while performing the search.

15. A non-transitory computer readable medium, comprising instructions for:

receiving a multilingual document;

determining a set of fragments of a content of the multilingual document, wherein each of the set of fragments comprises a portion of the content of the multilingual document;

determining an index language associated with each fragment of the determined set of fragments of the multilingual document;

based on the index language, indexing each fragment in a multilingual object index in association with the multilingual document, wherein the multilingual object index includes tokens from the multilingual document in multiple languages determined according to their respective index language;

receiving a search query; and

performing the search query by:

determining a set of search fragments of the search query, wherein each of the set of search fragments comprises a portion of the search query;

determining a search language associated with each search fragment of the determined set of search fragments; and

based on the search language associated with each search fragment, indexing each search fragment in a multilingual search index in association with the search query; and

performing a search of the multilingual object index according to the search query based on the multilingual search index.

16. The non-transitory computer readable medium of claim 15 , wherein indexing each fragment is performed using a language specific model for the index language.

17. The non-transitory computer readable medium of claim 16 , wherein the indexing of each search fragment is preformed using a language specific model for the search language.

18. The non-transitory computer readable medium of claim 17 , wherein the language specific model for the index language is different than then language specific model for the search language.

19. The non-transitory computer readable medium of claim 15 , wherein each fragment is indexed in a single field.

20. The non-transitory computer readable medium of claim 15 , wherein the multilingual search index comprises tokens of the search fragments according to their respective associated search language.

21. The non-transitory computer readable medium of claim 20 , wherein the multilingual search index is temporarily stored while performing the search.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 12, 2022
From: OBBARD, GEOFFREY MICHAEL
To: OPEN TEXT CORPORATION
Reel/Frame 061392/0656 →
Continuity (1)
Related Publication 20240119076A1 · Apr 11, 2024
References Cited (22)
US 6658377B1 · Anward et al. · 2003 [cited by applicant]
US 6842730B1 · Ejerhed et al. · 2005 [cited by applicant]
US 10762139B1 · Dai · 2020 [cited by examiner]
US 20110131212A1 · Shikha · 2011 [cited by examiner]
US 20120095748A1 · Li et al. · 2012 [cited by applicant]
US 20130339378A1 · Zheng · 2013 [cited by examiner]
US 20170078199A1 · Mosko · 2017 [cited by examiner]
US 20170364510A1 · Huang et al. · 2017 [cited by applicant]
US 20180330012A1 · Hopkins · 2018 [cited by examiner]
US 20190108279A1 · Moore et al. · 2019 [cited by applicant]
US 20210334299A1 · Sonntag et al. · 2021 [cited by applicant]
US 20240119070A1 · Obbard · 2024 [cited by applicant]
EP 2807535 · 2019 [cited by applicant]
Eric Brill., “A Simple Rule-Based Part of Speech Tagge” ANLC '92: Proceedings of the third conference on Applied natural language processing, Mar. 1992, pp. 152-155. [cited by applicant]
Nguyen et al., “RDRPOSTagger: A Ripple Down Rules-based Part-Of-Speech Tagger”, Proceedings of the Demonstrations at the 14th Conference of the European Chapter of the Association for Computational Linguistics, pp. 17-2… [cited by applicant]
International Search Report and Written Opinion issued by the Canadian Intellectual Property Office as the International Searching Authority (CA/ISA) for International PCT Application No. PCT/IB2023/060075, mailed Dec. … [cited by applicant]
LemmatizerME.txt, Apache Software Foundation, retrieved at <<opennlp/opennlp-tools/src/main/java/opennlp/tools/lemmatizer/ LemmatizerME.java at main⋅apache/opennlp⋅GitHub>>, 7 pages. [cited by applicant]
Dat Quoc Nguyen et al., “Ripple Down Rules for Part-of-Speech Tagging,” In Proc. of 12th CICLing—vol. Part I, Feb. 2011, pp. 190-201. [cited by applicant]
Dat Quoc Nguyen, et al., “RDRPOSTAGGER: Ripple Down Rules-Based Part-of-Speech Tagger,” The Demonstrations at the14 [cited by applicant]
Eric Brill, “A Simple Rule-Based Part of Speech Tagger,” ANLC '92: Proceedings of the third conference on Applied natural language processing, Mar. 1992, pp. 152-155. [cited by applicant]
Grzegorz Chrupała, “Towards a Machine-Learning Architecture for Lexical Functional Grammar Parsing,” Dublin City University, Apr. 2008, 136 pages. [cited by applicant]
Office Action issued by the United States Patent and Trademark Office (USPTO) for U.S. Appl. No. 17/962,177, mailed Sep. 29, 2024, 14 pages. [cited by applicant]
Cited By (1)
US 12,591,601