IP Library Granted Patent US 12,591,601
Granted Patent B2
US 12,591,601 · App. 17/962,177 · Granted Mar 31, 2026

System and method for hybrid multilingual search indexing

Inventor: Geoffrey Michael Obbard (Waterloo, CA)
Assignee: OPEN TEXT CORPORATION
G06F16/316G06F40/58
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,591,601
App. No.
17/962,177
Granted
Mar 31, 2026
Kind
B2
Abstract

System and method for the indexing and searching of multilingual documents are disclosed.

Claims (41)

1 . A system, comprising:

a processor; and

a computer readable medium storing instructions translatable by the processor to implement a multilingual search engine, comprising instructions for:

receiving a multilingual object, the multilingual object comprising a search query;

determining a set of fragments of the text of the multilingual object, where each of the set of fragments comprises a portion of the text of the multilingual object determined based on a multilingual sentence boundary detection model;

determining a language associated with each fragment of the determined set of fragments of the multilingual object;

determining a set of tokens for each fragment of the determined set of fragments, where determining the set of tokens for a fragment comprises analyzing the portion of the text of that fragment based on the language associated with that fragment;

indexing the set of tokens determined for each fragment in an index in association with the multilingual object, such that the index is a multilingual index including tokens from the multilingual object in multiple languages determined according to the language associated with each fragment; and

performing a search of a search index using the determined set of tokens.

2 . The system of claim 1 , wherein the set of tokens are associated with a single field of the multilingual object.

3 . The system of claim 1 , wherein the indexed set of tokens are associated with a single field of the multilingual object.

4 . The system of claim 1 , where each of the set of fragments is a sentence.

5 . The system of claim 4 , wherein determining the set of fragments is done using a machine learning model trained on sentence markers for multiple languages.

6 . The system of claim 1 , further comprising presplitting the text of the multilingual object based upon one or more markers before determining the set of fragments.

7 . The system of claim 1 , wherein analyzing the portion of the text of that fragment based on the language associated with that fragment comprises selecting a language model of a set of pluggable language models.

8 . A method, comprising:

receiving a multilingual object, the multilingual object comprising a search query;

determining a set of fragments of the text of the multilingual object, where each of the set of fragments comprises a portion of the text of the multilingual object determined based on a multilingual sentence boundary detection model;

determining a language associated with each fragment of the determined set of fragments of the multilingual object;

determining a set of tokens for each fragment of the determined set of fragments, where determining the set of tokens for a fragment comprises analyzing the portion of the text of that fragment based on the language associated with that fragment;

indexing the set of tokens determined for each fragment in an index in association with the multilingual object such that the index is a multilingual index including tokens from the multilingual object in multiple languages determined according to the language associated with each fragment; and

performing a search of a search index using the determined set of tokens.

9 . The method of claim 8 , wherein the set of tokens are associated with a single field of the multilingual object.

10 . The method of claim 8 , wherein the indexed set of tokens are associated with a single field of the multilingual object.

11 . The method of claim 8 , where each of the set of fragments is a sentence.

12 . The method of claim 11 , wherein determining the set of fragments is done using a machine learning model trained on sentence markers for multiple languages.

13 . The method of claim 8 , further comprising presplitting the text of the multilingual object based upon one or more markers before determining the set of fragments.

14 . The method of claim 8 , wherein analyzing the portion of the text of that fragment based on the language associated with that fragment comprises selecting a language model of a set of pluggable language models.

15 . A non-transitory computer readable medium, comprising instructions for:

receiving a multilingual object, the multilingual object comprising a search query;

determining a set of fragments of the text of the multilingual object, where each of the set of fragments comprises a portion of the text of the multilingual object determined based on a multilingual sentence boundary detection model;

determining a language associated with each fragment of the determined set of fragments of the multilingual object;

determining a set of tokens for each fragment of the determined set of fragments, where determining the set of tokens for a fragment comprises analyzing the portion of the text of that fragment based on the language associated with that fragment;

indexing the set of tokens determined for each fragment in an index in association with the multilingual object such that the index is a multilingual index including tokens from the multilingual object in multiple languages determined according to the language associated with each fragment; and

performing a search of a search index using the determined set of tokens.

16 . The non-transitory computer readable medium of claim 15 , wherein the set of tokens are associated with a single field of the multilingual object.

17 . The non-transitory computer readable medium of claim 15 , wherein the indexed set of tokens are associated with a single field of the multilingual object.

18 . The non-transitory computer readable medium of claim 15 , where each of the set of fragments is a sentence.

19 . The non-transitory computer readable medium of claim 18 , wherein determining the set of fragments is done using a machine learning model trained on sentence markers for multiple languages.

20 . The non-transitory computer readable medium of claim 15 , further comprising presplitting the text of the multilingual object based upon one or more markers before determining the set of fragments.

21 . The non-transitory computer readable medium of claim 15 , wherein analyzing the portion of the text of that fragment based on the language associated with that fragment comprises selecting a language model of a set of pluggable language models.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 17, 2022
From: OBBARD, GEOFFREY MICHAEL
To: OPEN TEXT CORPORATION
Reel/Frame 061437/0107 →
Continuity (1)
Related Publication 20240119070A1 · Apr 11, 2024
References Cited (42)
US 6658377B1 · Anward et al. · 2003 [cited by applicant]
US 6842730B1 · Ejerhed et al. · 2005 [cited by applicant]
US 10762139B1 · Dai et al. · 2020 [cited by applicant]
US 12254032B2 · Obbard · 2025 [cited by applicant]
US 20060184516A1 · Ellis · 2006 [cited by applicant]
US 20070078654A1 · Moore · 2007 [cited by examiner]
US 20110131212A1 · Shikha · 2011 [cited by examiner]
US 20120095748A1 · Li et al. · 2012 [cited by applicant]
US 20120239378A1 · Parfentieva · 2012 [cited by examiner]
US 20130339378A1 · Zheng · 2013 [cited by examiner]
US 20170078199A1 · Mosko · 2017 [cited by examiner]
US 20170364503A1 · Anisimovich · 2017 [cited by examiner]
US 20170364510A1 · Huang et al. · 2017 [cited by applicant]
US 20180330012A1 · Hopkins · 2018 [cited by examiner]
US 20190102390A1 · Antunes · 2019 [cited by applicant]
US 20190108279A1 · Moore et al. · 2019 [cited by applicant]
US 20190147109A1 · Offer · 2019 [cited by applicant]
US 20190332619A1 · De Sousa Webber · 2019 [cited by applicant]
US 20190362003A1 · Zhang · 2019 [cited by applicant]
US 20210334299A1 · Sonntag et al. · 2021 [cited by applicant]
US 20240119076A1 · Obbard · 2024 [cited by applicant]
US 20240126795A1 · Zhong · 2024 [cited by applicant]
US 20240394942A1 · Shankhdhar · 2024 [cited by applicant]
US 20250131021A1 · Obbard · 2025 [cited by applicant]
US 20260004064A1 · Obbard · 2026 [cited by applicant]
EP 2807535 · 2019 [cited by applicant]
Eric Brill., “A Simple Rule-Based Part of Speech Tagge” ANLC '92: Proceedings of the third conference on Applied natural language processing, Mar. 1992, pp. 152-155. [cited by applicant]
Nguyen et al., “RDRPOSTagger: A Ripple Down Rules-based Part-Of-Speech Tagger”, Proceedings of the Demonstrations at the 14th Conference of the European Chapter of the Association for Computational Linguistics, pp. 17-2… [cited by applicant]
International Search Report and Written Opinion issued by the Canadian Intellectual Property Office as the International Searching Authority (CA/ISA) for International PCT Application No. PCT/IB2023/060075, mailed Dec. … [cited by applicant]
Office Action issued by the United States Patent and Trademark Office (USPTO) for U.S. Appl. No. 17/962,157, mailed Dec. 21, 2023, 13 pages. [cited by applicant]
LemmatizerME.Txt, Apache Software Foundation, retrieved at <<opennlp/opennlp-tools/src/main/java/opennlp/tools/lemmatizer/ LemmatizerME.java at main.apache/opennlp⋅GitHub>>, 7 pages. [cited by applicant]
Dat Quoc Nguyen et al., “Ripple Down Rules for Part-of-Speech Tagging,” In Proc. of 12th CICLing—vol. Part I, Feb. 2011, pp. 190-201. [cited by applicant]
Dat Quoc Nguyen, et al., “RDRPOSTAGGER: Ripple Down Rules-Based Part-of- Speech Tagger,” The Demonstrations at the14 [cited by applicant]
Eric Brill, “A Simple Rule-Based Part of Speech Tagger,” ANLC '92: Proceedings of the third conference on Applied natural language processing, Mar. 1992, pp. 152-155. [cited by applicant]
Grzegorz Chrupała, “Towards a Machine-Learning Architecture for Lexical Functional Grammar Parsing,” Dublin City University, Apr. 2008, 136 pages. [cited by applicant]
Office Action issued by the United States Patent and Trademark Office (USPTO) for U.S. Appl. No. 17/962,157, mailed Jun. 12, 2024, 13 pages. [cited by applicant]
Notice of Allowance issued by the United States Patent and Trademark Office (USPTO) for U.S. Appl. No. 17/962,157, mailed Sep. 27, 2024, 11 pages. [cited by applicant]
International Preliminary Report on Patentability issued by the International Bureau of WIPO for International PCT Application No. PCT/IB2023/060075, mailed Apr. 17, 2025, 7 pages. [cited by applicant]
Office Action issued for U.S. Appl. No. 18/990,446, mailed Sep. 17, 2025, 18 pages. [cited by applicant]
Office Action issued by the United States Patent and Trademark Office (USPTO) for U.S. Appl. No. 18/755,364, mailed Jan. 29, 2026, 18 pages. [cited by applicant]
Nguyen, Dat Quoc, et al., A Robust Transformation-Based Learning Approach using Ripple Down Rules for Part-of-Speech Tagging, Al Communications, Dec. 19, 2015, 14 pgs. [cited by applicant]
Kwok, Rex BH, “Translations of Ripple Down Rules into Logic Formalisms, ” Int'l Conference on Knowledge Engineering and Knowledge Management, Berlin, Heidelberg: Springer Berlin Heidelberg, 2000, 14 pgs. [cited by applicant]