IP Library Granted Patent US 12,596,879
Granted Patent B2
US 12,596,879 · App. 18/044,890 · Granted Apr 7, 2026

Method and system for identifying citations within regulatory content

Inventors: Mahdi Ramezani (Vancouver, CA); Elijah Solomon Krag (Vancouver, CA); Amir Abbas Tahmasbi (Vancouver, CA); Margery Moore (Salt Spring Island, CA)
Assignee: Intelex Technologies, ULC
G06F40/30G06F18/2148G06F40/284G06V10/764G06V10/82G06V30/147G06V30/19147G06V30/19173G06V30/413G06V30/414G06V30/416G06V30/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,596,879
App. No.
18/044,890
Granted
Apr 7, 2026
Kind
B2
Abstract

A computer-implemented system and method for identifying citations within regulatory content is disclosed. The method involves receiving image data representing a format and layout of the regulatory content, receiving a language embedding including a plurality of tokens representing words or characters in the regulatory content, and generating a token mapping associating each of the tokens with a portion of the image data. The method also involves receiving the plurality of tokens and token mapping at an input of a citation classifier, the citation classifier having been trained to generate a classification output for each token based on the language embedding and the token mapping, the classification output identifying a plurality of citation tokens within the plurality of tokens. The method further involves processing the plurality of citation tokens to determine a hierarchical relationship between citation tokens, the hierarchical relationship being established based at least in part on the token mapping for the citation tokens.

Claims (41)

1 . A computer-implemented method for identifying citations within regulatory content based on image data representing the regulatory content, the method comprising:

receiving a language embedding including a plurality of tokens representing words or characters in the regulatory content;

generating a token mapping associating each of the tokens with a portion of the image data representing the regulatory content;

receiving the plurality of tokens and the token mappings at an input of a citation classifier, the citation classifier having been trained to generate a classification output for each of the tokens based on the language embedding and the token mapping of that token, wherein the classification output identifies a plurality of citation tokens from the tokens; and

processing the tokens based on the classification output of each of the tokens to determine a hierarchical relationship between the tokens, the hierarchical relationship being established based at least in part on the token mapping for at least some of the tokens,

wherein processing the classification output of each of the tokens comprises processing the plurality of citation tokens to determine the hierarchical relationship between the plurality of citation tokens; and

wherein processing the plurality of citation tokens comprises receiving pairwise combinations of citation tokens at an input of a sibling classifier, the sibling classifier having been trained to output a probability indicative of whether each pairwise combination of citation tokens has a common hierarchical level.

2 . The method of claim 1 further comprising receiving text data representing characters in the regulatory content and generating the language embedding based on the received text data.

3 . The method of claim 2 wherein generating the language embedding comprises generating the language embedding using a pretrained language model to process the received text data, the pretrained language model being operably configured to generate the tokens.

4 . The method of claim 2 wherein generating the language embedding comprises:

processing the image data to identify regions of interest within the regulatory content, each region of interest including a plurality of characters; and

wherein generating the language embedding comprises generating the language embedding only for text data associated with the regions of interest.

5 . The method of claim 4 wherein processing the image data to identify the regions of interest comprises receiving the image data at an input of a region of interest classifier, the region of interest classifier having been trained to identify regions within the image data that are associated with regulatory content that should not be processed to identify citation tokens.

6 . The method of claim 5 wherein the region of interest classifier is pre-trained using a plurality of images of portions of the regulatory content that should not be processed to identify citation tokens.

7 . The method of claim 1 wherein the image data comprises a plurality of image pages and generating the token mapping comprises generating a tensor for each of the image pages, the tensor having first and second dimensions corresponding to image pixels in that image page and a third dimension in which language embedding values are associated with portions of that image page, the tensor providing the input to the citation classifier.

8 . The method of claim 1 wherein receiving the image data comprises receiving text data including format data representing a layout of the regulatory content and generating an image data representation of the text data formatted with the format data.

9 . The method of claim 1 wherein the image data comprises a plurality of page images and wherein generating the token mapping comprises:

sizing each of the page images to match a common page size;

for each of the tokens, establishing a bounding box within an applicable page image of the page images that identifies a portion of the applicable page image corresponding to that token; and

determining at least a bounding box location for each of the tokens, the bounding box location being indicative of an indentation or position of that token within the applicable page image.

10 . The method of claim 9 further comprising processing the portion of the applicable page image within the bounding box to determine at least one of a font size associated with the token or a formatting associated with the token, and wherein the citation classifier is trained to generate the classification output based on at least one of the determined font size or the determined formatting associated with the token.

11 . The method of claim 1 further comprising:

generating a similarity matrix including a plurality of rows corresponding to the plurality of citation tokens and a plurality of columns corresponding to the plurality of citation tokens, the similarity matrix being populated with the probabilities determined by the sibling classifier for each of the pairwise combinations of citation tokens; and

processing the similarity matrix to generate a hierarchical level for each of the plurality of citation tokens.

12 . The method of claim 1 wherein the plurality of citation tokens each include either a citation number or citation title, and wherein a plurality of remaining tokens not identified as the plurality of citation tokens include body text, and further comprising associating the body text between two consecutive citation tokens with a first citation token of the two consecutive citation tokens.

13 . A system for identifying citations within regulatory content based on image data representing the regulatory content, the system comprising:

a token mapper operably configured to:

receive a language embedding including a plurality of tokens representing words or characters in the regulatory content; and

generate a token mapping associating each of the tokens with a portion of the image data representing the regulatory content;

a citation classifier operably configured to receive the plurality of tokens and the token mappings at an input of a citation classifier, the citation classifier having been trained to generate a classification output for each of the tokens based on the language embedding and the token mapping of that token, wherein the classification output identifies a plurality of citation tokens within the tokens;

a hierarchical relationship classifier operably configured to process the tokens based on the classification output of each of the tokens to determine a hierarchical relationship between the tokens, the hierarchical relationship being established based at least in part on the token mapping for at least some of the tokens,

wherein the hierarchical relationship classifier is operably configured to process the tokens based on the classification output of each of the tokens by processing the plurality of citation tokens to determine the hierarchical relationship between the plurality of citation tokens; and

wherein the hierarchical relationship classifier comprises a sibling classifier and is operably configured to process the plurality of citation tokens by receiving pairwise combinations of citation tokens at an input of the sibling classifier, the sibling classifier having been trained to output a probability indicative of whether each pairwise combination of citation tokens have a common hierarchical level; and

one or more processor circuits having a memory for storing codes, the codes being operable to direct the one or more processor circuits to implement each of the token mapper, the citation classifier, and the hierarchical relationship classifier.

14 . The system of claim 13 wherein the image data comprises a plurality of page images and wherein the token mapper is further operably configured to:

size each of the page images to match a common page size;

for each of the tokens, establish a bounding box within an applicable page image of the page images that identifies a portion of the applicable page image corresponding to that token; and

determine at least a bounding box location for each of the tokens, the bounding box location being indicative of an indentation or position of that token within the applicable page image.

15 . The system of claim 13 wherein the hierarchical relationship classifier is further operably configured to:

generate a similarity matrix including a plurality of rows corresponding to the plurality of citation tokens and a plurality of columns corresponding to the plurality of citation tokens, the similarity matrix being populated with the probabilities determined by the sibling classifier for each of the pairwise combinations of citation tokens; and

process the similarity matrix to generate a hierarchical level for each of the plurality of citation tokens.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 1, 2024
From: MOORE & GASPERECZ GLOBAL INC.
To: INTELEX TECHNOLOGIES, ULC
Reel/Frame 066619/0902 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: RAMEZANI, MAHDI; KRAG, ELIJAH SOLOMON; TAHMASBI, AMIR ABBAS; MOORE, MARGERY
To: MOORE & GASPERECZ GLOBAL INC.
Reel/Frame 064086/0264 →
Continuity (2)
Continuation 17017406 · Sep 10, 2020
Related Publication 20240013005A1 · Jan 11, 2024
References Cited (189)
US 7571149B1 · Yanosy, Jr. · 2009 [cited by applicant]
US 7778954B2 · Rhoads et al. · 2010 [cited by applicant]
US 7823120B2 · Kazakov et al. · 2010 [cited by applicant]
US 8156010B2 · Gopalakrishnan · 2012 [cited by applicant]
US 8266184B2 · Liu et al. · 2012 [cited by applicant]
US 8306819B2 · Liu et al. · 2012 [cited by applicant]
US 8447788B2 · Liu et al. · 2013 [cited by applicant]
US 8510345B2 · Liu et al. · 2013 [cited by applicant]
US 8676733B2 · Heisele · 2014 [cited by examiner]
US 8693790B2 · Jiang et al. · 2014 [cited by applicant]
US 8788523B2 · Martin et al. · 2014 [cited by applicant]
US 8897563B1 · Welling et al. · 2014 [cited by applicant]
US 9053179B2 · Zhang et al. · 2015 [cited by applicant]
US 9292483B2 · Deyab et al. · 2016 [cited by applicant]
US 9324025B2 · Byron et al. · 2016 [cited by applicant]
US 9460076B1 · Barba · 2016 [cited by applicant]
US 9811518B2 · Leidner et al. · 2017 [cited by applicant]
US 9953281B2 · Wiig et al. · 2018 [cited by applicant]
US 9972016B2 · Ghaisas et al. · 2018 [cited by applicant]
US 10013655B1 · Clark · 2018 [cited by applicant]
US 10516902B1 · Manoria et al. · 2019 [cited by applicant]
US 10565502B2 · Scholtes · 2020 [cited by applicant]
US 10853696B1 · Luo et al. · 2020 [cited by applicant]
US 10956673B1 · Ramezani · 2021 [cited by examiner]
US 11005888B2 · Wang et al. · 2021 [cited by applicant]
US 11042692B1 · Greisen et al. · 2021 [cited by applicant]
US 11080486B2 · Hao et al. · 2021 [cited by applicant]
US 11194784B2 · Brisimi et al. · 2021 [cited by applicant]
US 11194963B1 · Schafer et al. · 2021 [cited by applicant]
US 11232358B1 · Ramezani et al. · 2022 [cited by applicant]
US 11275936B2 · Neumann · 2022 [cited by applicant]
US 11314922B1 · Ramezani et al. · 2022 [cited by applicant]
US 11763321B2 · Gasperecz et al. · 2023 [cited by applicant]
US 11823477B1 · Ramezani et al. · 2023 [cited by applicant]
US 20020080196A1 · Bornstein et al. · 2002 [cited by applicant]
US 20040039594A1 · Narasimhan et al. · 2004 [cited by applicant]
US 20050160263A1 · Naizhen et al. · 2005 [cited by applicant]
US 20050188072A1 · Lee et al. · 2005 [cited by applicant]
US 20050198098A1 · Levin et al. · 2005 [cited by applicant]
US 20060088214A1 · Handley et al. · 2006 [cited by applicant]
US 20060117012A1 · Rizzolo et al. · 2006 [cited by applicant]
US 20070092140A1 · Handley · 2007 [cited by applicant]
US 20070226356A1 · Levin et al. · 2007 [cited by applicant]
US 20080216169A1 · Naizhen et al. · 2008 [cited by applicant]
US 20080263505A1 · StClair et al. · 2008 [cited by applicant]
US 20100153473A1 · Gopalan · 2010 [cited by applicant]
US 20100228548A1 · Liu et al. · 2010 [cited by applicant]
US 20110258182A1 · Singh · 2011 [cited by examiner]
US 20110289550A1 · Nakae · 2011 [cited by applicant]
US 20120011455A1 · Subramanian et al. · 2012 [cited by applicant]
US 20130236110A1 · Barrus · 2013 [cited by applicant]
US 20140160528A1 · Bloch et al. · 2014 [cited by applicant]
US 20160042251A1 · Cordova-Diba et al. · 2016 [cited by applicant]
US 20160063322A1 · Déjean et al. · 2016 [cited by applicant]
US 20160350885A1 · Clark · 2016 [cited by applicant]
US 20170235848A1 · Van Dusen et al. · 2017 [cited by applicant]
US 20170264643A1 · Bhuiyan et al. · 2017 [cited by applicant]
US 20180075554A1 · Clark et al. · 2018 [cited by applicant]
US 20180137107A1 · Buccapatnam Tirumala et al. · 2018 [cited by applicant]
US 20180197111A1 · Crabtree et al. · 2018 [cited by applicant]
US 20180225471A1 · Goyal et al. · 2018 [cited by applicant]
US 20190098054A1 · Ramachandran et al. · 2019 [cited by applicant]
US 20190147388A1 · Alexander · 2019 [cited by applicant]
US 20190251397A1 · Tremblay et al. · 2019 [cited by applicant]
US 20190279111A1 · Merrill et al. · 2019 [cited by applicant]
US 20190377785A1 · N et al. · 2019 [cited by applicant]
US 20200004877A1 · Ghafourifar et al. · 2020 [cited by applicant]
US 20200019767A1 · Porter et al. · 2020 [cited by applicant]
US 20200050620A1 · Clark et al. · 2020 [cited by applicant]
US 20200074515A1 · Ghatage et al. · 2020 [cited by applicant]
US 20200082204A1 · Beaver · 2020 [cited by applicant]
US 20200111023A1 · Pondicherry Murugappan et al. · 2020 [cited by applicant]
US 20200125659A1 · Brisimi et al. · 2020 [cited by applicant]
US 20200213365A1 · Ramachandran et al. · 2020 [cited by applicant]
US 20200233862A1 · Shaked et al. · 2020 [cited by applicant]
US 20200279271A1 · Gasperecz et al. · 2020 [cited by applicant]
US 20200311201A1 · Sainani et al. · 2020 [cited by applicant]
US 20200320349A1 · Yu et al. · 2020 [cited by applicant]
US 20200337625A1 · Aimone et al. · 2020 [cited by applicant]
US 20200349920A1 · Al Bawab et al. · 2020 [cited by applicant]
US 20210089767A1 · Ashek et al. · 2021 [cited by applicant]
US 20210117621A1 · Sharpe et al. · 2021 [cited by applicant]
US 20210124800A1 · Williams · 2021 [cited by applicant]
US 20210209358A1 · Alikhani et al. · 2021 [cited by applicant]
US 20210319173A1 · Gerber et al. · 2021 [cited by applicant]
US 20220147814A1 · Ramezani et al. · 2022 [cited by applicant]
US 20230177281A1 · Kamath et al. · 2023 [cited by applicant]
US 20230419110A1 · Ramezani et al. · 2023 [cited by applicant]
CA 2256408C · 1997 [cited by applicant]
CA 2381460A1 · 2001 [cited by applicant]
CA 2410881A1 · 2001 [cited by applicant]
CA 2419377A1 · 2002 [cited by applicant]
CA 2699397A1 · 2009 [cited by applicant]
CA 2699644A1 · 2009 [cited by applicant]
CA 3124358A1 · 2021 [cited by applicant]
CA 3210419A1 · 2023 [cited by applicant]
CA 3124358C · 2023 [cited by applicant]
CA 3210419C · 2024 [cited by applicant]
CN 110705223A · 2020 [cited by applicant]
EP 2993617A1 · 2016 [cited by applicant]
RU 2628431C1 · 2017 [cited by applicant]
WO 2022051838A1 · 2022 [cited by applicant]
WO 2022094723A1 · 2022 [cited by applicant]
WO 2022094724A1 · 2022 [cited by applicant]
Denk T.I. et al., “BERTgrid: Contextualized Embedding for 2D Document Representation and Understanding.,” arXiv: Computation and Language, pp. 1-4, 2019. [cited by applicant]
Locke, Daniel et al., Automatic cited decision retrieval: Working notes of lelab for FIRE Legal Track Precedence Retrieval Task, Queensland University of Technology, 2017, pp. 1-2. [cited by applicant]
Locke, Daniel et al., Towards Automatically Classifying Case Law Citation Treatment Using Neural Networks. The University of Queensland, Dec. 2019, pp. 1-8. [cited by applicant]
Sadeghian, Ali et al., “Automatic Semantic Edge Labeling Over Legal Citation Graphs” , Springer Science+Business Media B.V., Mar. 1, 2018, pp. 128-144. [cited by applicant]
Young T. et al., “Recent Trends in Deep Learning Based Natural Language Processing [Review Article],” IEEE Computational Intelligence Magazine, vol. 13, No. 3, pp. 1-32, 2018. [cited by applicant]
Radford A. et al., “Improving Language Understanding by Generative Pre-Training”, pp. 1-12, 2018. [cited by applicant]
Devlin, J. et al. “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, arXiv: Computation and Language, pp. 1-17, 2018. [cited by applicant]
Lan, Z. et al., “Albert: A Lite BERT for Self-Supervised Learning of Language Representations”, arXiv: Computation and Language, pp. 1-17, 2019. [cited by applicant]
Liu, Y. et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach”, arXiv: Computation and Language, pp. 2-13, 2019. [cited by applicant]
Sanh, V. et al., “DistilBERT, a Distilled Version of BERT: smaller, faster, cheaper and lighter” , arXiv: Computation and Language, pp. 1-5, 2019. [cited by applicant]
Lee, J. et al., “BioBERT: a pre-trained biomedical language representation model for biomedical text mining” , Bioinformatics, pp. 1-7, 2019. [cited by applicant]
Peters, M. E. et al., “To Tune or Not to Tune? Adapting Pretrained Representations to Diverse Tasks” , arXiv: Computation and Language, pp. 1-8, 2019. [cited by applicant]
Siblini, C. et al., “Multilingual Question Answering from Formatted Text applied to Conversational Agents”, arXiv: Computation and Language, pp. 1-10, 2019. [cited by applicant]
Lai, A. et al., “Natural Language Inference from Multiple Premises”, arXiv: Computation and Language, pp. 1-10, 2017. [cited by applicant]
Elwany, E. et al.; “BERT Goes to Law School: Quantifying the Competitive Advantage of Access to Large Legal Corpora in Contract Understanding”, arXiv: Computation and Language, 2019, pp. 1-4. [cited by applicant]
Nguyen, T. et al.; “Recurrent neural network-based models for recognizing requisite and effectuation parts in legal texts”, Springer Science+Business Media B.V., 2018, pp. 169-199. [cited by applicant]
Chalkidis, I. et al.; “Deep learning in law: early adaptation and legal word embeddings trained on large corpora”, Springer Nature B.V. 2018, pp. 171-198. [cited by applicant]
Tang, G. et al.; “Matching law cases and reference law provision with a neural attention model”, 2016, pp. 1-4. [cited by applicant]
Kingston, J.; “Using artificial intelligence to support compliance with the general data protection regulation”, Springer Science+Business Media B.V., 2017 pp. 429-443. [cited by applicant]
Fatema et al.; “A Semi-Automated Methodology for Extracting Access Control Rules from the European Data Protection Directive”, IEE Computer Society, 2016, pp. 25-32. [cited by applicant]
Mandal, S. et al.; “Modular Norm Models: A Lightweight Approach for Modeling and Reasoning about Legal Compliance”, IEE Computer Society, 2017, pp. 657-662. [cited by applicant]
Hashmi M.; “A Methodology for Extracting Legal Norms from Regulatory Documents”, IEEE Computer Society, 2015, pp. 41-50. [cited by applicant]
Kiyavitskaya, N. et al.; “Automating the Extraction of Rights and Obligations for Regulatory Compliance”, Conceptual Modeling, 2008, pp. 154-168. [cited by applicant]
Islam, M. et al. “RuleRS: a rule-based architecture for decision support systems”, Spring Science+Business Media B. V., 2018, pp. 315-344. [cited by applicant]
Yang X. et al., “Learning to Extract Semantic Structure from Documents Using Multimodal Fully Convolutional Neural Networks,” The Pennsylvania State University, pp. 1-16, 2017. [cited by applicant]
Katti A. et al.; “Chargrid: Towards Understanding 2D Documents,” arXiv: Computation and Language, 2018, pp. 1-11. [cited by applicant]
Soviany P. et al., “Optimizing the Trade-off between Single-Stage and Two-Stage Object Detectors using Image Difficulty Prediction,” arXiv: Computer Vision and Pattern Recognition, pp. 2-6, 2018. [cited by applicant]
Tsai H. et al., “Small and Practical BERT Models for Sequence Labeling,” pp. 9-11, 2019. [cited by applicant]
Zhong, Haoxi, et al; “How does NLP benefit legal system: A summary of legal artificial intelligence.” arXiv preprint arXiv: 2004.12158v5, 2020. [cited by applicant]
Bach, Ngo Xuan, et al.; “Reference extraction from Vietnamese legal documents”, Proceedings of the Tenth International Symposium on Information and Communication Technology, 2019. pp. 486-493. [cited by applicant]
Chalkidis, Ilias; Ion Androutsopoulos and Achilleas Michos; “Extracting contract elements”, Proceedings of the 16th edition of the International Conference on Artificial Intelligence and Law, 2017. [cited by applicant]
Chalkidis, Ilias; Ion Androutsopoulos and Achilleas Michos; “Obligation and prohibition extraction using hierarchical Rnns”, arXiv preprint arXiv:1805.03871v1, 2018. [cited by applicant]
International Search Report and Written Opinion issued by the Canadian Intellectual Property Office in connection with International Patent Application No. PCT/CA2021/051130 dated Nov. 1, 2021, 10 pages. [cited by applicant]
International Search Report and Written Opinion issued by the Canadian Intellectual Property Office in connection with International Patent Application No. PCT/CA2021/051586 dated Jan. 25, 2022, 10 pages. [cited by applicant]
Non-Final Office Action issued by the U.S. Patent and Trademark Office on Jul. 11, 2022 in connection with U.S. Appl. No. 16/562,589, 39 pages. [cited by applicant]
Yin, Pengcheng et al.; “TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data”, arXiv: 2005.08314v1 [cs.CL] May 17, 2020. [cited by applicant]
Deng, Xiang et al.; “TURL: Table Understanding through Representation Learning”, arXiv:2006.14806v2 [cs.IR] Dec. 3, 2020. [cited by applicant]
Du, Lun et al.; “TabularNet: A Neural Network Architecture for Understanding Semantic Structures of Tabular Data”, arXiv:2106.03096v2 [cs.LG] Jun. 16, 2021. [cited by applicant]
Herzig, Jonathan et al.; “TAPAS: Weakly Supervised Table Parsing via Pre-training”, arXiv:2004.02349v2 [cs.IR] Apr. 21, 2020. [cited by applicant]
Yang, Jingfeng et al.; “TABLEFORMER: Robust Transformer Modeling for Table-Text Encoding”, arXiv:2203.00274v1 [cs.CL] Mar. 1, 2022. [cited by applicant]
Herzig, Jonathan et al.; “Open Domain Question Answering over Tables via Dense Retrieval”, arXiv:2103.12011v2 [cs. CL] Jun. 9, 2021. [cited by applicant]
Majumder, Bodhisattwa Prasad et al.; “Representation Learning for Information Extraction from Form-like Documents”, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 6495-6504,… [cited by applicant]
Eisenschlos, Julian Martin et al.; “Understanding tables with intermediate pre-training”, arXiv:2010.00571v2 [cs.CL] Oct. 5, 2020. [cited by applicant]
Erickson, Nick et al.; “AutoGluon-Tabular: Robust and Accurate AutoML for Structured Data”, arXiv:2003.06505v1 [stat.ML] Mar. 13, 2020. [cited by applicant]
Logan IV, Robert L. et al.; “Multimodal Attribute Extraction”, arXiv:1711.11118v1 [cs.CL] Nov. 29, 2017. [cited by applicant]
Rebuffel, Clément et al., “A Hierarchical Model for Data-to-Text Generation”, arXiv:1912.10011v1 [cs.CL] Dec. 20, 2019. [cited by applicant]
Ma, Shuming et al.; “Key Fact as Pivot: A Two-Stage Model for Low Resource Table-to-Text Generation”, arXiv:1908.03067v1 [cs.CL] Aug. 8, 2019. [cited by applicant]
Zhang, Ziqi; “Towards Efficient and Effective Semantic Table Interpretation”, P. Mika et al. (Eds.) ISWC 2014, Part I, LNCS 8796, pp. 487-505, 2014, Springer International Publishing Switzerland. [cited by applicant]
Bhagavatula, Chandra Sekhar et al.; “TabEL: Entity Linking in Web Tables”, Northwestern University Evanston IL 60201, USA. [cited by applicant]
International Search Report and Written Opinion issued by the Canadian Intellectual Property Office in connection with International Patent Application No. PCT/CA2021/051585 dated Feb. 16, 2022, 7 pages. [cited by applicant]
Zhang, Shuo et al.; “Web Table Extraction, Retrieval and Augmentation: A Survey”, arXiv:2002.00207v2 [cs. IR] Feb. 5, 2020. [cited by applicant]
Ritze, Dominique et al.; Matching HTML Tables to Dbpedia; Wims 2015 Limassol, Cyprus. [cited by applicant]
Efthymiou, Vasilis et al.; “Matching Web Tables with Knowledge Base Entities: From Entity Lookups to Entity Embeddings”, pp. 1-16. [cited by applicant]
Jiménez-Ruiz, Ernesto et al.; “SemTab 2019: Resources to Benchmark Tabular Data to Knowledge Graph Matching Systems”, A Harth et al. (Eds.): ESWC 2020, LNCS 12123, pp. 514-530, 2020, Springer Nature Switzerland AG. [cited by applicant]
Zhang, Ziqi; “Effective and Efficient Semantic Table Interpretation using TableMiner”, Semantic Web tbd (2016) pp. 1-39, IOS Press. [cited by applicant]
Chen, Shuang et al.; “LinkingPark: An Integrated Approach for Semantic Table Interpretation”, CEUR-WS.org/Vol-2775/paper7, pp. 1-10. [cited by applicant]
Shigarov A. et al.; “TabbyXL: Software platform for rule-based spreadsheet data extraction and transformation”, Elsevier, Software X 10 (2019) 100270, pp. 1-6. [cited by applicant]
Shigarov A. et al.; “Rule-based spreadsheet data transformation from arbitrary to relational tables”, Elsevier, Informational Systems 71 (2017) pp. 123-136. [cited by applicant]
Xu, Yiheng et al.; “LayoutLM: Pre-training of Text and Layout for Document Image Understanding”, arXiv:1912.13318v5 [cs.CL] Jun. 16, 2020. [cited by applicant]
Dorodnykh, Nikita O. et al.; “Towards A Unviersal Approach for Semantic Interpretation of Spreadsheets Data”, IDEAS 2020, Aug. 12-14, 2020, Seoul, Republic of Korea. [cited by applicant]
Non-Final Office Action issued by the U.S. Patent and Trademark Office on Nov. 16, 2022, in connection with U.S. Appl. No. 17/898,788, 11 pages. [cited by applicant]
Non-Final Office Action issued by the U.S. Patent and Trademark Office on Mar. 27, 2023, in connection with U.S. Appl. No. 17/898,788, 27 pages. [cited by applicant]
International Preliminary Report on Patentability issued by the Canadian Intellectual Property Office in connection with International Patent Application No. PCT/CA2021/051130, Mar. 7, 2023, 7 pages. [cited by applicant]
International Preliminary Report on Patentability issued by the Canadian Intellectual Property Office in connection with International Patent Application No. PCT/CA2021/051585, May 8, 2023, 5 pages. [cited by applicant]
International Preliminary Report on Patentability issued by the Canadian Intellectual Property Office in connection with International Patent Application No. PCT/CA2021/051586, May 8, 2023, 7 pages. [cited by applicant]
“360factors introducing new innovative compliance activities management solution.”, https://www.360factors.com/regulatory-compliance-aba-conference-nashville/ (accessed on Jun. 18, 2024), 3 pages. [cited by applicant]
“Confusion matrix”, Wikipedia, https://en.wikipedia.org/wiki/Confusion_matrix (accessed on Jun. 18, 2024), 7 pages. [cited by applicant]
“Decision tree learning”, Wikipedia, https://en.wikipedia.org/wiki/Decision_tree_learning (accessed Jun. 18, 2024), 13 pages. [cited by applicant]
“Feature selection”, Wikipedia, https://en.wikipedia.org/wiki/Feature_selection (accessed on Jun. 18, 2024), 16 pages. [cited by applicant]
“Gradient boosting”, Wikipedia, https://en.wikipedia.org/wiki/Gradient_boosting (accessed on Jun. 18, 2024), 10 pages. [cited by applicant]
“Linguistic annotations”, https://spacy.io/usage/spacy-101#annotations (accessed on Jun. 18, 2024), 1 page. [cited by applicant]
“Named Entities”, https://spacy.io/usage/spacy-101#annotations-ner (accessed on Jun. 18, 2024), 1 page. [cited by applicant]
“Part-of-speech tags and dependencies”, https://spacy.io/usage/spacy-101#annotations-pos-deps (accessed on Jun. 18, 2024), 1 page. [cited by applicant]
“Pipelines”, https://spacy.io/usage/spacy-101#pipelines (accessed on Jun. 18, 2024), 1 page. [cited by applicant]
“Tokenization”, https://spacy.io/usage/spacy-101#annotations-token (accessed on Jun. 18, 2024), 2 pages. [cited by applicant]
“European Application Serial No. 21865414.3, Extended European Search Report mailed Jun. 20, 2024”, Intelex Technologies, ULC, 6 pages. [cited by applicant]
“European Application Serial No. 21887966.6, Extended European Search Report mailed Oct. 7, 2024”, Intelex Technologies, ULC, 10 pages. [cited by applicant]
Chalkidis, et al., “LEGAL-BERT: The Muppets straight out of Law School”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853,, Oct. 6, 2020, 7 pages. [cited by applicant]
Devlin, et al., “BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding”, May 24, 2019, 16 pages. [cited by applicant]
Hamdaqa, et al., “An approach based on citation analysis to support effective handling of regulatory compliance”, Future Generation Computer Systems, Elsevier Science Publishers. Amsterdam, NL, vol. 27, No. 4, Apr. 1, 2… [cited by applicant]
Kiperwasser, et al., “Simple and accurate dependency parsing using bidirectional LSTM feature representation”, Trans, of the Assn, for Computational Linguistics, vol. 4, 2016, pp. 313-327. [cited by applicant]
Le, et al., “Requirement text detection from contract packages to support project definition determination”, Advances in Informatics and Computing in Civil and Construction Engineering: Proc. of the 35th CIB W78 Confere… [cited by applicant]
Shaheen, et al., “Large Scale Legal Text Classification Using Transformer Models”, Oct. 24, 2020, 11 pages. [cited by applicant]
Zhang, et al., “A machine learning-based method for building code requirement hierarchy extraction”, CSCE Annual Conference, 2019, 10 pages. [cited by applicant]
Zhou, et al., “Onotology-based automated information extraction from building energy conservation codes”, Automation in Construction vol. 74 ., 2017, pp. 103-117. [cited by applicant]