IP Library › Granted Patent US 12,518,111
Granted Patent B2
US 12,518,111 · App. 18/452,803 · Granted Jan 6, 2026

Automating large-scale data collection

Inventors: Paria Jamshid Lou (Sydney, AU); Gioacchino Tangari (Sydney, AU); Jason Black (Redmond, WA); Bhagya Gayathri Hettige (Melbourne, AU); Xu Zhong (Melbourne, AU); Poorya Zaremoodi (Melbourne, AU); Thanh Long Duong (Seabrook, AU); Mark Edward Johnson (Sydney, AU)
Assignee: Oracle International Corporation
G06F40/40G06F40/284G06F40/289G10L15/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,518,111
App. No.
18/452,803
Granted
Jan 6, 2026
Kind
B2
Abstract

Obtaining collections of sentences in different languages that are usable for training models in various applications of artificial intelligence is provided. A method is provided that obtains, from text corpus, webpages in a plurality of languages, each of the webpages corresponding to an URL; obtains annotations for each of the webpages based on its URL, to obtain annotated data entries corresponding to the webpages, each of the annotated data entries including a classification label corresponding to a sub-topic of one of a plurality of topics, where each of the plurality of topics includes a corresponding plurality of sub-topics; filters the annotated data entries to obtain topic-specific content in a target language based on the classification labels, the topic-specific content corresponding to one or more sub-topics; performs post-processing on the topic-specific content to obtain result data; and outputs the result data for the topic.

Claims (79)

1 . A computer-implemented method comprising:

(a) obtaining, by a dataset generation system from a text corpus, a plurality of webpages in a plurality of languages, each of the plurality of webpages corresponding to a respective universal resource locator (URL);

(b) annotating, by the dataset generation system, each of the plurality of webpages based on the respective URL, to obtain annotated data entries, each of the annotated data entries corresponding to one of the plurality of webpages and comprising a classification label corresponding to a sub-topic of one of a plurality of topics, wherein each of the plurality of topics comprises a corresponding plurality of sub-topics;

(c) performing, by the dataset generation system, a first filtering for the annotated data entries to obtain topic-specific content in a target language among the plurality of languages by:

comparing one or more sub-topics among the corresponding plurality of sub-topics with classification labels correspondingly included in the annotated data entries, wherein the one or more sub-topics are associated with a topic among the plurality of topics, and

obtaining, from the annotated data entries, the topic-specific content corresponding to the one or more sub-topics based on the comparing; and

(d) performing, by the dataset generation system, post-processing on the topic-specific content to obtain result data, by performing at least one from among a second filtering and a normalizing on sentences included in the topic-specific content,

wherein (c) and (d) are performed for a plurality of different target languages among the plurality of languages, the target language being a first target language of the plurality of different target languages;

outputting, by the dataset generation system, the result data for each of the plurality of different target languages as a plurality of model training datasets for the topic for which (c) and (d) were performed;

transmitting, by the dataset generation system to a model training system, at least a first model training dataset among the plurality of model training datasets, the first model training dataset corresponding to the first target language;

training, by the model training system, a first machine learning model using the first model training dataset in the first target language, the training comprising inputting, into the first machine learning model, training examples from the first model training dataset, to obtain a set of values for model parameters for a first speech recognition model; and

generating, by the model training system, the first speech recognition model in the first target language, for the topic for which (c) and (d) were performed, wherein the first speech recognition model is configured with the set of values for a set of predetermined sentiments or a set of predetermined intents,

wherein the first speech recognition model is configured to, based on one or more sentences provided as an input by a user in the first target language, identify a certain sentiment from the set of predetermined sentiments or a certain intent from the set of predetermined intents, and output, for the one or more sentences, a prediction of the certain sentiment or a prediction of the certain intent.

2 . The computer-implemented method of claim 1 , wherein the performing the first filtering further comprises:

comparing keywords of a predetermined lexicon for the topic with words in sentences included in one from among the topic-specific content and the annotated data entries; and

obtaining the sentences containing the keywords, to generate topic-specific, lexicon-specific content,

wherein the performing the post-processing on the topic-specific content comprises performing the post-processing on the topic-specific, lexicon-specific content.

3 . The computer-implemented method of claim 1 , wherein the performing the first filtering further comprises:

obtaining, from the topic-specific content, variation-specific content corresponding to a variation of the target language, by searching URLs of webpages corresponding to the topic-specific content for an indication of a geographic locality corresponding to the variation,

wherein the performing the post-processing on the topic-specific content comprises performing the post-processing on the variation-specific content.

4 . The computer-implemented method of claim 1 , wherein the obtaining the plurality of webpages further comprises:

detecting a language of the plurality of webpages, respectively;

computing a language confidence score for the detected language; and

discarding, from the plurality of webpages, webpages whose language confidence score is less than a predetermined threshold, prior to the annotating.

5 . The computer-implemented method of claim 1 , wherein:

the obtaining the plurality of webpages further comprises tokenizing raw content included in each of the plurality of webpages into sentences, and

the sentences included in the topic-specific content are tokenized sentences.

6 . The computer-implemented method of claim 1 , wherein the performing the post-processing further comprises:

performing the second filtering on the sentences included in the topic-specific content to remove, from the topic-specific content, at least one from among an incomplete sentence and a grammatically incorrect sentence, to obtain a set of high-quality sentences; and

normalizing the set of high-quality sentences by removing, from at least one sentence of the set of high-quality sentences, at least one from among a bullet point and an emoji.

7 . The computer-implemented method of claim 1 , further comprising:

prior to the performing the post-processing, verifying the topic-specific content by extracting, from the topic-specific content, webpages corresponding to at least one from among a verified source domain, a verified publication, and a verified public source, to obtain the topic-specific content that is verified.

8 . The computer-implemented method of claim 1 , wherein:

each of the annotated data entries comprises a raw content section comprising raw content and a sentences section comprising sentences obtained by tokenizing the raw content, and

the annotating further comprises:

inserting tags in each of the annotated data entries, wherein the tags identify, for a webpage among the plurality of webpages, at least two from among the respective URL, the classification label for the URL, the target language, a target language confidence score, a webpage title, the raw content section, and the sentences section.

9 . The computer-implemented method of claim 1 , wherein:

(c) and (d) are performed for the plurality of topics in the plurality of different target languages, and

the outputting the result data further comprises providing sets of result data that are obtained for each of the plurality of topics in the plurality of languages as a plurality of training datasets for building speech recognition models in the plurality of different target languages for the plurality of topics.

10 . The computer-implemented method of claim 1 , wherein the performing the post-processing further comprises:

performing the second filtering on the sentences of the topic-specific content to identify sentences for a task-specific requirement,

wherein the task-specific requirement is a sentiment analysis.

11 . The computer-implemented method of claim 1 , further comprising:

prior to the annotating, receiving a first user selection input providing an identification of the target language; and

prior to the performing the first filtering, receiving a second user selection input providing an identification of the topic for which (c) and (d) are to be performed.

12 . The computer-implemented method of claim 11 , wherein the first user selection input and the second user selection input are provided by a customer of a cloud service provider.

13 . A system comprising:

a model training system; and

a dataset generation system comprising one or more processors, and one or more computer-readable media storing instructions which, when executed by the one or more processors, cause the system to perform operations including:

(a) obtaining, from a text corpus, a plurality of webpages in a plurality of languages, each of the plurality of webpages corresponding to a respective universal resource locator (URL),

(b) annotating each of the plurality of webpages based on the respective URL, to obtain annotated data entries, each of the annotated data entries corresponding to one of the plurality of webpages and comprising a classification label corresponding to a sub-topic of one of a plurality of topics, wherein each of the plurality of topics comprises a corresponding plurality of sub-topics,

(c) performing a first filtering for the annotated data entries to obtain topic-specific content in a target language among the plurality of languages by:

comparing one or more sub-topics among the corresponding plurality of sub-topics with classification labels correspondingly included in the annotated data entries, wherein the one or more sub-topics are associated with a topic among the plurality of topics, and obtaining, from the annotated data entries, the topic-specific content corresponding to the one or more sub-topics based on the comparing, and

(d) performing post-processing on the topic-specific content to obtain result data, by performing at least one from among a second filtering and a normalizing on sentences included in the topic-specific content,

wherein (c) and (d) are performed for a plurality of different target languages among the plurality of languages, the target language being a first target language of the plurality of different target languages,

outputting the result data for each of the plurality of different target languages as a plurality of model training datasets for the topic for which (c) and (d) were performed, and

transmitting, to the model training system, at least a first model training dataset among the plurality of model training datasets, the first model training dataset corresponding to the first target language, wherein the model training system is configured to:

training a first machine learning model using the first model training dataset in the first target language, the training including inputting, into the first machine learning model, training examples from the first model training dataset, to obtain a set of values for model parameters for a first speech recognition model, and

generating the first speech recognition model in the first target language, for the topic for which (c) and (d) were performed, wherein the first speech recognition model is configured with the set of values for a set of predetermined sentiments or a set of predetermined intents,

wherein the first speech recognition model is configured to, based on one or more sentences provided as an input by a user in the first target language, identify a certain sentiment from the set of predetermined sentiments or a certain intent from the set of predetermined intents, and output, for the one or more sentences, a prediction of the certain sentiment or a prediction of the certain intent.

14 . The system of claim 13 , wherein:

(c) and (d) are performed for the plurality of topics in the plurality of different target languages, and

the outputting the result data further includes providing sets of result data that are obtained for each of the plurality of topics in the plurality of languages as a plurality of training datasets for building speech recognition models in the plurality of different target languages for the plurality of topics.

15 . One or more non-transitory computer-readable media storing instructions which, when executed by one or more processors, cause a computer system comprising a model training system and a dataset generation system to perform operations including:

(a) obtaining, using the dataset generation system from a text corpus, a plurality of webpages in a plurality of languages, each of the plurality of webpages corresponding to a respective universal resource locator (URL),

(b) annotating, using the dataset generation system, each of the plurality of webpages based on the respective URL, to obtain annotated data entries, each of the annotated data entries corresponding to one of the plurality of webpages and comprising a classification label corresponding to a sub-topic of one of a plurality of topics, wherein each of the plurality of topics comprises a corresponding plurality of sub-topics,

(c) performing, using the dataset generation system, a first filtering for the annotated data entries to obtain topic-specific content in a target language among the plurality of languages by:

comparing one or more sub-topics among the corresponding plurality of sub-topics with classification labels correspondingly included in the annotated data entries, wherein the one or more sub-topics are associated with a topic among the plurality of topics, and

obtaining, from the annotated data entries, the topic-specific content corresponding to the one or more sub-topics based on the comparing, and

(d) performing, using the dataset generation system, post-processing on the topic-specific content to obtain result data, by performing at least one from among a second filtering and a normalizing on sentences included in the topic-specific content,

wherein (c) and (d) are performed for a plurality of different target languages among the plurality of languages, the target language being a first target language of the plurality of different target languages,

outputting, using the dataset generation system, the result data for each of the plurality of different target languages as a plurality of model training datasets for the topic for which (c) and (d) were performed,

transmitting, from the dataset generation system to the model training system, at least a first model training dataset among the plurality of model training datasets, the first model training dataset corresponding to the first target language,

training, using the model training system, a first machine learning model using the first model training dataset in the first target language, the training including inputting, into the first machine learning model, training examples from the first model training dataset, to obtain a set of values for model parameters for a first speech recognition model, and

generating, using the model training system, the first speech recognition model in the first target language, for the topic for which (c) and (d) were performed, wherein the first speech recognition model is configured with the set of values for a set of predetermined sentiments or a set of predetermined intents,

wherein the first speech recognition model is configured to, based on one or more sentences provided as an input by a user in the first target language, identify a certain sentiment from the set of predetermined sentiments or a certain intent from the set of predetermined intents, and output, for the one or more sentences, a prediction of the certain sentiment or a prediction of the certain intent.

16 . The one or more non-transitory computer-readable media of claim 15 , wherein:

(c) and (d) are performed for the plurality of topics in the plurality of different target languages, and

the outputting the result data further includes providing sets of result data that are obtained for each of the plurality of topics in the plurality of languages as a plurality of training datasets for building speech recognition models in the plurality of different target languages for the plurality of topics.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 21, 2023
From: LOU, PARIA JAMSHID; TANGARI, GIOACCHINO; BLACK, JASON; HETTIGE, BHAGYA GAYATHRI; ZHONG, XU; ZAREMOODI, POORYA; DUONG, THANH LONG; JOHNSON, MARK EDWARD
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 064654/0001 →
Continuity (2)
Provisional Application 63384468 · Nov 21, 2022
Related Publication 20240169161A1 · May 23, 2024
References Cited (43)
US 9235634B2 · Georgakis · 2016 [cited by examiner]
US 9836455B2 · Martens · 2017 [cited by examiner]
US 10282425B2 · Shaw · 2019 [cited by examiner]
US 11321615B1 · Genkin et al. · 2022 [cited by applicant]
US 20100241633A1 · Lotito · 2010 [cited by examiner]
US 20100306144A1 · Scholz · 2010 [cited by examiner]
US 20110301938A1 · Agrawal · 2011 [cited by examiner]
US 20180246972A1 · Shukla · 2018 [cited by examiner]
US 20200151245A1 · Maratta · 2020 [cited by examiner]
US 20220293107A1 · Leaman · 2022 [cited by examiner]
Ye et al. (Multilingual taxonomic web page classification for contextual targeting at yahoo)—Proceedings of the 28th . . . , 2022—dl.acm.org (Year: 2022). [cited by examiner]
Baykan et al. (A Comprehensive Study of Techniques for URL-Based Web Page Language Classification), ACM Transactions on the Web (TWEB), vol. 7, Issue 1 Article No. 3, pp. 1-37—https://doi.org/10.1145/2435215.2435218 (Ye… [cited by examiner]
Approximate K-NN Search, Approximate Search-Open Distro Documentation, Available Online at: https://opendistro.github.io/for-elasticsearch-docs/docs/knn/approximate-knn/, Accessed from Internet on Aug. 23, 2023, 5 pages. [cited by applicant]
Benchmarking Results, ANN-Benchmarks, Available Online at: https://ann-benchmarks.com/, Accessed from Internet on Aug. 23, 2023, 27 pages. [cited by applicant]
DMOZ, Creative Commons Attribution 3.0 Unported, Open Directory License, Available online at: https://en.wikipedia.org/wiki/DMOZ, Jun. 5, 1998, 14 pages. [cited by applicant]
DMOZ Internet Directory, Presented by DMOZLive.com, Available online at: http://dmozlive.com/, Aug. 17, 2023, 1 page. [cited by applicant]
Examples Using Common Crawl Data, Common Crawl, Available online at: https://commoncrawl.org/the-data/examples/, Aug. 17, 2023, 21 pages. [cited by applicant]
MLRun, GitHub, Machine Learning automation and tracking, Available Online at: https://github.com/mlrun/mlrun, Accessed from Internet on Aug. 23, 2023, 4 pages. [cited by applicant]
Parallelized ML Inference, Aws-Samples, GitHub, Available Online at: https://github.com/mlrun-samples/parallelize-ml-inference/blob/master/src/ml-inference/parallize_inference_pool.py, Accessed from Internet on Aug. 23,… [cited by applicant]
URL Classification DMOZ, Available online at: https://www.kaggle.com/datasets/revanthrex/url-classification, Aug. 17, 2023, 1 page. [cited by applicant]
WARCannon—Catastrophically Powerful Parallel WARC Processing, Available online at: https://github.com/c6fc/warcannon, Sep. 7, 2022, 12 pages. [cited by applicant]
Girard, Parse Petabytes of Data from CommonCrawl in Seconds, Available online at: https://www.primates.dev/parse-petabytes-of-data-from-commoncrawl-in-seconds/, Jan. 21, 2020, 6 pages. [cited by applicant]
Grave et al., Learning Word Vectors for 157 Languages, Proceedings of the Eleventh International Conference on Language Resources and Evaluation, Available Online at: https://aclanthology.org/L18-1550.pdf, May 2018, pp.… [cited by applicant]
Hoek, Web Classification Using DMOZ, Bachelor Thesis Computing Science, Available online at: https://www.cs.ru.nl/bachelors-theses/2021/Lisa_Hoek_1009553_Web_classification_using_DMOZ.pdf, Jan. 17, 2021, 47 pages. [cited by applicant]
Mackenzie et al., CC-News-En: A Large English News Corpus, CIKM Conference, Available Online at: https://people.eng.unimelb.edu.au/ammoffat/abstracts/cikm20ccnews.pdf, Oct. 19-23, 2020, 8 pages. [cited by applicant]
Manning et al., The Stanford CoreNLP Natural Language Processing Toolkit, Proceedings of 52nd Annual Meeting of the Association for Computational Linguistics: System Demonstrations, Association for Computational Linguis… [cited by applicant]
Mukhopadhyay et al., Domain-Specific Crawler Design, Web Searching and Mining, Available online at: https://www.researchgate.net/publication/333747680, Jul. 31, 2019, pp. 85-112. [cited by applicant]
Remus et al., Domain-Specific Corpus Expansion with Focused Webcrawling, Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), Available online at: https://aclanthology.org/L1… [cited by applicant]
Rettig et al., Fusing Vector Space Models for Domain-Specific Applications, Conference: 2019 IEEE 31st International Conference on Tools with Artificial Intelligence (ICTAI), Available online at: https://arxiv.org/pdf/1… [cited by applicant]
Shang et al., Automated Phrase Mining from Massive Text Corpora, IEEE Transactions on Knowledge and Data Engineering, Available online at: https://arxiv.org/pdf/1702.04457.pdf, Mar. 11, 2017, 14 pages. [cited by applicant]
Srivastava et al., Time and Domain Specific Twitter Data Mining for Plastic Ban based on Public Opinion, Proceedings of the Second International Conference on Innovative Mechanisms for Industry Applications (ICIMIA 2020… [cited by applicant]
Suarez et al., A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Available online at: https://a… [cited by applicant]
Suarez et al., Asynchronous Pipeline for Processing Huge Corpora on Medium to Low Resource Infrastructures, HAL Open Science, Available Online at: https://inria.hal.science/hal-02148693/file/Asynchronous_Pipeline_for_Pr… [cited by applicant]
Wahed et al., SAUCE: Truncated Sparse Document Signature Bit-Vectors for Fast Web-Scale Corpus Expansion, Computation and Language, Available online at: https://arxiv.org/pdf/2108.11948.pdf, Aug. 26, 2021, 11 pages. [cited by applicant]
Wang , Parallelizing Across Multiple Cpu/gpus to Speed Up Deep Learning Inference at the Edge, AWS Machine Learning Blog, Available Online at: https://aws.amazon.com/blogs/machine-learning/parallelizing-across-multiple-… [cited by applicant]
Wenzek et al., CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data, Available Online at: https://arxiv.org/pdf/1911.00359.pdf, Nov. 15, 2019, 9 pages. [cited by applicant]
Zhao et al., Understanding Data Storage and Ingestion for Large-Scale Deep Recommendation Model Training, In The 49th Annual International Symposium on Computer Architecture (ISCA 2022), Available online at: https://arx… [cited by applicant]
Chollet, Multi-GPU and Distributed Training, Keras, Available online at https://keras.io/guides/distributed_training, Apr. 29, 2020, 8 pages. [cited by applicant]
Dkeras-project/dkeras, Available online at https://github.com/dkeras-project/dkeras, retrieved Aug. 28, 2023, 5 pages. [cited by applicant]
Model Interference Using TensorFlow Keras API, Databricks on AWS, Available online at https://docs.databricks.com/en/machine-learning/model-inference/resnet-model-inference-keras.html, Jul. 11, 2023, 1 page. [cited by applicant]
Facebookresarch/cc_net, Available online at https://github.com/facebookresearch/cc net, retrieved Aug. 17, 2023, 5 pages. [cited by applicant]
“Data Collection for Artificial Intelligence and Machine”, Available online at: https://www.tagxdata.com/datacollection/, Accessed from Internet on Feb. 21, 2023, 4 pages. [cited by applicant]
“Data Collection for Machine Learning”, Mindy Support Outsourcing, Available online at: https://mindy-support.com/services-post/data-collection/, Jul. 2022, 5 pages. [cited by applicant]