IP Library › Granted Patent US 12,210,831
Granted Patent B2
US 12,210,831 · App. 17/493,819 · Granted Jan 28, 2025

Knowledge base with type discovery

Inventors: Elena Pochernina (London, GB); John Winn (Cambridge, GB); Matteo Venanzi (London, GB); Ivan Korostelev (London, GB); Pavel Myshkov (London, GB); Samuel Alexander Webster (Cambridge, GB); Yordan Kirilov Zaykov (Cambridge, GB); Nikita Voronkov (Bothell, WA); Dmitriy Meyerzon (Bellevue, WA); Marius Alexandru Bunescu (Redmond, WA); Alexander Armin Spengler (Cambridge, GB); Vladimir Gvozdev (Sammamish, WA); Thomas P. Minka (Cambridge, GB); Anthony Arnold Wieser (Fen Ditton, GB); Sanil Rajput (San Francisco, CA); John Guiver (Saffron Walden, GB)
Assignee: Microsoft Technology Licensing, LLC.
G06F40/30G06F16/9024G06F16/90335
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,210,831
App. No.
17/493,819
Filed
Oct 4, 2021
Granted
Jan 28, 2025
Kind
B2
Art Unit
2655
USPC
704/9
Abstract

In various examples there is a computer-implemented method of database construction. The method comprises storing a knowledge graph comprising nodes connected by edges, each node representing a topic. Accessing a topic type hierarchy comprising a plurality of types of topics, the topic type hierarchy having been computed from a corpus of text documents. One or more text documents are accessed and the method involves labelling a plurality of the nodes with one or more labels, each label denoting a topic type from the topic type hierarchy, by, using a deep language model; or for an individual one of the nodes representing a given topic, searching the accessed text documents for matches to at least one template, the template being a sequence of words and containing the given topic and a placeholder for a topic type; and storing the knowledge graph comprising the plurality of labelled nodes.

Claims (42)

1. A computer-implemented method of database construction comprising:

accessing a topic type hierarchy from text documents, the topic type hierarchy comprising types of topics;

labelling a first node representing a topic with a first topic type and a second topic type from the topic type hierarchy, wherein labelling the first node comprises:

searching a text document for a match to at least one template, the template being a sequence of words and containing a stored topic and a placeholder for a stored topic type, wherein a template match fills the placeholder for the stored topic type such that contents of the filled placeholder comprise a candidate topic type;

receiving, for the first and second topic types, a probability that the topic type applies to the first node from a topic type correctness model that has as input the candidate topic type and has as output an estimate of a likelihood of the candidate topic type being a correct topic type; and

storing a knowledge graph comprising the labelled first node.

2. The method of claim 1 comprising receiving a query comprising a topic, searching the knowledge graph to identify a node representing a topic similar to the query, outputting a topic of the identified node, and outputting a topic type of the identified node.

3. The method of claim 1 comprising receiving a selection of a topic type and filtering identified nodes to include only identified nodes having a selected topic type.

4. The method of claim 1 comprising receiving a query comprising a topic type, searching the knowledge graph to identify a node having a topic type label corresponding to the topic type of the query, outputting a topic of the identified node.

5. The method of claim 1 comprising receiving a query comprising a topic, searching the knowledge graph to identify a node within a specified number of hops away from a node representing the topic of the query, outputting a topic of an identified node and outputting a topic type of the identified node.

6. The method of claim 1 comprising filtering identified nodes to include only identified nodes having a same topic type as a topic type of a received query topic, and outputting the topics of the filtered identified nodes.

7. The method of claim 1 comprising computing the types of the topic type hierarchy from a corpus of text documents and using a plurality of seed types.

8. The method of claim 7 comprising: searching for topics in the corpus of text documents to identify topics having one of the seed types.

9. The method of claim 1 comprising outputting the contents of the placeholder as the candidate topic type.

10. The method of claim 1 comprising filtering candidate topic types to retain a specified number of most frequently occurring candidate topic types.

11. The method of claim 10 comprising using the retained candidate topic types as seed types and repeating a process of searching for topics in a same corpus of text documents to identify topics having one of the seed types, and for each identified topic, searching text near to the identified topic for matches to the at least one template, and when a template match is found which fills the placeholder for topic type, outputting contents of the placeholder as the candidate topic type.

12. The method of claim 10 comprising using the retained candidate topic types as seed types and repeating a process of searching for topics in a different corpus of text documents to identify topics having one of the seed types, and for each identified topic, searching text near to the identified topic for matches to the at least one template, and when a template match is found which fills the placeholder for topic type, outputting contents of the placeholder as the candidate topic type.

13. The method of claim 1 comprising selecting at least one of the first or second topic types of the first node for display to a user based on a comparison of the probabilities of the first and second topic types.

14. The method of claim 1 further comprising selecting both the first topic type and the second topic type of the first node for display to a user.

15. A database construction apparatus comprising:

a processor;

a memory storing instructions that, when executed by the processor, perform a method for:

accessing a topic type hierarchy from text documents, the topic type hierarchy comprising types of topics;

labelling nodes with labels denoting a topic type from the topic type hierarchy, by,

for an individual one of the nodes representing a stored topic,

searching a text document for a match to at least one template, the template being a sequence of words and containing the stored topic and a placeholder for a stored topic type, wherein a template match fills the placeholder for the stored topic type such that contents of the filled placeholder comprise a candidate topic type,

wherein labelling the nodes further comprises:

labelling a first node with a first topic type and a second topic type; and

receiving, for the first and second topic types, a probability that the topic type applies to the first node from a topic type correctness model that has as input the candidate topic type and has as output an estimate of a likelihood of the candidate topic type being a correct topic type; and

storing a knowledge graph comprising the labelled nodes.

16. The database construction apparatus of claim 15 wherein the instructions are for receiving a query comprising a topic type, searching the knowledge graph to identify nodes having topic type labels corresponding to the topic type of the query, and outputting topics of the identified nodes.

17. The database construction apparatus of claim 15 wherein the instructions are for receiving a query comprising a topic, searching the knowledge graph to identify nodes within a specified number of hops away from a node representing the topic of the query, outputting a topic of at least one of the identified nodes and outputting a topic type of the at least one identified node.

18. A database construction apparatus comprising:

a processor; and

a memory storing instructions that, when executed by the processor, perform a method for:

accessing a topic type hierarchy from text documents, the topic type hierarchy comprising types of topics;

labelling a first node with a first topic type by,

searching a text document for a match to at least one template, the template being a sequence of words and containing a stored topic and a placeholder for a stored topic type, wherein a template match fills the placeholder for the stored topic type such that contents of the filled placeholder comprise a candidate topic type;

receiving a probability that the first topic type applies to the first node from a topic type correctness model that has as input the candidate topic type and has as output an estimate of a likelihood of the candidate topic type being a correct topic type; and

storing a knowledge graph comprising the labelled first node.

19. The database construction apparatus of claim 15 wherein the instructions are for filtering candidate topic types and retaining a specified number of most frequently occurring candidate topic types.

20. The database construction apparatus of claim 15 wherein the instructions are for receiving a selection of a topic type, filtering identified nodes, and including only identified nodes having a selected topic type.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2021
From: POCHERNINA, ELENA; WINN, JOHN; VENANZI, MATTEO; KOROSTELEV, IVAN; MYSHKOV, PAVEL; WEBSTER, SAMUEL ALEXANDER; ZAYKOV, YORDAN KIRILOV; VORONKOV, NIKITA; MEYERZON, DMITRIY; BUNESCU, MARIUS ALEXANDRU; SPENGLER, ALEXANDER ARMIN; GVOZDEV, VLADIMIR; MINKA, THOMAS P.; WIESER, ANTHONY ARNOLD; RAJPUT, SANIL; GUIVER, JOHN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 057710/0695 →
Continuity (2)
Continuation In Part 17460123 · Aug 27, 2021
Related Publication 20230076773A1 · Mar 9, 2023
References Cited (115)
US 6006242A · Poole et al. · 1999 [cited by applicant]
US 6591258B1 · Stier et al. · 2003 [cited by applicant]
US 6601055B1 · Roberts et al. · 2003 [cited by applicant]
US 7082430B1 · Danielsen · 2006 [cited by examiner]
US 7096210B1 · Kramer et al. · 2006 [cited by applicant]
US 7502770B2 · Hillis et al. · 2009 [cited by applicant]
US 7756810B2 · Nelken et al. · 2010 [cited by applicant]
US 8103598B2 · Minka et al. · 2012 [cited by applicant]
US 8275737B2 · Kupershmidt et al. · 2012 [cited by applicant]
US 9251467B2 · Winn et al. · 2016 [cited by applicant]
US 9842166B1 · Leviathan · 2017 [cited by examiner]
US 9864795B1 · Halevy et al. · 2018 [cited by applicant]
US 10504198B1 · Ward · 2019 [cited by examiner]
US 10679008B2 · Dubey et al. · 2020 [cited by applicant]
US 10783162B1 · Montague · 2020 [cited by examiner]
US 11244113B2 · Adderly · 2022 [cited by examiner]
US 11809460B1 · Rausch · 2023 [cited by examiner]
US 20030177115A1 · Stern et al. · 2003 [cited by applicant]
US 20040059966A1 · Chan et al. · 2004 [cited by applicant]
US 20040095374A1 · Jojic et al. · 2004 [cited by applicant]
US 20040260692A1 · Brill et al. · 2004 [cited by applicant]
US 20050086222A1 · Wang et al. · 2005 [cited by applicant]
US 20050154690A1 · Nitta · 2005 [cited by examiner]
US 20100077324A1 · Harrington et al. · 2010 [cited by applicant]
US 20110119050A1 · Deschacht et al. · 2011 [cited by applicant]
US 20110251984A1 · Nie et al. · 2011 [cited by applicant]
US 20110307435A1 · Overell et al. · 2011 [cited by applicant]
US 20120203752A1 · Ha-thuc et al. · 2012 [cited by applicant]
US 20140201126A1 · Lotfi · 2014 [cited by applicant]
US 20140250046A1 · Winn · 2014 [cited by examiner]
US 20140280353A1 · Delaney · 2014 [cited by examiner]
US 20140337306A1 · Gramatica · 2014 [cited by examiner]
US 20150006501A1 · Talmon · 2015 [cited by examiner]
US 20150248222A1 · Stickler · 2015 [cited by examiner]
US 20160012020A1 · Yonghoon et al. · 2016 [cited by applicant]
US 20160012122A1 · Soares et al. · 2016 [cited by applicant]
US 20160063061A1 · Meyerzon · 2016 [cited by examiner]
US 20160140445A1 · Adderly · 2016 [cited by examiner]
US 20160323411A1 · Lee · 2016 [cited by examiner]
US 20170024652A1 · Kipersztok · 2017 [cited by applicant]
US 20170199928A1 · Zhang et al. · 2017 [cited by applicant]
US 20170235816A1 · Livshits · 2017 [cited by examiner]
US 20170316322A1 · Perincherry et al. · 2017 [cited by applicant]
US 20180060306A1 · Starostin · 2018 [cited by examiner]
US 20180075359A1 · Brennan · 2018 [cited by examiner]
US 20190081983A1 · Teal · 2019 [cited by applicant]
US 20190155947A1 · Chu et al. · 2019 [cited by applicant]
US 20190213484A1 · Winn · 2019 [cited by examiner]
US 20200089802A1 · Ronen · 2020 [cited by examiner]
US 20210026846A1 · Subramanya · 2021 [cited by examiner]
US 20210104234A1 · Zhang · 2021 [cited by examiner]
US 20210157858A1 · Stevens · 2021 [cited by examiner]
US 20210191949A1 · Sato · 2021 [cited by examiner]
US 20210209500A1 · Hu · 2021 [cited by examiner]
US 20210312919A1 · Sato · 2021 [cited by examiner]
US 20210326519A1 · Lin · 2021 [cited by examiner]
US 20220019740A1 · Meyerzon · 2022 [cited by examiner]
US 20220261545A1 · Lauber · 2022 [cited by examiner]
US 20220300544A1 · Potter · 2022 [cited by examiner]
US 20230030086A1 · Martinez Ayala · 2023 [cited by examiner]
US 20230067688A1 · Pochernina · 2023 [cited by examiner]
US 20230076773A1 · Pochernina · 2023 [cited by examiner]
US 20230342629A1 · Panda · 2023 [cited by examiner]
“Final Office Action Issued in U.S. Appl. No. 15/898,211”, Mailed Date: Oct. 26, 2022, 51 Pages. [cited by applicant]
“Final Office Action Issued in U.S. Appl. No. 15/898,211”, Mailed Date: Aug. 18, 2023, 42 Pages. [cited by applicant]
“Non Final Office Action Issued in U.S. Appl. No. 15/898,211”, Mailed Date: May 16, 2022, 44 Pages. [cited by applicant]
Chu, et al., “KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing”, In Proceedings of the ACM SIGMOD International Conference on Management of Data, May 27, 2015, 15 Pages. [cited by applicant]
Zhang, et al., “DeepDive: Declarative Knowledge Base Construction”, In Journal of Communications of the ACM, vol. 60, Issue 5, Apr. 24, 2017, pp. 93-102. [cited by applicant]
Niu, et al., “DeepDive: Web-scale Knowledge-base Construction using Statistical Learning and Inference”, In Journal of VLDS, vol. 12, Aug. 31, 2012, 4 Pages. [cited by applicant]
“Non Final Office Action Issued in U.S. Appl. No. 15/898,211”, Mailed Date: Feb. 27, 2023, 54 Pages. [cited by applicant]
Conneau, et al., “Supervised Learning of Universal Sentence Representations from Natural Language Inference Data”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Sep. 7, 2017, pp. … [cited by applicant]
Jordan, et al., ““Machine learning: Trends, perspectives, and prospects””, In Journal of Science , vol. 349, Issue 6245, 2015, pp. 255-260. [cited by applicant]
Carlson, et al., “Toward an Architecture for Never-Ending Language Learning”, In Proceedings of Twenty-Fourth AAAI Conference on Artificial Intelligence, Jul. 5, 2010, pp. 1306-1313. [cited by applicant]
Devlin, et al., “BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding”, In Proceedings of Conference of the North American Chapter of the Association for Computational Linguistics: Human Lang… [cited by applicant]
Dong, et al., “Knowledge Vault: A Web-Scale Approach to Probabilistic Knowledge Fusion”, In Proceedings of 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Aug. 24, 2014, pp. 601-610. [cited by applicant]
Klimt, et al., “The Enron Corpus: A New Dataset for Email Classification Research”, In Publication of Springer, Sep. 20, 2004, pp. 217-226. [cited by applicant]
Liu, et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach”, In Repository of arXiv:1907.11692v1, Jul. 26, 2019, 13 Pages. [cited by applicant]
Loshin, David, “Enterprise Knowledge Management: The Data Quality Approach—Chapter 1”, In Publication of Morgan Kaufmann, 2001, 24 Pages. [cited by applicant]
Maedche, et al., “Ontologies for Enterprise Knowledge Management”, In Journal of IEEE Intelligent Systems, vol. 18, Issue 2, Mar. 2003, pp. 26-33. [cited by applicant]
“Infer.NET 0.3”, Retrieved from: https:/web.archive.org/web/20181013050802/https://dotnet.github.io/infer/, Oct. 13, 2018, 2 pages. [cited by applicant]
Radicati, et al., “Email Statistics Report, 2015-2019”, Retrieved from: https://www.radicati.com/wp/wp-content/uploads/2015/02/Email-Statistics-Report-2015-2019-Executive-Summary.pdf, 2019, 4 Pages. [cited by applicant]
Sang, et al., “Introduction To The CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition”, In Proceedings of Seventh Conference on Natural Language Learning at HLT-NAACL, Jun. 12, 2003, 6 Pages. [cited by applicant]
Szekely, et al., “Building and Using a Knowledge Graph to Combat Human Trafficking”, In Publication of Springer, Oct. 11, 2015, pp. 205-221. [cited by applicant]
Voigt, et al., “The EU General Data Protection Regulation (GDPR)”, In Publication of Springer, Aug. 10, 2017, pp. 201-217. [cited by applicant]
Winn, et al., “Alexandria: Unsupervised High-Precision Knowledge Base Construction using a Probabilistic Program”, In Proceedings of Automated Knowledge Base Construction, Nov. 17, 2018, 20 Pages. [cited by applicant]
Winn, et al., “Enterprise Alexandria: Online High-Precision Enterprise Knowledge Base Construction with Typed Entities”, In Proceedings of 3rd Conference on Automated Knowledge Base Construction, Jun. 22, 2021, 13 Pages. [cited by applicant]
Yamada, et al., “LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attention”, In Proceedings of Conference on Empirical Methods in Natural Language Processing, Nov. 16, 2020, pp. 6442-6454. [cited by applicant]
Yangel, et al., “Belief Propagation with Strings”, In Technical Report MSR-TR-2017-11, Feb. 2017, 9 Pages. [cited by applicant]
Zhang, CE, “DeepDive: A Data Management System for Automatic Knowledge Base Construction”, A dissertation Submitted in Partial Fulfillment of the Requirements for the Degree of Doctor of Philosophy (Computer Sciences) a… [cited by applicant]
Zhang, et al., “ERNIE: Enhanced Language Representation with Informative Entities”, In Proceedings of 57th Annual Meeting of the Association for Computational Linguistics, Jul. 28, 2019, pp. 1441-1451. [cited by applicant]
Rajput, et al., “Alexandria in Microsoft Viva Topics: from Big Data to Big Knowledge”, Retrieved from: https://www.microsoft.com/en-us/research/blog/alexandria-in-microsoft-viva-topics-from-big-data-to-big-knowledge/, A… [cited by applicant]
Weikum, et al., “From Information to Knowledge: Harvesting Entities and Relationships from Web Sources”, In Proceedings of Twenty-Ninth ACM SIGMOD-SIGACT-SIGART Symposium on Principles of Database Systems, Jun. 6, 2010,… [cited by applicant]
Surdeanu, et al., “Overview of the English Slot Filling Track at the TAC2014 Knowledge Base Population Evaluation”, In Proceedings of 7th Text Analysis Conference, Nov. 17, 2014, 15 Pages. [cited by applicant]
Schmitz, et al., “Open Language Learning for Information Extraction”, In Proceedings of the Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning, Jul. 12, 2012… [cited by applicant]
Richardson, et al., “Markov Logic Networks”, In Journal of Machine Learning, vol. 62, Issue 1, Jan. 27, 2006, pp. 107-136. [cited by applicant]
Reinanda, R., “Entity Associations for Search”, A PhD Thesis Submitted in Dutch Research School for Information and Knowledge Systems, May 11, 2017, 184 Pages. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US2019/012237”, Mailed Date: Mar. 28, 2019, 13 Pages. [cited by applicant]
Niu, et al., “Elementary: Large-scale Knowledge-base Construction via Machine Learning and Statistical Inference”, In International Journal on Semantic Web and Information Systems, vol. 8, Issue 3, Jul. 1, 2012, 23 Page… [cited by applicant]
Mitchell, et al., “Never-Ending Learning”, In Proceedings of Twenty-Ninth AAAI Conference on Artificial Intelligence, Jan. 25, 2015, 9 Pages. [cited by applicant]
Minka, Thomas P. , “Expectation Propagation for Approximate Bayesian Inference”, In Proceedings of 17th Conference in Uncertainty in Artificial Intelligence, Aug. 2, 2001, pp. 362-369. [cited by applicant]
Hoffart, et al., “YAGO2: A Spatially and Temporally Enhanced Knowledge Base from Wikipedia”, In Journal of Artificial Intelligence, vol. 194, Jan. 2013, pp. 28-61. [cited by applicant]
Gupta, et al., “Biperpedia: An Ontology for Search Applications”, In Proceedings of VLDB Endowment, vol. 7, Issue 7, Mar. 2014, pp. 505-516. [cited by applicant]
Fader, et al., “Identifying Relations for Open Information Extraction”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Jul. 27, 2011, pp. 1535-1545. [cited by applicant]
Chaiken, et al., “SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets”, In Proceedings of VLDB Endowment, vol. 1, Issue 2, Aug. 23, 2008, pp. 1265-1276. [cited by applicant]
“Probabilistic-Programming.org”, Retrieved from; https://web.archive.org/web/20171121223322/probabilistic-programming.org/wiki/Home, Retrieved on: Nov. 21, 2017, 4 Pages. [cited by applicant]
“Probabilistic Programming for Advancing Machine Learning”, In Publication of Defense Advanced Research Projects Agency, Apr. 1, 2013, 47 Pages. [cited by applicant]
“Infer.NET 2.6”, Retrieved from: http://infernet.azurewebsites.net, Nov. 25, 2014, 1 Page. [cited by applicant]
Non-Final Office Action mailed on Dec. 21, 2023, in U.S. Appl. No. 17/460,123, 52 pages. [cited by applicant]
Communication pursuant to Article 94(3) EPC Received for European Application No. 19701418.6, mailed on Mar. 25, 2024, 8 pages. [cited by applicant]
Reinanda. R., “UvA-DARE (Digital Academic Repository).” accessed on URL: https://pure.uva.nl/ws/files/12382976/Reinanda_Thesis_complete.pdf, May 11, 2017, 185 pages. [cited by applicant]
Nguyen, et al., “Query-driven on-the-fly knowledge base construction.” Proceedings of the VLDB Endowment, vol. 11, Issue No. 1, 2017, pp. 66-79. [cited by applicant]
Non-Final Office Action mailed on May 31, 2024, in U.S. Appl. No. 15/898,211, 29 pages. [cited by applicant]
Final Office Action mailed on Dec. 5, 2024, in U.S. Appl. No. 15/898,211, 35 pages. [cited by applicant]
Qi, et al., “Measuring conflict and agreement between two prioritized knowledge bases in possibilistic logic”, Fuzzy Sets and Systems, vol. 161, Issue 14, Jul. 16, 2010, pp. 1906-1925. [cited by applicant]
Summons to attend oral proceedings pursuant to Rule 115(1) received in European Application No. 19701418.6, mailed on Nov. 25, 2024, 10 pages. [cited by applicant]