IP Library › Granted Patent US 12,602,370
Granted Patent B1
US 12,602,370 · App. 18/987,974 · Granted Apr 14, 2026

Detecting conflicting knowledge in data sources

Inventors: Roy Eisenstadt (Tel Aviv, IL); Alexander Ostrikov (Herzliya, IL); Alexander Tsvetkov (Tel Aviv, IL)
Assignee: Microsoft Technology Licensing, LLC
G06F16/2365G06F16/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,370
App. No.
18/987,974
Granted
Apr 14, 2026
Kind
B1
Abstract

A knowledge source is analyzed using machine learning models to detect conflicting or contradictory data in the documents of the knowledge source. The documents in a knowledge source are partitioned into non-overlapping, contiguous text segments containing a unique topic. Questions are generated for each text segment from a generative language model. Factoid answers are generated for each question from a question answering language model. Similar pairs of questions are identified from the answers of each question. A natural language inference model determines whether the answers to the similar questions are contradictory. A remedy is generated to address the documents having the contradictory or conflicting segments.

Claims (58)

1 . A system comprising:

a processor; and

a memory that stores a program that is configured to be executed by the processor, wherein the program comprises instructions to perform acts that:

obtain a plurality of documents of a knowledge source;

partition each document of the plurality of documents of the knowledge source into a plurality of segments, wherein a segment comprises a non-overlapping, contiguous block of natural language text that pertains to a particular topic;

cause a language model to generate at least one question for each segment of the plurality of segments;

cause the language model to generate a factoid answer for the at least one question of each segment of the plurality of segments;

find a pair of similar questions from the at least one question of each segment of the plurality of segments;

cause the language model to determine whether the factoid answers for the pair of similar questions are contradictory, whereby the pair of similar questions are a contradictory pair of similar questions; and

upon the language model determining that the pair of similar questions are a contradictory pair of similar questions, eliminate a select document associated with a select one of the contradictory pair of similar questions from the knowledge source.

2 . The system of claim 1 , wherein find a pair of similar questions from the at least one question for each segment of the plurality of segments is based on matching an embedding of each question of each segment of the plurality of segments with an embedding of all other questions of each segment of the plurality of segments.

3 . The system of claim 1 , wherein find a pair of similar questions from the at least one question for each segment of the plurality of segments is based on an edit distance between each question of each segment of the plurality of segments.

4 . The system of claim 1 , wherein the program comprises instructions to perform acts that:

rank each document associated with the contradictory pair of similar questions in accordance with a priority rule; and

select one document associated with the contradictory pair of similar questions to eliminate from the knowledge source based on the priority rule.

5 . The system of claim 4 , wherein the priority rule is based on a recency of the documents of the contradictory pair of similar questions.

6 . The system of claim 4 , wherein the priority rule is based on a type of the documents of the contradictory pair of similar questions or a type of a source of the documents of the contradictory pair of similar questions.

7 . The system of claim 4 , wherein the priority rule is based on a page rank of a website from which the documents of the contradictory pair of similar questions originated.

8 . A computer-implemented method comprising:

accessing a knowledge source comprising a plurality of documents;

generating a plurality of segments from the plurality of documents, wherein a segment comprises a non-overlapping, contiguous block of natural language text that pertains to a select topic;

causing a language model to generate a plurality of questions for each segment of the plurality of segments;

causing the language model to generate a factoid answer for each of the plurality of questions of each segment of the plurality of segments;

generating pairs of similar questions from the plurality of questions for each segment of the plurality of segments based on matching words in select ones of the plurality of questions;

causing the language model to determine whether the factoid answers for the pairs of similar questions are contradictory based on a non-existence of a logical relationship between the answers for each pair of similar questions, whereby each pair of similar questions is a contradictory pair of similar questions; and

selecting one of the documents of a contradictory pair of similar questions to remedy in the knowledge source.

9 . The computer-implemented method of claim 8 , comprising:

ranking documents of the contradictory pair of similar questions based on recency; and

eliminating a select one of the documents of the contradictory pair of similar questions having a longest recency.

10 . The computer-implemented method of claim 8 , comprising:

ranking the documents of the contradictory pair of similar questions based on a priority rule; and

eliminating a select one of the documents of the contradictory pair of similar questions based on the priority rule.

11 . The computer-implemented method of claim 10 , wherein the priority rule comprises type of source of the documents of the contradictory pair of similar questions or type of the documents of the contradictory pair of similar questions.

12 . The computer-implemented method of claim 8 , comprising:

computing a page rank for a website source of each document of the contradictory pair of similar questions; and

eliminating a select one of the documents of the contradictory pair of similar questions having a lowest page rank score.

13 . The computer-implemented method of claim 8 , wherein generating a plurality of segments from the plurality of documents further comprising:

identifying a segment from boundaries between text blocks based on topic shifts and lexical cohesion.

14 . The computer-implemented method of claim 8 , wherein generating pairs of similar questions from the plurality of questions for each segment of the plurality of segments is based on matching an embedding of a first question of a pair of similar questions with an embedding of a second question of the pair of similar questions.

15 . The computer-implemented method of claim 8 , further comprising:

updating the knowledge source to remedy the documents of the contradictory pair of similar questions; and

generating a training dataset from the updated knowledge source, wherein the training dataset is used to train the language model.

16 . A hardware storage device having stored thereon computer executable instructions that are structured to be executable by a processor of a computing device to thereby cause the computing device to perform actions that:

access a knowledge source comprising a plurality of documents;

identify at least one pair of documents of the plurality of documents containing conflicting information by performing actions that:

partition the plurality of documents into a plurality of segments, wherein a segment comprises a non-overlapping contiguous portion of text belonging to a topic;

cause a language model to generate a plurality of questions for each segment of the plurality of segments;

cause the language model to generate an answer for each of the plurality of questions for each segment of the plurality of segments;

identify at least one pair of similar questions from the plurality of questions for each segment of the plurality of segments, wherein a pair of similar questions is based on matching embeddings of each question of the plurality of questions or based on an edit distance of changes between pairs of questions of the plurality of questions; and

cause the language model to determine that the answers to the at least one pair of similar questions produce results that are conflicting, whereby each of the at least one pair of similar questions is a contradictory pair of similar questions; and

alter the knowledge source to eliminate a select document associated with a question of a contradictory pair of similar questions.

17 . The hardware storage device of claim 16 , wherein identify the at least one pair of similar questions from the plurality of questions for each segment of the plurality of segments is based on matching an embedding of each question of the plurality of questions for each segment of the plurality of segments with all other questions for the plurality of questions for each segment of the plurality of segments.

18 . The hardware storage device of claim 16 , having stored thereon computer executable instructions that are structured to be executable by a processor of a computing device to thereby cause the computing device to perform actions that:

choose the select document associated with a question of the contradictory pair of similar questions to eliminate from the knowledge source based on a page rank of the web site from which the select document originated.

19 . The hardware storage device of claim 16 , having stored thereon computer executable instructions that are structured to be executable by a processor of a computing device to thereby cause the computing device to perform actions that:

select the select document associated with a question of the contradictory pair of similar questions to eliminate from the knowledge base based on recency of the select document, source of the select document, or type of the select document.

20 . The hardware storage device of claim 16 , having stored thereon computer executable instructions that are structured to be executable by a processor of a computing device to thereby cause the computing device to perform actions that:

choose the select document associated with a question of the contradictory pair of similar questions to eliminate from the knowledge source based on a type of the documents of the contradictory pair of similar questions or a type of a source of the documents of the contradictory pair of similar questions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 22, 2024
From: EISENSTADT, ROY; OSTRIKOV, ALEXANDER; TSVETKOV, ALEXANDER
To: MICROSOFT TECHNOLOGY LICENSING, LLC.
Reel/Frame 069660/0454 →
References Cited (19)
US 10762438B1 · Zhang · 2020 [cited by examiner]
US 20170091170A1 · Cardillo · 2017 [cited by examiner]
US 20180349754A1 · Kumar · 2018 [cited by examiner]
US 20210406735A1 · Nahamoo · 2021 [cited by examiner]
US 20230205824A1 · Jablokov · 2023 [cited by examiner]
US 20240202448A1 · Suppa · 2024 [cited by examiner]
US 20240265041A1 · Rennie · 2024 [cited by examiner]
US 20250342188A1 · Vahdat · 2025 [cited by examiner]
Bachina, et al., “Ensemble ALBERT and RoBERTa for Span Prediction in Question Answering,” Proceedings of the 1st Workshop on Document-grounded Dialogue and Conversational Question Answering, 2021, pp. 1-6. [cited by applicant]
Hearst, Marti A., “TextTiling: Segmenting Text into Multi-paragraph Subtopic Passages,” in Computational Linguistics, vol. 23, Issue 1, Mar. 1997, pp. 33-64. [cited by applicant]
Khanna, et al., “Transformer-based Language Models for Factoid Question Answering at BioASQ9b,” arxiv.org, Sep. 15, 2021, 11 pages. [cited by applicant]
Reimers, et al., “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks”, In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on … [cited by applicant]
Sherman, et al., “Using Hidden Markov Models for Topic Segmentation of Meeting Transcripts,” IEEE Spoken Language Technology Workshop, 2008, pp. 1-4. [cited by applicant]
Weissenborn, et al., “FastQA: A Simple and Efficient Neural Architecture for Question Answering,” arxiv.org, Mar. 14, 2017, 10 pages. [cited by applicant]
Weissenborn, et al., “Making Neural QA as Simple as Possible but not Simpler,” Proceedings of the 21st Conference on Computational Natural Language Learning, Aug. 2017, pp. 271-280. [cited by applicant]
“Cosine Similarity—Wikipedia”, Retrieved from: https://en.wikipedia.org/wiki/Cosine_similarity, May 24, 2025, 08 Pages. [cited by applicant]
“Levenshtein Distance—Wikipedia”, Retrieved from: https://en.wikipedia.org/wiki/Levenshtein_distance, Jun. 28, 2025, 08 Pages. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2025/051169, Mar. 5, 2026, 13 pages. [cited by applicant]
Pan et al., “On the Risk of Misinformation Pollution with Large Language Models”, In the Findings of the Association for Computational Linguistics: EMNLP, Dec. 2023, pp. 1389-1403. [cited by applicant]