IP Library › Granted Patent US 12,265,561
Granted Patent B2
US 12,265,561 · App. 17/898,173 · Granted Apr 1, 2025

Machine learning for similarity scores between different document schemas

Inventors: Liviu-Sebastian Matei (Bucharest, RO); Filip Trojan (Beron, CZ)
Assignee: ORACLE INTERNATIONAL CORPORATION
G06F16/3326G06F16/3346
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,265,561
App. No.
17/898,173
Granted
Apr 1, 2025
Kind
B2
Abstract

A document repository may be searched for documents that are similar to a source document. Multiple queries may be generated based on a type of the source document, and the results may be combined in a unified response. User behavior may then be monitored, and implicit and explicit feedback may be gathered to evaluate the performance of the search. The gathered feedback may indicate how relevant each of the result documents are in comparison to the original source document. This feedback may then be used to adjust search parameters for the source document type, such that the performance of subsequent searches may be improved. A model may also be trained to classify implicit feedback using explicit feedback received from users.

Claims (45)

1. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

providing a plurality of documents having similarity scores that are calculated in response to receiving a first document, wherein the similarity scores indicate similarities between the first document and the plurality of documents;

receiving feedback from a user for a second document in the plurality of documents, wherein the feedback indicates whether the second document is relevant to the first document;

adjusting parameters of a configuration in response to the feedback, wherein the configuration is associated with documents having a same schema as the first document;

receiving a third document that has the same schema as the first document;

selecting the configuration based on the third document having the same schema as the first document; and

using adjusted parameters of the configuration to generate a plurality of queries, and executing the plurality of queries into a document repository to identify documents that are similar to the third document.

2. The non-transitory computer-readable medium of claim 1 , wherein the feedback comprises implicit feedback recorded as the user views the second document.

3. The non-transitory computer-readable medium of claim 2 , wherein the implicit feedback comprises a dwell time.

4. The non-transitory computer-readable medium of claim 2 , wherein the operations further comprise providing the implicit feedback to a classification model that determines whether the implicit feedback indicates that the second document is relevant to the first document.

5. The non-transitory computer-readable medium of claim 4 , wherein the feedback further comprises explicit feedback in addition to the implicit feedback, and the operations further comprise training the classification model using the implicit feedback as training data and the explicit feedback as a label for the training data.

6. The non-transitory computer-readable medium of claim 5 , wherein the explicit feedback comprises a user input indicating whether the first document is similar to the second document.

7. The non-transitory computer-readable medium of claim 4 , wherein the classification model determines whether a dwell time indicates that the second document is relevant to the first document.

8. The non-transitory computer-readable medium of claim 1 , wherein the operations further comprise generating a pop-up window with a display of the second document.

9. The non-transitory computer-readable medium of claim 8 , wherein the pop-up window comprises a control that allows the user to provide explicit feedback for the second document.

10. The non-transitory computer-readable medium of claim 1 , wherein the parameters of the configuration comprise a parameter that indicates a number of shingles used when generating the plurality of queries.

11. The non-transitory computer-readable medium of claim 1 , wherein the parameters of the configuration comprise a parameter that indicates a number of queries generated when generating the plurality of queries.

12. The non-transitory computer-readable medium of claim 1 , wherein the parameters of the configuration comprise a parameter that indicates weights applied to the plurality of queries.

13. The non-transitory computer-readable medium of claim 1 , wherein the parameters of the configuration comprise a learning parameter that controls how much the parameters of the configuration are adjusted in response to the feedback.

14. The non-transitory computer-readable medium of claim 1 , wherein the operations further comprise:

generating new similarity scores for the plurality of documents after adjusting the parameters of the configuration.

15. The non-transitory computer-readable medium of claim 14 , wherein the operations further comprise:

comparing the new similarity scores to the similarity scores to determine whether adjusting the parameters of the configuration improves the new similarity scores.

16. The non-transitory computer-readable medium of claim 15 , wherein determining whether adjusting the parameters of the configuration improves the new similarity scores comprises:

evaluating an objective function representing a search error.

17. The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise:

reverting back to original parameters of the configuration when adjusting the parameters of the configuration does not improve the new similarity scores.

18. The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise:

using adjusted parameters of the configuration when performing subsequent searches of the plurality of documents for documents of a same type as the first document when adjusting the parameters of the configuration improves the new similarity scores.

19. A system comprising:

one or more processors; and

one or more memory devices comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

providing a plurality of documents having similarity scores that are calculated in response to receiving a first document wherein the similarity scores indicate similarities between the first document and the plurality of documents;

receiving feedback from a user for a second document in the plurality of documents, wherein the feedback indicates whether the second document is relevant to the first document;

adjusting parameters of a configuration in response to the feedback, wherein the configuration is associated with documents having a same schema as the first document;

receiving a third document that has the same schema as the first document;

selecting the configuration based on the third document having the same schema as the first document; and

using adjusted parameters of the configuration to generate a plurality of queries, and executing the plurality of queries into a document repository to identify documents that are similar to the third document.

20. A method of using feedback to improve similarity scores for documents, the method comprising:

providing a plurality of documents having similarity scores that are calculated in response to receiving a first document wherein the similarity scores indicate similarities between the first document and the plurality of documents;

receiving feedback from a user for a second document in the plurality of documents, wherein the feedback indicates whether the second document is relevant to the first document;

adjusting parameters of a configuration in response to the feedback, wherein the configuration is associated with documents having a same schema as the first document;

receiving a third document that has the same schema as the first document;

selecting the configuration based on the third document having the same schema as the first document; and

using adjusted parameters of the configuration to generate a plurality of queries, and executing the plurality of queries into a document repository to identify documents that are similar to the third document.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2022
From: MATEI, LIVIU-SEBASTIAN; TROJAN, FILIP
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 060933/0894 →
Continuity (2)
Provisional Application 63239798 · Sep 1, 2021
Related Publication 20230068342A1 · Mar 2, 2023
References Cited (13)
US 6269368B1 · Diamond · 2001 [cited by examiner]
US 7272593B1 · Castelli · 2007 [cited by examiner]
US 7912701B1 · Gray · 2011 [cited by examiner]
US 8005823B1 · Marshall · 2011 [cited by examiner]
US 11983244B1 · Zhdanov · 2024 [cited by examiner]
US 20070208730A1 · Agichtein · 2007 [cited by examiner]
US 20120278341A1 · Ogilvy · 2012 [cited by examiner]
US 20130290339A1 · LuVogt · 2013 [cited by examiner]
US 20150169584A1 · Kwok · 2015 [cited by examiner]
US 20170091319A1 · Legrand · 2017 [cited by examiner]
US 20190057095A1 · Chakravarti · 2019 [cited by examiner]
US 20220004553A1 · Corvinelli · 2022 [cited by examiner]
US 20220215446A1 · Jiang · 2022 [cited by examiner]