Systems and methods for evaluating page content
Systems, methods, and non-transitory computer-readable media can determine a set of candidate values for a field in a page. The set of candidate values can be evaluated for accuracy based at least in part on a machine learning model, wherein the machine learning model outputs a respective score for each candidate value that measures an accuracy of the candidate value for the field in the page. A best scoring candidate value can be determined from the set of candidate values. The field in the page can be associated with the best scoring candidate value.
1. A computer-implemented method comprising:
determining, by a computing system, a set of candidate values for a field in a page;
evaluating, by the computing system, the set of candidate values for accuracy based at least in part on a machine learning model, wherein the evaluating further comprises:
obtaining, by the computing system, a score for a candidate value from the machine learning model based on a feature vector that represents the candidate value, wherein the feature vector is associated with at least one feature that represents a data pipeline endorsement of the candidate value by at least one data pipeline, wherein the data pipeline endorsement is weighted based on a consensus score associated with the at least one data pipeline, wherein the consensus score measures a rate at which the at least one data pipeline has historically provided data pipeline endorsements for candidate values that were endorsed by a threshold amount of users;
determining, by the computing system, a best scoring candidate value from the set of candidate values; and
associating, by the computing system, the field in the page with the best scoring candidate value.
2. The computer-implemented method of claim 1 , wherein the field corresponds to at least one of: a page category field, a website field, a phone number field, an hours of operation field, and a physical address field.
3. The computer-implemented method of claim 1 , further comprising:
causing, by the computing system, the field in the page to be populated with the best scoring candidate value.
4. The computer-implemented method of claim 1 , further comprising:
providing, by the computing system, the best scoring candidate value as a recommendation for populating the field in the page.
5. The computer-implemented method of claim 1 , wherein the feature vector includes a feature representing a set of weighted user endorsements for the candidate value, wherein a user endorsement is weighted based on a credibility score associated with the user, the credibility score measuring a credibility of the user.
6. The computer-implemented method of claim 5 , wherein the feature representing the set of weighted user endorsements for the candidate value is determined based at least in part on an activation function or a logit function.
7. The computer-implemented method of claim 1 , wherein the feature vector includes a feature representing a set of weighted data pipeline endorsements for the candidate value, wherein the set of weighted data pipeline endorsements includes the at least one data pipeline endorsement.
8. The computer-implemented method of claim 7 , wherein the feature representing the set of weighted data pipeline endorsements is determined based at least in part on an activation function or a logit function.
9. The computer-implemented method of claim 1 , further comprising:
causing, by the computing system, one or more users to be polled to confirm or improve an accuracy of the best scoring candidate value.
10. A system comprising:
at least one processor; and
a memory storing instructions that, when executed by the at least one processor, cause the system to perform:
determining a set of candidate values for a field in a page;
evaluating the set of candidate values for accuracy based at least in part on a machine learning model, wherein the evaluating further comprises:
obtaining a score for a candidate value from the machine learning model based on a feature vector that represents the candidate value, wherein the feature vector is associated with at least one feature that represents a data pipeline endorsement of the candidate value by at least one data pipeline, wherein the data pipeline endorsement is weighted based on a consensus score associated with the at least one data pipeline, wherein the consensus score measures a rate at which the at least one data pipeline has historically provided data pipeline endorsements for candidate values that were endorsed by a threshold amount of users;
determining a best scoring candidate value from the set of candidate values; and
associating the field in the page with the best scoring candidate value.
11. The system of claim 10 , wherein the field corresponds to at least one of: a page category field, a website field, a phone number field, an hours of operation field, and a physical address field.
12. The system of claim 10 , wherein the instructions further cause the system to perform:
causing the field in the page to be populated with the best scoring candidate value.
13. The system of claim 10 , wherein the instructions further cause the system to perform:
providing the best scoring candidate value as a recommendation for populating the field in the page.
14. A non-transitory computer-readable storage medium including instructions that, when executed by at least one processor of a computing system, cause the computing system to perform a method comprising:
determining a set of candidate values for a field in a page;
evaluating the set of candidate values for accuracy based at least in part on a machine learning model, wherein the evaluating further comprises:
obtaining a score for a candidate value from the machine learning model based on a feature vector that represents the candidate value, wherein the feature vector is associated with at least one feature that represents a data pipeline endorsement of the candidate value by at least one data pipeline, wherein the data pipeline endorsement is weighted based on a consensus score associated with the at least one data pipeline, wherein the consensus score measures a rate at which the at least one data pipeline has historically provided data pipeline endorsements for candidate values that were endorsed by a threshold amount of users;
determining a best scoring candidate value from the set of candidate values; and
associating the field in the page with the best scoring candidate value.
15. The non-transitory computer-readable storage medium of claim 14 , wherein the field corresponds to at least one of: a page category field, a website field, a phone number field, an hours of operation field, and a physical address field.
16. The non-transitory computer-readable storage medium of claim 14 , wherein the instructions further cause the computing system to perform:
causing the field in the page to be populated with the best scoring candidate value.
17. The non-transitory computer-readable storage medium of claim 14 , wherein the instructions further cause the computing system to perform:
providing the best scoring candidate value as a recommendation for populating the field in the page.