IP Library › Patent Application 16668544
Patent Application
App. No. 16/668,544

EXTRACTING UNSTRUCTURED DEMOGRAPHIC INFORMATION FROM A DATA SOURCE IN A STRUCTURED MANNER

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/668,544
Abstract

The present disclosure is directed to systems and methods for identifying demographic information in a marked up document. The method may include: detecting a plurality of fields representing demographic information in a marked up document; extracting a set of features based the detected demographic information; based at least in part on the set of features, determining whether the plurality of fields represent demographic information of a single provider using a machine learning model trained according to other marked up documents; and when the plurality of fields are determined to represent demographic information of a single provider, associating the plurality of fields to represent demographic information of the single provider.

Claims (52)

1 . A method of identifying demographic information in a marked up document, comprising:

(a) detecting a plurality of fields representing demographic information in a marked up document;

(b) extracting a set of features based the detected demographic information;

(c) based at least in part on the set of features, determining whether the plurality of fields represent demographic information of a single provider using a machine learning model trained according to other marked up documents; and

(d) when the plurality of fields are determined to represent demographic information of a single provider, associating the plurality of fields to represent demographic information of the single provider.

2 . The method of claim 1 , the determining (c) comprises:

rendering the marked up document;

determining where, in the rendered marked-up document, each of the plurality of fields is located; and

calculating a geometric distance between the respective locations of the plurality of fields in the rendered marked-up document, and

wherein the determining (d) occurs based at least in part on the calculated geometric distance.

3 . The method of claim 1 , the determining (c) comprises:

representing the marked up document in a document object model including a plurality of interconnected nodes;

determining where, in the document object model, each of the plurality of fields is located; and

calculating a number of hops between the respective locations of the plurality of fields in the rendered marked-up document, and

wherein the determining (d) occurs based on the number of hops.

4 . The method of claim 1 , the determining (c) comprises:

(e) determining how many a plurality of fields representing demographic information of a particular type are in a marked up document,

wherein determining occurs based at least in part on the number of fields of the particular type determined in (e).

5 . The method of claim 4 , wherein the particular type is at least one of an address or phone number.

6 . The method of claim 1 , further comprising:

(e) training the machine learning model using a sample set of pages and corresponding locations of identified fields.

7 . The method of claim 6 , further comprising:

(f) identifying fields in the sample set of pages based on known tags.

8 . The method of claim 1 , wherein the training the model comprises training the model using one or more of a support vector machines algorithm, linear regression algorithm, logistic regression algorithm, naive Bayes algorithm, linear discriminant analysis algorithm, decision trees algorithm, k-nearest neighbor algorithm, neural networks algorithm, and a similarity learning algorithm.

9 . The method of claim 1 , wherein the extracting the set of features comprises identifying a distance between two or more types of demographic information from among the plurality of types of demographic information.

10 . The method of claim 1 , wherein the extracting the set of features comprises identifying a number pairs of demographic information on a given site of a data source.

11 . A non-transitory program storage device having instructions stored thereon that, when executed by at least one computing device, causes the at least one computing device to perform a method, the method comprising:

(a) detecting a plurality of fields representing demographic information in a marked up document;

(b) extracting a set of features based the detected demographic information;

(c) based at least in part on the set of features, determining whether the plurality of fields represent demographic information of a single provider using a machine learning model trained according to other marked up documents; and

(d) when the plurality of fields are determined to represent demographic information of a single provider, associating the plurality of fields to represent demographic information of the single provider.

12 . The non-transitory program storage device of claim 11 , the determining (c) comprises:

rendering the marked up document;

determining where, in the rendered marked-up document, each of the plurality of fields is located; and

calculating a geometric distance between the respective locations of the plurality of fields in the rendered marked-up document, and

wherein the determining (d) occurs based at least in part on the calculated geometric distance.

13 . The non-transitory program storage device of claim 11 , the determining (c) comprises:

representing the marked up document in a document object model including a plurality of interconnected nodes;

determining where, in the document object model, each of the plurality of fields is located; and

calculating a number of hops between the respective locations of the plurality of fields in the rendered marked-up document, and

wherein the determining (d) occurs based on the number of hops.

14 . The non-transitory program storage device of claim 11 , the determining (c) comprises:

(e) determining how many a plurality of fields representing demographic information of a particular type are in a marked up document,

wherein determining occurs based at least in part on the number of fields of the particular type determined in (e).

15 . The non-transitory program storage device of claim 14 , wherein the particular type is at least one of an address or phone number.

16 . The non-transitory program storage device of claim 11 , the method further comprising:

(e) training the machine learning model using a sample set of pages and corresponding locations of identified fields.

17 . The non-transitory program storage device of claim 16 , the method further comprising:

(f) identifying fields in the sample set of pages based on known tags.

18 . The non-transitory program storage device of claim 11 , wherein the training the model comprises training the model using one or more of a support vector machines algorithm, linear regression algorithm, logistic regression algorithm, naive Bayes algorithm, linear discriminant analysis algorithm, decision trees algorithm, k-nearest neighbor algorithm, neural networks algorithm, and a similarity learning algorithm.

19 . The non-transitory program storage device of claim 11 , wherein the extracting the set of features comprises identifying a distance between two or more types of demographic information from among the plurality of types of demographic information.

20 . The non-transitory program storage device of claim 11 , wherein the extracting the set of features comprises identifying a number pairs of demographic information on a given site of a data source from among the plurality of data sources.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2026
From: VEDA DATA SOLUTIONS, INC
To: H1 INSIGHTS, INC.
Reel/Frame 073623/0895 →
RELEASE OF SECURITY INTEREST Recorded Jun 4, 2025
From: COMERICA BANK
To: VEDA DATA SOLUTIONS, INC.
Reel/Frame 071309/0392 →
SECURITY INTEREST Recorded Nov 27, 2023
From: VEDA DATA SOLUTIONS, INC.
To: COMERICA BANK
Reel/Frame 065668/0675 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2019
From: VERO-CIRO, CARLOS; LINDNER, ROBERT RAYMOND
To: VEDA DATA SOLUTIONS, INC., SUITE 200
Reel/Frame 050883/0214 →