IP Library Granted Patent US 10,824,688
Granted Patent B2
US 10,824,688 · App. 15/397,114 · Granted Nov 3, 2020

System, method, and computer program product for generation of local content corpus

Inventors: Roger Castillo (Palo Alto, CA); Thomas Jack (Austin, TX)
Assignee: Groupon, Inc.
G06F16/9537G06F16/22G06F16/248G06F16/285G06F16/29G06F16/58
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,824,688
App. No.
15/397,114
Granted
Nov 3, 2020
Kind
B2
Abstract

Various methods for generating a content corpus populated with content related to a particular geographic area are provided herein. One example method comprises, for each document in an initial local content corpus, applying a first set of heuristic filters to the raw content of each document, identifying at least a second term, applying a second set of heuristic filters to the raw content of each document, the second set of heuristic filters associated with the second term, iteratively performing the identification of additional terms and application of an additional set of heuristic filters associated with the additional terms until each identifiable term is extracted, determining a level on a geographic containment hierarchy indicative of a location to which each document from the set of documents is local, and for each place in a gazette, and for each document, determining a set of points in polygons indicative of its locality.

Claims (57)

1. A computer-implemented method, comprising:

for each document in an initial local content corpus, the content corpus comprising a plurality of documents, each document comprised of raw content, applying a first set of heuristic filters to the raw content of each document;

identifying, based on the application of the first set of heuristic filters, at least a second term;

applying a second set of heuristic filters to the raw content of each document, the second set of heuristic filters associated with the second term;

iteratively performing the identification of additional terms and application of an additional set of heuristic filters associated with the additional terms until each identifiable term is extracted;

determining a level on a geographic containment hierarchy indicative of a location to which each document from the set of documents is local, the level on the geographic containment hierarchy to which each document is determined to be local being dependent on a configurable resolution setting; and

for each place in a gazette, and for each document in the set of documents, determining a set of points in polygons indicative of its locality.

2. The computer-implemented method according to claim 1 , wherein the iterative performance of the identification of additional terms and the application of the additional set of heuristic filters associated with the additional terms is configured to generate a set of local terms referencing at least people, places, and organizations.

3. The computer-implemented method according to claim 1 , further comprising:

determining a presence of geospatial information within a particular document; and

utilizing the geospatial information in conjunction with particular term frequencies to determine to eliminate the document from the initial local content corpus.

4. The computer-implemented method according to claim 1 , further comprising:

for each place in the gazette, generating the initial set of local terms, generating the initial set of local terms comprising: accessing a document, and calculating an initial weighting for at least a portion of the initial set of local terms; and

creating the initial local content corpus utilizing the initial set of local terms, the initial local content corpus containing documents that are semantically related to each other with respect to the place, the local content corpus configured for providing the system with relevant and unranked documents.

5. The computer-implemented method according to claim 1 , further comprising:

for each place in a gazette, and for each document in the set of documents, determine a set of weightings associated with the terms in the document.

6. The computer-implemented method according to claim 1 , wherein the first set of heuristics comprises at least a topological filter with respect to the place.

7. The computer-implemented method according to claim 1 , further comprising:

obtaining the content from a plurality of disparate sources.

8. A computer program product comprising at least one non-transitory computer readable medium storing instructions translatable by one or more server machines to perform:

for each document in an initial local content corpus, the content corpus comprising a plurality of documents, each document comprised of raw content, apply a first set of heuristic filters to the raw content of each document;

identify, based on the application of the first set of heuristic filters, at least a second term;

apply a second set of heuristic filters to the raw content of each document, the second set of heuristic filters associated with the second term;

iteratively perform the identification of additional terms and application of an additional set of heuristic filters associated with the additional terms until each identifiable term is extracted;

determine a level on a geographic containment hierarchy indicative of a location to which each document from the set of documents is local, the level on the geographic containment hierarchy to which each document is determined to be local being dependent on a configurable resolution setting; and

for each place in a gazette, and for each document in the set of documents, determine a set of points in polygons indicative of its locality.

9. The computer program product according to claim 8 , wherein the iterative performance of the identification of additional terms and the application of the additional set of heuristic filters associated with the additional terms is configured to generate a set of local terms referencing at least people, places, and organizations.

10. The computer program product according to claim 8 , wherein the instructions are further translatable by the one or more server machines to perform:

determining a presence of geospatial information within a particular document; and

utilizing the geospatial information in conjunction with particular term frequencies to determine to eliminate the document from the initial local content corpus.

11. The computer program product according to claim 8 , wherein the instructions are further translatable by the one or more server machines to perform:

for each place in the gazette, generating the initial set of local terms, generating the initial set of local terms comprising: accessing a document, and calculating an initial weighting for at least a portion of the initial set of local terms; and

creating the initial local content corpus utilizing the initial set of local terms, the initial local content corpus containing documents that are semantically related to each other with respect to the place, the local content corpus configured for providing the system with relevant and unranked documents.

12. The computer program product according to claim 8 , wherein the instructions are further translatable by the one or more server machines to perform:

for each place in a gazette, and for each document in the set of documents, determine a set of weightings associated with the terms in the document.

13. The computer program product according to claim 8 , wherein the first set of heuristics comprises at least a topological filter with respect to the place.

14. The computer program product according to claim 8 , wherein the instructions are further translatable by the one or more server machines to perform:

obtaining the content from a plurality of disparate sources.

15. A system, comprising:

one or more server machines; and

at least one non-transitory computer readable medium storing instructions translatable by the one or more server machines to perform:

for each document in an initial local content corpus, the content corpus comprising a plurality of documents, each document comprised of raw content, apply a first set of heuristic filters to the raw content of each document;

identify, based on the application of the first set of heuristic filters, at least a second term;

apply a second set of heuristic filters to the raw content of each document, the second set of heuristic filters associated with the second term;

iteratively perform the identification of additional terms and application of an additional set of heuristic filters associated with the additional terms until each identifiable term is extracted;

determine a level on a geographic containment hierarchy indicative of a location to which each document from the set of documents is local, the level on the geographic containment hierarchy to which each document is determined to be local being dependent on a configurable resolution setting; and

for each place in a gazette, and for each document in the set of documents, determine a set of points in polygons indicative of its locality.

16. The system according to claim 15 , wherein the iterative performance of the identification of additional terms and the application of the additional set of heuristic filters associated with the additional terms is configured to generate a set of local terms referencing at least people, places, and organizations.

17. The system according to claim 15 , wherein the instructions are further translatable by the one or more server machines to perform:

determining a presence of geospatial information within a particular document; and

utilizing the geospatial information in conjunction with particular term frequencies to determine to eliminate the document from the initial local content corpus.

18. The system according to claim 15 , wherein the instructions are further translatable by the one or more server machines to perform:

for each place in the gazette, generating the initial set of local terms, generating the initial set of local terms comprising: accessing a document, and calculating an initial weighting for at least a portion of the initial set of local terms; and

creating the initial local content corpus utilizing the initial set of local terms, the initial local content corpus containing documents that are semantically related to each other with respect to the place, the local content corpus configured for providing the system with relevant and unranked documents.

19. The system according to claim 15 , wherein the instructions are further translatable by the one or more server machines to perform:

for each place in a gazette, and for each document in the set of documents, determine a set of weightings associated with the terms in the document.

20. The system according to claim 15 , wherein the first set of heuristics comprises at least a topological filter with respect to the place.

Assignments (6)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 12, 2024
From: GROUPON, INC.
To: BYTEDANCE INC.
Reel/Frame 068833/0811 →
RELEASE OF SECURITY INTEREST Recorded Feb 26, 2024
From: JPMORGAN CHASE BANK, N.A.
To: GROUPON, INC.; LIVINGSOCIAL, LLC (F/K/A LIVINGSOCIAL, INC.)
Reel/Frame 066676/0001 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RIGHTS Recorded Feb 26, 2024
From: JPMORGAN CHASE BANK, N.A.
To: GROUPON, INC.; LIVINGSOCIAL, LLC (F/K/A LIVINGSOCIAL, INC.)
Reel/Frame 066676/0251 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 11, 2020
From: CASTILLO, ROGER H.; JACK, THOMAS
To: BORROWED SUGAR, INC.
Reel/Frame 053455/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 11, 2020
From: BORROWED SUGAR, INC.
To: GROUPON, INC.
Reel/Frame 053456/0593 →
SECURITY INTEREST Recorded Jul 23, 2020
From: GROUPON, INC.; LIVINGSOCIAL, LLC
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 053294/0495 →
Continuity (3)
Continuation 13444691 · Apr 11, 2012
Provisional Application 61474095 · Apr 11, 2011
Related Publication 20170344565A1 · Nov 30, 2017