IP Library › Granted Patent US 9,514,216
Granted Patent B2
US 9,514,216 · App. 14/480,528 · Granted Dec 6, 2016

Automatic classification of segmented portions of web pages

Inventors: Lei Duan (San Jose, CA); Fan Li (Redwood City, CA); Srinivas Vadrevu (Milpitas, CA); Emre Velipasaoglu (San Francisco, CA); Swapnil Hajela (Fremont, CA); Deepayan Chakrabarti (Austin, TX)
Assignee: Yahoo! Inc.
G06F17/30598G06F15/18G06F17/30873G06K9/6256G06N5/04G06N99/005G06Q10/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,514,216
App. No.
14/480,528
Granted
Dec 6, 2016
Kind
B2
Abstract

Exemplary methods and apparatuses are provided which may be used for classifying and indexing segmented portions of web pages and providing related information for use in information extraction and/or information retrieval systems. In an embodiment, an index of segmented portions may be used by a search engine to respond to a search query. In an embodiment, one or more machine learned models may be used to identify one or more feature properties of a plurality of segmented portions within one or more files, or otherwise inferable from the one or more files. In an embodiment, one or more machine learned models may be used to classify one or more of a plurality of segmented portions as being at least one of a plurality of segment types.

Claims (48)

1. A method comprising:

with one or more special purpose computing devices coupled to a memory:

accessing a plurality of segmented portions of at least one of a plurality of displayable web pages represented by one or more digital signals of one or more files stored in a memory, wherein a particular displayable web page of the plurality of displayable web pages comprises at least two of the plurality of segmented portions;

using one or more machine learned models for:

identifying one or more feature properties of the plurality of segmented portions within the one or more files, or otherwise inferable from the one or more files,

classifying the at least two of the plurality of segmented portions as being at least one of a plurality of segment types based, at least in part, on the one or more identified feature properties, the one or more identified feature properties comprising at least language feature properties of a language model of content to be displayed in one or more of the at least two of the plurality of segmented portions, and

determining content quality scores for at least two of the plurality of segmented portions of at least the particular displayable web page; and

storing one or more digital signals in the memory as part of an index for the plurality of segmented portions, the index being based, at least in part, on the segment type, the index indicating the content quality scores.

2. The method as recited in claim 1 , wherein the classifying of the at least two of the plurality of segmented portions further comprises:

establishing segmented portion key-value content for at least one of the one or more machine learned models.

3. The method as recited in claim 2 , wherein the segmented portion key-value content comprises a segment portion score.

4. The method as recited in claim 2 , further comprising:

with the one or more special purpose computing devices, indexing the at least one of the plurality of segmented portions based, at least in part, on at least a portion of the segmented portion key-value content.

5. The method as recited in claim 1 , further comprising:

with the one or more special purpose computing devices, combining the at least two of the plurality of segmented portions into a single segmented portion based, at least in part, on at least one feature for the two or more of the plurality of segmented portions.

6. The method as recited in claim 1 , further comprising:

with the one or more special purpose computing devices, training at least one of the one or more machine learned models based, at least in part, on editorial input for a sample set of segmented portions.

7. The method as recited in claim 1 , wherein at least one of the one or more machine learned models operates in an unsupervised mode.

8. The method as recited in claim 7 , wherein the one or more machine learned models operating in an unsupervised mode identifies one or more digital signals representing a vector space representation as one of the feature properties.

9. The method a recited in claim 1 , wherein at least one of the plurality of segmented portions comprises one or more digital signals representing at least one document object model (DOM) node.

10. The method as recited in claim 1 , further comprising:

with the one or more special purpose computing devices:

accessing the one or more files for the at least one of the plurality of displayable web pages from the memory; and

identifying one or more digital signals representing the plurality of segmented portions based, at least in part, on an initial set of properties identifiable in one or more digital signals representing the at least one file.

11. An apparatus comprising:

a memory having stored therein one or more digital signals to represent at least one file for a particular displayable web page to comprise at least two of a plurality of segmented portions;

at least one processing unit coupled to the memory and programmed with instructions to:

access the plurality of segmented portions of the at least one displayable web page, and use one or more machine learned models to:

identify one or more feature properties of the plurality of segmented portions within the one or more files, or otherwise to be inferable from the one or more files,

classify the at least two of the plurality of segmented portions as at least one of a plurality of segment types to be based, at least in part, on the one or more feature properties to be identified, the one or more feature properties to be identified are to comprise at least language feature properties of language model of content to be displayed in one or more of the at least two of the plurality of segmented portions, and

determine content quality scores for at least two of the plurality of segmented portions of at least the particular displayable web page; and

establish an index in the memory, the index to be established for the plurality of segmented portions and to be based, at least in part, on the segment type, the index to indicate the content quality scores.

12. The apparatus as recited in claim 11 , wherein the at least one processing unit is to be programmed with instructions to establish segmented portion key-value content for at least one of the one or more machine learned models, for the at least one of the plurality of segmented portions, and to index the at least one of the plurality of segmented portions to be based, at least in part, on at least a portion of the segmented portion key-value content.

13. The apparatus as recited in claim 11 , wherein the at least one processing unit is to be programmed with instructions to combine the at least two of the plurality of segmented portions into a single segmented portion to be based, at least in part, on at least one feature for the two or more of the plurality of segmented portions.

14. The apparatus as recited in claim 11 , wherein at least one of the one or more machine learned models is to operate in an unsupervised mode and is to identify one or more digital signals to represent a vector space representation as one of the feature properties.

15. The apparatus as recited in claim 11 , wherein the at least one processing unit is to be programmed with instructions to identify the plurality of segmented portions to be based, at least in part, on an initial set of properties to be identifiable in the at least one file.

16. An article comprising:

a non-transitory computer readable medium having computer implementable instructions stored thereon to be implemented by one or more processing units in a computing device to transform the computing device into a special purpose device to:

access a plurality of segmented portions of at least one of a plurality of web pages to be displayable by one or more digital signals of one or more files stored in a memory, wherein a particular web page of the plurality of web pages is to comprise at least two of the plurality of segmented portions;

use one or more machine learned models to:

identify one or more feature properties of the plurality of segmented portions within the one or more files, or otherwise to be inferable from the one or more files,

classify the at least two of the plurality of segmented portions as at least one of a plurality of segment types to be based, at least in part, on the one or more identified feature properties, the one or more feature properties to be identified are to comprise at least language feature properties of language model of content to be displayed in one or more of the at least two of the plurality of segmented portions, and

determine content quality scores for at least two of the plurality of segmented portions of at least the particular web page; and

establish one or more digital signals to represent an index within a memory to be coupled to the one or more processing units, the index to be established for the plurality of segmented portions and to be based, at least in part, on the segment type, the index to indicate the content quality scores.

17. The article as recited in claim 16 , wherein the computer implementable instructions to be implemented by the one or more processing units in the computing device are to operatively transform the computing device into the special purpose device to establish one or more digital signals to represent segmented portion key-value content for at least one of the one or more machine learned models, for the at least one of the plurality of segmented portions, and to index at least one of the plurality of segmented portions to be based, at least in part, on at least a portion of the segmented portion key-value content.

18. The article as recited in claim 16 , wherein the computer implementable instructions to be implemented by the one or more processing units in the computing device are to operatively transform the computing device into the special purpose device to combine the at least two of the plurality of segmented portions into a single segmented portion to be based, at least in part, on at least one feature for the two or more of the plurality of segmented portions.

19. The article as recited in claim 16 , wherein at least one of the one or more machine learned models is to operate in an unsupervised mode and is to identify one or more digital signals to represent a vector space representation as one of the feature properties.

20. The article as recited in claim 16 , wherein the computer implementable instructions to be implemented by the one or more processing units in the computing device are to operatively transform the computing device into the special purpose device to identify the plurality of segmented portions to be based, at least in part, on one or more digital signals to represent an initial set of properties to be identifiable in the at least one file.

Assignments (9)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE ASSIGNOR NAME PREVIOUSLY RECORDED AT REEL: 052853 FRAME: 0153. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Mar 29, 2021
From: R2 SOLUTIONS LLC
To: STARBOARD VALUE INTERMEDIATE FUND LP, AS COLLATERAL AGENT
Reel/Frame 056832/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 12, 2021
From: EXCALIBUR IP, LLC
To: R2 SOLUTIONS LLC
Reel/Frame 055283/0483 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED ON REEL 053654 FRAME 0254. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST GRANTED PURSUANT TO THE PATENT SECURITY AGREEMENT PREVIOUSLY RECORDED. Recorded Dec 30, 2020
From: STARBOARD VALUE INTERMEDIATE FUND LP
To: R2 SOLUTIONS LLC
Reel/Frame 054981/0377 →
RELEASE OF SECURITY INTEREST IN PATENTS Recorded Jul 8, 2020
From: STARBOARD VALUE INTERMEDIATE FUND LP
To: ACACIA RESEARCH GROUP LLC; AMERICAN VEHICULAR SCIENCES LLC; BONUTTI SKELETAL INNOVATIONS LLC; CELLULAR COMMUNICATIONS EQUIPMENT LLC; INNOVATIVE DISPLAY TECHNOLOGIES LLC; LIFEPORT SCIENCES LLC; LIMESTONE MEMORY SYSTEMS LLC; MOBILE ENHANCEMENT SOLUTIONS LLC; MONARCH NETWORKING SOLUTIONS LLC; NEXUS DISPLAY TECHNOLOGIES LLC; PARTHENON UNIFIED MEMORY ARCHITECTURE LLC; R2 SOLUTIONS LLC; SAINT LAWRENCE COMMUNICATIONS LLC; STINGRAY IP SOLUTIONS LLC; SUPER INTERCONNECT TECHNOLOGIES LLC; TELECONFERENCE SYSTEMS LLC; UNIFICATION TECHNOLOGIES LLC
Reel/Frame 053654/0254 →
PATENT SECURITY AGREEMENT Recorded Jun 5, 2020
From: ACACIA RESEARCH GROUP LLC; AMERICAN VEHICULAR SCIENCES LLC; BONUTTI SKELETAL INNOVATIONS LLC; CELLULAR COMMUNICATIONS EQUIPMENT LLC; INNOVATIVE DISPLAY TECHNOLOGIES LLC; LIFEPORT SCIENCES LLC; LIMESTONE MEMORY SYSTEMS LLC; MERTON ACQUISITION HOLDCO LLC; MOBILE ENHANCEMENT SOLUTIONS LLC; MONARCH NETWORKING SOLUTIONS LLC; NEXUS DISPLAY TECHNOLOGIES LLC; PARTHENON UNIFIED MEMORY ARCHITECTURE LLC; R2 SOLUTIONS LLC; SAINT LAWRENCE COMMUNICATIONS LLC; STINGRAY IP SOLUTIONS LLC; SUPER INTERCONNECT TECHNOLOGIES LLC; TELECONFERENCE SYSTEMS LLC; UNIFICATION TECHNOLOGIES LLC
To: STARBOARD VALUE INTERMEDIATE FUND LP, AS COLLATERAL AGENT
Reel/Frame 052853/0153 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2016
From: YAHOO! INC.
To: EXCALIBUR IP, LLC
Reel/Frame 038950/0592 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2016
From: EXCALIBUR IP, LLC
To: YAHOO! INC.
Reel/Frame 038951/0295 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2016
From: YAHOO! INC.
To: EXCALIBUR IP, LLC
Reel/Frame 038383/0466 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 26, 2014
From: DUAN, LEI; LI, FAN; VADREVU, SRINIVAS; VELIPASAOGLU, EMRE; HAJELA, SWAPNIL; CHAKRABARTI, DEEPAYAN
To: YAHOO! INC.
Reel/Frame 034271/0979 →
Continuity (2)
Continuation 12538776 · Aug 10, 2009
Related Publication 20150066934A1 · Mar 5, 2015