IP Library Granted Patent US 11,170,759
Granted Patent B2
US 11,170,759 · App. 16/724,960 · Granted Nov 9, 2021

System and method for discriminating removing boilerplate text in documents comprising structured labelled text elements

Inventors: David Alexander Sim (Cambridge, GB); David Paul Austen Ryland (Cambridge, GB)
Assignee: Verint Systems UK Limited
G10L13/08G10L13/047H04L67/02H04L67/42
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,170,759
App. No.
16/724,960
Granted
Nov 9, 2021
Kind
B2
Abstract

A method, system, and computer program product for discriminating boilerplate text in documents, such as web pages. An example method includes: receiving documents structured as labelled text elements; generating a local language model for each labelled text element of the received documents; comparing local language models for different labelled text elements that have the same label; for each comparison of local language models, deriving a similarity indicator, and using the similarity indicators of all the comparisons to derive a similarity score for that label; using the similarity scores to determine labels associated with text elements comprising boilerplate text; and providing the textual content of the labelled text elements to a receiving computer system; and identifying the textual content of labelled text elements that include boilerplate text.

Claims (45)

1. A computer implemented method comprising:

i) receiving from a server, documents structured as labelled text elements;

ii) generating a local language model for each labelled text element of the received documents;

iii) comparing local language models for different labelled text elements that have the same label;

iv) for each comparison of local language models in iii), deriving a similarity indicator, and using the similarity indicators of all the comparisons to derive a similarity score for that label;

v) using similarity scores to determine labels associated with text elements comprising boilerplate text; and

vi) providing textual content of the labelled text elements to a receiving computer system, the determination made in v) being used to: provide the textual content of selected labelled text elements to the receiving computer system and/or to provide information to the receiving computer system identifying the textual content of labelled text elements that comprise boilerplate text.

2. The computer implemented method according to claim 1 wherein the receiving computer system comprises an information retrieval system.

3. The computer implemented method according to claim 2 wherein the textual content provided in vi) and received by the information retrieval system is used to populate an index of the information retrieval system.

4. The computer implemented method according to claim 1 wherein the receiving computer system is a text to speech system.

5. The computer implement method according to claim 1 further comprising:

vii) generating a global language model from combined text of all text elements having the same label; and

viii) repeating vii) for each label to provide a global language model for each label; and wherein iii) comprises: comparing the local language model for each text element with the global language model for that label to derive a similarity indicator for each local language model.

6. The computer implemented method according to claim 1 , wherein iii) comprises: comparing the local language model for each text element with other local language models for other text elements with the same label to derive, for each comparison, a similarity indicator.

7. The computer implemented method according to claim 1 where iv) comprises: averaging the similarity indicators for a label to derive the similarity score for the label.

8. The computer implemented method according to claim 1 wherein the similarity score for a label is indicative of similarity of text between the elements of that label.

9. The computer implemented method according to claim 1 wherein v) comprises comparing the similarity scores for labels with a threshold.

10. The computer implemented method according to claim 1 comprising comparing similarity scores of different labels to determine one or more labels which from relative similarity score compared to other labels, are associated with text elements that comprise text which is relatively similar to other elements of that label.

11. The computer implemented method according to claim 1 wherein the local language model is a probabilistic model.

12. The computer implemented method according to claim 1 wherein the documents are web pages of a website.

13. The computer implemented method according to claim 12 , wherein the documents are a minimum of 2000 web pages of a website.

14. A system comprising:

a computer, and a computer readable medium coupled to the computer having instructions stored thereon which, when executed by the computer cause the computer to perform operations comprising:

i) receiving documents structured as labelled text elements from a server;

ii) generating a local language model for each text element of the received documents;

iii) comparing local language models for different text elements having the same label to derive similarity indicators, and using the similarity indictors to derive a similarity score for that label;

iv) using the similarity score to determine labels whose elements comprise boilerplate text; and

v) passing to an information retrieval system indicators of the labels determined in iv) above whose elements comprise boilerplate text.

15. A computer program product for a processing system comprised of a computer, the computer program product comprising a non-transitory computer readable storage medium having computer readable program code embodied therewith, the computer readable program code including computer instructions, where a processor, responsive to executing the computer instructions, performs operations comprising:

i) receiving documents structured as labelled text elements from a server;

ii) generating a local language model for each text element of the received documents;

iii) comparing local language models for different text elements having the same label to derive similarity indicators, and using the similarity indictors to derive a similarity score for that label;

iv) using the similarity score to determine labels whose elements comprise boilerplate text; and

v) passing to an information retrieval system indicators of thorn the labels determined in iv) above whose elements comprise boilerplate text.

16. The computer program product of claim 15 wherein the operations further comprise:

vi) generating a global language model from combined text of all text elements having the same label; and

vii) repeating vii) for each label to provide a global language model for each label; and

wherein iii) comprises: comparing the local language model for each text element with the global language model for that label to derive a similarity indicator for each local language model.

17. The computer program product of claim 15 wherein the operations further comprise:

comparing the local language model for each text element with other local language models for other text elements with the same label to derive, for each comparison, a similarity indicator.

18. The computer program product of claim 15 wherein the operations further comprise:

averaging the similarity indicators for a label to derive the similarity score for the label.

19. The computer program product of claim 15 wherein iv) comprises comparing similarity scores for labels with a threshold.

20. The computer program product of claim 15 wherein the operations further comprise:

comparing similarity scores of different labels to determine one or more labels which from relative similarity score compared to other labels, are associated with text elements that comprise text which is relatively similar to other elements of that label.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 23, 2019
From: SIM, DAVID ALEXANDER; RYLAND, DAVID PAUL AUSTEN
To: VERINT SYSTEMS UK LIMITED
Reel/Frame 051355/0541 →
Priority Claims (1)
GB 1821327 · Dec 31, 2018 · national
Continuity (1)
Related Publication 20200219481A1 · Jul 9, 2020
Cited By (4)
US 12,190,059 US 12,242,806 US 12,626,058 US 12,705,420