IP Library Granted Patent US 11,593,564
Granted Patent B2
US 11,593,564 · App. 16/901,677 · Granted Feb 28, 2023

Systems and methods for extracting patent document templates from a patent corpus

Inventors: Ian C. Schick (Ojai, CA); Kevin Knight (Marina del Rey, CA)
Assignee: Specifio, Inc.
G06F40/30G06F40/137G06F40/284G06F40/56G06Q10/10G06Q50/184G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,593,564
App. No.
16/901,677
Granted
Feb 28, 2023
Kind
B2
Abstract

Systems, methods, and storage media for extracting patent document templates from a patent corpus are disclosed. Exemplary implementations may: obtain a patent corpus; receive one or more parameters; determine one or more subsets of the patent corpus by filtering the patent corpus based on the one or more parameters; identify one or more document clusters within individual ones of the one or more subsets of the patent corpus; obtain a patent document template corresponding to the first document cluster; and/or perform other operations.

Claims (44)

1. A method for extracting patent document templates from a patent corpus, the method comprising:

obtaining a patent corpus, the patent corpus including a plurality of patent documents;

receiving one or more parameters, the one or more parameters including a first parameter;

determining one or more subsets of the patent corpus by filtering the patent corpus based on the one or more parameters, the one or more subsets of the patent corpus including a first subset of the patent corpus;

identifying one or more document clusters within individual ones of the one or more subsets of the patent corpus, the one or more document clusters including a first document cluster within the first subset of the patent corpus, wherein the first document cluster includes a plurality of patent documents sharing common text, wherein the identifying the one or more document clusters includes comparing some or all combinations of pairs of patent documents contained in a given subset of the patent corpus, and wherein the comparing the some or all combinations of the pairs of patent documents contained in the given subset of the patent corpus includes:

identifying one or more specific patent document sections in individual patent documents included in the one or more subsets of the patent corpus where related patent documents frequently share common spans of text, wherein the specific patent document sections include one or more of a first portion of a summary section, a last portion of a summary section, a first portion of a brief description of drawing section, a last portion of a brief description of drawings section, a first portion of a detailed description section, or a last portion of a detailed description section;

obtaining the spans of text included in the one or more of the specific patent document sections of the individual patent documents included in the one or more subsets of the patent corpus; and

comparing, for individual ones of the pairs of patent documents, the spans of text obtained from the individual patent documents of the pairs of patent documents to determine if they are common text; and

obtaining a patent document template corresponding to the first document cluster, the patent document template including the common text of the plurality of patent documents sharing common text.

2. The method of claim 1 , wherein the individual patent documents of the plurality of patent documents include one or both of published patents or published patent applications.

3. The method of claim 1 , wherein the plurality of patent documents corresponds to a specific patent jurisdiction, wherein the patent corpus is provided by a patent office, and wherein the patent corpus is in a public domain.

4. The method of claim 1 , wherein the plurality of patent documents corresponds to a publication date range.

5. The method of claim 1 , wherein the individual patent documents are in an electronic form.

6. The method of claim 1 , wherein a given one of the one or more parameters include one or more of a patent assignee, a name of a competitor of a patent assignee, an inventor name, a name of a law firm that prepared a corresponding patent application, a name of an attorney who prepared a corresponding patent application, a name of a law firm that filed a corresponding patent application, a name of an attorney who filed a corresponding patent application, a name of a law firm handling prosecution of a corresponding patent application, a name of an attorney prosecuting a corresponding patent application, an examiner associated with examination of a corresponding patent application, a patent application filing date, a patent application filing date range, a patent application publication date, a patent application publication date range, a patent issuance date, a patent issuance date range, a patent classification, a range of patent classifications, or an identifier of a cited prior art reference corresponding to a patent application.

7. The method of claim 1 , wherein the first subset of the patent corpus includes a plurality of subset documents, the plurality of subset documents including patent documents associated with a specific patent assignee and a specific law firm responsible for filing underlying patent applications associated with the plurality of subset documents.

8. The method of claim 1 , wherein the spans of text are determined to be the common text if they are similar or identical text, wherein the spans of text that are similar or identical include a first span.

9. The method of claim 8 , wherein the first span includes one or more of a sentence, a paragraph, or a group of adjacent paragraphs.

10. The method of claim 8 , wherein the common text included in the plurality of patent documents sharing common text includes one or more of boilerplate language, a stock description, a stock description of a stock drawing figure, or a stock definition.

11. The method of claim 1 , wherein the identifying the one or more document clusters includes encoding spans such that individual spans are represented by unique encodings.

12. The method of claim 11 , wherein the encoding the spans includes applying one or more of a hash function, character encoding, or semantics encoding to the individual spans.

13. The method of claim 11 , wherein the unique encodings enable rapid comparison between patent documents contained in a given document cluster.

14. The method of claim 1 , wherein the patent document template is a basis for a new patent application.

15. The method of claim 1 , wherein one or more of the identifying the one or more of the specific patent document sections in the individual patent documents, the obtaining the spans of text included in the one or more of the specific patent document sections, or the comparing the spans of text, are performed using an operation based on a machine learning model.

16. A system configured for extracting patent document templates from a patent corpus, the system comprising:

one or more hardware processors configured by machine-readable instructions to:

obtain a patent corpus, the patent corpus including a plurality of patent documents;

receive one or more parameters, the one or more parameters including a first parameter;

determine one or more subsets of the patent corpus by filtering the patent corpus based on the one or more parameters, the one or more subsets of the patent corpus including a first subset of the patent corpus;

identify one or more document clusters within individual ones of the one or more subsets of the patent corpus, the one or more document clusters including a first document cluster within the first subset of the patent corpus, wherein the first document cluster includes a plurality of patent documents sharing common text, wherein identifying the one or more document clusters includes comparing some or all combinations of pairs of patent documents contained in a given subset of the patent corpus, and wherein comparing the some or all combinations of the pairs of patent documents contained in the given subset of the patent corpus includes:

identifying one or more specific patent document sections in individual patent documents included in the one or more subsets of the patent corpus where related patent documents frequently share common spans of text, wherein the specific patent document sections include one or more of a first portion of a summary section, a last portion of a summary section, a first portion of a brief description of drawing section, a last portion of a brief description of drawings section, a first portion of a detailed description section, or a last portion of a detailed description section;

obtaining the spans of text included in the one or more of the specific patent document sections of the individual patent documents included in the one or more subsets of the patent corpus; and

comparing, for individual ones of the pairs of patent documents, the spans of text obtained from the individual patent documents of the pairs of patent documents to determine if they are common text; and

obtain a patent document template corresponding to the first document cluster, the patent document template including the common text of the plurality of patent documents sharing common text.

17. The system of claim 16 , wherein the one or more hardware processors are further configured by to machine-readable instructions to implement an operation based on a machine learning model to identify the one or more of the specific patent document sections in the individual patent documents, obtain the spans of text included in the one or more of the specific patent document sections, and/or compare the spans of text.

18. A non-transient computer-readable storage medium having instructions embodied thereon, the instructions being executable by one or more processors to perform a method for extracting patent document templates from a patent corpus, the method comprising:

obtaining a patent corpus, the patent corpus including a plurality of patent documents;

receiving one or more parameters, the one or more parameters including a first parameter;

determining one or more subsets of the patent corpus by filtering the patent corpus based on the one or more parameters, the one or more subsets of the patent corpus including a first subset of the patent corpus;

identifying one or more document clusters within individual ones of the one or more subsets of the patent corpus, the one or more document clusters including a first document cluster within the first subset of the patent corpus, wherein the first document cluster includes a plurality of patent documents sharing common text, wherein the identifying the one or more document clusters includes comparing some or all combinations of pairs of patent documents contained in a given subset of the patent corpus, and wherein the comparing the some or all combinations of the pairs of patent documents contained in the given subset of the patent corpus includes:

identifying one or more specific patent document sections in individual patent documents included in the one or more subsets of the patent corpus where related patent documents frequently share common spans of text, wherein the specific patent document sections include one or more of a first portion of a summary section, a last portion of a summary section, a first portion of a brief description of drawing section, a last portion of a brief description of drawings section, a first portion of a detailed description section, or a last portion of a detailed description section;

obtaining the spans of text included in the one or more of the specific patent document sections of the individual patent documents included in the one or more subsets of the patent corpus; and

comparing, for individual ones of the pairs of patent documents, the spans of text obtained from the individual patent documents of the pairs of patent documents to determine if they are common text; and

obtaining a patent document template corresponding to the first document cluster, the patent document template including the common text of the plurality of patent documents sharing common text.

19. The non-transient computer-readable storage medium of claim 18 , wherein one or more of the identifying the one or more of the specific patent document sections in the individual patent documents, the obtaining the spans of text included in the one or more of the specific patent document sections, or the comparing the spans of text, are performed using an operation based on a machine learning model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2024
From: SPECIFIO, INC.
To: PAXIMAL, INC.
Reel/Frame 066266/0834 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2020
From: SCHICK, IAN C.; KNIGHT, KEVIN
To: SPECIFIO, INC.
Reel/Frame 052941/0214 →
Continuity (35)
Continuation In Part 16814335 · Mar 10, 2020
Continuation In Part 16739655 · Jan 10, 2020
Continuation 16510074 · Jul 12, 2019
Continuation 16901677
Continuation In Part 16025720 · Jul 2, 2018
Continuation In Part 16025687 · Jul 2, 2018
Continuation In Part 15994756 · May 31, 2018
Continuation 15936239 · Mar 26, 2018
Continuation 15892679 · Feb 9, 2018
Provisional Application 62590274 · Nov 23, 2017
Provisional Application 62564210 · Sep 27, 2017
Provisional Application 62561876 · Sep 22, 2017
Provisional Application 62553096 · Aug 31, 2017
Provisional Application 62546743 · Aug 17, 2017
Provisional Application 62539014 · Jul 31, 2017
Provisional Application 62534793 · Jul 20, 2017
Provisional Application 62528907 · Jul 5, 2017
Provisional Application 62526316 · Jun 28, 2017
Provisional Application 62526314 · Jun 28, 2017
Provisional Application 62523262 · Jun 22, 2017
Provisional Application 62523257 · Jun 22, 2017
Provisional Application 62523258 · Jun 22, 2017
Provisional Application 62523260 · Jun 22, 2017
Provisional Application 62519852 · Jun 14, 2017
Provisional Application 62519847 · Jun 14, 2017
Provisional Application 62519850 · Jun 14, 2017
Provisional Application 62516640 · Jun 7, 2017
Provisional Application 62515096 · Jun 5, 2017
Provisional Application 62479136 · Mar 30, 2017
Provisional Application 62459199 · Feb 15, 2017
Provisional Application 62459235 · Feb 15, 2017
Provisional Application 62459208 · Feb 15, 2017
Provisional Application 62459357 · Feb 15, 2017
Provisional Application 62459246 · Feb 15, 2017
Related Publication 20200311351A1 · Oct 1, 2020
Cited By (1)
US 12,662,012