IP Library Granted Patent US 12705270
Granted Patent B2
US 12705270 · App. 18/853,645 · Granted Aug 11, 2026

Method for visualizing patent documents through similarity assessment based on natural language processing and device for providing the same

Inventor: Inkyung Choi (Seoul, KR)
Assignee: TANALYSIS CO., LTD.
G06F16/338G06F16/345G06F16/358G06F2216/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12705270
App. No.
18/853,645
Granted
Aug 11, 2026
Kind
B2
Abstract

The present disclosure relates to a method and device for providing patent document information based on natural language processing of patent documents. A method for visualizing patent documents implemented by a computer according to the present disclosure includes: receiving information about a target patent; and allowing a user interface to be displayed, wherein the user interface includes a first panel for classifying cores defining at least one component information extracted from the input target patent according to determined colors, and a second panel for classifying similarity assessment results of the target patent with respect to similar patents according to the determined colors. According to the present disclosure, by providing, in the form of a GUI, an interface for inputting patent information to be analyzed, users can easily request analysis of patent documents.

Claims (60)

1 . A method implemented by a computer, comprising:

receiving target patent information;

extracting a first embedding vector of a target patent and first embedding vectors of prior patent documents from a first embedding vector database that stores document-level embedding vectors for each patent document;

calculating a first similarity between the first embedding vector of the target patent and the first embedding vectors of the prior patent documents, and extracting a candidate patent list comprising the prior patent documents having the first similarity greater than or equal to a threshold value;

extracting second embedding vectors of the target patent and the prior patent documents in the candidate patent list from a second embedding vector database that stores sentence-level embedding vectors for each patent document, and calculating a second similarity based on the second embedding vectors for each of cores of the target patent to generate similar patent document information; and

displaying a user interface that includes a first panel which distinguishes cores defining at least one component information extracted from the input target patent according to a determined color, and a second panel which distinguishes similarity assessment results of the target patent with respect to the similar patent document information according to the determined color,

wherein the target patent information comprises at least one sentence,

wherein each of the cores corresponds to at least one received sentence of the target patent, the first panel being configured to visually distinguish the cores from one another, and

wherein the similarity is calculated according to a predetermined criterion using result data from patent examinations or trials in which the validity of patents has been examined or determined, and the target patent is provided in the second panel with the calculated similarity.

2 . The method of claim 1 , wherein the first panel includes extraction criteria for each core for extracting the similar patents.

3 . The method of claim 2 , wherein the first panel includes weights of the similarity assessment with the similar patents.

4 . The method of claim 2 , wherein the second panel includes results according to the extraction criteria for each core for extracting the similar patents.

5 . The method of claim 1 , wherein the interface further includes a third panel that provides corresponding paragraph information for each core of the target patent or the similar patents.

6 . The method of claim 1 , wherein the interface further includes a third panel that provides a validity determination result according to a comparison for each core of the similar patents extracted from the target patent.

7 . The method of claim 6 , wherein the third panel provides a statistical validity score of the target patent.

8 . The method of claim 7 , wherein the third panel provides a distribution location of a validity score of a patent whose validity has been determined in the past or a validity score of a patent whose validity has been determined to be invalid, based on the validity score of the target patent.

9 . The method of claim 1 , wherein the interface includes a fourth panel that maps feature vectors of the target patent or the similar patents into a feature space and provides the feature vectors.

10 . The method of claim 9 , wherein a first feature vector of the target patent and a second feature vector of the similar patents have a distance in the feature space corresponding to a similarity calculated for the target patent and the similar patents.

11 . A non-transitory computer-readable recording medium storing a program executable by a processor for performing a method implemented by a computer, the method comprising:

receiving target patent information;

extracting a first embedding vector of a target patent and first embedding vectors of prior patent documents from a first embedding vector database that stores document-level embedding vectors for each patent document;

calculating a first similarity between the first embedding vector of the target patent and the first embedding vectors of the prior patent documents, and extracting a candidate patent list comprising the prior patent documents having the first similarity greater than or equal to a threshold value;

extracting second embedding vectors of the target patent and the prior patent documents in the candidate patent list from a second embedding vector database that stores sentence-level embedding vectors for each patent document, and calculating a second similarity based on the second embedding vectors for each of cores of the target patent to generate similar patent document information; and

displaying a user interface that includes a first panel which distinguishes cores defining at least one component information extracted from the input target patent according to a determined color, and a second panel which distinguishes similarity assessment results of the target patent with respect to the similar patent document information according to the determined color,

wherein the target patent information comprises at least one sentence,

wherein each of the cores corresponds to at least one received sentence of the target patent, the first panel being configured to visually distinguish the cores from one another, and

wherein the similarity is calculated according to a predetermined criterion using result data from patent examinations or trials in which the validity of patents has been examined or determined, and the target patent is provided in the second panel with the calculated similarity.

12 . A method for visualizing patent documents implemented by a computer, comprising:

displaying first and second cores of a target patent;

displaying a first similar document similar to the target patent based on a calculated similarity; and

displaying 1-1th core mapping information and 1-2th core mapping information of the first similar document,

wherein the 1-1th core mapping information is information generated based on a 1-1th text similar to the first core among texts of the first similar document, and

the 1-2th core mapping information is generated based on a 1-2th text similar to the second core among the texts of the first similar document,

wherein target patent information is received by the computer,

a first embedding vector of the target patent and first embedding vectors of prior patent documents are extracted from a first embedding vector database that stores document-level embedding vectors for each patent document,

a first similarity is calculated between the first embedding vector of the target patent and the first embedding vectors of the prior patent documents, and a candidate patent list, comprising the prior patent documents having the first similarity greater than or equal to a threshold value, is extracted,

second embedding vectors of the target patent and the prior patent documents is extracted in the candidate patent list from a second embedding vector database that stores sentence-level embedding vectors for each patent document, and a second similarity is calculated based on the second embedding vectors for each of cores of the target patent to generate similar patent document information, and

a user interface, that includes a first panel which distinguishes cores defining at least one component information extracted from the input target patent according to a determined color, and a second panel which distinguishes similarity assessment results of the target patent with respect to the similar patent document information according to the determined color, is displayed,

wherein the target patent information comprises at least one sentence,

wherein each of the cores corresponds to at least one received sentence of the target patent, and the cores are visually distinguished from one another, and

wherein the similarity is calculated according to a predetermined criterion using result data from patent examinations or trials in which the validity of patents has been examined or determined, and the target patent is provided in the second panel with the calculated similarity.

13 . The method of claim 12 , wherein the 1-1th text has a similarity to the first core among the texts of the first similar document that is greater than or equal to a first threshold value.

14 . The method of claim 12 , wherein the 1-2th text has a similarity to the second core among the texts of the first similar document that is greater than or equal to a second threshold value smaller than the first threshold value.

15 . The method of claim 12 , further comprising:

displaying a third core of the target patent; and

displaying 1-3th core mapping information of the first similar document,

wherein the 1-3th core mapping information is information generated based on a 1-3th text similar to the third core among the texts of the first similar document.

16 . The method of claim 12 , wherein the 1-1th core mapping information is displayed at a first position corresponding to the first core, and

the 1-2th core mapping information is displayed at a second position corresponding to the second core.

17 . The method of claim 16 , wherein the first and second cores and the 1-1th and 1-2th core mapping information are arranged in a matrix form,

the first and second cores are arranged in a first row, and

the 1-1th and 1-2th core mapping information is arranged in a second row.

18 . The method of claim 12 , wherein the first text is provided in plurality, and

the 1-1th core mapping information includes at least one of the number of the plurality of first texts, an average similarity of the 1-1th texts to the first core, and a similarity of the 1-1th text most similar to the first core among the plurality of 1-1th texts.

19 . The method of claim 18 , wherein a color of the 1-1th core mapping information is determined based on at least one of the number of the plurality of first texts, the average similarity of the 1-1th texts to the first core, and a maximum similarity of the 1-1th text most similar to the first core among the plurality of 1-1th texts.

20 . The method of claim 12 , further comprising:

displaying a second similar document similar to the target patent; and

displaying 2-1th core mapping information and 2-2th core mapping information of the second similar document,

wherein the 2-1th core mapping information is information generated based on a 2-1th text similar to the second core among texts of the second similar document, and

the 2-2th core mapping information is information generated based on a 2-2th text similar to the second core among the texts of the second similar document.