IP Library › Granted Patent US 12,488,011
Granted Patent B2
US 12,488,011 · App. 18/635,181 · Granted Dec 2, 2025

Document search system, document search method, program, and non-transitory computer readable storage medium

Inventors: Kazuki Higashi (Atsugi, JP); Junpei Momo (Sagamihara, JP)
Assignee: Semiconductor Energy Laboratory Co., Ltd.
G06F16/24578G06F16/93G06F40/268G06F40/279G06N3/08G06N20/00G06F2216/11G06Q10/10G06Q50/184
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,011
App. No.
18/635,181
Granted
Dec 2, 2025
Kind
B2
Abstract

A highly accurate document search, particularly a search for a document relating to intellectual property, is achieved with an easy input method. A document search system includes a processing portion. The processing portion has a function of extracting a keyword included in text data, a function of extracting a related term of the keyword from words included in a plurality of pieces of first reference text analysis data, a function of giving a weight to each of the keyword and the related term, a function of giving a score to each of a plurality of pieces of second reference text analysis data on the basis of the weight, a function of ranking the plurality of pieces of second reference text analysis data on the basis of the score to generate ranking data, and a function of outputting the ranking data.

Claims (22)

1 . A document search device, for searching a document related or similar to an input document, comprising a processing portion that is configured to:

obtain a first distributed representation vector and a first weight of the first distributed representation vector from the input document;

extract related terms of a keyword extracted from words included in the input document based on a similarity degree with between the first distributed representation vector and a second distributed representation vector of words included in a plurality of pieces of reference text analysis data;

obtain a second weight of the related terms, the second weight comprising a product of the first weight by the similarity degree;

output the first weight and the second weight;

receive a change to value of one or more of the first weight and the second weight; and

execute the search of the document by using the first weight and the second weight that were compiled.

2 . The document search device according to claim 1 , wherein the first distributed representation vector and the second distributed representation vector are obtained through a machine learning of distributed representation of a word included in the plurality of pieces of reference text analysis data.

3 . The document search device according to claim 2 , wherein the plurality of pieces of the reference text analysis data is generated by performing a morphological analysis on reference text data.

4 . The document search device according to claim 1 , wherein the first distributed representation vector is obtained from a keyword extracted from the input document.

5 . The document search device according to claim 2 , wherein the machine learning uses a neural network.

6 . A document search device comprising:

a memory portion having instructions stored thereon that, when executed by a processing portion, cause the processing portion to perform operations for searching a document in view of an input document, the operations comprising:

obtaining a first distributed representation vector and a first weight of the first distributed representation vector from the input document,

extracting related terms of a keyword extracted from words included in the input document based on a similarity degree with between the first distributed representation vector and a second distributed representation vector of the words included in a plurality of pieces of reference text analysis data,

obtaining a second weight of the related terms, the second weight comprising a product of the first weight and the similarity degree,

outputting the first weight and the second weight, and

executing a search of the document using a changed value of one or more of the first weight and the second weight, the changed value being provided through compiling of keyword data and related term data.

7 . The document search device according to claim 6 , wherein the first distributed representation vector and the second distributed representation vector are obtained through a machine learning of distributed representation of a word included in the plurality of pieces of reference text analysis data.

8 . The document search device according to claim 7 , wherein the plurality of pieces of the reference text analysis data is generated by performing a morphological analysis on reference text data.

9 . The document search device according to claim 7 , wherein the machine learning uses a neural network.

10 . The document search device according to claim 6 , wherein the first distributed representation vector is obtained from a keyword extracted from the input document.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 17, 2024
From: HIGASHI, KAZUKI; MOMO, JUNPEI
To: SEMICONDUCTOR ENERGY LABORATORY CO., LTD.
Reel/Frame 068012/0822 →
Priority Claims (1)
JP 2018-055934 · Mar 23, 2018 · national
Continuity (3)
Continuation 17064871 · Oct 7, 2020
Continuation 16979197
Related Publication 20240273108A1 · Aug 15, 2024
References Cited (39)
US 8200695B2 · Cha et al. · 2012 [cited by applicant]
US 10141069B2 · Ikeda. et al. · 2018 [cited by applicant]
US 10699794B2 · Ikeda et al. · 2020 [cited by applicant]
US 11086857B1 · Ganu et al. · 2021 [cited by applicant]
US 11182806B1 · Arfa et al. · 2021 [cited by applicant]
US 20080168288A1 · Jia et al. · 2008 [cited by applicant]
US 20140067846A1 · Edwards et al. · 2014 [cited by applicant]
US 20140324808A1 · Sandhu et al. · 2014 [cited by applicant]
US 20160328388A1 · Cao et al. · 2016 [cited by applicant]
US 20160343452A1 · Ikeda et al. · 2016 [cited by applicant]
US 20180032606A1 · Tolman et al. · 2018 [cited by applicant]
US 20190155915A1 · Huang et al. · 2019 [cited by applicant]
US 20190221204A1 · Zhang et al. · 2019 [cited by applicant]
US 20200176069A1 · Ikeda et al. · 2020 [cited by applicant]
CN 101055580A · 2007 [cited by applicant]
CN 103886063A · 2014 [cited by applicant]
CN 105631009A · 2016 [cited by applicant]
CN 106776713A · 2017 [cited by applicant]
CN 107247780A · 2017 [cited by applicant]
CN 109002473A · 2018 [cited by applicant]
JP 03172966A · 1991 [cited by applicant]
JP 08263521A · 1996 [cited by applicant]
JP 2007065745A · 2007 [cited by applicant]
JP 2011039639A · 2011 [cited by applicant]
JP 2014106665A · 2014 [cited by applicant]
JP 2015041239A · 2015 [cited by applicant]
JP 2016219011A · 2016 [cited by applicant]
JP 2017130195A · 2017 [cited by applicant]
JP 2017134675A · 2017 [cited by applicant]
WO WO2017107566 · 2017 [cited by applicant]
Takano.A et al., “Development of the generic association engine for processing large corpora”, Mar. 30, 2007, IPA. [cited by applicant]
International Search Report (Application No. PCT/IB2019/052022) Dated Jun. 18, 2019. [cited by applicant]
Written Opinion (Application No. PCT/IB2019/052022) Dated Jun. 18, 2019. [cited by applicant]
word2vec, https://github.com/imsky/word2vec, Jan. 31, 2015. [cited by applicant]
Nakayama. Y, “Feature extraction and TF-IDF”, https://qiita.com/ynakayama/items/300460aa718363abc85c, Jun. 23, 2014. [cited by applicant]
Dwaipayan.R et al., “Using Word Embeddings for Automatic Query Expansion”, Jul. 21, 2016. [cited by applicant]
Chinese Office Action (Application No. 201980033402.2) dated Apr. 25, 2024. [cited by applicant]
Chinese Office Action (Application No. 201980033402.2) dated Apr. 3, 2025. [cited by applicant]
Chinese Office Action (Application No. 201980033402.2) dated Nov. 12, 2024. [cited by applicant]