IP Library Granted Patent US 9,128,907
Granted Patent B2
US 9,128,907 · App. 14/446,540 · Granted Sep 8, 2015

Language model generating device, method thereof, and recording medium storing program thereof

Inventors: Kazuhiro Arai (Tokyo, JP); Tadashi Emori (Tokyo, JP)
Assignee: NEC INFORMATEC SYSTEMS, LTD.
G06F17/21G06F17/2755G10L15/063G10L15/183Y10S707/99934
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,128,907
App. No.
14/446,540
Granted
Sep 8, 2015
Kind
B2
Abstract

A text in a corpus including a set of world wide web (web) pages is analyzed. At least one word appropriate for a document type set according to a voice recognition target is extracted based on an analysis result. A word set is generated from the extracted at least one word. A retrieval engine is caused to perform a retrieval process using the generated word set as a retrieval query of the retrieval engine on the Internet, and a link to a web page from the retrieval result is acquired. A language model for voice recognition is generated from the acquired web page.

Claims (90)

1. A word retrieval device, comprising:

an analyzer configured to carry out morphological analysis of text;

an extractor configured to extract at least one word representing a feature of the text from words of a morphological analysis result by the analyzer, using word information quantity on the words; and

a retrieval device configured to retrieve text related to the at least one word from a web page using the at least one word as a retrieval query,

wherein the word information quantity is represented by I x , where T x represents a power of an appearance frequency of each word and I x is defined as follows:

I

x

=

T

x

x

=

t

T

x

×

100.

2. The word retrieval device according to claim 1 , the morphological analysis includes detection of information of notation or the notation and a part of speech for each word in the text.

3. The word retrieval device according to claim 1 ,

further comprising a selector which sets a character string defining a document type,

wherein the extractor compares each word obtained from the morphological analysis result with the character string and extracts the word when the word corresponds to the character string.

4. The word retrieval device according to claim 3 ,

wherein the character string includes information of notation and a part of speech of the character string, and

the extractor compares the notation or the notation and part of speech of each word in the text, with the notation or the notation and part of speech of the character string and extracts a word corresponding to the notation or the notation and part of speech of the character string.

5. The word retrieval device according to claim 4 ,

wherein the extractor determines whether or not a part of speech of a word not corresponding to the character string is a noun and excludes the word from an extraction target when the part of speech of the word is not the noun.

6. A word retrieval method of a word retrieval device, the method comprising:

carrying out morphological analysis of text;

extracting at least one word representing a feature of the text from words of a morphological analysis result by the analyzer, using word information quantity on the words; and

retrieving text related to the at least one word from a web page using the at least one word as a retrieval query,

wherein the word information quantity is represented by I x , where T x represents a power of an appearance frequency of each word and I x is defined as follows:

I

x

=

T

x

x

=

t

T

x

×

100.

7. The word retrieval method according to claim 6 , the morphological analysis includes detection of information of notation or the notation and a part of speech for each word in the text.

8. The word retrieval method according to claim 6 , further comprising:

setting a character string defining a document type; and

comparing each word obtained from the morphological analysis result with the character string and extracting the word when the word corresponds to the character string.

9. The word retrieval method according to claim 8 , wherein the character string includes information of notation and a part of speech of the character string,

the comparing being comparing of the notation or the notation and part of speech of each word in the text, with the notation or the notation and part of speech of the character string,

the extracting being extracting of a word corresponding to the notation or the notation and part of speech of the character string.

10. The word retrieval method according to claim 9 , further comprising:

determining whether or not a part of speech of a word not corresponding to the character string is a noun, and

excluding the word from an extraction target when the part of speech of the word is not the noun.

11. A non-transitory computer-readable recording medium storing a word retrieving program used in a computer of a word retrieval device and causing the computer to execute a method comprising:

carrying out morphological analysis of text;

extracting at least one word representing a feature of the text from words of a morphological analysis result by the analyzer, using word information quantity on the words; and

retrieving text related to the at least one word from a web page using the at least one word as a retrieval query,

wherein the word information quantity is represented by I x , where T x represents a power of an appearance frequency of each word and I x is defined as follows:

I

x

=

T

x

x

=

t

T

x

×

100.

12. The non-transitory computer-readable recording medium according to claim 11 , the morphological analysis includes detection of information of notation or the notation and a part of speech for each word in the text.

13. The non-transitory computer-readable recording medium according to claim 11 , further comprising:

setting a character string defining a document type; and

comparing each word obtained from the morphological analysis result with the character string and extracting the word when the word corresponds to the character string.

14. The non-transitory computer-readable recording medium according to claim 13 ,

wherein the character string includes information of notation and a part of speech of the character string,

the comparing being comparing of the notation or the notation and part of speech of each word in the text, with the notation or the notation and part of speech of the character string,

the extracting being extracting of a word corresponding to the notation or the notation and part of speech of the character string.

15. The non-transitory computer-readable recording medium according to claim 14 , further comprising:

determining whether or not a part of speech of a word not corresponding to the character string is a noun, and

excluding the word from an extraction target when the part of speech of the word is not the noun.

Assignments (2)
MERGER AND CHANGE OF NAME Recorded Jun 26, 2017
From: NEC INFORMATEC SYSTEMS, LTD.; NEC SOFT, LTD.
To: NEC SOLUTION INNOVATORS, LTD.
Reel/Frame 042817/0787 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 30, 2014
From: ARAI, KAZUHIRO; EMORI, TADASHI
To: NEC INFORMATEC SYSTEMS, LTD.
Reel/Frame 033420/0485 →
Priority Claims (1)
JP 2010-229526 · Oct 12, 2010 · national
Continuity (2)
Division 13271424 · Oct 12, 2011
Related Publication 20140343926A1 · Nov 20, 2014