IP Library Granted Patent US 8,831,945
Granted Patent B2
US 8,831,945 · App. 13/271,424 · Granted Sep 9, 2014

Language model generating device, method thereof, and recording medium storing program thereof

Inventors: Kazuhiro Arai (Tokyo, JP); Tadashi Emori (Tokyo, JP)
Assignee: NEC Informatec Systems, Ltd.
G10L15/063G10L15/183Y10S707/99934
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,831,945
App. No.
13/271,424
Granted
Sep 9, 2014
Kind
B2
Abstract

A text in a corpus including a set of world wide web (web) pages is analyzed. At least one word appropriate for a document type set according to a voice recognition target is extracted based on an analysis result. A word set is generated from the extracted at least one word. A retrieval engine is caused to perform a retrieval process using the generated word set as a retrieval query of the retrieval engine on the Internet, and a link to a web page from the retrieval result is acquired. A language model for voice recognition is generated from the acquired web page.

Claims (74)

1. A language model generating device implemented by hardware, comprising:

a corpus analyzer which analyzes text in a corpus including a set of world wide web (web) pages;

an extractor which extracts at least one word appropriate for a document type set according to a voice recognition target based on an analysis result by the corpus analyzer;

a word set generator which generates at least one word set from the at least one word extracted by the extractor;

a hardware-implemented web page acquiring device which causes a retrieval engine to perform a retrieval process using the word set generated by the word set generator as a retrieval query of the retrieval engine on the Internet and acquires a web page related to the word set from the retrieval result; and

a language model generator which generates a language model for voice recognition from the web page,

wherein the word set generator calculates a word information quantity, representing similarity to the corpus on each word extracted by the extractor and generates the at least one word set from at least one word whose word information quantity is greater than or equal to the predetermined value, the word information quantity being represented by I x , where T x represents a power of an appearance frequency of each word and I x is defined as follows:

I

x

=

T

x

x

=

t

T

x

×

100

,

wherein at least one of the analyzing, extracting, generating, and retrieving is performed by a processor.

2. The language model generating device according to claim 1 ,

wherein the word set generator generates the predetermined number of word sets, each of which includes the predetermined number of words, are generated randomly from words whose word information quantity values are respectively greater than or equal to a predetermined value.

3. A language model generating method, comprising:

analyzing a text in a corpus including a set of world wide web (web) pages;

extracting at least one word appropriate for a document type set according to a voice recognition target based on an analysis result;

generating at least one word set from the at least one extracted word;

causing a retrieval engine to perform a retrieval process using the generated word set as a retrieval query of the retrieval engine on the Internet and acquiring a web page related to the word set from the retrieval result; and

generating a language model for voice recognition from the acquired web page,

wherein the at least one word set is generated such that a word information quantity representing similarity to the corpus is calculated on each extracted word, and the at least one word set is generated from at least one word whose word information quantity is greater than or equal to a predetermined value, the word information quantity being represented by I x , where T x represents a power of an appearance frequency of each word and I x is defined as follows:

I

x

=

T

x

x

=

t

T

x

×

100.

4. The language model generating method according to claim 3 ,

wherein the word set is generated such that the predetermined number of word sets, each of which includes the predetermined number of words, are generated randomly from words whose word information quantity values are respectively greater than or equal to a predetermined value.

5. A non-transitory computer-readable recording medium storing a language model generating program used in a computer of a language model generating device and causing the computer to execute a method comprising:

analyzing a text in a corpus including a set of world wide web (web) pages;

extracting at least one word appropriate for a document type set according to a voice recognition target based on an analysis result;

generating at least one word set from at least one extracted word;

causing a retrieval engine to perform a retrieval process using the generated word set as a retrieval query of the retrieval engine on the Internet and acquiring a web page related to the word set from the retrieval result; and

generating a language model for voice recognition from the acquired web page,

wherein in the step of generating the at least one word set, a word information quantity representing similarity to the corpus is calculated on each extracted word, and the at least one word set is generated from at least one word whose word information quantity is greater than or equal to a predetermined value, the word information quantity being represented by I x , where T x represents a power of an appearance frequency of each word and I x is defined as follows:

I

x

=

T

x

x

=

t

T

x

×

100.

6. A non-transitory computer-readable recording medium according to claim 5 ,

wherein the word set is generated such that the predetermined number of word sets, each of which includes the predetermined number of words, are generated randomly from words whose word information quantity values are respectively greater than or equal to a predetermined value.

Assignments (2)
MERGER AND CHANGE OF NAME Recorded Jun 26, 2017
From: NEC INFORMATEC SYSTEMS, LTD.; NEC SOFT, LTD.
To: NEC SOLUTION INNOVATORS, LTD.
Reel/Frame 042817/0787 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2011
From: ARAI, KAZUHIRO; EMORI, TADASHI
To: NEC INFORMATEC SYSTEMS, LTD.
Reel/Frame 027263/0445 →
Priority Claims (1)
JP 2010-229526 · Oct 12, 2010 · national
Continuity (1)
Related Publication 20120089397A1 · Apr 12, 2012