IP Library Patent Application 13363401
Patent Application
App. No. 13/363,401

Selection of Language Model Training Data

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
13/363,401
Abstract

An intelligent selection system selects language model training data to obtain in-domain training datasets. The selection is accomplished by estimating a cross-entropy difference for each candidate text segment from a generic language dataset. The cross-entropy difference is a difference between the cross-entropy of the text segment according to the in-domain language model and the cross-entropy of the text segment according to a language model trained on a random sample of the data source from which the text segment is drawn. If the difference satisfies a threshold condition, the text segment is added as an in-domain text segment to a training dataset.

Claims (28)

1 . A method comprising:

determining an in-domain cross-entropy of a data segment from a domain-specific dataset according to an in-domain language model;

determining a non-domain-specific cross-entropy of the data segment according to a non-domain-specific language model;

determining a difference between the in-domain cross-entropy and the non-domain-specific cross-entropy; and

adding the data segment to a training dataset for the in-domain language model, if the difference satisfies a threshold condition.

2 . The method of claim 1 wherein the data segment is a text segment.

3 . The method of claim 1 wherein the in-domain language model is a language model used for machine translation.

4 . The method of claim 1 wherein the in-domain language model is at least one of (1) a language model used for speech recognition and (2) a search algorithm related language model.

5 . The method of claim 1 wherein the non-domain-specific language model is a language model trained on a random sample of the non-domain-specific dataset.

6 . One or more computer-readable storage media encoding computer-executable instructions for executing on a computer system a computer process, the computer process comprising:

scoring a data segment from a non-domain-specific dataset based on a difference between a cross-entropy of the data segment according to an in-domain language model and a cross-entropy of the data segment according to a non-domain-specific language model.

7 . The one or more computer-readable storage media of claim 6 wherein the computer process further comprising adding the data segment to an in-domain training dataset for the in-domain language model, if the difference satisfies a threshold condition.

8 . The one or more computer-readable storage media of claim 6 wherein the data segment is a text segment.

9 . The one or more computer-readable storage media of claim 6 wherein the data segment is a segment of a biological sequence.

10 . The one or more computer-readable storage media of claim 6 wherein the in-domain language model is a language model used for machine translation.

11 . The one or more computer-readable storage media of claim 6 wherein the in-domain language model is an n-gram language model.

12 . The one or more computer-readable storage media of claim 6 wherein the non-domain-specific language model is a language model trained on a random sample of the non-domain-specific dataset.

13 . The one or more computer-readable storage media of claim 6 wherein the computer process further comprising partitioning the non-domain-specific dataset into the data segments, each data segment being a sentence.

14 . The one or more computer-readable storage media of claim 6 wherein the computer process further comprising determining the difference in a log domain.

15 . The one or more computer-readable storage media of claim 7 wherein the non-domain-specific dataset comprising a first component in a first language and a second component in a second language and wherein scoring the data segment from the non-domain-specific dataset further comprising scoring the first component.

16 . The one or more computer-readable storage media of claim 15 wherein adding the data segment to the in-domain training dataset for the in-domain language model further comprising adding the first component and the second component to the in-domain training dataset for the in-domain language model.

17 . A system comprising:

a selection engine configured to select a text segment from a non-domain-specific dataset;

a determination engine configured to determine an in-domain cross-entropy of the text segment according to an in-domain language model and to determine a non-domain-specific cross-entropy of the text segment according to a non-domain-specific language model; and

a differentiator configured to determine a difference between the in-domain cross-entropy and the non-domain-specific cross-entropy.

18 . The system of claim 17 further comprising a comparator configured to compare the difference with a threshold.

19 . The system of claim 18 wherein the comparator is further configured to add the data segment to a training dataset for the in-domain language model, if the difference satisfies a threshold condition.

20 . The system of claim 17 wherein the non-domain-specific language model is a language model trained on a random sample of the non-domain-specific dataset.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034544/0541 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2012
From: MOORE, ROBERT CARTER; LEWIS, WILLIAM DUNCAN
To: MICROSOFT CORPORATION
Reel/Frame 027629/0689 →