IP Library Granted Patent US 12,430,519
Granted Patent B2
US 12,430,519 · App. 18/030,275 · Granted Sep 30, 2025

Data processing device, data processing method, and data processing program

Inventors: Atsunori Ogawa (Musashino, JP); Naohiro Tawara (Musashino, JP); Marc Delcroix (Musashino, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G06F40/56G06F40/253
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,519
App. No.
18/030,275
Granted
Sep 30, 2025
Kind
B2
Abstract

A data processing device includes processing circuitry configured to extract a second word corresponding to a first word included in first text from among a plurality of words belonging to a predetermined domain, repeat processing in the extraction for all words included in the first text to generate a confusion network that expresses a plurality of sentence possibilities with one network configuration and is an expression format of a word sequence, and search for a grammatically correct word string in the confusion network using a language model that evaluates grammatical correctness of the word string, and select a word string to be output.

Claims (47)

1. A data processing device comprising:

processing circuitry configured to:

use a replacement model, that estimates a domain-dependent word probability distribution at a position where a word appears in a sentence, to estimate a word probability distribution obtained by replacing every first word of a plurality of first words included in text of a first domain with a second word of a second domain;

sample values corresponding to a vocabulary size from a Gumbel distribution;

add the values to the word probability distribution to generate a new word probability distribution;

repeat the generation of the new word probability distribution a plurality of times for all of the plurality of first words included in the text of the first domain, select a plurality of the second words from a top of the new word probability distribution for every first word of the plurality of first words, and arrange the selected second words to be ordered by position to generate a confusion network that expresses a plurality of sentence possibilities with one network configuration and is an expression format of a word sequence;

search for a grammatically correct word string in the confusion network using a language model that evaluates grammatical correctness of the word string; and

select a word string to be output,

wherein the language model is trained based on learning data,

wherein the data processing device adds the selected word string to the learning data to enrich the learning data,

wherein the language model is further trained using the enriched learning data, and

wherein the processing circuitry is configured to use the further trained language model to search for the grammatically correct word string in the confusion network.

2. The data processing device according to claim 1 ,

wherein the processing circuitry is further configured to preferentially output either the grammatically correct word string or a word string having diversity according to a preset first parameter indicating a degree of priority between the grammatically correct word string and the word string having diversity.

3. The data processing device according to claim 2 , wherein a value of the first parameter is adjusted according to a second parameter that sets a number of words indicated in each segment of the confusion network.

4. A data processing method executed by a data processing device, the data processing method comprising:

using a replacement model, that estimates a domain-dependent word probability distribution at a position where a word appears in a sentence, to estimate a word probability distribution obtained by replacing every first word of a plurality of first words included in text of a first domain with a second word of a second domain, sampling values corresponding to a vocabulary size from a Gumbel distribution, and adding the values to the word probability distribution to generate a new word probability distribution;

repeating the generation of the new word probability distribution a plurality of times for all of the plurality of first words included in the text of the first domain, select a plurality of the second words from a top of the new word probability distribution for every first word of the plurality of first words, and arrange the selected second words to be ordered by position to generate a confusion network that expresses a plurality of sentence possibilities with one network configuration and is an expression format of a word sequence; and

searching for a grammatically correct word string in the confusion network using a language model that evaluates grammatical correctness of the word string; and

selecting a word string to be output,

wherein the language model is trained based on learning data,

wherein the method further comprises adding the selected word string to the learning data to enrich the learning data,

wherein the language model is further trained using the enriched learning data, and

wherein the method further comprises using the further trained language model to search for the grammatically correct word string in the confusion network.

5. A non-transitory computer-readable recording medium storing therein a data processing program that causes a computer to execute a process comprising:

using a replacement model, that estimates a domain-dependent word probability distribution at a position where a word appears in a sentence, to estimate a word probability distribution obtained by replacing every first word of a plurality of first words included in text of a first domain with a second word of a second domain, sampling values corresponding to a vocabulary size from a Gumbel distribution, and adding the values to the word probability distribution to generate a new word probability distribution;

repeating the generation of the new word probability distribution a plurality of times for all of the plurality of first words included in the text of the first domain, select a plurality of the second words from a top of the new word probability distribution for every first word of the plurality of first words, and arrange the selected second words to be ordered by position to generate a confusion network that expresses a plurality of sentence possibilities with one network configuration and is an expression format of a word sequence; and

searching for a grammatically correct word string in the confusion network using a language model that evaluates grammatical correctness of the word string, and

selecting a word string to be output,

wherein the language model is trained based on learning data,

wherein the process further comprises adding the selected word string to the learning data to enrich the learning data,

wherein the language model is further trained using the enriched learning data, and

wherein the process further comprises using the further trained language model to search for the grammatically correct word string in the confusion network.

6. The data processing device according to claim 1 , wherein the first domain is different from the second domain.

7. The data processing device according to claim 6 , wherein

the first domain includes text in a formal style, and

the second domain includes text in a conversational style.

8. The data processing device according to claim 1 , wherein

the first domain includes text in a formal style, and

the second domain includes text in a colloquial style.

9. The data processing device according to claim 1 , wherein domain corresponds to a category defined by a content of a sentence, a style of the sentence, or a combination thereof.

10. The data processing device according to claim 1 , wherein the replacement model is a bidirectional long short-term memory (LSTM) model configured to estimate the domain-dependent word probability distribution by processing a forward partial word string of the text of the first domain and a backward partial word string of the text of the first domain.

11. The data processing device according to claim 10 , wherein the replacement model generates a domain-dependent hidden state vector by concatenating hidden state vectors from a forward LSTM layer, a backward LSTM layer, and a binary domain label indicating the second domain.

12. The data processing device according to claim 1 , wherein the processing circuitry is configured to select the plurality of the second words from the top of the new word probability distribution such that a cumulative probability of the selected second words does not exceed a preset threshold.

13. The data processing device according to claim 12 , wherein the processing circuitry is configured to normalize probabilities of the selected second words such that a total probability equals one for each position in the confusion network.

14. The data processing device according to claim 1 , wherein the replacement model is trained using learning data from both the first domain and the second domain.

15. The data processing device according to claim 1 , wherein the processing circuitry is configured to select the plurality of the second words from a top K words of the new word probability distribution, where K is a preset parameter determining a number of words per segment of the confusion network.

Assignments (2)
CHANGE OF NAME Recorded Aug 20, 2025
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 072556/0180 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 5, 2023
From: OGAWA, ATSUNORI; TAWARA, NAOHIRO; DELCROIX, MARC
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 063225/0570 →
Continuity (1)
Related Publication 20240005104A1 · Jan 4, 2024
References Cited (17)
US 8326598B1 · Macherey · 2012 [cited by examiner]
US 20130018649A1 · Deshmukh · 2013 [cited by examiner]
US 20130054224A1 · Jiang · 2013 [cited by examiner]
US 20190286709A1 · Okura et al. · 2019 [cited by applicant]
CN 102650988A · 2012 [cited by applicant]
JP 2019159743A · 2019 [cited by applicant]
JP 2020112915A · 2020 [cited by applicant]
Rosti et al., Improved Word-Level System Combination for Machine Translation, 2007, Proceedings of the 45th Annual Meeting of the Association of Computational Linguistics, pp. 312-219 (Year: 2007). [cited by examiner]
Prasov et al., Fusing Eye Gaze with Speech Recognition Hypotheses to Resolve Exophoric References in Situated Dialogue, 2010, Proceedings of the 2010 Conference on Empirical Methods in Natural Language Processing, pp. 4… [cited by examiner]
Stolcke, “SRILM—An Extensible Language Modeling Toolkit”, International Conference on Spoken Language Processing, ICSLP, Proceedings, 2002, 4 pages. [cited by applicant]
HSU, “Generalized Linear Interpolation of Language Models”, ASRU, 2007, pp. 136-140. [cited by applicant]
Mangu et al., “Finding Consensus in Speech Recognition: Word Error Minimization and Other Applications of Confusion Networks”, Computer Speech and Language, vol. 14, No. 4, arXiv:cs/0010012v1 [cs.CL], Oct. 7, 2000, 35 p… [cited by applicant]
Kobayashi, “Contextual Augmentation: Data Augmentation by Words with Paradigmatic Relations”, Proc. NAACL-HLT, 2018, 6 pages. [cited by applicant]
Maekawa, “Corpus of Spontaneous Japanese: Its Design and Evaluation”, Proc. Workshop on Spontaneous Speech Processing and Recognition (SSPR), 2003, 6 pages. [cited by applicant]
Hori et al., “Low-Latency Real-Time Meeting Recognition and Understanding Using Distant Microphones and Omni-Directional Camera”, IEEE Transactions on Audio, Speech, and Language Processing, vol. 20, No. 2, Feb. 2012, p… [cited by applicant]
Vaswani et al., “Attention Is All You Need”, 31st Conference on Neural Information Processing Systems (NIPS), 2017, 15 pages. [cited by applicant]
Ogawa et al., “Unsupervised Domain Transfer for Training Data Augmentation of Language Modeling”, Reports of the 2020 Autumm Meeting of the Acoustical Society of Japan 1-2-9, 2020, 5 pages including English Translation. [cited by applicant]