IP Library Granted Patent US 12,308,021
Granted Patent B2
US 12,308,021 · App. 17/995,529 · Granted May 20, 2025

Punctuation mark delete model training device, punctuation mark delete model, and determination device

Inventor: Taku Katou (Chiyoda-ku, JP)
Assignee: NTT DOCOMO, INC.
G10L15/1822G06F16/33G06F40/232G10L15/22G10L15/26G10L25/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,308,021
App. No.
17/995,529
Granted
May 20, 2025
Kind
B2
Abstract

A punctuation mark delete model learning device is a device that generates, through machine learning, a punctuation mark delete model, and comprises a first learning data generation unit that generates first learning data consisting of a pair of an input sentence including a punctuation mark, a preceding sentence that is a sentence with the punctuation mark assigned at an end of the sentence, and a subsequent sentence following the punctuation mark, and a label indicating whether or not the assignment of the punctuation mark is correct, on the basis of a first text corpus consisting of text obtained by speech recognition processing, and a model learning unit that updates parameters of the punctuation mark delete model on the basis of an error between a probability obtained by inputting the input sentences of the first learning data to the punctuation mark delete model and the label.

Claims (35)

1. A punctuation mark delete model learning device for generating, through machine learning, a punctuation mark delete model for determining whether or not a punctuation mark assigned to text obtained by speech recognition processing is correct,

wherein the punctuation mark delete model receives two consecutive sentences of a first sentence with a punctuation mark assigned at an end of the sentence and a second sentence following the first sentence, and outputs a probability indicating whether the punctuation mark assigned at an end of the first sentence is correct, and

the punctuation mark delete model learning device comprises circuitry configured to:

generate first learning data consisting of a pair of an input sentence including a preceding sentence, the preceding sentence being a sentence with a punctuation mark assigned at an end of the sentence, and a subsequent sentence, the subsequent sentence being a sentence following the punctuation mark in text constituting a first text corpus, and a label indicating whether or not the assignment of the punctuation mark is correct on the basis of the first text corpus, the first text corpus being text including of one or more sentences obtained by speech recognition processing and having a punctuation mark assigned thereto on the basis of information obtained by speech recognition processing; and

update parameters of the punctuation mark delete model on the basis of an error between the probability obtained by inputting the input sentences of the first learning data to the punctuation mark delete model and the label associated with the input sentence.

2. The punctuation mark delete model learning device according to claim 1 , wherein the circuitry assigns the label of the first learning data on the basis of the presence or absence of a punctuation mark at an end of a sentence corresponding to preceding sentence included in the input sentence in a second text corpus, the second text corpus being text consisting of the same text as each first text corpus and having a punctuation mark assigned at the end of the sentence included in the text.

3. The punctuation mark delete model learning device according to claim 2 , wherein the first text corpus is text with a punctuation mark inserted at an end of each voice section divided by a silent section having a predetermined length or longer.

4. The punctuation mark delete model learning device according to claim 1 , wherein the first text corpus is text with a punctuation mark inserted at an end of each voice section divided by a silent section having a predetermined length or longer.

5. The punctuation mark delete model learning device according to claim 1 , wherein the circuitry is further configured to

generate second learning data consisting of a pair of an input sentence including a preceding sentence with a punctuation mark assigned at an end of the sentence and a subsequent sentence following the punctuation mark in the text constituting a fourth text corpus, and a label indicating whether or not the assignment of the punctuation mark is correct, on the basis of a third text corpus consisting of text including a sentence with a punctuation mark legitimately assigned at an end of the sentence, and the fourth text corpus consisting of text obtained by randomly inserting a punctuation mark into the text constituting the third text corpus, the label of the second learning data being assigned on the basis of presence or absence of a punctuation mark at an end of the sentence corresponding to the preceding sentence included in the input sentence of the second learning data in the third text corpus, and

wherein the circuitry updates the parameters of the punctuation mark delete model on the basis of the error between the probability obtained by inputting the input sentences of the first learning data and the second learning data to the punctuation mark delete model and the label associated with the input sentence.

6. The punctuation mark delete model learning device according to claim 1 ,

wherein the punctuation mark delete model comprises a bidirectional long short-term storage network including respective forward and backward long short-term storage networks, and

wherein the circuitry

inputs a word string included in the preceding sentence and a word string included in the subsequent sentence to the long short-term storage network in the forward direction according to an arrangement order in the input sentence from the word at the beginning of the preceding sentence,

inputs the word string included in the preceding sentence and the word string included in the subsequent sentence to the long short-term storage network in the backward direction in an order reverse to the arrangement order in the input sentence from the word at the end of the subsequent sentence, and

acquires the probability on the basis of the output of the long short-term storage network in the forward direction and the output of the long short-term storage network in the backward direction.

7. The punctuation mark delete model learning device according to claim 6 ,

wherein the punctuation mark delete model comprises

a hidden layer configured to combine an output of the long short-term storage network in the forward direction with an output of the long short-term storage network in the backward direction; and

an output layer configured to generate the probability on the basis of an output of the hidden layer.

8. The punctuation mark delete model learning device according to claim 7 ,

wherein the model learning unit circuitry

inputs a word string from a beginning of the preceding sentence to an end of the preceding sentence to the long short-term storage network in the forward direction according to an arrangement order in the input sentence, and

inputs a word string from an end of the subsequent sentence to a beginning of the subsequent sentence to the long short-term storage network in the backward direction according to an order reverse to the arrangement order in the input sentence.

9. A determination device for determining whether or not a punctuation mark assigned to text obtained by speech recognition processing is correct, the determination device comprising circuitry configured to:

extract a determination input sentence consisting of two consecutive sentences including a first sentence with a punctuation mark assigned at an end of the sentence and a second sentence following the first sentence from determination target text, the determination target text being text serving as a determination target including one or more sentences obtained by speech recognition processing,

input the determination input sentence to a punctuation mark delete model and determine whether or not the punctuation mark assigned at the end of the sentence of the first sentence is correct; and

delete the punctuation mark determined to be incorrect by the circuitry from the determination target text and output the determination target text with the corrected punctuation mark,

wherein the punctuation mark delete model

is a model learned by machine learning for causing a computer to function, and

is constructed by machine learning for

receiving two consecutive sentences of a first sentence with a punctuation mark assigned at an end of the sentence and a second sentence following the first sentence, and outputting a probability indicating whether the punctuation mark assigned at an end of the first sentence is correct,

setting a pair of an input sentence including a preceding sentence, the preceding sentence being a sentence with a punctuation mark assigned at an end of the sentence, and a subsequent sentence, the subsequent sentence being a sentence following the punctuation mark in text constituting a first text corpus, and a label indicating whether or not the assignment of the punctuation mark is correct, as first learning data, on the basis of the first text corpus, the first text corpus being text including of one or more sentences obtained by speech recognition processing and having a punctuation mark assigned thereto on the basis of information obtained by speech recognition processing; and

updating parameters of the punctuation mark delete model on the basis of an error between the probability output by inputting the input sentences included in the first learning data to the punctuation mark delete model and the label associated with the input sentence.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 5, 2022
From: KATOU, TAKU
To: NTT DOCOMO, INC.
Reel/Frame 061321/0218 →
Priority Claims (1)
JP 2020-074788 · Apr 20, 2020 · national
Continuity (1)
Related Publication 20230223017A1 · Jul 13, 2023
References Cited (21)
US 11138519B1 · Giri · 2021 [cited by examiner]
US 11250452B2 · Kulkarni · 2022 [cited by examiner]
US 11269665B1 · Podgorny · 2022 [cited by examiner]
US 20140297267A1 · Spencer · 2014 [cited by examiner]
US 20150262209A1 · Orsini · 2015 [cited by examiner]
US 20150317069A1 · Clements · 2015 [cited by examiner]
US 20220019737A1 · Choi · 2022 [cited by examiner]
US 20230072015A1 · Ihori · 2023 [cited by examiner]
US 20230223017A1 · Katou · 2023 [cited by examiner]
CN 109887492A · 2019 [cited by examiner]
CN 111062204A · 2020 [cited by examiner]
CN 113326350A · 2021 [cited by examiner]
CN 113704396A · 2021 [cited by examiner]
CN 113609843B · 2022 [cited by examiner]
CN 115081455A · 2022 [cited by examiner]
JP 2014164575A · 2014 [cited by applicant]
KR 101691327B1 · 2016 [cited by examiner]
KR 102705393B1 · 2024 [cited by examiner]
WO WO2021215262A1 · 2021 [cited by examiner]
International Search Report mailed on Jul. 13, 2021 in PCT/JP2021/014931 filed on Apr. 8, 2021, total 3 pages. [cited by applicant]
International Preliminary Report and Written Opinion issued Nov. 3, 2022 in PCT/JP2021/014931 (submitting English translation only), 5 pages. [cited by applicant]