IP Library Granted Patent US 9,978,371
Granted Patent B2
US 9,978,371 · App. 14/980,400 · Granted May 22, 2018

Text conversion method and device

Inventors: Lin Ma (Hong Kong, HK); Weibin Zhang (Hong Kong, HK); Pascale Fung (Hong Kong, HK)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G10L15/26G06F17/2755G06F17/2775G10L15/02G10L15/193
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,978,371
App. No.
14/980,400
Granted
May 22, 2018
Kind
B2
Abstract

The method includes acquiring a target spoken text, where the target spoken text includes a non-spoken morpheme and a spoken morpheme; determining, from a target weighted finite-state transducer (WFST) model database, a target WFST model corresponding to the target spoken text, where output of a state that is corresponding to the spoken morpheme and that is in the target WFST model is empty, and output and input of a state that is corresponding to the non-spoken morpheme and that is in the target WFST model are the same; and determining, according to the target WFST model, a written text corresponding to the target spoken text, where the written text includes the non-spoken morpheme and does not include the spoken morpheme.

Claims (38)

1. A text conversion method, comprising:

acquiring a target spoken text that comprises a non-spoken morpheme and a spoken morpheme, wherein the spoken morpheme is a type of morpheme that has one or more spoken morpheme characteristics, and wherein the spoken morpheme characteristics comprise an inserted morpheme, a repeated morpheme, and an amending morpheme;

determining an initial weighted finite-state transducer (WFST) model database according to a text training database by means of statistical learning, wherein the initial WFST model database comprises N initial spoken WFST models corresponding to N spoken texts, wherein the N initial spoken WFST models include first states, wherein each spoken text of the N spoken texts comprises the type of morpheme, and wherein an output of each state of the first states that corresponds to the type of morpheme is not empty;

determining a spoken morpheme characteristic WFST model database according to a spoken morpheme training database and the spoken morpheme characteristics and by means of statistical learning, wherein an output of each state in the spoken morpheme characteristic WFST model database that corresponds to the type of morpheme is empty;

modifying the N initial spoken WFST models in the initial WFST model database according to the spoken morpheme characteristic WFST model database to determine N modified spoken WFST models, wherein an output of each state of the N modified WFST models that corresponds to the type of morpheme is empty;

determining a target WFST model database, wherein the target WFST model database comprises the N modified spoken WFST models;

determining, from the target WFST model database, a target WFST model corresponding to the target spoken text, wherein an output of each state in the target WFST model that corresponds to the type of morpheme is empty, and wherein, for each state in the target WFST model that corresponds to the type of morpheme, an output of the state is the same as an input of the state; and

converting the target spoken text into a written text corresponding to the target spoken text by processing the target spoken text using the target WFST model, wherein the written text comprises the non-spoken morpheme and does not comprise the spoken morpheme.

2. The method according to claim 1 , wherein modifying the N initial spoken WFST models in the initial WFST model database according to the spoken morpheme characteristic WFST model database to determine N modified spoken WFST models comprises:

identifying first morphemes in the N initial spoken WFST models by identifying each morpheme in the N initial spoken WFST models that corresponds to the type of morpheme;

determining, from the spoken morpheme characteristic WFST model database, a spoken morpheme characteristic WFST model for each of the first morphemes; and

combining each initial spoken WFST model with a corresponding spoken morpheme characteristic WFST model determined for the initial spoken WFST model.

3. A text conversion device, comprising:

a memory: and

a computer processor coupled to the memory and configured to:

acquire a target spoken text that comprises a spoken morpheme and a non-spoken morpheme, wherein the spoken morpheme is a type of morpheme that has one or more spoken morpheme characteristics, and wherein the spoken morpheme characteristics comprise an inserted morpheme, a repeated morpheme, and an amending morpheme;

determine an initial weighted finite-state transducer (WFST) model database according to a text training database by means of statistical learning, wherein the initial WFST model database comprises N initial spoken WFST models corresponding to N spoken texts, wherein the N initial spoken WFST models include first states, wherein each spoken text of the N spoken texts comprises the type of morpheme, and wherein an output of each state of the first states that corresponds to the type of morpheme is not empty;

determine a spoken morpheme characteristic WFST model database according to a spoken morpheme training database and the spoken morpheme characteristics and by means of statistical learning, wherein an output of each state in the spoken morpheme characteristic WFST model database that corresponds to the type of morpheme is empty;

modify the N initial spoken WFST models in the initial WFST model database according to the spoken morpheme characteristic WFST model database to determine N modified spoken WFST models, wherein an output of each state of the N modified WFST models that corresponds to the type of morpheme is empty;

determine a target WFST model database, wherein the target WFST model database comprises the N modified spoken WFST models;

determine, from the target WFST model database, a target WFST model corresponding to the target spoken text, wherein an output of each state in the target WFST model that corresponds to the type of morpheme is empty, and wherein, for each state in the target WFST model that corresponds to the type of morpheme, output of the state is the same as an input of the state; and

convert the target spoken text into a written text corresponding to the target spoken text by processing the target spoken text using the target WFST model, wherein the written text comprises the non-spoken morpheme and does not comprise the spoken morpheme.

4. The device according to claim 3 , wherein, to determine the N modified spoken WFST models, the computer processor is configured to:

identify first morphemes in the N initial spoken WFST models by identifying each morpheme in the N initial spoken WFST models that corresponds to the type of morpheme;

determine, from the spoken morpheme characteristic WFST model database, a spoken morpheme characteristic WFST model for each spoken morpheme of the first morphemes; and

combine each initial spoken WFST model with a corresponding spoken morpheme characteristic WFST model determined for the initial spoken WFST model.

5. A text conversion device, comprising:

a non-transitory computer-readable medium configured to store a text training database and a spoken morpheme training database; and

a computer processor coupled to the non-transitory computer-readable medium and configured to:

determine an initial weighted finite-state transducer (WFST) model database according to the text training database by means of statistical learning, wherein the initial WFST model database comprises N initial spoken WFST models corresponding to N spoken texts, each spoken text of the N spoken texts comprises a spoken morpheme, and output of a state of the spoken morpheme in each initial spoken WFST model of the N initial spoken WFST models is not empty;

determine a spoken morpheme characteristic WFST model database according to the spoken morpheme training database and characteristics of the spoken morpheme and by means of statistical learning, wherein an output of each state in the spoken morpheme characteristic WFST model database that corresponds to an inserted morpheme, a repeated morpheme, or an amending morpheme is empty;

modify the N initial spoken WFST models in the initial WFST model database according to the spoken morpheme characteristic WFST model database to determine N modified spoken WFST models, wherein an output of each state of the N modified WFST models that corresponds to the inserted morpheme, the repeated morpheme, or the amending morpheme is empty;

determine a target WFST model database, wherein the target WFST model database comprises the N modified spoken WFST models; and

convert an acquired target spoken text into a written text corresponding to the acquired target spoken text by processing the acquired target spoken text using a target WFST model from the target WFST model database.

6. The device according to claim 5 , wherein, to determine the N modified spoken WFST models, the computer processor is configured to:

determine the spoken morpheme in each initial spoken WFST model;

determine, from the spoken morpheme characteristic WFST model database, a spoken morpheme characteristic WFST model corresponding to the spoken morpheme in each initial spoken WFST model; and

combine each initial spoken WFST model with a corresponding spoken morpheme characteristic WFST model determined for the initial spoken WFST model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2016
From: MA, LIN; ZHANG, WEIBIN; FUNG, PASCALE
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 039117/0758 →
Priority Claims (1)
CN 2015 1 0017057 · Jan 13, 2015 · national
Continuity (1)
Related Publication 20160203819A1 · Jul 14, 2016