SPEECH OUTPUT METHOD, DEVICE AND MEDIUM
Embodiments of the present disclosure disclose a speech output method and apparatus, a device and a medium, and relate to speech processing technologies. Examples of the method include: determining a target text to be processed; matching the target text with a local text database to determine a preset text corresponding to the target text; and determining, based on the preset text, output speech of the target text from a local speech database to output the output speech; wherein the local speech database is pre-configured based on a correspondence between a text and speech.
1 . A speech output method, comprising:
determining a target text to be processed;
determining a preset text corresponding to the target text by matching the target text with a local text database; and
determining, based on the preset text, output speech of the target text from a local speech database to output the output speech;
wherein the local speech database is pre-configured based on a correspondence between a text and speech.
2 . The method of claim 1 , wherein determining the preset text corresponding to the target text by matching the target text with the local text database comprises:
in response to failing to determine the preset text corresponding to the target text by matching the target text as a whole with the local text database, splitting the target text to obtain at least two target keywords; and
matching the at least two target keywords with the local text database respectively to determine preset keywords corresponding to the target keywords; and
determining, based on the preset text, the output speech of the target text from the local speech database comprises:
determining, based on the preset keywords, the output speech of the target text from the local speech database.
3 . The method of claim 2 , wherein determining, based on the preset keywords, the output speech of the target text from the local speech database comprises:
determining, based on the preset keywords, speech segments corresponding to the target keywords from the local speech database; and
splicing the speech segments based on a sequence of the target keywords in the target text, to obtain the output speech of the target text.
4 . The method of claim 3 , wherein determining, based on the preset keywords, the output speech of the target text from the local speech database comprises:
for a specific keyword that fails to match with a preset keyword from the local text database in the at least two target keywords, determining a synthesized speech segment corresponding to the specific keyword by adopting offline text to speech; and
splicing, based on the sequence of the target keywords in the target text, the synthesized speech segment and the speech segment determined from the local speech database to obtain the output speech of the target text.
5 . The method of claim 1 , which is applied to an offline navigation scene,
wherein the local speech database comprises navigation terms.
6 . An electronic device, comprising:
at least one processor; and
a storage device communicatively connected to the at least one processor; wherein,
the storage device stores an instruction executable by the at least one processor, and when the instruction executed by the at least one processor, the processor implements a speech output method, and the speech output method comprises:
determining a target text to be processed;
determining a preset text corresponding to the target text by matching the target text with a local text database; and
determining, based on the preset text, output speech of the target text from a local speech database to output the output speech;
wherein the local speech database is pre-configured based on a correspondence between a text and speech.
7 . The electronic device of claim 6 , wherein determining the preset text corresponding to the target text by matching the target text with the local text database comprises:
in response to failing to determine the preset text corresponding to the target text by matching the target text as a whole with the local text database, splitting the target text to obtain at least two target keywords; and
matching the at least two target keywords with the local text database respectively to determine preset keywords corresponding to the target keywords; and
determining, based on the preset text, the output speech of the target text from the local speech database comprises:
determining, based on the preset keywords, the output speech of the target text from the local speech database.
8 . The electronic device of claim 7 , wherein determining, based on the preset keywords, the output speech of the target text from the local speech database comprises:
determining, based on the preset keywords, speech segments corresponding to the target keywords from the local speech database; and
splicing the speech segments based on a sequence of the target keywords in the target text, to obtain the output speech of the target text.
9 . The electronic device of claim 8 , wherein determining, based on the preset keywords, the output speech of the target text from the local speech database comprises:
for a specific keyword that fails to match with a preset keyword from the local text database in the at least two target keywords, determining a synthesized speech segment corresponding to the specific keyword by adopting offline text to speech; and
splicing, based on the sequence of the target keywords in the target text, the synthesized speech segment and the speech segment determined from the local speech database to obtain the output speech of the target text.
10 . The electronic device of claim 6 , wherein the local speech database comprises navigation terms.
11 . A non-transitory computer-readable storage medium having a computer instruction stored thereon, wherein the computer instruction is configured to make a computer implement a speech output method, and the speech output method comprises:
determining a target text to be processed;
determining a preset text corresponding to the target text by matching the target text with a local text database; and
determining, based on the preset text, output speech of the target text from a local speech database to output the output speech;
wherein the local speech database is pre-configured based on a correspondence between a text and speech,
12 . The storage medium of claim 11 , wherein determining the preset text corresponding to the target text by matching the target text with the local text database comprises:
in response to failing to determine the preset text corresponding to the target text by matching the target text as a whole with the local text database, splitting the target text to obtain at least two target keywords; and
matching the at least two target keywords with the local text database respectively to determine preset keywords corresponding to the target keywords; and
determining, based on the preset text, the output speech of the target text from the local speech database comprises:
determining, based on the preset keywords, the output speech of the target text from the local speech database.
13 . The storage medium of claim 12 , wherein determining, based on the preset keywords, the output speech of the target text from the local speech database comprises:
determining, based on the preset keywords, speech segments corresponding to the target keywords from the local speech database; and
splicing the speech segments based on a sequence of the target keywords in the target text, to obtain the output speech of the target text.
14 . The storage medium of claim 13 , wherein determining, based on the preset keywords, the output speech of the target text from the local speech database comprises:
for a specific keyword that fails to match with a preset keyword from the local text database in the at least two target keywords, determining a synthesized speech segment corresponding to the specific keyword by adopting offline text to speech; and
splicing, based on the sequence of the target keywords in the target text, the synthesized speech segment and the speech segment determined from the local speech database to obtain the output speech of the target text.
15 . The storage medium of claim 11 , wherein the local speech database comprises navigation terms.