IP Library Granted Patent US 10,827,067
Granted Patent B2
US 10,827,067 · App. 15/730,462 · Granted Nov 3, 2020

Text-to-speech apparatus and method, browser, and user terminal

Inventor: Xiang Liu (Guangzhou, CN)
Assignee: Guangzhou UCWeb Computer Technology Co., Ltd.
H04M3/4938G10L13/02G10L13/08G10L15/08G10L15/22H04L29/0809G10L15/1822G10L15/26G10L15/30H04L67/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,827,067
App. No.
15/730,462
Granted
Nov 3, 2020
Kind
B2
Abstract

A text-to-speech method includes outputting an instruction according to voice information entered by a user; obtaining text information according to the instruction; converting the text information to audio; and playing the audio. According to the embodiments of the present invention, news or other text content in a browser can be played by voice, which liberates hands and eyes of a user. The user can use the browser in some scenarios where the user cannot easily use the browser, such as driving a car, thereby improving user experience.

Claims (71)

1. A text-to-speech method, implemented by a user terminal having a browser, wherein the text-to-speech method comprises:

receiving first voice information entered by a user, wherein the first voice information comprises a participle of a verb;

determining a function requested in the first voice information entered by the user by analyzing the participle of the verb;

obtaining a set of prestored instructions based on the function requested in the first voice information entered by the user;

obtaining a first prompt voice based on the set of prestored instructions;

feeding back the first prompt voice to prompt the user for second voice information comprising an adjective corresponding to the set of prestored instructions;

receiving the second voice information comprising the adjective that is entered by the user in response to the first prompt voice;

searching a memory of the user terminal for a prestored instruction from the set of prestored instructions which matches the second voice information comprising the adjective entered by the user;

outputting the prestored instruction according to the second voice information comprising the adjective entered by the user;

obtaining text information according to the prestored instruction, by:

determining whether the text information is prestored in the user terminal;

in response to determining that the text information is prestored in the user terminal, obtaining the text information from the user terminal according to the prestored instruction; and

in response to determining that the text information is not stored in the user terminal, obtaining the text information from a server according to the prestored instruction;

obtaining prestored audio information according to the text information; and

playing the prestored audio information.

2. The text-to-speech method according to claim 1 , wherein the text-to-speech method further comprises loading the text information to a web page of the browser.

3. The text-to-speech method according to claim 1 , wherein the prestored audio information was obtained by:

partitioning the text information into words or phrases;

searching an audio library for multiple audio segments corresponding to the words or the phrases;

synthesizing the multiple audio segments into audio information; and

storing the audio information.

4. The text-to-speech method according to claim 1 , wherein the text-to-speech method further comprises playing a second prompt voice after playing of the audio is completed.

5. The text-to-speech method of claim 1 , wherein the obtaining prestored audio information according to the text information comprises obtaining the prestored audio information without converting the text information to an audio.

6. The text-to-speech method of claim 1 , further comprising:

in response to determining that the prestored audio information is not available, converting the text information to an audio.

7. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor of a computer system, cause the computer system to perform a method comprising:

receiving first voice information entered by a user, wherein the first voice information comprises a participle of a verb;

determining a function requested in the first voice information entered by the user by analyzing the participle of the verb;

obtaining a set of prestored instructions based on the function requested in the first voice information entered by the user;

obtaining a first prompt voice based on the set of prestored instructions;

feeding back the first prompt voice to prompt the user for second voice information comprising an adjective corresponding to the set of prestored instructions;

receiving the second voice information comprising the adjective that is entered by the user in response to the first prompt voice;

searching a memory of the user terminal for a prestored instruction from the set of prestored instructions which matches the second voice information comprising the adjective entered by the user;

outputting the prestored instruction according to the second voice information comprising the adjective entered by the user;

obtaining text information according to the prestored instruction, by:

determining whether the text information is prestored in the user terminal;

in response to determining that the text information is prestored in the user terminal, obtaining the text information from the user terminal according to the prestored instruction; and

in response to determining that the text information is not stored in the user terminal, obtaining the text information from a server according to the prestored instruction;

obtaining prestored audio information according to the text information; and

playing the prestored audio information.

8. The non-transitory computer-readable storage medium according to claim 7 , wherein the method further comprises loading the text information to a web page of a browser.

9. The non-transitory computer-readable storage medium according to claim 7 , wherein the method further comprises:

obtaining the prestored audio information by:

partitioning the text information into words or phrases;

searching an audio library for multiple audio segments corresponding to the words or the phrases;

synthesizing the multiple audio segments into audio information; and

storing the audio information.

10. A user terminal, comprises:

a processor; and

a memory storing executable instructions that, when executed by the processor, cause the processor to:

receive first voice entered by a user, wherein the first voice information comprises a participle of a verb;

determining a function requested in the first voice information entered by the user by analyzing the participle of the verb;

obtain a set of prestored instructions based on the function requested in the first voice information entered by the user;

obtain a first prompt voice based on the set of prestored instructions;

provide the first prompt voice to prompt the user for second voice information comprising an adjective corresponding to the set of prestored instructions;

receive the second voice information comprising the adjective that is entered by the user in response to the first prompt voice;

search the memory of the user terminal for a prestored instruction from the set of prestored instructions which matches the second voice information comprising the adjective entered by the user;

obtain the prestored instruction according to the second voice information comprising the adjective entered by the user;

obtain text information according to the prestored instruction, by:

determining whether the text information is prestored in the user terminal;

in response to determining that the text information is prestored in the user terminal, obtaining the text information from the user terminal according to the prestored instruction; and

in response to determining that the text information is not stored in the user terminal, obtaining the text information from a server according to the prestored instruction;

obtain prestored audio information according to the text information; and

play the prestored audio information.

11. The user terminal according to claim 10 , wherein the memory further stores executable instructions that, when executed by the processor, cause the processor to: load the text information to a web page of a browser.

12. The user terminal according to claim 10 , wherein the memory further stores executable instructions that, when executed by the processor, cause the processor to:

obtain the prestored audio information by:

partitioning the text information into words or phrases;

searching an audio library for multiple audio segments corresponding to the words or the phrases; and

synthesizing the multiple audio segments into the audio.

13. The user terminal according to claim 10 , wherein the memory further stores executable instructions that, when executed by the processor, cause the processor to: play a second prompt voice after playing of the audio is completed.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 29, 2020
From: GUANGZHOU UCWEB COMPUTER TECHNOLOGY CO., LTD.
To: ALIBABA GROUP HOLDING LIMITED
Reel/Frame 053601/0565 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2017
From: LIU, XIANG
To: GUANGZHOU UCWEB COMPUTER TECHNOLOGY CO., LTD.
Reel/Frame 043841/0821 →
Priority Claims (1)
CN 2016 1 0894538 · Oct 13, 2016 · national
Continuity (1)
Related Publication 20180109677A1 · Apr 19, 2018
Cited By (1)
US 12,387,725