IP Library › Granted Patent US 8,935,165
Granted Patent B2
US 8,935,165 · App. 13/617,222 · Granted Jan 13, 2015

Method for displaying words and processing device and computer program product thereof

Inventors: Yu-Chen Huang (Tao Yuan Shien, TW); Che-Kuang Lin (Tao Yuan Shien, TW)
Assignee: Quanta Computer Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,935,165
App. No.
13/617,222
Granted
Jan 13, 2015
Kind
B2
Abstract

The disclosure provides a method for displaying words. In the method, a speech signal is received. A pitch contour and an energy contour of the speech signal are extracted. Speech recognition is performed on the speech signal to recognize a plurality of words corresponding to the speech signal and determine time alignment information of each of the plurality of words. At least one display parameter of each of the plurality of words is determined according to the pitch contour, the energy contour and the time alignment information of each of the plurality of words. Thus, the plurality of words is integrated into a sentence according to the at least one display parameter of each of the plurality of words. Then, the sentence is displayed on at least one display device.

Claims (65)

1. A method for displaying words, comprising:

receiving a speech signal;

extracting a pitch contour of the speech signal;

extracting an energy contour of the speech signal;

performing speech recognition on the speech signal to recognize a plurality of words corresponding to the speech signal and determine time alignment information of each of the plurality of words;

determining at least one display parameter of each of the plurality of words according to the pitch contour, the energy contour and the time alignment information of each of the plurality of words;

integrating the plurality of words into a sentence according to the at least one display parameter of each of the plurality of words; and

outputting the sentence to be displayed on at least one display device.

2. The method as claimed in claim 1 , wherein the at least one display parameter comprises a position parameter, a size parameter and a distance parameter.

3. The method as claimed in claim 2 , further comprising:

capturing a facial image via a video camera;

determining facial expression intensity according to the facial image; and

determining whether to insert at least one first emoticon into the sentence according to the facial expression intensity.

4. The method as claimed in claim 3 , further comprising:

calculating a plurality of Mel-scale frequency cepstral coefficients of the speech signal;

calculating energy of the speech signal according to the plurality of Mel-scale frequency cepstral coefficients to obtain the energy contour; and

performing the speech recognition on the speech signal according to the plurality of Mel-scale frequency cepstral coefficients to recognize the plurality of words and determine the time alignment information of each of the plurality of words.

5. The method as claimed in claim 4 , wherein the time alignment information of each of the plurality of words comprises a starting time and an ending time in the speech signal of each of the plurality of words.

6. The method as claimed in claim 5 , further comprising:

determining the distance parameter of each of the plurality of words according to a time difference between the starting time and the ending time of each of the plurality of words.

7. The method as claimed in claim 6 , further comprising:

calculating an average energy of each of the plurality of words according to the energy contour between the starting time and the ending time of each of the plurality of words; and

determining the size parameter of each of the plurality of words according to the average energy of each of the plurality of words.

8. The method as claimed in claim 7 , further comprising:

calculating a regression line of each of the plurality of words according to the pitch contour between the starting time and the ending time of each of the plurality of words; and

determining the position parameter of each of the plurality of words according to a slope of the regression line of each of the plurality of words.

9. The method as claimed in claim 8 , further comprising:

determining whether to insert at least one second emoticon into a position in the sentence near at least one word the plurality of words according to the average energy of the at least one word and the slope of the regression line of the at least one word; and

if so, determining the at least one second emoticon according to the average energy of the at least one word and the slope of the regression line of the at least one word.

10. A processing device, comprising:

a speech input unit, receiving a speech signal;

a processor, comprising:

a pitch extracting module, extracting a pitch contour of the speech signal;

an energy calculating module, extracting an energy contour of the speech signal;

a speech recognition engine, performing speech recognition on the speech signal to recognize a plurality of words corresponding to the speech signal and determine time alignment information of each of the plurality of words; and

a text processing module, determining at least one display parameter of each of the plurality of words according to the pitch contour, the energy contour and the time alignment information of each of the plurality of words and integrating the plurality of words into a sentence according to the at least one display parameter of each of the plurality of words; and

a text output unit, outputting the sentence to be displayed on at least one display device.

11. The processing device as claimed in claim 10 , wherein the at least one display parameter comprises a position parameter, a size parameter and a distance parameter.

12. The processing device as claimed in claim 11 further comprising:

an image input unit, capturing an image,

wherein the processor further comprises:

a face recognition module, performing face recognition on the image to retrieve a facial image;

a facial feature extracting module, extracting a facial feature of the facial image; and

an expression parameter module, determining facial expression intensity according to the facial feature, and

wherein the text processing module further determines whether to insert at least one first emoticon into the sentence according to the facial expression intensity.

13. The processing device as claimed in claim 12 , wherein the processor further comprises:

a Mel-scale frequency cepstral coefficient (MFCC) module, calculating a plurality of Mel-scale frequency cepstral coefficients of the speech signal,

wherein the energy calculating module calculates energy of the speech signal according to the plurality of Mel-scale frequency cepstral coefficients to obtain the energy contour, and

wherein the speech recognition engine recognizes the plurality of words and determines the time alignment information of each of the plurality of words according to the plurality of Mel-scale frequency cepstral coefficients.

14. The processing device as claimed in claim 13 , wherein the time alignment information of each of the plurality of words comprises a starting time and an ending time in the speech signal of each of the plurality of words.

15. The processing device as claimed in claim 14 , wherein the text processing module determines the distance parameter of each of the plurality of words according to a time difference between the starting time and the ending time of each of the plurality of words.

16. The processing device as claimed in claim 15 , wherein the text processing module calculates an average energy of each of the plurality of words according to the energy contour between the starting time and the ending time of each of the plurality of words and determines the size parameter of each of the plurality of words according to the average energy of each of the plurality of words.

17. The processing device as claimed in claim 16 , wherein the text processing module calculates a regression line of each of the plurality of words according to the pitch contour between the starting time and the ending time of each of the plurality of words and determines the position parameter of each of the plurality of words according to a slope of the regression line of each of the plurality of words.

18. The processing device as claimed in claim 17 , wherein the text processing module determines whether to insert at least one second emoticon into a position in the sentence near at least one word of the plurality of words according to the average energy of the at least one word and the slope of the regression line of the at least one word, and if so, the text processing module determines the at least one second emoticon according to the average energy of the at least one word and the slope of the regression line of the at least one word.

19. A computer program product embodied in a non-transitory computer-readable storage medium, wherein the computer program product is loaded into and executed by an electronic device for performing a method for displaying words, the computer program product comprising:

a first code for receiving a speech signal;

a second code for extracting a pitch contour of the speech signal;

a third code for extracting an energy contour of the speech signal;

a fourth code for performing speech recognition on the speech signal to recognize a plurality of words corresponding to the speech signal and determine time alignment information of each of the plurality of words;

a fifth code for determining at least one display parameter of each of the plurality of words according to the pitch contour, the energy contour and the time alignment information of each of the plurality of words; and

a sixth code for integrating the plurality of words into a sentence according to the at least one display parameter of each of the plurality of words and outputting the sentence to be displayed on at least one display device.

20. The computer program product as claimed in claim 19 , further comprising:

a seventh code for capturing a facial image via a video camera;

an eighth code for determining facial expression intensity according to the facial image; and

a ninth code for determining whether to insert at least one first emoticon into the sentence according to the facial expression intensity.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 14, 2012
From: HUANG, YU-CHEN; LIN, CHE-KUANG
To: QUANTA COMPUTER INC.
Reel/Frame 028974/0292 →
Priority Claims (1)
TW 101120062 A · Jun 5, 2012 · national
Continuity (1)
Related Publication 20130325464A1 · Dec 5, 2013