IP Library › Granted Patent US 9,691,381
Granted Patent B2
US 9,691,381 · App. 13/400,585 · Granted Jun 27, 2017

Voice command recognition method and related electronic device and computer-readable medium

Inventors: Yiou-Wen Cheng (Hsinchu, TW); Liang-Che Sun (Taipei, TW); Chao-Ling Hsu (Hsinchu, TW); Hsi-Kang Tsao (Taipei, TW); Jyh-Horng Lin (Hsinchu, TW)
Assignee: MEDIATEK INC.
G10L15/22G10L2015/223G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,691,381
App. No.
13/400,585
Granted
Jun 27, 2017
Kind
B2
Abstract

An electronic device for browsing a document is disclosed. The document being browsed includes a plurality of command-associated text strings. First, a text string selector of the electronic device selects a plurality of candidate text strings from the command-associated text strings. Afterward, an acoustic string provider of the electronic device prepares a candidate acoustic string for each of the candidate text strings. Thereafter, a microphone of the electronic device receives a voice command. Next, a speech recognizer of the electronic device searches the candidate acoustic strings for a target acoustic string that matches the voice command, wherein the target acoustic string corresponds to a target text string of the candidate text strings. Finally, a document browser of the electronic device executes a command associated with the target text string.

Claims (35)

1. A voice command recognition method for an electronic device, comprising:

utilizing a text string selector to dynamically select a plurality of command-associated text strings of a document being browsed as candidate text strings, wherein the candidate text strings are displayed to a user;

utilizing an acoustic string provider to prepare a candidate acoustic string for each of the candidate text strings, respectively;

utilizing a microphone to receive a voice command;

utilizing a speech recognizer to search the candidate acoustic strings for a target acoustic string that matches the voice command, the target acoustic string corresponding to a target text string of the candidate text strings; and

utilizing the electronic device to execute a command associated with the target text string;

wherein each of the candidate acoustic strings represents the pronunciation of a symbol additionally displayed next to a corresponding candidate text string to represent that corresponding text string.

2. The method of claim 1 , wherein the additionally displayed symbols are a sequence of numbers and each of the candidate acoustic strings represents the pronunciation of a number of a corresponding candidate text string among the plurality of candidate text strings.

3. The method of claim 1 , wherein the step of selection comprises:

selecting the command-associated text strings in a displayed region of the document to be the candidate text strings.

4. The method of claim 1 , wherein the step of selection comprises:

selecting the command-associated text strings in a region specified by a user browsing the document to be the candidate text strings.

5. The method of claim 1 , wherein the step of selection comprises:

selecting the command-associated text strings in a gazed region of the document to be the candidate text strings.

6. The method of claim 1 , wherein the step of selection comprises:

selecting the command-associated text strings in a gesture-specified region of the document to be the candidate text strings.

7. The method of claim 1 , wherein the step of selection comprises:

selecting the command-associated text strings in a subordinate level of a parent object specified by a user browsing the document to be the candidate text strings.

8. An electronic device, comprising:

a text string selector, configured to dynamically select a plurality of command-associated text strings of a document being browsed as candidate text strings, wherein the candidate text strings are displayed to a user;

an acoustic string provider, configured to prepare a candidate acoustic string for each of the candidate text strings, respectively;

a microphone, configured to receive a voice command; and

a speech recognizer, configured to search the candidate acoustic strings for a target acoustic string that matches the voice command, the target acoustic string corresponding to a target text string of the candidate text strings;

wherein each of the candidate acoustic strings represents the pronunciation of a symbol additionally displayed next to a corresponding candidate text string to represent that corresponding text string, and the electronic device is configured to execute a command associated with the target text string.

9. The electronic device of claim 8 , wherein the additionally displayed symbols are a sequence of numbers and each of the candidate acoustic strings represents the pronunciation of a number of a corresponding candidate text string among the plurality of candidate text strings.

10. The electronic device of claim 8 , wherein the text string selector is configured to select the command-associated text strings in a displayed region of the document to be the candidate text strings.

11. The electronic device of claim 8 , wherein the text string selector is configured to select the command-associated text strings in a gazed region of the document to be the candidate text strings, and the electronic device further comprises an eye tracker configured to assist in defining the gazed region by tracking a user's eyes.

12. The electronic device of claim 8 , wherein the text string selector is configured to select the command-associated text strings in a gesture-specified region of the document to be the candidate text strings, and the electronic device further comprises a gesture detector configured to assist in defining the gesture-specified region by detecting a user's gestures.

13. A non-transitory computer-readable medium comprising a microphone and storing at least one computer program which, when executed by an electronic device, causes the electronic device to perform operations comprising:

dynamically selecting a plurality of command-associated text strings of a document being browsed as candidate text strings, wherein the candidate text strings are displayed to a user;

preparing a candidate acoustic string for each of the candidate text strings, respectively;

receiving a voice command;

searching the candidate acoustic strings for a target acoustic string that matches the voice command, the target acoustic string corresponding to a target text string of the candidate text strings; and

executing a command associated with the target text string;

wherein each of the candidate acoustic strings represents the pronunciation of a symbol additionally displayed next to a corresponding candidate text string to represent that corresponding text string.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 21, 2012
From: CHENG, YIOU-WEN; SUN, LIANG-CHE; HSU, CHAO-LING; TSAO, HSI-KANG; LIN, JYH-HORNG
To: MEDIATEK INC.
Reel/Frame 027732/0847 →
Continuity (1)
Related Publication 20130218573A1 · Aug 22, 2013