IP Library Granted Patent US 11,211,048
Granted Patent B2
US 11,211,048 · App. 16/478,702 · Granted Dec 28, 2021

Method for sensing end of speech, and electronic apparatus implementing same

Inventors: Yong Ho Kim (Seoul, KR); Sourabh Pateriya (Bangalore, IN); Sunah Kim (Seongnam-si, KR); Gahyun Joo (Suwon-si, KR); Sang-Woong Hwang (Yongin-si, KR); Say Jang (Yongin-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G10L15/05G10L15/22G10L15/25G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,211,048
App. No.
16/478,702
Granted
Dec 28, 2021
Kind
B2
Abstract

Provided are an apparatus and a method, a variety of embodiments of the apparatus comprising a microphone, memory, and a processor functionally connected to the microphone or memory, wherein the processor is configured to: count end-point detection (EPD) time on the basis of a voice input; when the EPD time expires, determine whether the final word of the voice input corresponds to a previously configured word stored in memory; and, if the final word corresponds to the previously configured word, then extend the EPD time and wait for reception of a voice input. Additionally, other embodiments are possible.

Claims (59)

1. An electronic device comprising:

a microphone;

a camera or sensor arranged to detect a user's gesture;

a memory; and

a processor functionally connected with the microphone or the memory,

wherein the processor is configured to:

count an end point detection (EPD) time based on receiving a voice input,

determine whether a last word of the voice input corresponds to a predetermined word stored in the memory and a predetermined gesture is detected by the camera or sensor when the EPD time expires,

extend the EPD time after the reception of the voice input when the last word corresponds to the predetermined word and the predetermined gesture is detected, and

provide a service based on text converted from the voice input, wherein the processor is further configured to:

analyze a user's intent to end a speech in a method of calculating an index related to the user's intent to end the speech by giving a weight value or a point to whether the predetermined word is detected and whether the predetermined gesture is detected,

set an EPD extension time to a first time when the index is a first index and set the EPD extension time to a second time longer than the first time when the index is a second index higher than the first index, and

when the index is greater than or equal to a predetermined index, extend the EPD time based on the EPD extension time, and

wherein the predetermined gesture includes a specific gesture for a word that will be used to speak.

2. The electronic device of claim 1 , wherein the processor is further configured to, when the last word corresponds to a predetermined word comprising at least one of an empty word, a conjunction, or a waiting instruction, extend the EPD time.

3. The electronic device of claim 1 , wherein the processor is further configured to, when an additional voice input is detected before the EPD time expires, extend the EPD time.

4. The electronic device of claim 1 ,

wherein the predetermined word comprises a common word and a personal word, and

wherein the processor is further configured to:

determine similarity between a voice command recognized after a voice command failure and a previous voice command, and

collect the personal word based on a degree of the similarity.

5. The electronic device of claim 4 , wherein the processor is further configured to:

analyze changed text information between the voice command and the previous voice command, and

when the changed text information is detected a predetermined number of times or more, update the text information with the personal word.

6. The electronic device of claim 1 , wherein the processor is further configured to:

determine whether a sentence according to the voice input is completed when the EPD time expires, and

when it is determined that the sentence is not completed, extend the EPD time.

7. The electronic device of claim 6 , wherein the processor is further configured to determine whether to perform an operation of determining whether the sentence is completed, based on a type of a voice command according to the voice input.

8. The electronic device of claim 1 , wherein the processor is further configured to:

extend the EPD time according to a fixed value, or to change the EPD time to a value corresponding to context recognition, and

extend the EPD time according to the changed value.

9. The electronic device of claim 1 , wherein the processor is further configured to determine the EPD time or the EPD extension time, based on context information of the electronic device and characteristic information of the user.

10. The electronic device of claim 1 , wherein the processor is further configured to analyze the user's intent to end the speech based on at least one of context information of the electronic device, characteristic information of the user, whether an additional voice input is detected, or whether a sentence is completed.

11. The electronic device of claim 10 , wherein the processor is further configured to:

calculate the index related to the user's intent to end the speech by giving the weight value or the point to at least one of a silence detection time, or whether the sentence is completed.

12. The electronic device of claim 1 , wherein the service comprises mobile search, schedule management, calling, memo, and music play.

13. An operation method of an electronic device, the method comprising:

counting, by a processor of the electronic device, an end point detection (EPD) time, based on receiving a voice input through a microphone of the electronic device;

when the EPD time expires, determining, by the processor, whether a last word of the voice input corresponds to a predetermined word stored in a memory of the electronic device and a predetermined gesture is detected by a camera or sensor of the electronic device;

when the last word corresponds to the predetermined word and the predetermined gesture is detected, extending, by the processor, the EPD time after the reception of the voice input; and

providing, by the processor, a service based on text converted from the voice input,

wherein the method further comprises:

analyzing, by the processor, a user's intent to end a speech in a method of calculating an index related to the user's intent to end the speech by giving a weight value or a point to whether the predetermined word is detected and whether the predetermined gesture is detected,

setting, by the processor, an EPD extension time to a first time when the index is a first index and setting, by the processor, the EPD extension time to a second time longer than the first time when the index is a second index higher than the first index, and

when the index is greater than or equal to a predetermined index, extending, by the processor, the EPD time based on the EPD extension time, and

wherein the predetermined gesture includes a specific gesture for a word that will be used to speak.

14. The method of claim 13 ,

wherein the predetermined word comprises a common word and a personal word, and

wherein the method further comprises:

determining, by the processor, similarity between a voice command recognized after a voice command failure and a previous voice command; and

collecting, by the processor, the personal word based on a degree of the similarity.

15. The method of claim 14 , wherein collecting comprises:

analyzing, by the processor, changed text information between the voice command and the previous voice command; and

when the changed text information is detected a predetermined number of times or more, updating, by the processor, the text information with the personal word.

16. The method of claim 13 , further comprising:

when the EPD time expires, determining, by the processor, whether a sentence according to the voice input is completed; and

when it is determined that the sentence is not completed, extending, by the processor, the EPD time.

17. The method of claim 13 , further comprising determining, by the processor, the EPD time or the EPD extension time, based on context information of the electronic device and characteristic information of a user.

18. The method of claim 13 , further comprising analyzing, by the processor, the user's intent to end the speech based on at least one of context information of the electronic device, characteristic information of the user, whether an additional voice input is detected, or whether a sentence is completed.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 17, 2019
From: KIM, YONG HO; PATERIYA, SOURABH; KIM, SUNAH; JOO, GAHYUN; HWANG, SANG-WOONG; JANG, SAY
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 049780/0580 →
Priority Claims (1)
KR 10-2017-0007951 · Jan 17, 2017 · national
Continuity (1)
Related Publication 20190378493A1 · Dec 12, 2019
Cited By (18)
US 12,197,817 US 12,200,297 US 12,211,502 US 12,219,314 US 12,236,952 US 12,293,763 US 12,301,635 US 12,333,404 US 12,361,943 US 12,367,879 US 12,386,434 US 12,386,491 US 12,477,470 US 12,536,996 US 12,562,156 US 12,608,171 US 12,619,452 US 12,744,038