IP Library › Granted Patent US 12,272,352
Granted Patent B2
US 12,272,352 · App. 17/516,045 · Granted Apr 8, 2025

Electronic device and method for performing voice recognition thereof

Inventors: Ojun Kwon (Suwon-si, KR); Hyunjin Park (Suwon-si, KR); Kiyong Lee (Suwon-si, KR); Yoonju Lee (Suwon-si, KR); Jisup Lee (Suwon-si, KR); Jaeyung Yeo (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L15/18G10L13/02H04L67/125
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,272,352
App. No.
17/516,045
Granted
Apr 8, 2025
Kind
B2
Abstract

Various embodiments relate to an electronic device and a voice recognition performing method of an electronic device which are capable of receiving a voice input of a user and executing a function corresponding to a user command generated by the voice input. An electronic device according to various embodiments may include: a communication circuitry, a microphone, a display, and a processor, wherein the processor may be configured to: receive a voice utterance through the microphone, perform speech recognition on the received voice using a natural language platform for processing a command, determine whether to process the command based on an interaction with a server while performing the speech recognition, generate intermediate data corresponding to a state in which the speech recognition is performed based on determining processing based on the interaction with the server, control the communication circuitry to transmit the intermediate data to the server, receive a processing result of processing the command from the server based on the intermediate data, and control the display to display the processing result.

Claims (70)

1. An electronic device comprising:

a communication circuitry;

a microphone;

a display;

at least one processor comprising processing circuitry and operatively connected to the communication circuitry, the microphone, and the display; and

memory storing instructions which, when executed by the at least one processor, cause the electronic device to perform operations comprising:

receiving a voice utterance through the microphone,

providing the received voice utterance to a first natural language platform, comprising a plurality of processing modules, for performing speech recognition to process a command corresponding to the received voice utterance,

determining whether a failure occurs in processing the command in the performing of the speech recognition,

based on determining no occurrence of a failure, continuously performing speech recognition for processing the command using the first natural language platform of the electronic device,

based on determining an occurrence of a failure,

identify a result of determining compatibility between the electronic device and a server comprising a second natural language platform for the server to process intermediate data generated by the electronic device, based on an exchange of compatibility-related information between the electronic device and the server;

based on identifying a result of compatibility between the electronic device and the server, generating the intermediate data corresponding to a portion of the command processed by a processing module of the plurality of processing modules until the determined occurrence, controlling the communication circuitry to transmit the intermediate data to the server for continuously processing the command from the portion of the command by a processing module of the second natural language platform corresponding to the processing module of the first natural language platform associated with the determined occurrence, and

based on identifying a result of non-compatibility between the electronic device and the server, controlling the communication circuitry to transmit the voice utterance to the server for processing,

receiving a processing result of the command from the server, and

controlling the display to display the processing result.

2. The electronic device of claim 1 , wherein the processing modules of the first natural language platform comprise a speech recognition module, a natural language processing module, and a text-to-speech (TTS) module, and

wherein memory stores instructions which, when executed, cause the electronic device to perform operations comprising sequentially processing the command using the speech recognition module, the natural language processing module, and the TTS module.

3. The electronic device of claim 2 , wherein the intermediate data is related to one of the speech recognition module, the natural language processing module, or the TTS module.

4. The electronic device of claim 3 , wherein the intermediate data is different for each respective processing module based on a characteristic of the respective processing module.

5. The electronic device of claim 2 , wherein memory stores instructions which, when executed, cause the electronic device to perform operations comprising:

performing automatic speech recognition (ASR) on audio data corresponding to the voice utterance using the speech recognition module,

controlling the communication circuitry to transmit first intermediate data related to the ASR to the server based on occurrence of a failure for the ASR,

performing natural language processing (NLP) based on a result of the ASR, using the NLP module based on identifying no occurrence of a failure for the ASR, and

controlling the communication circuitry to transmit second intermediate data related to the NLP to the server based on occurrence of a failure for the NLP.

6. The electronic device of claim 5 , wherein memory stores instructions which, when executed, cause the electronic device to perform operations comprising:

performing TTS conversion using the TTS module based on no occurrence of a failure for the NLP, and

controlling the communication circuitry to transmit third intermediate data related to the TTS conversion to the server based on occurrence of a failure for the TTS conversion.

7. The electronic device of claim 1 , wherein memory stores instructions which, when executed, cause the electronic device to perform operations comprising:

determining an operation agent to provide the processing result, and

providing a user interface comprising the processing result and information about the operation agent.

8. The electronic device of claim 1 , wherein memory stores instructions which, when executed, cause the electronic device to perform operations comprising:

based on identifying a relationship between the voice utterance and content stored in the electronic device, controlling the communication circuitry to transmit content information comprising the relationship and information about the content and the intermediate data to the server.

9. A method of operating an electronic device, the method comprising:

receiving a voice utterance through a microphone;

providing the received voice utterance to a first natural language platform, comprising a plurality of processing modules, for performing speech recognition to process a command corresponding to the received voice utterance;

determining whether a failure occurs in processing the command in the performing of the speech recognition;

based on determining no occurrence of a failure, continuously performing speech recognition for processing the command using the first natural language platform of the electronic device;

based on determining an occurrence of a failure,

identifying a result of determining compatibility between the electronic device and a server comprising a second natural language platform for the server to process intermediate data generated by the electronic device, based on an exchange of compatibility-related information between the electronic device and the server;

based on identifying a result of compatibility between the electronic device and the server, generating the intermediate data corresponding to a portion of the command processed by a processing module of the plurality of processing modules until the determined occurrence, and transmitting the intermediate data to the server for continuously processing the command from the portion of the command by a processing module of the second natural language platform corresponding to the processing module of the first natural language platform associated with the determined occurrence through communication circuitry; and

based on identifying a result of non-compatibility between the electronic device and the server, controlling the communication circuitry to transmit the voice utterance to the server for processing,

receiving a processing result of the command from the server; and

displaying the processing result through a display.

10. The method of claim 9 , wherein the processing modules of the first natural language platform comprise a speech recognition module, a natural language processing module, and a text-to-speech (TTS) module, and

wherein the method comprises sequentially processing the command using the speech recognition module, the natural language processing module, and the TTS module.

11. The method of claim 10 , wherein the intermediate data is related to one of the speech recognition module, the natural language processing module, or the TTS module, and

wherein the intermediate data is different for each respective processing module based on a characteristic of the respective processing module.

12. The method of claim 10 , comprising:

performing automatic speech recognition (ASR) on audio data corresponding to the voice utterance using the speech recognition module,

transmitting first intermediate data related to the ASR to the server based on occurrence of a failure for the ASR,

performing natural language processing (NLP) based on a result of the ASR using the NLP module based on identifying no occurrence of a failure for the ASR,

transmitting second intermediate data related to the NLP to the server based on occurrence of a failure for the NLP,

performing TTS conversion using the TTS module based on no occurrence of a failure for the NLP, and

transmitting third intermediate data related to the TTS conversion to the server based on occurrence of a failure for the TTS conversion.

13. A non-transitory computer-readable recording medium storing a program which, when executed, causes an electronic device to perform operations comprising:

receiving a voice utterance through a microphone;

providing the received voice utterance to a first natural language platform, comprising a plurality of processing modules, for performing speech recognition to process a command corresponding to the received voice utterance;

determining whether a failure occurs in processing the command in the performing of the speech recognition;

based on determining no occurrence of a failure, continuously performing speech recognition for processing the command using the first natural language platform of the electronic device;

based on determining an occurrence of a failure,

identifying a result of determining compatibility between the electronic device and a server comprising a second natural language platform for the server to process intermediate data generated by the electronic device, based on an exchange of compatibility-related information between the electronic device and the server;

based on identifying a result of compatibility between the electronic device and the server, generating intermediate data corresponding to a portion of the command processed by a processing module of the plurality of processing modules until the determined occurrence, and transmitting the intermediate data to the server for continuously processing the command from the portion of the command by a processing module of the second natural language platform corresponding to the processing module of the first natural language platform associated with the determined occurrence through communication circuitry; and

based on identifying a result of non-compatibility between the electronic device and the server, controlling the communication circuitry to transmit the voice utterance to the server for processing,

receiving a processing result of the command from the server; and

displaying the processing result through a display.

14. An electronic device comprising the non-transitory computer-readable recording medium of claim 13 .

15. The electronic device of claim 1 , wherein the compatibility-related information comprises version information.

16. The method of claim 9 , wherein the compatibility-related information comprises version information.

17. The non-transitory computer-readable recording medium of claim 13 , wherein the compatibility-related information comprises version information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2021
From: KWON, OJUN; PARK, HYUNJIN; LEE, KIYONG; LEE, YOONJU; LEE, JISUP; YEO, JAEYUNG
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 058669/0955 →
Priority Claims (1)
KR 10-2020-0140539 · Oct 27, 2020 · national
Continuity (2)
Continuation PCTKR2021015187 · Oct 27, 2021
Related Publication 20220130377A1 · Apr 28, 2022
References Cited (20)
US 20020091527A1 · Shiau · 2002 [cited by examiner]
US 20120265528A1 · Gruber · 2012 [cited by examiner]
US 20120316878A1 · Singleton · 2012 [cited by examiner]
US 20130151250A1 · VanBlon · 2013 [cited by examiner]
US 20150262571A1 · Kaszczuk · 2015 [cited by examiner]
US 20160042748A1 · Jain et al. · 2016 [cited by applicant]
US 20160260430A1 · Panemangalore et al. · 2016 [cited by applicant]
US 20180197545A1 · Willett · 2018 [cited by examiner]
US 20190066674A1 · Jaygarl et al. · 2019 [cited by applicant]
US 20200051555A1 · Jaygarl et al. · 2020 [cited by applicant]
US 20200051560A1 · Yi et al. · 2020 [cited by applicant]
US 20200265840A1 · Kwon et al. · 2020 [cited by applicant]
US 20200286477A1 · Jaygarl et al. · 2020 [cited by applicant]
US 20210090555A1 · Mahmood · 2021 [cited by examiner]
KR 1020190023341 · 2019 [cited by applicant]
KR 1020200016774 · 2020 [cited by applicant]
KR 1020200017293 · 2020 [cited by applicant]
KR 1020200087497 · 2020 [cited by applicant]
KR 1020200101103 · 2020 [cited by applicant]
Search Report and Written Opinion issued Feb. 3, 2022 in counterpart International Patent Application No. PCT/KR2021/015187. [cited by applicant]