IP Library › Granted Patent US 12,198,675
Granted Patent B2
US 12,198,675 · App. 18/171,079 · Granted Jan 14, 2025

Electronic apparatus and method for controlling thereof

Inventors: Hosang Sung (Suwon-si, KR); Kyoungbo Min (Suwon-si, KR); Seonho Hwang (Suwon-si, KR); Doohwa Hong (Seoul, KR); Eunmi Oh (Suwon-si, KR); Jonghoon Jeong (Suwon-si, KR); Kihyun Choo (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L13/08G10L13/00G10L13/047G10L17/00G10L25/63
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,675
App. No.
18/171,079
Granted
Jan 14, 2025
Kind
B2
Abstract

An electronic apparatus which acquires input data to be input into a TTS module for outputting a voice through the TTS module, acquires a voice signal corresponding to the input data through the TTS module, detects an error in the acquired voice signal based on the input data, corrects the input data based on the detection result, and acquires a corrected voice signal corresponding to the corrected input data through the TTS module.

Claims (82)

1. An electronic apparatus comprising:

at least one processor; and

a memory configured to store at least one instruction which, when executed by the at least one processor, causes the electronic apparatus to:

acquire input data to be input into a text-to-speech (TTS) module for outputting a voice through the TTS module,

acquire a voice signal corresponding to the input data through the TTS module,

identify, by the at least one processor, a change in a frequency characteristic of the voice signal,

detect, by the at least one processor, an error in the voice signal based on the identified change in the frequency characteristic of the voice signal,

correct the input data based on a result of detecting the error, and

acquire a corrected voice signal corresponding to the input data corrected based on the result of detecting the error through the TTS module.

2. The electronic apparatus of claim 1 , wherein the at least one instruction further causes the electronic apparatus to:

based on the change in the frequency characteristic of the voice signal, identify, by the at least one processor, a change in a frequency, amplitude, and cycle of the voice signal,

based on the identified change in the frequency, amplitude, and cycle of the voice signal, identify, by the at least one processor, a change of a pitch of the voice, and

based on the identified change of the pitch of the voice, detect, by the at least one processor, the error in the voice signal.

3. The electronic apparatus of claim 2 , wherein the at least one instruction further causes the electronic apparatus to:

based on the identified change of the pitch of the voice, identify, by the at least one processor, at least one of an emotion, a voice tone, a style, and a prosody, and

based on the identified at least one of the emotion, the voice tone, the style, and the prosody, detect, by the at least one processor, an error in the voice signal.

4. The electronic apparatus of claim 1 , wherein the at least one instruction further causes the electronic apparatus to:

based on the detection of the error in the voice signal, adjust, by the at least one processor, a frequency pitch of the voice signal, and

correct the input data based on the adjusted frequency pitch of the voice signal.

5. The electronic apparatus of claim 1 , wherein the input data comprises first text data, and

wherein the at least one instruction further causes the electronic apparatus to:

convert the voice signal into second text data,

compare the first text data included in the input data and the second text data, and

detect the error in the voice signal based on a result of comparing the first text data and the second text data.

6. The electronic apparatus of claim 1 , wherein the at least one instruction further causes the electronic apparatus to:

compare a length of the voice signal and a length of text data included in the input data, and

detect the error in the voice signal based on a result of comparing the length of the voice signal and the length of the text data included in the input data.

7. The electronic apparatus of claim 1 , wherein the at least one instruction further causes the electronic apparatus to:

based on detecting the error in the voice signal, correct at least one of a spacing or a punctuation mark of text data included in information on the input data, and

input corrected input data having the at least one of the spacing or the punctuation mark of the text data into the TTS module.

8. The electronic apparatus of claim 1 , wherein the at least one instruction further causes the electronic apparatus to:

based on detecting the error in the voice signal, correct the input data by applying a speech synthesis markup language (SSML) to text data included in the input data, and

input corrected input data having the speech synthesis markup language (SSML) applied to the text data into the TTS module.

9. The electronic apparatus of claim 1 , wherein the at least one instruction further causes the electronic apparatus to:

convert a received user voice into text data by using a voice recognition module, analyze an intent of the text data, and

acquire response information corresponding to the received user voice as the input data.

10. The electronic apparatus of claim 1 , further comprising:

a speaker,

wherein the at least one instruction further causes the electronic apparatus to:

add an indicator indicating correction to the voice signal, and

output the voice signal having the indicator through the speaker.

11. The electronic apparatus of claim 1 , further comprising:

a speaker; and

a microphone,

wherein the at least one instruction further causes the electronic apparatus to:

output the voice signal through the speaker, and

based on the voice signal output through the speaker being received through the microphone, detect the error in the voice signal received through the microphone based on the input data.

12. The electronic apparatus of claim 11 , wherein the at least one instruction further causes the electronic apparatus to:

identify an identity of the voice signal received through the microphone,

based on the voice signal received through the microphone being the voice signal output through the speaker based on the identity, detect the error in the voice signal, and

based on the voice signal received through the microphone having been uttered by a user based on the identity, convert the voice signal into text data by using a voice recognition module, and analyze an intent of the text data and acquire response information corresponding to the received user voice as the input data.

13. The electronic apparatus of claim 1 , further comprising:

a communicator,

wherein the at least one instruction further causes the electronic apparatus to transmit the voice signal to an external apparatus through the communicator.

14. A method of controlling an electronic apparatus, the method comprising:

acquiring input data to be input into a text-to-speech (TTS) module for outputting a voice through the TTS module;

acquiring a voice signal corresponding to the input data through the TTS module;

identifying, by at least one processor of the electronic apparatus, a change in a frequency characteristic of the voice signal;

detecting, by the at least one processor, an error in the voice signal based on the identified change in the frequency characteristic of the voice signal;

correcting the input data based on a result of the detecting the error; and

acquiring a corrected voice signal corresponding to the input data corrected based on the result of detecting the error through the TTS module.

15. The method of claim 14 , wherein the detecting the error comprises:

based on the change in the frequency characteristic of the voice signal, identifying a change in a frequency, amplitude, and cycle of the voice signal;

based on the identified change in the frequency, amplitude, and cycle of the voice signal, identifying a change of a pitch of the voice; and

based on the identified change of the pitch of the voice, detecting an error in the voice signal.

16. The method of claim 15 , wherein the detecting the error comprises:

based on the identified change of the pitch of the voice, identify, by the at least one processor, at least one of an emotion, a voice tone, a style, and a prosody, and

based on the identified at least one of the emotion, the voice tone, the style, and the prosody, detect, by the at least one processor, an error in the voice signal.

17. The method of claim 14 , wherein the correcting the input data comprises:

based on the detection of the error in the voice signal, adjusting a frequency pitch of the voice signal; and

correcting the input data based on the adjusted frequency pitch of the voice signal.

18. The method of claim 14 , wherein the input data comprises first text data, and

wherein the detecting the error comprises:

converting the voice signal into second text data;

comparing the first text data included in the input data and the second text data; and

detecting the error in the voice signal based on a result of the comparing the first text data and the second text data.

19. The method of claim 14 , wherein the detecting the error comprises:

comparing a length of the voice signal and a length of text data included in the input data; and

detecting the error in the voice signal based on a result of comparing the length of the voice signal and the length of the text data included in the input data.

20. The method of claim 14 , wherein the correcting comprises:

based on detecting the error in the voice signal, correcting at least one of a spacing or a punctuation mark of text data included in the input data; and

inputting corrected input data having the at least one of the spacing or the punctuation mark of the text data into the TTS module.

Priority Claims (1)
KR 10-2019-0024192 · Feb 28, 2019 · national
Continuity (2)
Continuation 16788418 · Feb 12, 2020
Related Publication 20230206897A1 · Jun 29, 2023
References Cited (56)
US 5905970A · Aoyagi · 1999 [cited by examiner]
US 6081780A · Lumelsky · 2000 [cited by applicant]
US 6246672B1 · Lumelsky · 2001 [cited by applicant]
US 6477495B1 · Nukaga · 2002 [cited by examiner]
US 9245521B2 · Park et al. · 2016 [cited by applicant]
US 9275633B2 · Cath et al. · 2016 [cited by applicant]
US 9293129B2 · Zhao et al. · 2016 [cited by applicant]
US 9576574B2 · van Os · 2017 [cited by applicant]
US 9978359B1 · Kaszczuk · 2018 [cited by examiner]
US 10008196B2 · Maisonnier et al. · 2018 [cited by applicant]
US 10276149B1 · Liang · 2019 [cited by examiner]
US 10643600B1 · Aryal · 2020 [cited by examiner]
US 11587547B2 · Sung · 2023 [cited by examiner]
US 20010025242A1 · Joncour · 2001 [cited by applicant]
US 20020111794A1 · Yamamoto · 2002 [cited by examiner]
US 20050228672A1 · Lewis · 2005 [cited by examiner]
US 20060149546A1 · Runge · 2006 [cited by examiner]
US 20070027686A1 · Schramm · 2007 [cited by examiner]
US 20080091426A1 · Rempel et al. · 2008 [cited by applicant]
US 20080183473A1 · Nagano · 2008 [cited by examiner]
US 20090043583A1 · Agapi · 2009 [cited by examiner]
US 20090299733A1 · Agapi · 2009 [cited by examiner]
US 20100005362A1 · Ito et al. · 2010 [cited by applicant]
US 20120035917A1 · Kim · 2012 [cited by examiner]
US 20140019127A1 · Park et al. · 2014 [cited by applicant]
US 20140257815A1 · Zhao · 2014 [cited by examiner]
US 20150205779A1 · Bak et al. · 2015 [cited by applicant]
US 20160294739A1 · Stoehr et al. · 2016 [cited by applicant]
US 20170269975A1 · Wood · 2017 [cited by examiner]
US 20170270086A1 · Fume · 2017 [cited by examiner]
US 20170352346A1 · Paulik · 2017 [cited by examiner]
US 20180211652A1 · Mun et al. · 2018 [cited by applicant]
US 20180336882A1 · Reber · 2018 [cited by examiner]
US 20190295527A1 · Pore · 2019 [cited by examiner]
US 20200051545A1 · Iwase · 2020 [cited by examiner]
US 20200279551A1 · Sung · 2020 [cited by examiner]
US 20230206897A1 · Sung · 2023 [cited by examiner]
CN 1366659A · 2002 [cited by applicant]
CN 101183525A · 2008 [cited by applicant]
CN 103366731A · 2013 [cited by applicant]
CN 103546787A · 2014 [cited by applicant]
CN 108288464A · 2018 [cited by examiner]
CN 109461435A · 2019 [cited by examiner]
CN 109979427A · 2019 [cited by applicant]
CN 110600002A · 2019 [cited by applicant]
JP 2003108170A · 2003 [cited by applicant]
JP 2018109663A · 2018 [cited by applicant]
KR 100918644B1 · 2009 [cited by applicant]
KR 101221188A · 2013 [cited by applicant]
KR 1020160080711A1 · 2016 [cited by applicant]
Communication dated Nov. 3, 2023, issued by the China National Intellectual Property Administration in Chinese Application No. 202080017465.1. [cited by applicant]
Communication issued Mar. 28, 2023 by the Korean Intellectual Property Office in Korean Patent Application No. 10-2019-0024192. [cited by applicant]
Communication issued May 11, 2023 by the European Patent Office in European Patent Application No. 20763060.9. [cited by applicant]
Communication issued Oct. 25, 2021 by the European Patent Office in counterpart European Patent Application No. 20763060.9. [cited by applicant]
International Search Report dated May 13, 2020 issued by the International Searching Authority in counterpart International Application No. PCT/KR2020/001581 (PCT/ISA/210). [cited by applicant]
International Written Opinion dated May 13, 2020 issued by the International Searching Authority in counterpart International Application No. PCT/KR2020/001581 (PCT/ISA/237). [cited by applicant]