IP Library › Granted Patent US 12,266,343
Granted Patent B2
US 12,266,343 · App. 17/534,969 · Granted Apr 1, 2025

Electronic device and control method thereof

Inventors: Sangjun Park (Suwon-si, KR); Kihyun Choo (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L13/08G06N3/045G10L19/032
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,266,343
App. No.
17/534,969
Granted
Apr 1, 2025
Kind
B2
Abstract

The electronic device may include a communication interface; a memory configured to store a first neural network model; and a processor configured to: receive, from an external electronic device via the communication interface, compressed information related to an acoustic feature obtained based on a text; decompress the compressed information to obtain decompressed information; and obtain sound information corresponding to the text by inputting the decompressed information into the first neural network model. The first neural network model may be obtained by training a relationship between a plurality of sample acoustic features and a plurality of sample sounds corresponding to the plurality of sample acoustic features.

Claims (53)

1. An electronic device comprising:

a communication interface;

a memory configured to store a first neural network model; and

a processor connected to the communication interface and the memory,

wherein the processor is configured to:

receive, from an external electronic device via the communication interface, compressed information related to an acoustic feature obtained based on a text;

decompress the compressed information to obtain decompressed information; and

obtain sound information corresponding to the text by inputting the decompressed information into the first neural network model,

wherein the first neural network model is obtained by training a relationship between a first plurality of sample acoustic features and a first plurality of sample sounds corresponding to the first plurality of sample acoustic features, and

wherein the first neural network model is trained to obtain the first plurality of sample sounds based on the first plurality of sample acoustic features and noise; and

wherein the first plurality of sample acoustic features are acoustic features that are distorted by compressing and decompressing a plurality of original acoustic features.

2. The electronic device according to claim 1 , wherein the compressed information relates to a plurality of acoustic features based on the text,

wherein the processor is further configured to:

sequentially receive the compressed information from the external electronic device via the communication interface; and

based on a packet corresponding to a first acoustic feature of the plurality of acoustic features being damaged, obtain the sound information corresponding to the text by inputting a predetermined value to the first neural network model instead of the first acoustic feature, and

wherein the first neural network model is trained to output a sound based on at least one second acoustic feature adjacent to the predetermined value.

3. The electronic device according to claim 1 , wherein the compressed information relates to a plurality of acoustic features based on the text, and

wherein the processor is further configured to:

sequentially receive the compressed information from the external electronic device via the communication interface; and

based on a first acoustic feature of the plurality of acoustic features being damaged, predict the first acoustic feature based on at least a second acoustic feature adjacent to the first acoustic feature.

4. The electronic device according to claim 1 , wherein the noise comprises at least one of Gaussian noise or uniform noise.

5. The electronic device according to claim 1 , wherein the compressed information is obtained by quantizing the acoustic feature.

6. The electronic device according to claim 1 , wherein the compressed information is obtained by compressing the acoustic feature that is obtained by inputting the text into a second neural network model, and

wherein the second neural network model is obtained by training another relationship between a second plurality of sample texts and a second plurality of sample acoustic features corresponding to the second plurality of sample texts.

7. The electronic device according to claim 6 , wherein each second sample acoustic feature of the second plurality of sample acoustic features is obtained from a signal from which a frequency component equal to or greater than a threshold value is removed from a sound corresponding to a user voice.

8. An electronic system comprising:

a first electronic device configured to:

obtain an acoustic feature based on a text by inputting the text into a prosody neural network model,

obtain compressed information related to the acoustic feature, and

transmit the compressed information; and

a second electronic device configured to:

receive the compressed information transmitted by the first electronic device, decompress the compressed information to obtain decompressed information, and

obtain a sound corresponding to the text by inputting the decompressed information into a neural vocoder neural network model,

wherein the prosody neural network model is obtained by training a first relationship between a first plurality of sample texts and a first plurality of sample acoustic features corresponding to the first plurality of sample texts,

wherein the neural vocoder neural network model is obtained by training a relationship between a second plurality of sample acoustic features and a second plurality of sample sounds corresponding to the second plurality of sample acoustic features, and

wherein the neural vocoder neural network model is trained to obtain the second plurality of sample sounds based on the second plurality of sample acoustic features and noise, and

wherein the first plurality of sample acoustic features are acoustic features that are distorted by compressing and decompressing a plurality of original acoustic features.

9. The electronic system according to claim 8 , wherein the first electronic device is further configured to, based on a user voice input being received, obtain the text from the user voice input, and input the text into the prosody neural network model.

10. A method for controlling an electronic device, the method comprising:

receiving, from an external electronic device, compressed information related to an acoustic feature obtained based on a text;

decompressing the compressed information to obtain decompressed information; and

obtaining sound information corresponding to the text by inputting the decompressed information into a first neural network model,

wherein the first neural network model is obtained by training a relationship between a first plurality of sample acoustic features and a first plurality of sample sounds corresponding to the first plurality of sample acoustic features, and

wherein the first neural network model is trained to obtain the first plurality of sample sounds based on the first plurality of sample acoustic features and noise; and

wherein the first plurality of sample acoustic features are acoustic features that are distorted by compressing and decompressing a plurality of original acoustic features.

11. The method according to claim 10 , wherein the compressed information relates to a plurality of acoustic features based on the text,

wherein the receiving the compressed information comprises sequentially receiving the compressed information from the external electronic device,

wherein the obtaining the sound information comprises, based on a packet corresponding to a first acoustic feature of the plurality of acoustic features being damaged, obtaining the sound information corresponding to the text by inputting a predetermined value into the first neural network model instead of the first acoustic feature, and

wherein the first neural network model is trained to output a sound based on at least second acoustic feature adjacent to the predetermined value.

12. The method according to claim 10 , wherein the compressed information relates to a plurality of acoustic features based on the text,

wherein the receiving the compressed information comprises sequentially receiving the compressed information from the external electronic device, and

wherein the method further comprises based on a first acoustic feature of the plurality of acoustic features being damaged, predicting the first acoustic feature based on at least a second acoustic feature adjacent to the first acoustic feature.

13. The method according to claim 10 , wherein the noise comprises at least one of Gaussian noise or uniform noise.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 24, 2021
From: PARK, SANGJUN; CHOO, KIHYUN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 058207/0769 →
Priority Claims (1)
KR 10-2021-0024028 · Feb 23, 2021 · national
Continuity (2)
Continuation PCTKR2021012757 · Sep 17, 2021
Related Publication 20220270588A1 · Aug 25, 2022
References Cited (29)
US 5673362A · Matsumoto · 1997 [cited by applicant]
US 6810379B1 · Vermeulen et al. · 2004 [cited by applicant]
US 6985856B2 · Wang et al. · 2006 [cited by applicant]
US 9286885B2 · Sienel et al. · 2016 [cited by applicant]
US 9761219B2 · Xu · 2017 [cited by examiner]
US 10176809B1 · Piérard · 2019 [cited by examiner]
US 10573293B2 · Bengio et al. · 2020 [cited by applicant]
US 20040128128A1 · Wang et al. · 2004 [cited by applicant]
US 20060004577A1 · Nukaga et al. · 2006 [cited by applicant]
US 20200074985A1 · Clark et al. · 2020 [cited by applicant]
US 20210151029A1 · Gururani · 2021 [cited by examiner]
JP 5233565A · 1993 [cited by applicant]
JP 201148076A · 2011 [cited by applicant]
JP 5524131B2 · 2014 [cited by applicant]
KR 1020050091034A · 2005 [cited by applicant]
KR 100850419B1 · 2008 [cited by applicant]
KR 102044520B1 · 2019 [cited by applicant]
KR 1020200092501A · 2020 [cited by applicant]
WO 2019222591A1 · 2019 [cited by applicant]
WO 2020145472A1 · 2020 [cited by applicant]
WO 2021002967A1 · 2021 [cited by applicant]
Ding et al, A systemic review of the development of speech synthesis, 2023 the 8th internationl Conference on Computera nd Communication Systems, pp. 28-33 (Year: 2023). [cited by examiner]
Bollepalli et al, Normal to Lombard adaptation of speech synthesis using long short term memory recurrent neural networks, Speech Communication, Apr. 18, 2019, pp. 64-75 (Year: 2019). [cited by examiner]
Jean-Marc Valin et al., “A Real-Time Wideband Neural Vocoder at 1.6 kb/s Using LPCNet”, Interspeech, Sep. 29, 2019, pp. 1-5 (6 pages total). [cited by applicant]
Jan Skoglund et al., “Improving Opus Low Bit Rate Qualitiy with Neural Speech Synthesis”, Interspeech, Oct. 25-29, 2020, pp. 2847-2851 (5 pages total). [cited by applicant]
International Search Report (PCT/ISA/210) and Written Opinion (PCT/ISA/237) dated Dec. 20, 2021 issued by the International Searching Authority in International Application No. PCT/KR2021/012757. [cited by applicant]
Kwang Myung Jeon et al., “HMM-Based Distributed Text-To-Speech Synthesis Incorporating Speaker-Adaptive Training,” International Journal of Multimedia and Ubiquitous Engineering, 2014, pp. 107-120 (14 pages total). [cited by applicant]
Dario Rethage et al., “A Wavenet for Speech Denoising,” IEEE, 2018, pp. 5069-5073 (5 pages total). [cited by applicant]
Communication dated Feb. 19, 2025 issued by the Korean Patent Office for Korean Patent Application No. 10-2021-0024028. [cited by applicant]