IP Library › Granted Patent US 12,217,016
Granted Patent B2
US 12,217,016 · App. 17/746,238 · Granted Feb 4, 2025

Electronic apparatus for translating voice input using neural network models and method for controlling thereof

Inventors: Beomseok Lee (Suwon-si, KR); Sathish Indurthi (Suwon-si, KR); Mohd Abbas Zaidi (Suwon-si, KR); Nikhil Kumar (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06F40/58G06F40/284G06F40/44G10L13/02G10L15/16G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,016
App. No.
17/746,238
Granted
Feb 4, 2025
Kind
B2
Abstract

An electronic apparatus, including a microphone; a memory configured to store at least one instruction; and a processor configured to: acquire a first token corresponding to a first user voice input in a first language acquired through the microphone, acquire a first text in a second language by inputting the first token into a first neural network model, acquire a feature value corresponding to a predicted subsequent token, which is predicted to be uttered after the first token, by inputting the first text into a second neural network model, and based on a second token being acquired subsequent to the first token, acquire a second text in the second language by inputting the first token, the second token, the first text, and the feature value into the first neural network model.

Claims (39)

1. An electronic apparatus comprising:

a microphone;

a memory configured to store at least one instruction; and

a processor configured to:

acquire a first token corresponding to a first user voice input in a first language acquired through the microphone,

acquire a first text in a second language by inputting the first token into a first neural network model,

acquire a feature value corresponding to a predicted subsequent token, which is predicted to be uttered after the first token, by inputting the first text into a second neural network model, and

based on a second token being acquired subsequent to the first token, acquire a second text in the second language by inputting the first token, the second token, the first text, and the feature value into the first neural network model,

wherein the first neural network model acquires the second text based on a first feature related to a grammatical relationship between the first token and the second token, a second feature related to a semantic relationship between the first token and the second token and the feature value corresponding to the predicted subsequent token.

2. The electronic apparatus of claim 1 , wherein based on an input token of the first language being inputted into the first neural network model, the first neural network model is trained to acquire text of the second language corresponding to the input token, or to identify an additional input token in addition to the input token.

3. The electronic apparatus of claim 2 , wherein the first neural network model further comprises a first encoder and a first decoder, and

wherein, based on a context vector being acquired by inputting the input token into the first encoder, the first decoder is trained to acquire the text of the second language corresponding to the input token.

4. The electronic apparatus of claim 3 , wherein the first decoder is configured to, based on a probability value acquired based on the context vector being greater than a predetermined value, acquire the text of the second language corresponding to the input token, and based on the probability value acquired based on the context vector being less than the predetermined value, identify an additional token in addition to the input token.

5. The electronic apparatus of claim 1 , further comprising a display,

wherein the processor is further configured to control the display to display the first text and the second text.

6. The electronic apparatus of claim 1 , further comprising:

a speaker,

wherein the processor is further configured to control the speaker to output a voice message corresponding to the first text and the second text.

7. A method for controlling an electronic apparatus comprising:

acquiring a first token corresponding to a first user voice input in a first language;

acquiring a first text in a second language by inputting the first token into a first neural network model;

acquiring a feature value corresponding to a predicted subsequent token, which is predicted to be uttered after the first token, by inputting the first text into a second neural network model; and

based on a second token being acquired subsequent to the first token, acquiring a second text in the second language by inputting the first token, the second token, the first text, and the feature value into the first neural network model,

wherein the first neural network model acquires the second text based on a first feature related to a grammatical relationship between the first token and the second token, a second feature related to a semantic relationship between the first token and the second token and the feature value corresponding to the predicted subsequent token.

8. The method of claim 7 ,

wherein, based on an input token of the first language being inputted into the first neural network model, the first neural network model to acquire text of the second language corresponding to the input token or to identify an additional input token in addition to the input token.

9. The method of claim 8 ,

wherein the first neural network model further comprises a first encoder and a first decoder, and

wherein, based on a context vector being acquired by inputting the input token into the first encoder, the first decoder is configured to acquire the text of the second language corresponding to the input token.

10. The method of claim 9 ,

wherein the first decoder is configured to, based on a probability value acquired based on the context vector being greater than a predetermined value, acquire the text of the second language corresponding to the input token, and based on the probability value acquired based on the context vector being less than the predetermined value, identify an additional token in addition to the input token.

11. The method of claim 7 , further comprising displaying the first text and the second text.

12. The method of claim 7 , further comprising outputting a voice message corresponding to the first text and the second text.

13. A non-transitory computer-readable recording medium storing a program which, when executed by at least one processor, causes the at least one processor to:

acquire a first token corresponding to a first user voice input in a first language;

acquire a first text in a second language by inputting the first token into a first neural network model;

acquire a feature value corresponding to a predicted subsequent token predicted, which is predicted to be uttered after the first token, by inputting the first text into a second neural network model; and

based on a second token being acquired subsequent to the first token, acquire a second text in the second language by inputting the first token, the second token, the first text, and the feature value into the first neural network model,

wherein the first neural network model acquires the second text based on a first feature related to a grammatical relationship between the first token and the second token, a second feature related to a semantic relationship between the first token and the second token and the feature value corresponding to the predicted subsequent token.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2022
From: LEE, BEOMSEOK; INDURTHI, SATHISH; ZAIDI, MOHD ABBAS; KUMAR, NIKHIL
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 059934/0616 →
Priority Claims (2)
KR 10-2021-0063183 · May 17, 2021 · national
KR 10-2021-0180154 · Dec 15, 2021 · national
Continuity (1)
Related Publication 20220366157A1 · Nov 17, 2022
References Cited (23)
US 8527258B2 · Kim et al. · 2013 [cited by applicant]
US 11301625B2 · Roh et al. · 2022 [cited by applicant]
US 20170372693A1 · Rangarajan Sridhar et al. · 2017 [cited by applicant]
US 20200104371A1 · Ma et al. · 2020 [cited by applicant]
US 20200117715A1 · Lee · 2020 [cited by examiner]
KR 100713647B1 · 2007 [cited by applicant]
KR 101589433B1 · 2016 [cited by applicant]
KR 1020180000089A · 2018 [cited by applicant]
KR 1020200059625A · 2020 [cited by applicant]
KR 1020210121818A · 2021 [cited by applicant]
Wang, Mingxuan, et al. “Memory-enhanced decoder for neural machine translation.” arXiv preprint arXiv:1606.02003 (2016). [cited by examiner]
Gulcehre, Caglar, et al. “On integrating a language model into neural machine translation.” Computer Speech & Language 45 (2017): 137-148. [cited by examiner]
Werlen, Lesly Miculicich, et al. “Self-attentive residual decoder for neural machine translation.” arXiv preprint arXiv: 1709.04849 (2017). [cited by examiner]
Alinejad, Ashkan, Maryam Siahbani, and Anoop Sarkar. “Prediction improves simultaneous neural machine translation.” Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018. [cited by examiner]
Dalvi, Fahim, et al. “Incremental decoding and training methods for simultaneous translation in neural machine translation.” arXiv preprint arXiv:1806.03661 (2018). [cited by examiner]
Arivazhagan, Naveen, et al. “Monotonic infinite lookback attention for simultaneous machine translation.” arXiv preprint arXiv:1906.05218 (2019). [cited by examiner]
Schneider, Felix, and Alex Waibel. “Towards stream translation: Adaptive computation time for simultaneous machine translation.” Proceedings of the 17th International Conference on Spoken Language Translation. 2020. [cited by examiner]
Wu, Xueqing, et al. “Learning to Use Future Information in Simultaneous Translation.” (2020). [cited by examiner]
Zhang, Shaolei, Yang Feng, and Liangyou Li. “Future-Guided Incremental Transformer for Simultaneous Translation.” arXiv e-prints (2020): arXiv-2012. [cited by examiner]
Yi Ren et al., “SimulSpeech: End-to-End Simultaneous Speech to Text Translation”, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 2020, 10 pages total. [cited by applicant]
Mingbo Ma et al., “STACL: Simultaneous Translation with Implicit Anticipation and Controllable Latency using Prefix-to-Prefix Framework”, Proceedings of the 57th Annual Meeting of the Association for Computational Lingu… [cited by applicant]
Xutai Ma et al., “SimuIMT to SimuIST: Adapting Simultaneous Text Translation to End-to-End Simultaneous Speech Translation”, Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computati… [cited by applicant]
Xutai Ma et al., “Monotonic Multihead Attention”, arXiv:1909.12406v1 [cs.CL], Sep. 26, 2019, 11 pages total. [cited by applicant]