IP Library Granted Patent US 12,308,016
Granted Patent B2
US 12,308,016 · App. 17/666,900 · Granted May 20, 2025

Electronic device including speaker and microphone and method for operating the same

Inventors: Hoseon Shin (Gyeonggi-do, KR); Chulmin Lee (Gyeonggi-do, KR); Taegu Kim (Gyeonggi-do, KR); Kyounggu Woo (Gyeonggi-do, KR); Youngwoo Lee (Gyeonggi-do, KR)
Assignee: Samsung Electronics Co., Ltd
G10L15/08G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,308,016
App. No.
17/666,900
Granted
May 20, 2025
Kind
B2
Abstract

According to various embodiments, an electronic device is provided. The electronic device includes a communication circuit, a plurality of microphones, a speaker, and at least one processor. The at least one processor is configured to output audio through the speaker based on data received from an external electronic device through the communication circuit, identify an utterance including a specified keyword received through at least some of the plurality of microphones, based on identifying the utterance including the specified keyword, decrease the volume of the audio output through the speaker and perform an operation for providing a speech of a user of the electronic device and a speech of a person other than the user of the electronic device based on at least some of ambient sounds received through at least some of the plurality of microphones.

Claims (78)

1. An electronic device comprising:

a communication circuit;

a plurality of microphones;

at least one sensor configured to detect movement of the electronic device;

a speaker;

at least one processor; and

memory storing instructions,

wherein the instructions that, when executed by the at least one processor individually or collectively, cause the electronic device to:

output audio through the speaker based on data received from an external electronic device through the communication circuit;

identify an utterance including a specified keyword received through at least one of the plurality of microphones;

identify whether a user of the electronic device has turned their head based on a pattern of values received by the at least one sensor; and

based on identifying the utterance including the specified keyword and the user turning their head, decrease a volume of the output audio through the speaker and perform an operation for providing a speech of the user of the electronic device and a speech of a conversant with the user of the electronic device based on at least some of ambient sounds received through at least some of the plurality of microphones, from the user's preset mouth direction, forward direction, and a direction corresponding to an angle at which the head is turned,

wherein providing the speech of the user and the speech of the conversant with the user further comprises generating a new speech model by learning with a speaker embedding for the user, a speaker embedding for the conversant with the user, and at least one other speaker embedding for at least one other speaker contemporaneously speaking, the new speech model providing unique identifications for the conversant with the user, and each of the at least one other speaker.

2. The electronic device of claim 1 ,

wherein the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to:

store a plurality of keywords from the external electronic device or sounds corresponding to the plurality of keywords in the memory;

receive the utterance through the at least one of the plurality of microphones during outputting of the audio; and

identify whether the utterance includes the specified keyword based on the received utterance and at least one of the plurality of keywords or the sounds corresponding to the plurality of keywords.

3. The electronic device of claim 2 , wherein the plurality of keywords includes at least one first keyword and at least one second keyword, and

wherein the at least one first keyword is a name of the user, and the at least one second keyword is generated based on the at least one first keyword.

4. The electronic device of claim 1 , wherein the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to:

identify whether the utterance has been spoken by the user, based on identifying the utterance including the specified keyword, and

based on identifying that the utterance has been spoken by the user, decrease the volume of the audio output through the speaker and perform the operation for providing the speech of the user of the electronic device and the speech of the conversant with the user.

5. The electronic device of claim 4 , further comprising a sensor,

wherein the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to:

identify at least one specified value by using the sensor based on identifying the utterance including the specified keyword, the at least one specified value indicating that the utterance has been spoken by the user; and

identify that the utterance has been spoken by the user, based on identifying the at least one specified value.

6. The electronic device of claim 4 , wherein the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to:

generate first feature information about at least one speech of the user, before identifying the utterance including the specified keyword;

compare feature information about the utterance received through at least one of the plurality of microphones with the first feature information, based on identifying the utterance including the specified keyword; and

identify that the utterance has been spoken by the user based on a result of the comparison.

7. The electronic device of claim 1 , wherein the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to:

receive the speech of the conversant with the user through at least one of the plurality of microphones, based on identifying the utterance including the specified keyword;

when the speech of the conversant with the user has been received for a specified time period, generate a speaker embedding for the conversant with the user based on the speech of the conversant with the user received for the specified time period; and

obtain the speech of the user and the speech of the conversant with the user based on an obtained at least one speech model, and output the speech of the user and the speech of the conversant with the user.

8. The electronic device of claim 7 ,

wherein the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to:

compare at least one speaker embedding for at least one person pre-stored in the memory with the speaker embedding for the conversant with the user; and

when the conversant with the user is identified as different from the at least one person based on a result of the comparison, obtain the at least one speech model.

9. The electronic device of claim 8 , wherein the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to generate the at least one speech model by performing training by using the speaker embedding for the user and the speaker embedding for the conversant with the user as input data and using a first identifier corresponding to the user and a second identifier corresponding to the conversant with the user as output data, and

wherein the at least one speech model is configured to output the first identifier or the second identifier in response to input of the speaker embedding for the user or the speaker embedding for the conversant with the user.

10. The electronic device of claim 9 , wherein the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to:

receive the ambient sounds through at least one of the plurality of microphones after the at least one speech model is generated;

separate a plurality of speeches from the received ambient sounds, the plurality of speeches corresponding to a plurality of persons;

generate speaker embeddings for the plurality of persons based on the plurality of speeches;

obtain a plurality of identifiers corresponding to the plurality of speeches based on the generated speaker embeddings being input to the at least one speech model; and

obtain a speech of the user related to the first identifier and a speech of the conversant with the user related to the second identifier among the plurality of speeches.

11. The electronic device of claim 10 , wherein the at least one processor is configured to post-process the obtained speeches of the user and the conversant with the user and output the post-processed speeches of the user and the conversant with the user through the speaker.

12. The electronic device of claim 10 , wherein the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to:

further identify a third identifier of a person selected by the user;

further obtain a third speech corresponding to the third identifier among the plurality of speeches; and

output the speech of the user, the speech of the conversant with the user, and the third speech through the speaker.

13. The electronic device of claim 1 , wherein the instructions, when executed by the at least one processor individually or collectively, cause the electronic device to:

based on identifying the utterance including the specified keyword, stop output of anti-noise generated by an active noise cancellation circuit of the electronic device through the speaker.

14. A method of operating an electronic device, the method comprising:

outputting audio through a speaker based on data received from an external electronic device through a communication circuit of the electronic device;

identifying an utterance including a specified keyword received through at least some of a plurality of microphones of the electronic device;

identifying whether a user of the electronic device has turned their head based on a pattern of values received by at least one sensor; and

based on identifying the utterance including the specified keyword and the user turning their head, decreasing volume of the audio output through the speaker and performing an operation for providing a speech of the user of the electronic device and a speech of a conversant with the user of the electronic device based on at least some of ambient sounds received through at least one of the plurality of microphones, from the user's preset mouth direction, forward direction, and a direction corresponding to an angle at which the head is turned,

wherein providing the speech of the user and the speech of the conversant with the user further comprises generating a new speech model by learning with a speaker embedding for the user, a speaker embedding for the conversant with the user, and at least one other speaker embedding for at least one other speaker contemporaneously speaking, the new speech model providing unique identifications for the conversant with the user, and each of the at least one other speaker.

15. The method of claim 14 , further comprising:

storing a plurality of keywords from the external electronic device or sounds corresponding to the plurality of keywords in a memory of the electronic device;

receiving the utterance through the at least one of the plurality of microphones during outputting of the audio; and

identifying whether the utterance includes the specified keyword based on the received utterance and at least one of the plurality of keywords or the sounds corresponding to the plurality of keywords.

16. The method of claim 15 , wherein the plurality of keywords include at least one first keyword and at least one second keyword, and

wherein the at least one first keyword is a name of the user, and the at least one second keyword is generated based on the at least one first keyword.

17. The method of claim 14 , further comprising:

identifying whether the utterance has been spoken by the user, based on identifying the utterance including the specified keyword; and

based on identifying that the utterance has been spoken by the user, decreasing the volume of the audio output through the speaker and perform the operation for providing the speech of the user of the electronic device and the speech of the conversant with the user.

18. The method of claim 17 , further comprising:

identifying at least one specified value by using a sensor of the electronic device based on identifying the utterance including the specified keyword, the at least one specified value indicating that the utterance has been spoken by the user; and

identifying that the utterance has been spoken by the user, based on identifying the at least one specified value.

19. A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by at least one processor of an electronic device, cause the electronic device to perform operations, the operations comprising:

outputting audio through a speaker based on data received from an external electronic device through a communication circuit of the electronic device;

identifying an utterance including a specified keyword received through at least some of a plurality of microphones of the electronic device;

identifying whether a user of the electronic device has turned their head based on a pattern of values received by at least one sensor; and

based on identifying the utterance including the specified keyword and the user turning their head, decreasing volume of the audio output through the speaker and performing an operation for providing a speech of the user of the electronic device and a speech of a conversant with the user of the electronic device based on at least some of ambient sounds received through at least one of the plurality of microphones, from the user's preset mouth direction, forward direction, and a direction corresponding to an angle at which the head is turned,

wherein providing the speech of the user and the speech of the conversant with the user further comprises generating a new speech model by learning with a speaker embedding for the user, a speaker embedding for the conversant with the user, and at least one other speaker embedding for at least one other speaker contemporaneously speaking, the new speech model providing unique identifications for the conversant with the user, and each of the at least one other speaker.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 8, 2022
From: SHIN, HOSEON; LEE, CHULMIN; KIM, TAEGU; WOO, KYOUNGGU; LEE, YOUNGWOO
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 058927/0047 →
Priority Claims (1)
KR 10-2021-0021842 · Feb 18, 2021 · national
Continuity (2)
Continuation PCTKR2022001151 · Jan 21, 2022
Related Publication 20220261218A1 · Aug 18, 2022
References Cited (30)
US 10681453B1 · Meiyappan · 2020 [cited by examiner]
US 11412133B1 · Morton · 2022 [cited by examiner]
US 12020704B2 · Nguyen · 2024 [cited by examiner]
US 20020141599A1 · Trajkovic · 2002 [cited by examiner]
US 20160173049A1 · Mehta · 2016 [cited by applicant]
US 20170071822A1 · Tran et al. · 2017 [cited by applicant]
US 20170193978A1 · Goldman · 2017 [cited by applicant]
US 20170195787A1 · Ichimura · 2017 [cited by applicant]
US 20170215011A1 · Goldstein · 2017 [cited by examiner]
US 20190124436A1 · Usher · 2019 [cited by examiner]
US 20190228778A1 · Lesso · 2019 [cited by examiner]
US 20190341057A1 · Zhang · 2019 [cited by examiner]
US 20210014610A1 · Carrigan et al. · 2021 [cited by applicant]
US 20210043191A1 · Chao · 2021 [cited by examiner]
US 20210099787A1 · Yang · 2021 [cited by examiner]
US 20210168516A1 · Wexler · 2021 [cited by examiner]
US 20210183358A1 · Mao · 2021 [cited by examiner]
US 20210407532A1 · Salahuddin et al. · 2021 [cited by applicant]
US 20220014839A1 · Tartz · 2022 [cited by examiner]
US 20220139388A1 · Sharifi · 2022 [cited by examiner]
US 20220261218A1 · Shin · 2022 [cited by examiner]
CN 103414982A · 2013 [cited by applicant]
JP 2005192004A · 2005 [cited by applicant]
JP 2009300915A · 2009 [cited by applicant]
JP 201651915A · 2016 [cited by applicant]
KR 1020030009504A · 2003 [cited by applicant]
KR 101634133B1 · 2016 [cited by applicant]
KR 1020200113058A · 2020 [cited by applicant]
KR 1020200145219A · 2020 [cited by applicant]
International Search Report dated Apr. 26, 2022. [cited by applicant]