IP Library › Granted Patent US 12,370,680
Granted Patent B2
US 12,370,680 · App. 18/224,881 · Granted Jul 29, 2025

Robot for acquiring learning data and method for controlling thereof

Inventors: Yongkook Kim (Suwon-si, KR); Saeyoung Kim (Suwon-si, KR); Junghoe Kim (Suwon-si, KR); Hyeontaek Lim (Suwon-si, KR); Boseok Moon (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
B25J9/163B25J11/0005G10L15/22G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,370,680
App. No.
18/224,881
Granted
Jul 29, 2025
Kind
B2
Abstract

A robot transmits a command to control an external device around the robot based on pre-stored environment information while the robot is operating in a learning mode. The external device makes a noise as part of its operation. Also, the robot outputs user speech for learning while the external device is operating. The robot learns a speech recognition model based on the noise and speech of a user acquired through a microphone of the robot. The speech recognition model is then used by the robot or by another device to better understand the user when the user talks. The robot is then able to more accurately understand and properly execute speech commands from the user.

Claims (56)

1. A robot for acquiring learning data, the robot comprising:

a speaker;

a microphone;

a driver;

a communication interface;

a memory storing at least one instruction; and

at least one processor connected to the speaker, the microphone, the driver, the communication interface, and the memory for controlling the robot,

wherein the at least one processor, by executing the at least one instruction, is configured to:

control the communication interface so that the robot transmits a command to an external device around the robot based on pre-stored environment information while the robot is operating in a learning mode,

output first user speech for learning while the external device, responsive to the command, is outputting a noise, and

learn a speech recognition model based on the noise and the first user speech for learning acquired through the microphone.

2. The robot of claim 1 , wherein the at least one processor is further configured to:

acquire second user speech uttered by a user while the robot is operating in a speech recognition mode,

acquire environment information comprising first information about an ambient noise at a time when the second user speech is acquired and second information about the robot, and

store the environment information in the memory.

3. The robot of claim 2 , wherein the pre-stored environment information further comprises third information about a first place where the second user speech is uttered, and

wherein the at least one processor is further configured to determine a device to output the first user speech for learning based on the first place where the second user speech is uttered.

4. The robot of claim 3 , wherein the at least one processor is further configured to:

based on the first place where the second user speech is uttered and a second place where the robot is located when acquiring the second user speech being a same place, determine that the robot will output the first user speech for learning, and

based on the first place where the second user speech is uttered and the second place where the robot is located when acquiring the second user speech being different places, determine that a second external device located in the first place where the second user speech is uttered will output the first user speech for learning.

5. The robot of claim 3 , wherein the at least one processor is further configured to:

generate a text-to-speech (TTS) model based on the second user speech, and

generate the first user speech for learning based on the TTS model.

6. The robot of claim 5 , wherein the at least one processor is further configured to generate the first user speech for learning by inputting, into the TTS model, at least one of a predefined text and a text frequently used by the user.

7. The robot of claim 2 , wherein the environment information comprises first movement information of the robot when acquiring the second user speech, and

wherein the at least one processor is further configured to control the driver to drive the robot based on second movement information of the robot while the external device, responsive to the command, outputs the noise.

8. The robot of claim 1 , wherein the at least one processor is further configured to:

analyze an audio signal to identify a speech period and a non-speech period, wherein the noise and the first user speech for learning acquired through the microphone comprises the audio signal, and

determine a second speech recognition section for learning as having a second start time point and a second end time point based on a first start time point and a first end time point of the speech period of the first user speech for learning as a start time point and an end time point of a first speech recognition section of the robot.

9. The robot of claim 1 , wherein the at least one processor is further configured to operate in the learning mode, based on a preset event being detected, and

wherein the preset event comprises at least one of a first event of entering a time zone set by a user, a second event of entering a time zone at which learning data was acquired in the past, and a third event in which the user is detected as going outside.

10. The robot of claim 1 , wherein the at least one processor is further configured to control the communication interface to transmit the speech recognition model to an external device capable of recognizing speech.

11. A method of controlling a robot for acquiring learning data, the method comprising:

transmitting a command to an external device around the robot based on pre-stored environment information while the robot is operating in a learning mode;

outputting first user speech for learning while the external device, responsive to the command, is outputting a noise; and

learning a speech recognition model based on the noise and the first user speech for learning acquired through a microphone provided in the robot.

12. The method of claim 11 , further comprising:

acquiring second user speech uttered by a user while the robot is operating in a speech recognition mode;

acquiring environment information comprising first information about an ambient noise at a time when the second user speech is acquired and second information about the robot; and

storing the environment information.

13. The method of claim 12 , wherein the pre-stored environment information further comprises third information about a first place where the second user speech is uttered, and

wherein the method further comprises determining a device to output the first user speech for learning based on the first place where the second user speech is uttered.

14. The method of claim 13 , wherein the determining comprises:

based on the first place where the second user speech is uttered and a second place where the robot is located when acquiring the second user speech being a same place, determining that the robot will output the first user speech for learning; and

based on the first place where the second user speech is uttered and the second place where the robot is located when acquiring the second user speech being different places, determining that a second external device located in the first place where the second user speech is uttered will output the first user speech for learning.

15. The method of claim 13 , further comprising:

generating a text-to-speech (TTS) model based on the second user speech; and

generating the first user speech for learning based on the TTS model.

16. The method of claim 11 , further comprising driving the robot while the external device, responsive to the command, outputs a first noise.

17. The method of claim 16 , wherein the external device is an air conditioner.

18. The method of claim 17 , wherein the command specifies strong wind intensity.

19. The method of claim 11 , further comprising the robot issuing a second command to an air purifier, the second command specifying operation of the air purifier so that a second noise is generated.

20. A non-transitory computer readable medium storing a program to execute a control method of a robot, the control method comprising:

transmitting a command to an external device around the robot based on pre-stored environment information while the robot is operating in a learning mode;

outputting first user speech for learning while the external device, responsive to the command, is outputting noise; and

learning a speech recognition model based on the noise and the first user speech for learning acquired through a microphone provided in the robot.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 21, 2023
From: KIM, YONGKOOK; KIM, SAEYOUNG; KIM, JUNGHOE; LIM, HYEONTAEK; MOON, BOSEOK
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 064362/0259 →
Priority Claims (1)
KR 10-2022-0088584 · Jul 18, 2022 · national
Continuity (2)
Continuation PCTKR2023005199 · Apr 18, 2023
Related Publication 20240017406A1 · Jan 18, 2024
References Cited (49)
US 6842734B2 · Yamada et al. · 2005 [cited by applicant]
US 8185399B2 · Di Fabbrizio et al. · 2012 [cited by applicant]
US 9275638B2 · Meloney · 2016 [cited by examiner]
US 9302393B1 · Rosen · 2016 [cited by examiner]
US 9697822B1 · Naik · 2017 [cited by examiner]
US 10127905B2 · Lee et al. · 2018 [cited by applicant]
US 11037548B2 · Kim · 2021 [cited by applicant]
US 11055356B2 · Ritchey · 2021 [cited by examiner]
US 11164586B2 · Kim · 2021 [cited by examiner]
US 11250843B2 · Yun · 2022 [cited by applicant]
US 11282522B2 · Chae · 2022 [cited by examiner]
US 11443747B2 · Chae · 2022 [cited by examiner]
US 11457983B1 · Roh · 2022 [cited by examiner]
US 11501794B1 · Kim · 2022 [cited by examiner]
US 11551662B2 · Yun et al. · 2023 [cited by applicant]
US 11583998B2 · Moon · 2023 [cited by examiner]
US 20020158599A1 · Fujita · 2002 [cited by examiner]
US 20030097261A1 · Jeon et al. · 2003 [cited by applicant]
US 20160199977A1 · Breazeal · 2016 [cited by examiner]
US 20170076719A1 · Lee · 2017 [cited by examiner]
US 20190348021A1 · Trim · 2019 [cited by examiner]
US 20190385600A1 · Kim · 2019 [cited by examiner]
US 20190385614A1 · Kim · 2019 [cited by examiner]
US 20190392818A1 · Lee · 2019 [cited by applicant]
US 20200005766A1 · Kim · 2020 [cited by examiner]
US 20200258535A1 · Vatanparvar · 2020 [cited by examiner]
US 20200368616A1 · Delamont · 2020 [cited by examiner]
US 20200374269A1 · Lidman · 2020 [cited by examiner]
US 20210016431A1 · Kim · 2021 [cited by examiner]
US 20210043186A1 · Nagano · 2021 [cited by examiner]
US 20210043204A1 · Hwang · 2021 [cited by examiner]
US 20220165253A1 · Sharifi et al. · 2022 [cited by applicant]
US 20220375469A1 · Yang · 2022 [cited by examiner]
JP 2001215990A · 2001 [cited by applicant]
KR 100251003B1 · 2000 [cited by applicant]
KR 1020130123945A · 2013 [cited by applicant]
KR 1020180127100A · 2018 [cited by applicant]
KR 1020190096856A · 2019 [cited by applicant]
KR 1020190104490A · 2019 [cited by applicant]
KR 1020200046262A · 2020 [cited by applicant]
KR 102209689B1 · 2021 [cited by applicant]
KR 102213177B1 · 2021 [cited by applicant]
KR 1020210089347A · 2021 [cited by applicant]
KR 102321798B1 · 2021 [cited by applicant]
Search Report (PCT/ISA/210) dated Jul. 17, 2023 issued by the ISA for International Application No. PCT/KR2023/005199. [cited by applicant]
Written Opinion (PCT/ISA/237) dated Jul. 17, 2023 issued by the ISA for International Application No. PCT/KR2023/005199. [cited by applicant]
José Novoa et al., “DNN-HMM based Automatic Speech Recognition for HRI Scenarios”, HRI '18: Proceedings of the 2018 ACM/IEEE International Conference on Human-Robot Interaction, Feb. 2018, pp. 150-159, DOI: 10.1145/3171… [cited by applicant]
José Novoa et al., “Automatic Speech Recognition for Indoor HRI Scenarios”, ACM Transactions on Human-Robot Interaction (THRI), Mar. 2021, vol. 10, No. 2, Article No. 17, pp. 1-30, DOI: 10.1145/3442629, XP058493826. [cited by applicant]
Communication issued on Jun. 12, 2025 from the European Patent Office for European Patent Application No. 23843125.8. [cited by applicant]