IP Library › Granted Patent US 12,488,800
Granted Patent B2
US 12,488,800 · App. 18/112,736 · Granted Dec 2, 2025

Display apparatus and operating method thereof

Inventors: Seokjae Oh (Suwon-si, KR); Yeseul Park (Suwon-si, KR); Yuseong Jeon (Suwon-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G10L15/32G06F3/14G10L15/063G10L15/22G10L2015/228H04N21/42203H04N21/475H04N21/4788
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,800
App. No.
18/112,736
Granted
Dec 2, 2025
Kind
B2
Abstract

A method of operating a display apparatus includes: obtaining situation information for voice recognizer selection, selecting at least one of a plurality of voice recognizers based on the situation information, obtaining a voice recognition result from a voice signal, using the selected at least one voice recognizer, and obtaining a chat message from the voice recognition result.

Claims (56)

1 . A display apparatus comprising:

a memory including one or more instructions; and

a processor, comprising processing circuitry, configured to execute the one or more instructions stored in the memory to:

obtain situation information for voice recognizer selection,

select at least one of a plurality of voice recognizers based on the situation information,

obtain a voice recognition result from a voice signal using the selected at least one voice recognizer, and

obtain a chat message from the voice recognition result,

wherein each of the plurality of voice recognizers comprises a learning model configured to be trained with one or more different training data.

2 . The display apparatus of claim 1 , further comprising a display,

wherein the processor is further configured to execute the one or more instructions to control the display to display content and chat messages of a chat room related to the content,

wherein the situation information comprises at least one of content information about the content or chat information related to chatting.

3 . The display apparatus of claim 2 , wherein the chat information comprises information about at least one of a title of the chat room or content of the chat messages, and

the content information comprises at least one of subject of the content, a voice signal output together with the content, subtitles, a program name of the content, a content topic, a content type, a content genre, a channel type, a broadcasting station, a producer, a cast, a director, or a content broadcast time.

4 . The display apparatus of claim 1 ,

wherein the different training data comprise at least one of training data by language, training data by field, training data by program type, training data by program genre, training data by broadcasting station, training data by channel, training data by producer, training data by cast, training data by director, training data by region, personalized training data obtained based on user information, or training data obtained based on information about a group to which the user belongs.

5 . The display apparatus of claim 4 , wherein the user information comprises at least one of user profile information, viewing history information, or chat message content information input by the user, and

the information about the group to which the user belongs comprises at least one of profile information of people whose user information overlaps the user by a reference value or more, viewing history information of the people, or chat message content information input by the people.

6 . The display apparatus of claim 1 , wherein the plurality of voice recognizers are identified by label information indicating a type of training data used to train the learning model,

wherein the processor is further configured to execute the one or more instructions to select at least one of the plurality of voice recognizers based on a similarity between the situation information and the label information.

7 . The display apparatus of claim 1 , wherein the processor is further configured to, based on the selected voice recognizers being plural, obtain a plurality of voice recognition results from the voice signal using the selected plurality of voice recognizers.

8 . The display apparatus of claim 7 , further comprising a display,

wherein the processor is further configured to execute the one or more instructions to:

filter a specified number of or fewer voice recognition results, based on a weight matrix from among the plurality of voice recognition results,

obtain chat messages corresponding to the filtered voice recognition results, and

output the chat messages through the display.

9 . The display apparatus of claim 8 , wherein the processor is further configured to execute the one or more instructions to:

based on a plurality of chat messages being output through the display, transmit one chat message selected from among the plurality of chat messages to a chat server.

10 . The display apparatus of claim 9 , wherein the processor is further configured to execute the one or more instructions to update the weight matrix based on the selection.

11 . A method performed by a display apparatus comprising a processor comprising processing circuitry, and memory storing instructions, the method for of operating a display apparatus, the method and comprising:

obtaining, by the processing circuitry, situation information for voice recognizer selection;

selecting, by the processing circuitry, at least one of a plurality of voice recognizers based on the situation information;

obtaining, by the processing circuitry, a voice recognition result from a voice signal, using the selected at least one voice recognizer; and

obtaining a chat message from the voice recognition result, wherein each of the plurality of voice recognizers comprises a learning model trained with one or more different training data.

12 . The method of claim 11 , further comprising displaying content and chat messages of a chat room related to the content,

wherein the situation information comprises at least one of content information about the content or chat information related to chatting.

13 . The method of claim 12 , wherein the chat information comprises at least one of title information of the chat room or content information of the chat messages, and

the content information comprises at least one of subject of the content, a voice signal output together with the content, subtitles, a program name of the content, a content topic, a content type, a content genre, a channel type, a broadcasting station, a producer, a cast, a director, or a content broadcast time.

14 . The method of claim 11 ,

wherein the different training data comprise at least one of training data by language, training data by field, training data by program type, training data by program genre, training data by broadcasting station, training data by channel, training data by producer, training data by cast, training data by director, training data by region, personalized training data obtained based on user information, or training data obtained based on information about a group to which the user belongs.

15 . The method of claim 14 , wherein the user information comprises at least one of user profile information, viewing history information of the user, or chat message content information input by the user, and

the information about the group to which the user belongs comprises at least one of profile information of people whose user information overlaps the user by a reference value or more, viewing history information of the people, or chat message content information input by the people.

16 . The method of claim 11 , wherein the plurality of voice recognizers are identified by label information indicating a type of training data used to train the learning model,

wherein the selecting of the at least one of the plurality of voice recognizers comprises selecting at least one of the plurality of voice recognizers, based on a similarity between the situation information and the label information.

17 . The method of claim 11 , wherein the obtaining of the voice recognition result comprises,

based on the selected voice recognizers being plural, obtaining a plurality of voice recognition results from the user's voice signal using the plurality of selected voice recognizers.

18 . The method of claim 17 , wherein the obtaining of the chat message comprises:

filtering a specified number of or fewer voice recognition results based on a weight matrix from among the plurality of voice recognition results; and

obtaining chat messages corresponding to the filtered voice recognition results,

wherein the method further comprises outputting the chat messages.

19 . The method of claim 18 , further comprising, based on a plurality of chat messages being output, transmitting one chat message selected from among the plurality of chat messages to a chat server.

20 . A non-transitory computer-readable recording medium having recorded thereon a program which when executed by a processor of a display apparatus, causes the display to perform operations comprising:

obtaining situation information for voice recognizer selection;

selecting at least one of a plurality of voice recognizers based on the situation information;

obtaining a voice recognition result from a voice signal, using the selected at least one voice recognizer; and

obtaining a chat message from the voice recognition result,

wherein each of the plurality of voice recognizers comprises a learning model configured to be trained with one or more different training data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 22, 2023
From: OH, SEOKJAE; PARK, YESEUL; JEON, YUSEONG
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 062768/0432 →
Priority Claims (1)
KR 10-2022-0023209 · Feb 22, 2022 · national
Continuity (2)
Continuation PCTKR2023001837 · Feb 8, 2023
Related Publication 20230267934A1 · Aug 24, 2023
References Cited (64)
US 8862467B1 · Casado · 2014 [cited by examiner]
US 9070366B1 · Mathias et al. · 2015 [cited by applicant]
US 9570070B2 · Baldwin et al. · 2017 [cited by applicant]
US 10217455B2 · Cho · 2019 [cited by examiner]
US 10629196B2 · Park · 2020 [cited by examiner]
US 10950228B1 · Tan · 2021 [cited by examiner]
US 11068667B2 · Choi et al. · 2021 [cited by applicant]
US 11189282B2 · Jeong · 2021 [cited by examiner]
US 11218592B2 · Kim · 2022 [cited by examiner]
US 11289074B2 · Lee · 2022 [cited by applicant]
US 11587571B2 · Choi et al. · 2023 [cited by applicant]
US 12027168B2 · Cho · 2024 [cited by examiner]
US 20050159833A1 · Giaimo · 2005 [cited by examiner]
US 20090030687A1 · Cerra · 2009 [cited by examiner]
US 20090030688A1 · Cerra · 2009 [cited by examiner]
US 20090204409A1 · Mozer · 2009 [cited by examiner]
US 20090320076A1 · Chang · 2009 [cited by examiner]
US 20100145694A1 · Ju · 2010 [cited by examiner]
US 20110054895A1 · Phillips · 2011 [cited by examiner]
US 20110066634A1 · Phillips · 2011 [cited by examiner]
US 20130179168A1 · Bae · 2013 [cited by examiner]
US 20150039318A1 · Shin · 2015 [cited by examiner]
US 20150073801A1 · Shin · 2015 [cited by examiner]
US 20150088524A1 · Shin · 2015 [cited by examiner]
US 20150189362A1 · Lee · 2015 [cited by examiner]
US 20150189390A1 · Sirpal · 2015 [cited by examiner]
US 20150228279A1 · Biadsy · 2015 [cited by examiner]
US 20150279363A1 · Furumoto · 2015 [cited by examiner]
US 20160027440A1 · Gelfenbeyn · 2016 [cited by examiner]
US 20160088333A1 · Bhatia · 2016 [cited by examiner]
US 20160188150A1 · Abida · 2016 [cited by examiner]
US 20170300831A1 · Gelfenbeyn et al. · 2017 [cited by applicant]
US 20180053502A1 · Biadsy · 2018 [cited by examiner]
US 20180166076A1 · Higuchi · 2018 [cited by examiner]
US 20180182383A1 · Kim · 2018 [cited by examiner]
US 20180261220A1 · Higbie · 2018 [cited by examiner]
US 20180374476A1 · Lee et al. · 2018 [cited by applicant]
US 20190130901A1 · Kato · 2019 [cited by examiner]
US 20190237085A1 · Ryu · 2019 [cited by examiner]
US 20190279638A1 · Kwon · 2019 [cited by examiner]
US 20200074990A1 · Kim et al. · 2020 [cited by applicant]
US 20200099785A1 · Kim · 2020 [cited by examiner]
US 20200184989A1 · Jang · 2020 [cited by examiner]
US 20200302935A1 · Choi · 2020 [cited by examiner]
US 20210043204A1 · Wang et al. · 2021 [cited by applicant]
US 20210065718A1 · Choi · 2021 [cited by applicant]
US 20210152870A1 · Lee · 2021 [cited by examiner]
US 20210217407A1 · Mohapatra · 2021 [cited by examiner]
US 20210342555A1 · Choi et al. · 2021 [cited by applicant]
US 20230267934A1 · Oh · 2023 [cited by examiner]
US 20230350707A1 · Kim · 2023 [cited by examiner]
US 20240267580A1 · Lee · 2024 [cited by examiner]
KR 1020180067977A · 2018 [cited by applicant]
KR 1020180108400A · 2018 [cited by applicant]
KR 1020190001434 · 2019 [cited by applicant]
KR 101949731 · 2019 [cited by applicant]
KR 1020200092166 · 2020 [cited by applicant]
KR 1020210017392 · 2021 [cited by applicant]
KR 1020210027991A · 2021 [cited by applicant]
KR 102225984 · 2021 [cited by applicant]
KR 1020210039049A · 2021 [cited by applicant]
International Search Report dated May 9, 2023 issued in International Application No. PCT/KR2023/001837 (9 pages). [cited by applicant]
Miranda et al., “How to Build Domain Specific Automatic Speech Recognition Models on GPUs” (Dec. 2019), https://developer.nvidia.com/blog/how-to-build-domain-specific-automatic-speech-recognition-models-on-gpus/, 4 page… [cited by applicant]
Taubenheim et al., “GPU-Accelerated Speech to Text with Kaldi: A Tutorial on Getting Started” (Oct. 2019), https://developer.nvidia.com/blog/gpu-accelerated-speech-to-text-with-kaldi-a-tutorial-on-getting-started/, 5 pa… [cited by applicant]