IP Library Granted Patent US 12,536,992
Granted Patent B2
US 12,536,992 · App. 17/575,195 · Granted Jan 27, 2026

Electronic device and method for providing voice recognition service

Inventors: Hojung Lee (Suwon-si, KR); Hyungtak Choi (Suwon-si, KR); Munjo Kim (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L15/08G06F3/04842G06F3/167G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,992
App. No.
17/575,195
Granted
Jan 27, 2026
Kind
B2
Abstract

A method and an electronic device for providing a voice recognition service are provided. The method includes: receiving a voice command of a user while one or more objects are displayed on a screen of the electronic device; based on the receiving the voice command, identifying the one or more objects displayed on the screen; interpreting text converted from the voice command, based on types of the one or more objects; and performing an operation related to an object selected from among the one or more objects, based on a result of interpreting the text, wherein the types of the one or more objects are identified based on whether the one or more objects are selectable by a user input to the electronic device.

Claims (72)

1 . A method, performed by an electronic device, of providing a voice recognition service, the method comprising:

receiving a voice command of a user while one or more objects are displayed on a screen of the electronic device;

based on the receiving the voice command, identifying the one or more objects displayed on the screen;

obtaining text from the voice command;

determining, from the text obtained from the voice command, a type of an utterance intent of the user as actionable or accessible, or not actionable or accessible;

identifying types of the one or more objects as text or interpretable, and as selectable or non-selectable by a user input to the electronic device;

based on the type of the utterance intent being actionable or accessible, interpreting the text obtained from the voice command in a first mode by performing natural language understanding with respect to the text based on a data structure generated based on the types of the one or more objects displayed on the screen, and by:

based on the type of the utterance intent being actionable, preferentially referring to the one or more objects that are selectable; and

based on the type of the utterance intent being accessible, preferentially referring to the one or more objects that are text;

selecting a selectable object from the one or more objects displayed on the screen, based on a result of the interpreting of the text;

performing an operation related to a selected object;

receiving a second voice command of the user while the one or more objects are displayed on the screen of the electronic device;

based on the receiving the second voice command, identifying the one or more objects displayed on the screen;

obtaining text from the second voice command;

determining, from the text obtained from the second voice command, the type of the utterance intent of the user as actionable or accessible, or not actionable or accessible;

identifying types of the one or more objects as text or interpretable, and as selectable or non-selectable by a user input to the electronic device;

based on the type of the utterance intent being not actionable or not accessible, interpreting the text obtained from the second voice command in a second mode by performing the natural language understanding on the text without using the data structure generated based on the types of the one or more objects; and

generating a response message in a form of at least one of voice, text, or video, to the result of interpreting the text from the second voice command.

2 . The method of claim 1 , further comprising:

generating the data structure by:

identifying the types of the one or more objects through character recognition based on image processing with respect to the one or more objects or through meta data reading with respect to an application that provides the one or more objects;

determining priorities of the one or more objects, based on the types of the one or more objects; and

generating the data structure in a tree form indicating a relationship between the one or more objects and terms related to the one or more objects.

3 . The method of claim 2 , wherein the terms related to the one or more objects are obtained from at least one of text information included in the one or more objects or attribute information of the one or more objects included in the meta data of the application that provides the one or more objects.

4 . The method of claim 2 , wherein the one or more objects comprise at least one of image information or text information,

the image information is displayed on an object layer of the screen,

the text information is displayed on a text layer of the screen, and

the types of the one or more objects are identified based on the object layer and the text layer.

5 . The method of claim 2 , wherein the identifying the types of the one or more objects further comprises identifying whether the one or more objects include text information.

6 . The method of claim 1 , further comprising:

generating the data structure,

wherein the identifying the one or more objects displayed on the screen comprises:

determining priorities of the one or more objects of the data structure, based on the type of the utterance intent and the types of the one or more objects.

7 . The method of claim 1 , wherein the determining whether to operate in the first mode or in the second mode is also based on whether an activation word is included in the text.

8 . The method of claim 1 , wherein the result of the interpreting the text comprises information about at least one of the utterance intent of the user, the selected object, or a function to be executed by the electronic device in relation to the selected object.

9 . The method of claim 1 , wherein the operation related to the selected object comprises at least one of playing a video related to the selected object, enlarging and displaying an image or text related to the selected object, or outputting an audio based on the text included in the selected object.

10 . The method of claim 1 , wherein the performing the operation related to the selected object comprises generating and outputting the response message related to the selected object, based on the voice command of the user.

11 . An electronic device for providing a voice recognition service, the electronic device comprising:

a display;

a microphone;

a memory storing one or more instructions; and

at least one processor configured to execute one or more instructions stored in the memory to:

receive a voice command of a user via the microphone while one or more objects are displayed on a screen of the display;

based on receipt of the voice command, identify the one or more objects displayed on the screen;

obtain text from the voice command;

determine, from the text obtained from the voice command, a type of an utterance intent of the user as (i) actionable or accessible or (ii) not actionable or not accessible;

identify types of the one or more objects as text or interpretable, and as selectable or non-selectable by a user input to the electronic device;

based on the type of the utterance intent being actionable or accessible, interpret the text obtained from the voice command in a first mode by performing natural language understanding with respect to the text based on a data structure generated based on the types of the one or more objects displayed on the screen, and by:

based on the type of the utterance intent being actionable, preferentially referring to the one or more objects that are selectable; and

based on the type of the utterance intent being accessible, preferentially referring to the one or more objects that are text;

based on the type of the utterance intent being not actionable or not accessible, interpret the text obtained from the voice command in a second mode by performing the natural language understanding on the text without using the data structure generated based on the types of the one or more objects;

select a selectable object from the one or more objects displayed on the screen, based on a result of the interpreting of the text; and

perform an operation related to a selected object.

12 . The electronic device of claim 11 , wherein the at least one processor is further configured to execute the one or more instructions to:

generate the data structure;

identify the types of the one or more objects through character recognition based on image processing with respect to the one or more objects or through meta data reading with respect to an application that provides the one or more objects; and

determine priorities of the one or more objects, based on the types of the one or more objects, to generate the data structure in a tree form indicating a relationship between the one or more objects and terms related to the one or more objects.

13 . A server for providing a voice recognition service through an electronic device comprising a display, the server comprising:

a communication interface configured to communicate with the electronic device;

a memory storing one or more instructions; and

at least one processor configured to execute one or more instructions stored in the memory to:

receive information about a voice command of a user received through the electronic device while one or more objects are displayed on a screen of the electronic device;

identify the one or more objects displayed on the screen, based on receipt of the voice command;

obtain text from the voice command;

determine, from the text obtained from the voice command, a type of an utterance intent of the user as actionable or accessible;

identify types of the one or more objects as text or interpretable, and as selectable or non-selectable by a user input to the electronic device;

based on the type of the utterance intent being actionable or accessible, interpret the text obtained from the voice command in a first mode by performing natural language understanding with respect to the text based on a data structure generated based on the types of the one or more objects displayed on the screen, and by:

based on the type of the utterance intent being actionable, preferentially referring to the one or more objects that are selectable; and

based on the type of the utterance intent being accessible, preferentially referring to the one or more objects that are text;

based on the type of the utterance intent being not actionable or not accessible, interpret the text obtained from the voice command in a second mode by performing the natural language understanding on the text without using the data structure generated based on the types of the one or more objects;

select a selectable object from the one or more objects displayed on the screen, based on a result of the interpreting of the text; and

control the electronic device to perform an operation related to a selected object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2022
From: LEE, HOJUNG; CHOI, HYUNGTAK; KIM, MUNJO
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 058649/0632 →
Priority Claims (1)
KR 10-2021-0034865 · Mar 17, 2021 · national
Continuity (2)
Continuation PCTKR2021016882 · Nov 17, 2021
Related Publication 20220301549A1 · Sep 22, 2022
References Cited (35)
US 8126715B2 · Paek · 2012 [cited by applicant]
US 9129011B2 · Yang · 2015 [cited by examiner]
US 9171542B2 · Gandrabur et al. · 2015 [cited by applicant]
US 9761225B2 · Gandrabur et al. · 2017 [cited by applicant]
US 9927949B2 · Gray et al. · 2018 [cited by applicant]
US 10068574B2 · Zhang · 2018 [cited by examiner]
US 10331312B2 · Napolitano · 2019 [cited by examiner]
US 10600413B2 · Zhang et al. · 2020 [cited by applicant]
US 10757148B2 · Nelson · 2020 [cited by examiner]
US 10884701B2 · Thangarathnam · 2021 [cited by examiner]
US 11355098B1 · Zhong · 2022 [cited by examiner]
US 11392688B2 · Lewis · 2022 [cited by examiner]
US 11429344B1 · Fernandez · 2022 [cited by examiner]
US 11720324B2 · Kim · 2023 [cited by examiner]
US 11726806B2 · Kim · 2023 [cited by examiner]
US 20140270258A1 · Wang · 2014 [cited by examiner]
US 20150317979A1 · Yang et al. · 2015 [cited by applicant]
US 20150348551A1 · Gruber · 2015 [cited by examiner]
US 20200260127A1 · Chung et al. · 2020 [cited by applicant]
US 20210118463A1 · Chung et al. · 2021 [cited by applicant]
US 20210280195A1 · Srinivasan · 2021 [cited by examiner]
US 20210311701A1 · Cubukcu · 2021 [cited by examiner]
US 20210383794A1 · Kim · 2021 [cited by examiner]
US 20220254333A1 · Aliev · 2022 [cited by examiner]
CN 109545223A · 2019 [cited by applicant]
JP 2016519377A · 2016 [cited by applicant]
KR 1020150072608A · 2015 [cited by applicant]
KR 1020150072625A · 2015 [cited by applicant]
KR 1020150125464A · 2015 [cited by applicant]
KR 101660269B1 · 2016 [cited by applicant]
KR 1020180061691A · 2018 [cited by applicant]
KR 102009316A · 2019 [cited by applicant]
International Search Report and Written Opinion (PCT/ISA/210, PCT/ISA/220, and PCT/ISA/237) issued Feb. 21, 2022 by the International Searching Authority in International Application No. PCT/KR2021/016882. [cited by applicant]
Junki Ohmura et al., “Context-Aware Dialog Re-Ranking for Task-Oriented Dialog Systems”, arXiv:1811.11430v1, Nov. 28, 2018, 8 pages total. [cited by applicant]
Yue Weng et al., “Joint Contextual Modeling for ASR Correction and Language Understanding”, arXiv:2002.00750v1, IEEE, Jan. 28, 2020, 7 pages total. [cited by applicant]