IP Library Granted Patent US 11,893,813
Granted Patent B2
US 11,893,813 · App. 17/294,594 · Granted Feb 6, 2024

Electronic device and control method therefor

Inventors: Jeongho Mok (Suwon-si, KR); Heejun Song (Suwon-si, KR); Sanghyuk Yoon (Suwon-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G06V30/153G06N3/08G06V20/00G10L15/00G06V30/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,893,813
App. No.
17/294,594
Granted
Feb 6, 2024
Kind
B2
Abstract

An electronic device and a control method therefor are provided. The present electronic device comprises: a communication interface including a circuit, a memory for storing at least one instruction, and a processor for executing the at least one instruction, wherein the processor acquires contents through the communication interface, acquires information about a text included in an image of the contents, and acquires, on the basis of the information about the text included in the image of the contents, caption data of the contents by performing voice recognition for voice data included in the contents.

Claims (31)

1. An electronic device comprising:

a communication interface comprising circuitry;

a memory storing at least one instruction; and

a processor configured to execute the at least one instruction,

wherein the processor is configured to:

obtain a content via the communication interface,

obtain information on a text included in an image of the content,

obtain caption data of the content by performing speech recognition for speech data included in the content based on the information on the text included in the image of the content, and

perform the speech recognition for the speech data by applying a weight to each of an appearance time of the text, an appearance position of the text and a size of the text included in the image of the content obtained by analyzing image data included in the content.

2. The device of claim 1 , wherein the processor is further configured to obtain the information on the text included in the image of the content through optical character reader (OCR) for image data included in the content.

3. The device of claim 1 , wherein the processor is further configured to perform the speech recognition for speech data corresponding to a first screen by applying a weight to a text included in the first screen while performing the speech recognition for the speech data corresponding to the first screen of the image of the content.

4. The device of claim 1 , wherein the processor is further configured to perform the speech recognition for the speech data by applying a high weight to a text with a long appearance time or a large number of times of appearance among texts included in the image of the content obtained by analyzing image data included in the content.

5. The device of claim 1 , wherein the processor is further configured to perform the speech recognition for the speech data by applying a high weight to a text displayed at a fixed position among texts included in the image of the content obtained by analyzing image data included in the content.

6. The device of claim 1 , wherein the processor is further configured to:

determine a type of the content by analyzing the content, and

perform the speech recognition for the speech data by applying a weight to a text related to the determined type of the content.

7. The device of claim 6 , wherein the processor is further configured to determine the type of the content by analyzing metadata included in the content.

8. The device of claim 6 , wherein the processor is further configured to:

obtain information on the content by inputting image data included in the content to an artificial intelligence model trained for scene understanding, and

determine the type of the content based on the obtained information on the content.

9. A method for controlling an electronic device, the method comprising:

obtaining a content;

obtaining information on a text included in an image of the content; and

obtaining caption data of the content by performing speech recognition for speech data included in the content based on the information on the text included in the image of the content,

wherein the obtaining of the caption data comprises:

performing the speech recognition for the speech data by applying a weight to each of an appearance time of the text, an appearance position of the text and a size of the text included in the image of the content obtained by analyzing image data included in the content.

10. The method of claim 9 , wherein the obtaining of the information on the text comprises obtaining the information on the text included in the image of the content through optical character reader (OCR) for image data included in the content.

11. The method of claim 9 , wherein the obtaining of the caption data comprises performing the speech recognition for speech data corresponding to a first screen by applying a weight to a text included in the first screen while performing the speech recognition for the speech data corresponding to the first screen of the image of the content.

12. The method of claim 9 , wherein the obtaining of the caption data comprises performing the speech recognition for the speech data by applying a high weight to a text with a long appearance time or a large number of times of appearance among texts included in the image of the content obtained by analyzing image data included in the content.

13. The method of claim 9 , wherein the obtaining of the caption data comprises performing the speech recognition for the speech data by applying a high weight to a text displayed at a fixed position among texts included in the image of the content obtained by analyzing image data included in the content.

14. The method of claim 9 , wherein the obtaining of the caption data comprises performing the speech recognition for the speech data by applying a weight based on at least one of an appearance position of the text and a size of the text included in the image of the content obtained by analyzing image data included in the content.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2021
From: MOK, JEONGHO; SONG, HEEJUN; YOON, SANGHYUK
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 056262/0751 →
Priority Claims (1)
KR 10-2019-0013965 · Feb 1, 2019 · national
Continuity (1)
Related Publication 20220012520A1 · Jan 13, 2022