IP Library › Granted Patent US 12,526,485
Granted Patent B1
US 12,526,485 · App. 18/449,306 · Granted Jan 13, 2026

Content aware graphical subtitles

Inventors: Hooman Mahyar (Mercer Island, WA); James C. Willeford (Seattle, WA); Arjun Cholkar (Bothell, WA); Xinyu Li (Sammamish, WA); Zhikang Zhang (Seattle, WA)
Assignee: Amazon Technologies, Inc.
H04N21/4884H04N21/4316H04N21/44008
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,526,485
App. No.
18/449,306
Granted
Jan 13, 2026
Kind
B1
Abstract

Systems, devices, and methods are provided for context aware graphical subtitles. Techniques described herein may involve determining a representation for a graphical subtitle. The graphical subtitles may have various customizable properties that allow for creative expression. Representations of graphical subtitles may be determined using lower-level information extracted from detectors, such as audio detectors and/or visual detectors, as well as higher-level information determined by encoder-decoders.

Claims (57)

1 . A computer-implemented method, comprising:

determining a digital content item;

utilizing a plurality of detectors to determine semantic information for the digital content item collected from a sequence of visual frames of the digital content item;

determining subtitle text for the sequence of visual frames;

determining, based at least in part on the semantic information, an information map that encodes relationships between at least audio, visual, textual, and dialog information of the digital content item, wherein the information map comprises at least a first non-speaking visual event in the sequence of visual frames;

providing at least a portion of the information map to one or more encoder-decoders;

determining, by the one or more encoder-decoders and the at least portion of the information map, one or more annotations to the information map;

determining an updated information map based at least in part on the one or more annotations;

determining, based at least in part on the updated information map and the first non-speaking visual event, a visual effect; and

synchronizing presentation of the visual effect with the subtitle text.

2 . The computer-implemented method of claim 1 , wherein the plurality of detectors comprises one or more visual detectors and one or more audio detectors.

3 . The computer-implemented method of claim 1 , wherein:

the one or more encoder-decoders determines a macrogenre for the portion of the digital content item; and

the one or more annotations indicate one or more events of the portion of the digital content item that contributes to the macrogenre being contextually relevant to the portion of the digital content item.

4 . The computer-implemented method of claim 1 , further comprising:

determining to present the subtitle text during a first portion of the sequence of visual frames in synchronization with the first non-speaking visual event; and

determining to omit presentation of the subtitle text during a second portion of the visual sequence of visual frames in synchronization with a second non-speaking visual event.

5 . The computer-implemented method of claim 1 , further comprising determining one or more active region and inactive regions for the sequence of visual frames.

6 . The computer-implemented method of claim 1 , further comprising selecting one or more text properties for the subtitle text.

7 . A system, comprising:

one or more processors; and

memory storing executable instructions that, as a result of execution by the one or more processors, cause the system to:

determine a set of semantic information for a digital content item, wherein the set of semantic information is extracted based at least in part on audio, visual, textual, and dialog information from a sequence of frames of the digital content item;

determine an information map that encodes relationships for the set of semantic information, wherein the information map comprises at least a first non-speaking visual event determined to be in the sequence of frames;

determine, based at least in part on the information map, one or more annotations that associate a portion of the information map to a macrogenre;

determine, based at least in part on the information map, the first non-speaking visual event, and the one or more annotations, a visual effect; and

synchronize presentation of the visual effect with a subtitle text.

8 . The system of claim 7 , wherein the executable instruction, as a result of execution by the one or more processors, further cause the system to:

use one or more encoder-decoders to determine the one or more annotations based at least in part on the information map.

9 . The system of claim 8 , wherein:

the one or more encoder-decoders determines the macrogenre for the portion of the digital content item; and

the one or more annotations indicate one or more events of the portion of the digital content item that contributes to the macrogenre being contextually relevant to the portion of the digital content item.

10 . The system of claim 7 , wherein the set of semantic information is determined based at least in part by a set of detectors comprising one or more visual detectors and one or more audio detectors.

11 . The system of claim 7 , wherein the visual effect comprises at least one of: a font color, font type, font size or background color based at least in part on one or more visual aspects of the digital content item.

12 . The system of claim 7 , wherein the instructions to synchronize the presentation include instructions that, as a result of execution by the one or more processors, cause the system to:

display the subtitle text in association with the first non-speaking visual event;

determine a second non-speaking visual event in the information map; and

omit display of the subtitle text in association with the second non-speaking visual event.

13 . The system of claim 7 , wherein the executable instruction, as a result of execution by the one or more processors, further causes the system to determine one or more active regions and inactive regions for the portion of the digital content item.

14 . A non-transitory computer-readable storage medium storing executable instructions that, as a result of being executed by one or more processors of a computer system, cause the computer system to at least:

determine a set of semantic information for a digital content item, wherein the set of semantic information is extracted based at least in part on audio, visual, textual, and dialog information from a sequence of frames of the digital content item;

determine an information map that encodes relationships for the set of semantic information, wherein the information map comprises at least a first non-speaking visual event determined to be in the sequence of frames;

determine, based at least in part on the information map, one or more annotations that associate a portion of the information map to a macrogenre;

determine, based at least in part on the information map, the first non-speaking visual event, and the one or more annotations, a visual effect; and

synchronize presentation of the visual effect with a subtitle text.

15 . The non-transitory computer-readable storage medium of claim 14 , wherein the executable instruction, as a result of execution by the one or more processors, further cause the computer system to:

use one or more encoder-decoders to determine the one or more annotations based at least in part on the information map.

16 . The non-transitory computer-readable storage medium of claim 15 , wherein:

the one or more encoder-decoders determines the macrogenre for the portion of the digital content item; and

the one or more annotations indicate one or more events of the portion of the digital content item that contributes to the macrogenre being contextually relevant to the portion of the digital content item.

17 . The non-transitory computer-readable storage medium of claim 14 , wherein the visual effect comprises at least one of: a font color, font type, font size or background color based at least in part on one or more visual aspects of the digital content item.

18 . The non-transitory computer-readable storage medium of claim 14 , wherein the executable instruction, as a result of execution by the one or more processors, further cause the computer system to:

display the subtitle text in association with the first non-speaking visual event;

determine a second non-speaking visual event in the information map; and

omit display of the subtitle text in association with the second non-speaking visual event.

19 . The non-transitory computer-readable storage medium of claim 14 , wherein the executable instruction, as a result of execution by the one or more processors, further cause the computer system to determine one or more active regions and inactive regions for the portion of the digital content item.

20 . The non-transitory computer-readable storage medium of claim 14 , wherein the set of semantic information is determined based at least in part by a set of detectors comprising one or more visual detectors and one or more audio detectors.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2023
From: MAHYAR, HOOMAN; WILLEFORD, JAMES C.; CHOLKAR, ARJUN; LI, XINYU; ZHANG, ZHIKANG
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 064679/0729 →
References Cited (5)
US 11211053B2 · Aronowitz · 2021 [cited by examiner]
US 20060064716A1 · Sull · 2006 [cited by examiner]
US 20140278370A1 · Chen · 2014 [cited by examiner]
US 20230007359A1 · Aher · 2023 [cited by examiner]
US 20230362451A1 · Candelore · 2023 [cited by examiner]