IP Library Granted Patent US 12,587,714
Granted Patent B2
US 12,587,714 · App. 18/208,147 · Granted Mar 24, 2026

Methods and systems to automate subtitle display based on user language profile and preferences

Inventors: Jean-Yves Couleaud (Mission Viejo, CA); Reda Harb (Issaquah, WA); Tao Chen (Palo Alto, CA)
Assignee: Adeia Guides Inc.
H04N21/4884G10L13/02G10L15/005H04N21/4755H04N21/4856
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,587,714
App. No.
18/208,147
Granted
Mar 24, 2026
Kind
B2
Abstract

Systems and methods for automatically activating subtitles during presentation of a media asset are disclosed. In an example, a media application stores a user profile including a language classification probability vector corresponding to the user profile. The media application generates for display a media asset via a user interface. The media application compares the language classification probability vector to audio of the media asset presented via the user interface, and in response to determining that the audio of the media asset is beyond a threshold distance from the language classification probability vector, activates subtitles for the media asset.

Claims (83)

1 . A method for selectively activating subtitles, the method comprising:

storing, in a user profile associated with a client device, a language classification probability vector (LCPV) corresponding to the user profile, the LCPV comprising a plurality of discrete values each corresponding to a language or an accent, wherein the LCPV for the user profile is available in real time; and

wherein the LCPV is at least one of generated at a server and transmitted to the client device as metadata or determined by the client device;

generating for display a media asset via a user interface, wherein the media asset comprises a plurality of audio portions each corresponding to a subject of a plurality of subjects depicted in a scene of the media asset;

for a first subject of the plurality of subjects depicted in the scene:

identifying a first LCPV corresponding to the first subject, wherein the first LCPV is available in real time;

based at least in part on a comparison between the first LCPV corresponding to the first subject and the LCPV corresponding to the user profile, determining that the first LCPV corresponding to the first subject is beyond a threshold distance from the LCPV corresponding to the user profile;

based at least in part on determining that the first LCPV corresponding to the first subject is beyond the threshold distance from the LCPV corresponding to the user profile, activating first subtitles for an audio portion corresponding to the first subject; and

for a second subject of the plurality of subjects depicted in the scene:

identifying a second LCPV corresponding to the second subject, wherein the second LCPV is available in real time;

based at least in part on a comparison between the second LCPV corresponding to the second subject and the LCPV corresponding to the user profile, determining that the second LCPV corresponding to the second subject is below the threshold distance from the LCPV corresponding to the user profile; and

based at least in part on determining that the second LCPV corresponding to the second subject is below the threshold distance from the LCPV corresponding to the user profile, maintaining second subtitles for an audio portion corresponding to the second subject in an off state.

2 . The method of claim 1 , further comprising:

receiving a selected language corresponding to the user profile;

receiving a speech input corresponding to the user profile; and

modifying the LCPV corresponding to the user profile based on the selected language and the speech input.

3 . The method of claim 1 , wherein audio of the media asset comprises a media asset language profile represented by metadata associated with the media asset, the method further comprising:

comparing the LCPV corresponding to the user profile to the media asset language profile represented by the metadata associated with the media asset.

4 . The method of claim 1 , wherein the LCPV corresponding to the user profile indicates the user profile is associated with a proficiency in both a primary language and a secondary language, the method further comprising:

identifying in audio of the media asset a first section of dialogue associated with the primary language;

identifying in the audio of the media asset a second section of dialogue associated with the secondary language; and

preventing activation of subtitles for the first section of dialogue and the second section of dialogue.

5 . The method of claim 1 , further comprising:

receiving an input to turn subtitles on or off for the media asset; and

modifying the LCPV corresponding to the user profile based on the input.

6 . The method of claim 1 , further comprising:

presenting a list of subjects of the plurality of subjects corresponding to the media asset;

receiving a selection of a selected subject from the list of subjects; and

activating subtitles for the media asset during portions of the media asset in which the selected subject speaks.

7 . The method of claim 6 , wherein the media asset is a first media asset comprising the selected subject, the method further comprising:

storing a reference to the selected subject in the user profile;

identifying a second media asset comprising the selected subject;

comparing an LCPV of a first language profile associated with the selected subject and the first media asset to an LCPV of a second language profile associated with the selected subject and the second media asset; and

in response to determining that the LCPV of the first language profile and the LCPV of the second language profile are within the threshold distance, activating subtitles for the second media asset comprising the selected subject during portions of the second media asset in which the selected subject speaks.

8 . The method of claim 1 , further comprising:

presenting a list of subjects of the plurality of subjects corresponding to the media asset;

receiving a selection of a selected subject from the list of subjects; and

replacing audio of the media asset corresponding to the selected subject with computer generated audio.

9 . The method of claim 1 , wherein the media asset is generated for display via a primary device, the method further comprising generating for display on a secondary device the first subtitles for the audio portion corresponding to the first subject.

10 . A system for selectively activating subtitles, the system comprising:

control circuitry configured to:

store, in a user profile associated with a client device, a language classification probability vector (LCPV) corresponding to the user profile, the LCPV comprising a plurality of discrete values each corresponding to a language or accent, wherein the LCPV for the user profile is available in real time; and

wherein the LCPV is at least one of generated at a server and transmitted to the client device as metadata or determined by the client device; and

input/output circuitry configured to display a media asset via a user interface, wherein the media asset comprises a plurality of audio portions each corresponding to a subject of a plurality of subjects depicted in a scene of the media asset; and

wherein the control circuitry is further configured to:

for a first subject of the plurality of subjects depicted in the scene:

identify a first LCPV corresponding to the first subject, wherein the first LCPV is available in real time;

based at least in part on a comparison between the first LCPV corresponding to the first subject and the LCPV corresponding to the user profile, determine that the first LCPV corresponding to the first subject is beyond a threshold distance from the LCPV corresponding to the user profile;

based at least in part on determining that the first LCPV corresponding to the first subject is beyond the threshold distance from the LCPV corresponding to the user profile, activate first subtitles for an audio portion corresponding to the first subject; and

for a second subject of the plurality of subjects depicted in the scene:

identify a second LCPV corresponding to the second subject, wherein the second LCPV is available in real time;

based at least in part on a comparison between the second LCPV corresponding to the second subject and the LCPV corresponding to the user profile, determine that the second LCPV corresponding to the second subject is below the threshold distance from the LCPV corresponding to the user profile; and

based at least in part on determining that the second LCPV corresponding to the second subject is below the threshold distance from the LCPV corresponding to the user profile, maintain second subtitles for an audio portion corresponding to the second subject in an off state.

11 . The system of claim 10 , wherein:

the input/output circuitry is further configured to:

receive a selected language corresponding to the user profile; and

receive a speech input corresponding to the user profile; and

the control circuitry is further configured to:

modify the LCPV corresponding to the user profile based on the selected language and the speech input.

12 . The system of claim 10 , wherein audio of the media asset comprises a media asset language profile represented by metadata associated with the media asset, and wherein the control circuitry is further configured to: compare the LCPV corresponding to the user profile to the media asset language profile represented by the metadata associated with the media asset.

13 . The system of claim 10 , wherein the LCPV corresponding to the user profile indicates the user profile is associated with a proficiency in both a primary language and a secondary language, wherein the control circuitry is further configured to:

identify in audio of the media asset a first section of dialogue associated with the primary language;

identify in the audio of the media asset a second section of dialogue associated with the secondary language; and

prevent activation of subtitles for the first section of dialogue and the second section of dialogue.

14 . The system of claim 10 , wherein:

the input/output circuitry is further configured to receive an input to turn subtitles on or off for the media asset; and

the control circuitry is further configured to modify the LCPV corresponding to the user profile based on the input.

15 . The system of claim 10 , wherein:

the input/output circuitry is further configured to:

present a list of subjects of the plurality of subjects corresponding to the media asset; and

receive a selection of a selected subject from the list of subjects; and

the control circuitry is further configured to activate subtitles for the media asset during portions of the media asset in which the selected subject speaks.

16 . The system of claim 15 , wherein the media asset is a first media asset comprising the selected subject, wherein the control circuitry is further configured to:

store a reference to the selected subject in the user profile;

identify a second media asset comprising the selected subject;

compare an LCPV of a first language profile associated with the selected subject and the first media asset to an LCPV of a second language profile associated with the selected subject and the second media asset; and

in response to determining that the LCPV of the first language profile and the LCPV of the second language profile are within the threshold distance, activate subtitles for the second media asset comprising the selected subject during portions of the second media asset in which the selected subject speaks.

17 . The system of claim 10 , wherein:

the input/output circuitry is further configured to:

present a list of subjects of the plurality of subjects corresponding to the media asset; and

receive a selection of a selected subject from the list of subjects; and

the control circuitry is further configured to replace audio of the media asset corresponding to the selected subject with computer generated audio.

18 . The system of claim 10 , wherein the media asset is generated for display via a primary device, and wherein the input/output circuitry is further configured to generate for display on a secondary device the first subtitles for the audio portion corresponding to the first subject.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 18, 2023
From: COULEAUD, JEAN-YVES; HARB, REDA; CHEN, TAO
To: ADEIA GUIDES INC.
Reel/Frame 064936/0278 →
Continuity (1)
Related Publication 20240414409A1 · Dec 12, 2024
References Cited (15)
US 9854324B1 · Panchaksharaiah · 2017 [cited by examiner]
US 10182266B2 · Panchaksharaiah et al. · 2019 [cited by applicant]
US 11665392B2 · Chandrashekar et al. · 2023 [cited by applicant]
US 20110246172A1 · Liberman et al. · 2011 [cited by applicant]
US 20170374423A1 · Anderson · 2017 [cited by examiner]
US 20200007946A1 · Olkha · 2020 [cited by examiner]
US 20230402033A1 · Heinzmann · 2023 [cited by examiner]
JP 2013201505A · 2013 [cited by applicant]
Hinsvark et al., “Accented Speech Recognition: A Survey,” [5] https://arxiv.org/abs/2104.10747 pp. 1-5 (2021). [cited by applicant]
https://aws.amazon.com/about-aws/whats-new/2021/08/amazon-chime-sdk-amazon-transcribe-amazon-transcribe-medical/ retrieved from internet Aug. 29, 2023. [cited by applicant]
https://meeting.tencent.com/support-doc-detail/1174/index.html retrieved from internet Aug. 29, 2023. [cited by applicant]
https://support.google.com/meet/answer/9300310?hl=en&co=GENIE.Platform%3DDesktop retrieved from the internet Aug. 29, 2023. [cited by applicant]
Najafian et al., “Automatic accent identification as an analytical tool for accent robust automatic speech recognition,” https://www.sciencedirect.com/science/article/abs/pii/S0167639317300043 (2020). [cited by applicant]
Radzikowski et al., “Accent modification for speech recognition of non-native speakers using neural style transfer,” EURASIP Journal on Audio, Speech, and Music Processing, pp. 1-10 (2021). [cited by applicant]
Timed Text Markup Language 1 (TTML1) (Third Edition) https://www.w3.org/TR/2018/REC-ttml1-20181108/ retrieved from internet Aug. 29, 2023. [cited by applicant]