IP Library › Granted Patent US 12,651,449
Granted Patent B2
US 12,651,449 · App. 18/354,066 · Granted Jun 9, 2026

Modular artificial intelligence platform for media generation and methods for use therewith

Inventors: Rory Donovan (Arroyo Grande, CA); Alan Salimov (Oakland, CA); Alexander Gluklick Braun (Oakland, MI); Kerim Doruk Karinca (San Francisco, CA)
Assignee: Virtuous AI, Inc.
G06V10/82G06V10/80G06V10/87
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,449
App. No.
18/354,066
Granted
Jun 9, 2026
Kind
B2
Abstract

A modular artificial intelligence (AI) platform operates by: receiving media input that includes image data and text data; generating encoded text data via a text encoder module that includes first language processing AI; generating encoded image data via an image encoder module that includes a plurality of neural networks and a long short-term memory; generating concept structure data via a concept identification module that includes graph-based learning AI; generating decoded text data via a text decoder module that includes language processing AI; generating decoded image data, via an image decoder module that includes a plurality of neural networks and a long short-term memory; and combining the decoded image data and the decoded text data to generate media output data.

Claims (37)

1 . A modular artificial intelligence (AI) platform comprises:

a network interface configured to communicate via a network;

at least one processor; and

a non-transitory machine-readable storage medium that stores operational instructions that, when executed by the at least one processor, cause the at least one processor to perform operations that include:

receiving, via the network interface, media input that includes image data and text data associated therewith;

generating encoded text data, based on the text data and via a text encoder module that includes first language processing AI;

generating encoded image data, based on the image data and via an image encoder module that includes a first plurality of neural networks and at least one first long short-term memory;

generating concept structure data, based on the encoded text data and the encoded image data and via a concept identification module that includes graph-based learning AI;

generating decoded text data, based on the concept structure data and the encoded text data and via a text decoder module that includes second language processing AI;

generating decoded image data, based on the concept structure data and the encoded image data and via an image decoder module that includes a second plurality of neural networks and at least one second long short-term memory; and

combining the decoded image data and the decoded text data to generate media output data.

2 . The modular AI platform of claim 1 , wherein the image decoder module and the text decoder module are trained based on the concept structure data.

3 . The modular AI platform of claim 1 , wherein the image encoder module and the text encoder module are trained based on the concept structure data.

4 . The modular AI platform of claim 1 , wherein the first language processing AI includes a Bidirectional Encoder Representations from Transformers (BERT) AI model.

5 . The modular AI platform of claim 1 , wherein the second language processing AI includes a Bidirectional Encoder Representations from Transformers (BERT) AI model.

6 . The modular AI platform of claim 1 , wherein the first plurality of neural networks includes k U-net models operating as k-experts.

7 . The modular AI platform of claim 6 , wherein the k-experts are trained independently on different subsets of the data, and outputs of k-experts are then combined using a gating mechanism that selects a most relevant one of the k-experts for each input image of the image data.

8 . The modular AI platform of claim 1 , wherein the concept identification module includes a third long short-term memory that processes the encoded image data for input to the graph-based learning AI.

9 . The modular AI platform of claim 1 , wherein the concept identification module includes a fourth long short-term memory that processes the encoded text data for input to the graph-based learning AI.

10 . The modular AI platform of claim 1 , wherein the graph-based learning AI operates based on a GraphSAGE model.

11 . A method comprises:

receiving, via a network interface, media input that includes image data and text data associated therewith;

generating encoded text data, based on the text data and via a text encoder module that includes first language processing AI;

generating encoded image data, based on the image data and via an image encoder module that includes a first plurality of neural networks and at least one first long short-term memory;

generating concept structure data, based on the encoded text data and the encoded image data and via a concept identification module that includes graph-based learning AI;

generating decoded text data, based on the concept structure data and the encoded text data and via a text decoder module that includes second language processing AI;

generating decoded image data, based on the concept structure data and the encoded image data and via an image decoder module that includes a second plurality of neural networks and at least one second long short-term memory; and

combining the decoded image data and the decoded text data to generate media output data.

12 . The method of claim 11 , wherein the image decoder module and the text decoder module are trained based on the concept structure data.

13 . The method of claim 11 , wherein the image encoder module and the text encoder module are trained based on the concept structure data.

14 . The method of claim 11 , wherein the first language processing AI includes a Bidirectional Encoder Representations from Transformers (BERT) AI model.

15 . The method of claim 11 , wherein the second language processing AI includes a Bidirectional Encoder Representations from Transformers (BERT) AI model.

16 . The method of claim 11 , wherein the first plurality of neural networks includes k U-net models operating as k-experts.

17 . The method of claim 16 , wherein the k-experts are trained independently on different subsets of the data, and outputs of k-experts are then combined using a gating mechanism that selects a most relevant one of the k-experts for each input image of the image data.

18 . The method of claim 11 , wherein the concept identification module includes a third long short-term memory that processes the encoded image data for input to the graph-based learning AI.

19 . The method of claim 11 , wherein the concept identification module includes a fourth long short-term memory that processes the encoded text data for input to the graph-based learning AI.

20 . The method of claim 11 , wherein the graph-based learning AI operates based on a GraphSAGE model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 16, 2023
From: DONOVAN, RORY; SALIMOV, ALAN; BRAUN, ALEXANDER GLUKLICK; KARINCA, KERIM DORUK
To: VIRTUOUS AI, INC.
Reel/Frame 064603/0923 →
Continuity (1)
Related Publication 20250029375A1 · Jan 23, 2025
References Cited (41)
US 8221238B1 · Shaw · 2012 [cited by applicant]
US 10671854B1 · Mahyar · 2020 [cited by applicant]
US 10905962B2 · Kaethler · 2021 [cited by applicant]
US 11052311B2 · Bleasdale · 2021 [cited by applicant]
US 20060121990A1 · O'Kelley · 2006 [cited by applicant]
US 20060247055A1 · O'Kelley · 2006 [cited by applicant]
US 20060287096A1 · O'Kelley · 2006 [cited by applicant]
US 20090234663A1 · McCann · 2009 [cited by applicant]
US 20090325709A1 · Shi · 2009 [cited by applicant]
US 20180373781A1 · Palrecha · 2018 [cited by applicant]
US 20190028336A1 · Coronado et al. · 2019 [cited by applicant]
US 20190213498A1 · Adjaoute · 2019 [cited by applicant]
US 20190238934A1 · Yun · 2019 [cited by applicant]
US 20190266912A1 · Barzman · 2019 [cited by applicant]
US 20190291008A1 · Cox · 2019 [cited by applicant]
US 20200078688A1 · Kaethler · 2020 [cited by applicant]
US 20200142999A1 · Pedersen · 2020 [cited by applicant]
US 20200250525A1 · Kumar Addepalli et al. · 2020 [cited by applicant]
US 20200364727A1 · Scott-Green · 2020 [cited by applicant]
US 20210004440A1 · Purnell · 2021 [cited by applicant]
US 20210038979A1 · Bleasdale · 2021 [cited by applicant]
US 20210168166A1 · Liu · 2021 [cited by applicant]
US 20220012296A1 · Marey · 2022 [cited by examiner]
US 20220198144A1 · Yang · 2022 [cited by examiner]
US 20220263860A1 · Crabtree · 2022 [cited by examiner]
US 20220377083A1 · Kim · 2022 [cited by applicant]
US 20240331235A1 · Smock · 2024 [cited by examiner]
US 20240386015A1 · Crabtree · 2024 [cited by examiner]
US 20250029375A1 · Donovan · 2025 [cited by examiner]
WO 2020092956A1 · 2020 [cited by applicant]
Zhang et al., “Hierarchical Cross-Modality Semantic Correlation Learning Model for Multimodal Summarization,” arXiv: 2112.12072v1 [cs.CV] Dec. 16, 2021 (Year: 2021). [cited by examiner]
Kapoor et al., “Underwater Moving Object Detection using an End-to-End Encoder-Decoder Architecture and GraphSage with Aggregator and Refactoring,” 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Wor… [cited by examiner]
Wang et al., “Learning to Compose Topic-Aware Mixture of Experts for Zero-Shot Video Captioning,” 1811.02765v2 [cs.CL] Nov. 23, 2018 (Year: 2018). [cited by examiner]
Zhibin Lu, “VGCN-BERT : Augmenting BERT with Graph Embedding for Text Classification—Application to offensive language detection,” Université de Montréal Thesis, 2020 (Year: 2020). [cited by examiner]
Hendrycks et al.; Aligning AI with Shared Human Values; arXiv:2008.02275v5, Jul. 2021; 30 pgs. [cited by applicant]
International Searching Authority; International Search Report and Written Opinion; International Application No. PCT/US2022/041024; Dec. 21, 2022; 11 pgs. [cited by applicant]
European Patent Office; Extended European Search Report; Application No. 24177224.3; Feb. 4, 2025; 12 pgs. [cited by applicant]
Kapoor et al., Underwater Moving Object Detection using an End-to-End Encoder-Decoder Architecture and GraphSage with Aggregator and Refactoring, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Works… [cited by applicant]
Wang et al., Learning to Compose Topic-Aware Mixture of Experts for Zero-Shot Video Captioning, arxiv.org, Cornell University Library, Nov. 7, 2018 (Nov. 7, 2018), XP081047116. [cited by applicant]
You et al., Implicit Anatomical Rendering for Medical Image Segmentation with Stochastic Experts, arxiv.org, Cornell University Library, Apr. 6, 2023 (Apr. 6, 2023), XP091478597. [cited by applicant]
Zhang et al., Hierarchical Cross-Modality Semantic Correlation Learning Model for Multimodal Summarization, arxiv.org, Cornell University Library, Dec. 16, 2021, (Dec. 16, 2021), XP091122580. [cited by applicant]