IP Library › Granted Patent US 12,400,122
Granted Patent B2
US 12,400,122 · App. 17/171,167 · Granted Aug 26, 2025

Narrative-based content discovery employing artificial intelligence

Inventors: Debajyoti Ray (Culver City, CA); Walter Kortschak (Culver City, CA)
Assignee: RIVETAI, INC.
G06N3/088G06N3/045G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,400,122
App. No.
17/171,167
Granted
Aug 26, 2025
Kind
B2
Abstract

Processor-based systems and/or methods of operation may generate queries and suggest legacy narrative content (e.g., video content, script content) for a narrative under development. An artificial neural network (ANN, e.g., autoencoder) is trained on pairs of video and text vectors to capture attributes or nuances beyond those typical of keyword searching. Query vector representations generated using an instance of the ANN may be matched against candidate vector representations, for instance generated using an instance of the ANN from legacy narratives. Such may query for missing video and/or text for a narrative under development. Matches may be returned, including scores or ranks. Feature vectors may be shared without jeopardizing source narrative content. Legacy source narrative content may remain secure behind a controlling entity's network security wall.

Claims (62)

1. A method of operation of a computational system that implements at least one artificial neural network, the method comprising:

comparing a query vector representation generated by at least a first instance of one autoencoder running on a first processor-based system, against a plurality of candidate vector representations generated by at least a second instance of the autoencoder running on a second processor-based system,

the query vector representation representing an inquiry regarding an incomplete narrative that lacks a portion of video or lacks a portion of text descriptors,

the plurality of candidate vector representations representing each of a plurality of candidate narratives,

the autoencoder trained on a training data set comprising a plurality of pairs of vectors, each pair of the vectors corresponding to a respective one of a plurality of training narratives and including a video vector and a text vector,

the video vector comprising a plurality of video descriptors extracted from a sequence of images of the corresponding training narrative and the text vector comprising a plurality of text descriptors extracted from a set of scene descriptions of the corresponding training narrative and extracted from at least a portion of a script of the corresponding training narrative, the video vector and the text vector of each pair aligned with one another; and

generating an indication of any matches from the candidate vector representations with the query vector representation including at least one of a score or a rank of the respective match.

2. The method of claim 1 , further comprising receiving the inquiry at the second processor-based system of a second entity, wherein the query vector representation was generated via the first instance of the autoencoder at the first processor-based system of a first entity, and the comparing occurs at the second processor-based system of the second entity.

3. The method of claim 2 , further comprising:

transmitting by the second processor-based system of the second entity to the first processor-based system of the first entity the indication of any matches from the candidate vector representations with the query vector representation including at least one of a score or a rank of the respective match.

4. The method of claim 2 , further comprising:

receiving at least the second instance of the autoencoder at the second processor-based system of the second entity.

5. The method of claim 2 , further comprising:

generating the candidate vector representations by at least the second instance of the autoencoder at the second processor-based system of the second entity from a set of source content including a plurality of source narratives stored behind a network wall of the second entity.

6. The method of claim 2 , further comprising:

receiving at least the first instance of the autoencoder at the first processor-based system of the first entity.

7. The method of claim 1 , further comprising:

transmitting at least the first instance of the autoencoder to the first processor-based system of a first entity; and

transmitting at least the second instance of the autoencoder to the second processor-based system of a second entity.

8. The method of claim 1 , further comprising:

receiving the inquiry at a third processor-based system of an intermediary entity, wherein the query vector representation was generated via at least the first instance of the autoencoder at the first processor-based system of a first entity; and

receiving the candidate vector representations at the intermediary entity, where the candidate vector representations were generated by at least the second instance of the autoencoder at the second processor-based system of a second entity.

9. The method of claim 8 wherein the comparing occurs at the third processor-based system of the intermediary, and further comprising:

transmitting the indication of any matches from the candidate vector representations with the query vector representation including at least one of a score or a rank of the respective match by the third processor-based system of the intermediary entity to the first processor-based system of the first entity.

10. The method of claim 1 , further comprising:

transmitting the inquiry from the first processor-based system of a first entity, the inquiry comprising the query vector representation, and wherein the query vector representation was generated via at least the first instance of the autoencoder and the candidate vector representations were generated by at least the second instance of the autoencoder; and

receiving by the first processor-based system of the first entity the indication of any matches from the candidate vector representations with the query vector representation including at least one of a score or a rank of the respective match.

11. The method of claim 1 , further comprising:

generating the query vector representation by the first processor-based system of a first entity via the first instance of the autoencoder.

12. The method of claim 11 wherein generating the query vector representation by the first processor-based system of the first entity via the first instance of the autoencoder comprises: generating the query vector representation from the incomplete narrative that lacks a portion of the video.

13. The method of claim 11 wherein generating the query vector representation by the first processor-based system of the first entity via the first instance of the autoencoder comprises: generating the query vector representation from the incomplete narrative that lacks a portion of a script.

14. The method of claim 11 wherein generating the query vector representation by the first processor-based system of the first entity via the first instance of the autoencoder comprises: providing a video vector and a text vector to the first instance of the autoencoder.

15. The method of claim 11 wherein generating the query vector representation by the first processor-based system of the first entity via the first instance of the autoencoder comprises: automatically extracting a video vector from at least a portion of a sequence of images of at least a portion of the incomplete narrative.

16. The method of claim 11 wherein generating the query vector representation by the first processor-based system of the first entity via the first instance of the autoencoder comprises: automatically extracting a text vector from at least a portion of a portion of a script of the incomplete narrative.

17. The method of claim 11 wherein generating the query vector representation by the first processor-based system of the first entity via the first instance of the autoencoder comprises: automatically extracting a text vector from at least a portion of a description of at least one scene of at least a portion of the incomplete narrative.

18. The method of claim 11 wherein generating the query vector representation by the first processor-based system of the first entity via the first instance of the autoencoder comprises:

automatically extracting one or more text descriptors of at least one scene of at least a portion of the incomplete narrative; and

updating one or more of the automatically extracted text descriptors based on user input.

19. A computational system that implements at least one artificial neural network, the computational system comprising:

at least one processor;

at least one nontransitory processor-readable medium communicatively coupled to the at least one processor and that stores processor-executable instructions which, when executed by the at least one processor, cause the at least one processor to:

compare a query vector representation generated by at least a first instance of one autoencoder running on a first processor-based system of a first entity against a plurality of candidate vector representations generated by at least a second instance of the autoencoder, the query vector representation representing an inquiry regarding an incomplete narrative that lacks a portion of video or lacks a portion of text descriptors, the plurality of candidate vector representations representing each of a plurality of candidate narratives, the autoencoder trained on a training data set comprising a plurality of pairs of vectors, each pair of the vectors corresponding to a respective one of a plurality of training narratives and including a video vector and a text vector, the video vector comprising a plurality of video descriptors extracted from a sequence of images of the corresponding training narrative and the text vector comprising a plurality of text descriptors extracted from a set of scene descriptions of the corresponding training narrative and extracted from at least a portion of a script of the corresponding training narrative, the video vector and the text vector of each pair aligned with one another; and

generate an indication of any matches from the candidate vector representations with the query vector representation including at least one of a score or a rank of the respective match.

20. The computational system of claim 19 wherein, when executed, the processor-executable instructions further cause the at least one processor further to:

generate the plurality of candidate vector representations from the plurality of candidate narratives using the second instance of the autoencoder at a second entity; and

receive the inquiry from the first processor-based system of the first entity, wherein the compare occurs at the second entity.

21. The computational system of claim 20 wherein, when executed, the processor-executable instructions further cause the at least one processor further to:

transmit to the first entity the indication of any matches from the candidate vector representations with the query vector representation including at least one of a score or a rank of the respective match.

22. The computational system of claim 20 wherein, when executed, the processor-executable instructions further cause the at least one processor further to:

transmit at least the second instance of the autoencoder to the processor-based system of the second entity.

23. The computational system of claim 20 wherein the plurality of candidate narratives are derived from a set of source content including a plurality of source narratives stored behind a network wall of the second entity.

24. The computational system of claim 19 wherein when executed, the processor-executable instructions further cause the at least one processor further to:

transmit at least the first instance of the trained autoencoder to the first processor-based system of a first entity; and

transmit at least the second instance of the autoencoder to a second processor-based system of an intermediary entity.

25. The computational system of claim 19 wherein, when executed, the processor-executable instructions further cause the at least one processor further to:

receive the inquiry from the first processor-based system of the first entity, the inquiry comprising the query vector representation, and wherein the query vector representation was generated via at least the first instance of the autoencoder and the candidate vector representations were generated by at least the second instance of the autoencoder; and

transmit, to the processor-based system of the first entity the indication of any matches from the candidate vector representations with the query vector representation including at least one of a score or a rank of the respective match.

26. The computational system of claim 19 wherein the first processor-based system provides a video vector and a text vector to the first instance of the autoencoder.

27. The computational system of claim 26 wherein the first processor-based system automatically extracts the video vector from at least a portion of a sequence of images of at least a portion of the incomplete narrative.

28. The computational system of claim 26 wherein the first processor-based system automatically extracts the text vector from at least a portion of a portion of a script of the incomplete narrative.

29. The computational system of claim 26 wherein the first processor-based system automatically extracts the text vector from at least a portion of a description of at least one scene of at least a portion of the incomplete narrative.

30. The computational system of claim 19 wherein the first processor-based system automatically extracts one or more text descriptors of at least one scene of at least a portion of the incomplete narrative; and updates one or more of the automatically extracted text descriptors based on user input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2021
From: RAY, DEBAJYOTI; KORTSCHAK, WALTER
To: RIVETAI, INC.
Reel/Frame 056214/0572 →
Continuity (1)
Related Publication 20220253715A1 · Aug 11, 2022
References Cited (51)
US 5604855A · Crawford · 1997 [cited by applicant]
US 7027974B1 · Busch et al. · 2006 [cited by applicant]
US 8094881B2 · Matsugu et al. · 2012 [cited by applicant]
US 8824861B2 · Gentile et al. · 2014 [cited by applicant]
US 9106812B1 · Price et al. · 2015 [cited by applicant]
US 10339922B2 · Kotri et al. · 2019 [cited by applicant]
US 10503775B1 · Ranzinger · 2019 [cited by examiner]
US 10558761B2 · Li · 2020 [cited by examiner]
US 10789288B1 · Ranzinger · 2020 [cited by applicant]
US 20050081159A1 · Gupta et al. · 2005 [cited by applicant]
US 20080221892A1 · Nathan et al. · 2008 [cited by applicant]
US 20090024963A1 · Lindley et al. · 2009 [cited by applicant]
US 20090094039A1 · MacDonald et al. · 2009 [cited by applicant]
US 20110135278A1 · Klappert · 2011 [cited by applicant]
US 20120173980A1 · Dachs · 2012 [cited by applicant]
US 20130124984A1 · Kuspa · 2013 [cited by applicant]
US 20130211841A1 · Ehsani et al. · 2013 [cited by applicant]
US 20140143183A1 · Sigal et al. · 2014 [cited by applicant]
US 20150186771A1 · Bhatt et al. · 2015 [cited by applicant]
US 20150199995A1 · Silverstein et al. · 2015 [cited by applicant]
US 20150205762A1 · Kulikowska · 2015 [cited by applicant]
US 20160147399A1 · Berajawala et al. · 2016 [cited by applicant]
US 20170039883A1 · Hunt et al. · 2017 [cited by applicant]
US 20170078621A1 · Sahay et al. · 2017 [cited by applicant]
US 20170357720A1 · Torabi · 2017 [cited by examiner]
US 20180121798A1 · Barkan · 2018 [cited by examiner]
US 20180136828A1 · Threewits · 2018 [cited by applicant]
US 20190213253A1 · Ray et al. · 2019 [cited by applicant]
US 20190287217A1 · Cooke · 2019 [cited by examiner]
US 20200125600A1 · Jo · 2020 [cited by examiner]
US 20200272695A1 · Dogan et al. · 2020 [cited by applicant]
US 20220161816A1 · Gyllenhammar · 2022 [cited by examiner]
US 20230188319A1 · Froelicher · 2023 [cited by examiner]
US 20230196769A1 · Trott · 2023 [cited by examiner]
EP 3012776A1 · 2016 [cited by applicant]
WO WO0150668 · 2001 [cited by applicant]
WO 2010081225A1 · 2010 [cited by applicant]
WO 2019140120A1 · 2019 [cited by applicant]
WO 2019140129A1 · 2019 [cited by applicant]
International Search Report and Written Opinion for PCT/US2022/015500, mailed May 16, 2022, 10 pages. [cited by applicant]
Pelin, Dogan et al., Label-Based Automatic Alignment of Video with Narrative Sentences, Computer Vision—ECCV 2016 Workshops, Sep. 18, 2016, 36 pages. [cited by applicant]
International Search Report and Written Opinion for PCT/US2019/013095, mailed Jun. 5, 2019, 15 pages. [cited by applicant]
International Search Report and Written Opinion for PCT/US2019/013105, mailed May 2, 2019, 14 pages. [cited by applicant]
Lopez, M., “Netflix's New App Aims to Simplify the Film Production Process,” Videoink, URL= https://www.thevideoink.com/2018/03/07/netflixs-new-app-aims-simplify-film-production-process/, Mar. 7, 2018, 1 page. [cited by applicant]
Roettgers, J. “Netflix's Newest App Isn't for Consumers, but the People Making Its Shows,” Variety, URL = http://variety.com/2018/digital/news/netflix-move-app-production-technology-1202720581/, Mar. 7, 2018, 3 pages. [cited by applicant]
Olah, Christopher , “Understanding LSTM Networks”, http://colah.github.io/posts/2015-08-Understanding-LSTMs/, Aug. 28, 2015, 15 pages. [cited by applicant]
Si, Mei , et al., Si et al., “Thespian: Using Multi-Agent Fitting to Craft Interactive Drama,” AMAS '05, copyright 2005 ACM, p. 21-28. (Year: 2005). [cited by applicant]
Tsai, Chia-Ming , et al., Tsai et al., “Scene-Based Movie Summarization Via Role-Community Networks” IEEE Transactions on Circuits and Systems for Video Technology, vol. 23, No. 11, Nov. 2013, p. 1927-1940 (Year 2013). [cited by applicant]
EP Search Report maiied Feb. 5, 2025 in App. No. 22753183.7-1203 / 4292021 PCT/US2022015500, 8 pages. [cited by applicant]
Michael Wray et al: “Fine-Grained Action Retrieval Through Multiple Parts-of-Speech Embeddings”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Aug. 9, 2019 (Aug. 9, 2019), … [cited by applicant]
Mithun Niluthpol Chowdhury et al: “Learning Joint Embedding with Multimodal Cues for Cross-Modal Video-Text, Retrieval”,Proceedings of the 14th ACM Web Science Conference 2022, ACMPUB27, New York, NY, USA, Jun. 5, 2018 … [cited by applicant]