IP Library › Granted Patent US 12,670,714
Granted Patent B2
US 12,670,714 · App. 17/551,407 · Granted Jun 30, 2026

Object interaction detection and inferences using semantic learning

Inventors: Hugo Latapie (Long Beach, CA); Ozkan Kilic (Long Beach, CA); Adam James Lawrence (Pasadena, CA); Gaowen Liu (Austin, TX); Andrew Albert Pletcher (Scotts Valley, CA)
Assignee: Cisco Technology, Inc.
G06V20/41G06T7/70G06V10/70G06T2207/10016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,714
App. No.
17/551,407
Filed
Dec 15, 2021
Granted
Jun 30, 2026
Kind
B2
Art Unit
2178
USPC
382/103
Abstract

In one embodiment, a device converts video data into a set of tracklets, each tracklet representing a different object depicted in the video data. The device identifies a particular object depicted in the video data as being an attractor or repulsor with respect to one or more other objects depicted in the video data, based on an analysis of their respective tracklets. The device makes, using a semantic reasoning engine, an inference about the video data, based in part on the particular object being identified as an attractor or repulsor. The device provides data based on the inference for display.

Claims (56)

1 . A method comprising:

converting, by a device, video data into a set of tracklets, each tracklet representing a different object depicted in the video data;

identifying, by the device and based on the set of tracklets, a particular object depicted in the video data as being an attractor or a repulsor with respect to one or more other objects depicted in the video data, the identifying including:

generating a first motion vector for a first tracklet associated with the particular object, the first motion vector indicative of a velocity and direction of motion of the particular object;

generating a second motion vector for a second tracklet associated with the one or more other objects, the second motion vector indicative of a velocity and direction of motion of the one or more other objects;

analyzing the first motion vector and the second motion vector to determine whether more than a threshold number of the one or more other objects are moving towards the particular object or away from the particular object;

determining that the particular object is an attractor when it is determined that the one or more objects are moving towards the particular object;

determining that the particular object is a repulsor when it is determined that the one or more objects are moving away from the particular object; and

dynamically applying and reassigning a focus of attention to the particular object over a plurality of frames of the video data based on a number of the one or more other objects moving toward or away from the particular object;

making, by the device and using a semantic reasoning engine that uses a knowledge base, an inference about the video data, based in part on the particular object being identified as an attractor or repulsor and an orientation of the particular object relative to the one or more other objects, the inference indicative of a dangerous or urgent situation, wherein the knowledge base represents concepts and relationships among the concepts that are domain-agnostic prior to runtime and instantiated with attractor-specific or repulsor-specific predicates during the inference, and has not been trained with the video data; and

providing, by the device, data for display based on the inference, wherein the data comprises both the identification of the attractor or the repulsor of the particular object and a predicted future convergence or divergence event between the particular object and the one or more other objects.

2 . The method as in claim 1 , wherein the video data is generated by a plurality of cameras deployed to a location.

3 . The method as in claim 2 , wherein identifying the particular object as being an attractor or repulsor comprises:

re-identifying the particular object across video data from different cameras in the plurality of cameras by correlating tracklets associated with the particular object.

4 . The method as in claim 1 , wherein identifying the particular object as being an attractor or repulsor comprises:

analyzing interactions between the first tracklet of the particular object and the second tracklet of the one or more other objects depicted in the video data.

5 . The method as in claim 1 , wherein making, by the device and using a semantic reasoning engine, the inference about the video data, based in part on the particular object being identified as an attractor or repulsor comprises:

using zero-shot learning to infer a condition of the particular object, without the device being trained to recognize the condition using sample data representative of that condition.

6 . The method as in claim 1 , wherein the semantic reasoning engine uses a knowledge graph comprising concepts and relationships, to make the inference about the video data.

7 . The method as in claim 6 , wherein the concepts of the knowledge graph represent at least one of: a medical condition, fighting, aggression, or a stampede.

8 . The method as in claim 1 , wherein the inference about the video data comprises a predicted future event.

9 . An apparatus, comprising:

a network interface to communicate with a computer network;

a processor coupled to the network interface; and

a memory configured to store one or more instructions, that when executed by the processor, configure the processor to:

convert video data into a set of tracklets, each tracklet representing a different object depicted in the video data;

identify a particular object depicted in the video data as being an attractor or a repulsor with respect to one or more other objects depicted in the video data by dynamically applying and reassigning a focus of attention to the particular object over a plurality of frames responsive to detecting relative motion patterns between the particular object and the one or more other objects, based on an analysis of their respective tracklets, wherein the particular object is identified as the attractor when the one or more other objects move towards it and is identified as the repulsor when the one or more other objects move away from it, wherein identification of the attractor or the repulsor is determined by analyzing motion vectors and spatial-temporal relationships between tracklets of the particular object and the one or more other objects without altering tracklet boundaries during tracking;

make, using a semantic reasoning engine that uses a knowledge base, an inference about the video data, based in part on the particular object being identified as an attractor or repulsor and an orientation of the particular object relative to the one or more other objects, the inference indicative of a dangerous or urgent situation, wherein the knowledge base represents concepts and relationships among the concepts that are domain-agnostic prior to runtime and instantiated with attractor-specific or repulsor-specific predicates during the inference, and has not been trained with the video data; and

provide data for display based on the inference, wherein the data comprises both the identification of the attractor or the repulsor of the particular object and a predicted future convergence or divergence event between the particular object and the one or more other objects.

10 . The apparatus as in claim 9 , wherein the video data is generated by a plurality of cameras deployed to a location.

11 . The apparatus as in claim 10 , wherein the apparatus identifies the particular object as being an attractor or repulsor by:

re-identifying the particular object across video data from different cameras in the plurality of cameras by correlating tracklets associated with the particular object.

12 . The apparatus as in claim 9 , wherein the apparatus identifies the particular object as being an attractor or repulsor by:

analyzing interactions between a tracklet of the particular object and those of the one or more other objects depicted in the video data.

13 . The apparatus as in claim 9 , wherein the apparatus makes, using a semantic reasoning engine, the inference about the video data, based in part on the particular object being identified as an attractor or repulsor comprises:

using zero-shot learning to infer a condition of the particular object, without the apparatus being trained to recognize the condition using sample data representative of that condition.

14 . The apparatus as in claim 9 , wherein the semantic reasoning engine uses a knowledge graph comprising concepts and relationships, to make the inference about the video data.

15 . The apparatus as in claim 14 , wherein the concepts of the knowledge graph represent at least one of: a medical condition, fighting, aggression, or a stampede.

16 . The apparatus as in claim 9 , wherein the inference about the video data comprises a predicted future event.

17 . A tangible, non-transitory, computer-readable medium storing program instructions that cause a device to execute a process comprising:

converting, by the device, video data into a set of tracklets, each tracklet representing a different object depicted in the video data;

identifying, by the device, a particular object depicted in the video data as being an attractor or a repulsor with respect to one or more other objects depicted in the video data by dynamically applying and reassigning a focus of attention to the particular object over a plurality of frames responsive to detecting relative motion patterns between the particular object and the one or more other objects, based on an analysis of their respective tracklets, wherein the particular object is identified as the attractor when the one or more other objects move towards it and is identified as the repulsor when the one or more other objects move away from it, wherein identification of the attractor or the repulsor is determined by analyzing motion vectors and spatial-temporal relationships between tracklets of the particular object and the one or more other objects without altering tracklet boundaries during tracking;

making, by the device and using a semantic reasoning engine that uses a knowledge base, an inference about the video data, based in part on the particular object being identified as an attractor or repulsor and an orientation of the particular object relative to the one or more other objects, the inference indicative of a dangerous or urgent situation, wherein the knowledge base represents concepts and relationships among the concepts that are domain-agnostic prior to runtime and instantiated with attractor-specific or repulsor-specific predicates during the inference, and has not been trained with the video data; and

providing, by the device, data for display based on the inference, wherein the data comprises both the identification of the attractor or the repulsor of the particular object and a predicted future convergence or divergence event between the particular object and the one or more other objects.

18 . The method as in claim 1 , wherein the identifying the particular object as being an attractor or a repulsor includes:

analyzing motion vectors for a set of objects that are within a predetermined distance to the particular object.

19 . The apparatus as in claim 9 , wherein the identifying the particular object as being an attractor or a repulsor includes:

identifying a set of other objects within a radius of the particular object based on their respective tracklets;

generating, for each of the set of other objects, a relative velocity vector between the particular object and each of the set of other objects;

generating, for each of the set of other objects, an inter-tracklet vector between the particular object and each of the set of other objects; and

determining, based on a product of the relative velocity vector and the inter-tracklet vector for each of the set of other objects, whether each of the set of other objects is moving toward or away from the particular object.

20 . The tangible, non-transitory, computer-readable medium as in claim 17 , wherein the identifying the particular object as being an attractor or a repulsor includes:

identifying a set of other objects within a radius of the particular object based on their respective tracklets;

generating, for each of the set of other objects, a relative velocity vector between the particular object and each of the set of other objects;

generating, for each of the set of other objects, an inter-tracklet vector between the particular object and each of the set of other objects; and

determining, based on a product of the relative velocity vector and the inter-tracklet vector for each of the set of other objects, whether each of the set of other objects is moving toward or away from the particular object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 15, 2021
From: LATAPIE, HUGO; KILIC, OZKAN; LAWRENCE, ADAM JAMES; LIU, GAOWEN; PLETCHER, ANDREW ALBERT
To: CISCO TECHNOLOGY, INC.
Reel/Frame 058395/0169 →
Continuity (1)
Related Publication 20230186626A1 · Jun 15, 2023
References Cited (86)
US 10887197B2 · Fenoglio et al. · 2021 [cited by applicant]
US 10965516B2 · Fenoglio et al. · 2021 [cited by applicant]
US 20090153661A1 · Cheng · 2009 [cited by examiner]
US 20140372348A1 · Lehmann et al. · 2014 [cited by applicant]
US 20160140984A1 · Cecchi et al. · 2016 [cited by applicant]
US 20180374233A1 · Zhou · 2018 [cited by examiner]
US 20200118682A1 · Villazón-Terrazas · 2020 [cited by examiner]
US 20210042532A1 · Latapie et al. · 2021 [cited by applicant]
US 20210045360A1 · Harvey · 2021 [cited by examiner]
US 20210174155A1 · Smith et al. · 2021 [cited by applicant]
US 20210279615A1 · Latapie et al. · 2021 [cited by applicant]
US 20230052573A1 · Gnanasambandam · 2023 [cited by examiner]
US 20230059673A1 · Latapie · 2023 [cited by examiner]
CN 107480578A · 2017 [cited by applicant]
CN 110472604A · 2021 [cited by applicant]
WO WO2021011992 · 2021 [cited by applicant]
Estevam et al. (Zero-Shot Action Recognition in Videos: A Survey, pub. 2020), (Year: 2020). [cited by examiner]
Patel et al. Video Representation and Suspicious Event Detection using Semantic Technologies, pp. 1-25, Pub. 2020 (Year: 2020). [cited by examiner]
Patel et al. (Video Representation and Suspicious Event Detection using Semantic Technologies, published Mar. 9, 2021, pp. 1-25 ) (Year: 2021). [cited by examiner]
Estevam et al. (Zero-Shot Action Recognition in Videos—A Survey, published Nov. 2020, pp. 1-24) (Year: 2020). [cited by examiner]
Agrawal, et al., “VQA: Visual Question Answering”, Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2015, 25 pages, arXiv:1505.00468v7 [cs.CL]. [cited by applicant]
Aleksander, Igor, “Machine consciousness” In Scholarpedia. 3(2):4162, Oct. 21, 2011, 7 pages. [cited by applicant]
Anderson, et al., “Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2018, pp. 6077-6086, IEEE, Salt Lake Cit… [cited by applicant]
Baudrillard, Jean, “Simulacra and Simulation”, 1981, 159 pages, Galilee. [cited by applicant]
Baz, et al., “Context-aware hybrid classification system for fine-grained retail product recognition”, 2016 IEEE 12th Image, Video, and Multidimensional Signal Processing Workshop (IVMSP), Jul. 2016, 5 pages, IEEE, Bord… [cited by applicant]
Bělohlávek, Radim, “Concept lattices and order in fuzzy logic”, Annals of Pure and Applied Logic 128 (2004) 277-298, Elsevier. [cited by applicant]
Box, G. E. P., “Science and Statistics”, In Journal of the American Statistical Association, 71(356), Dec. 1976, pp. 791-799. [cited by applicant]
Chalmers, David J., “The Conscious Mind: In Search of a Fundamental Theory”, 1996, 433 pages, Oxford University Press, New York. [cited by applicant]
Chella, et al., “A cognitive framework for imitation learning”, Robotics and Autonomous Systems 54, Mar. 2006, pp. 403-408, Elsevier. [cited by applicant]
Chella, et al., “Artificial Consciousness”, Chapter 20, In Perception-Action Cycle, 2011, pp. 637-671, Springer, New York. [cited by applicant]
Chella, et al., “Machine Consciousness: A Manifesto for Robotics”, In International Journal of Machine Consciousness, 1(1), Jun. 2009, pp. 33-51, World Scientific Publishing Company. [cited by applicant]
Cohen, Paul R., “Projections as Concepts”, Computer Science Department Faculty Publication Series (194), https://scholarworks.umass.edu/cs/_faculty/_pubs/194, 1997, 6 pages, University of Massachusetts, Amherst. [cited by applicant]
Cui, et al., “A survey on network embedding”, IEEE Transactions on Knowledge and Data Engineering, vol. 31, Issue: 5, May 1, 2019, pp. 833-852, IEEE. [cited by applicant]
De Bono, Edward, “The Mechanism of Mind”, 1967, 276 pages, Penguin Books. [cited by applicant]
Düntsch, et al., “Modal-style operators in qualitative data analysis”, 2002 IEEE International Conference on Data Mining, 2002. Proceedings, Dec. 2002, pp. 155-162, IEEE, Maebashi City, Japan. [cited by applicant]
Franco, et al., “Grocery product detection and recognition”, Expert Systems With Applications 81 (2017), pp. 163-176, Elsevier Ltd. [cited by applicant]
George, et al., “Recognizing Products: A Per-exemplar Multi-label Image Classification Approach”, ECCV 2014, Part II, LNCS 8690, 2014, pp. 440-455, Springer International Publishing Switzerland. [cited by applicant]
Goertzel, et al., “CogPrime Architecture for Embodied Artificial General Intelligence”, 2013 IEEE Symposium on Computational Intelligence for Human-like Intelligence (CIHLI), Apr. 2013, pp. 60-67, IEEE, Singapore. [cited by applicant]
Goertzel, Ben, “OpenCogPrime: A Cognitive Synergy Based Architecture for Artificial General Intelligence”, 2009 8th IEEE International Conference on Cognitive Informatics, Jun. 2009, pp. 60-68, IEEE, Hong Kong, China. [cited by applicant]
Gorban, et al., “Blessing of dimensionality: mathematical foundations of the statistical physics of data”, Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 376.2118, Ja… [cited by applicant]
Grover, et al., “node2vec: Scalable Feature Learning for Networks”, KDD '16: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Aug. 2016, pp. 855-864, Association for Co… [cited by applicant]
Hamilton, et al., “Representation Learning on Graphs: Methods and Applications”, Bulletin of the IEEE Computer Society Technical Committee on Data Engineering, 2017, 23 pages, IEEE. [cited by applicant]
Hammer, et al., “A Reasoning Based Model for Anomaly Detection in the Smart City Domain”, IntelliSys 2020, AISC 1251, pp. 144-159, 2021, Springer Nature Switzerland AG. [cited by applicant]
Hobbs, Jerry R., “Granularity”, In Proceedings of the Ninth International Joint Conference on Artificial Intelligence, 1985, pp. 432-435, Morgan Kaufmann. [cited by applicant]
Horowitz, Alexandra, “Smelling themselves: Dogs investigate their own odours longer when modified in an “olfactory mirror” test”, Behavioural Processes, 2017, 41 pages. [cited by applicant]
Johnson, Mark, “The Body in The Mind”, 1987, 268 pages, The University of Chicago Press. [cited by applicant]
Kiryati, et al., “A probabilistic Hough transform”, Pattern Recognition. 24(4), 1991, pp. 303-316, The Pattern Recognition Society. [cited by applicant]
Korzybski, Alfred, “Manhood Of Humanity, The Science and Art of Human Engineering”, 1921, 240 pages, E. P. Dutton & Company, New York, NY. [cited by applicant]
Korzybski, Alfred, “Science and Sanity: An Introduction to Non-Aristotelian Systems and General Semantics”, 5th Edition, 1994, 910 pages, Institute of General Semantics, New York, NY. [cited by applicant]
Korzybski, Alfred, “Videos—This Is Not That”, online: https://www.thisisnotthat.com/korzybski-videos/, accessed Nov. 18, 2021, 7 pages. [cited by applicant]
Lakoff, G., “Women, Fire, and Dangerous Things”, 1984, 631 pages, University of Chicago Press. [cited by applicant]
Latapie, et al., “A Metamodel and Framework for Artificial General Intelligence From Theory to Practice”, Journal of Artificial Intelligence and Consciousness, Feb. 12, 2021, 1:30, 24 pages, World Scientific Publishing … [cited by applicant]
Li, et al., “Concept learning via granular computing: A cognitive viewpoint”, Information Sciences 298 (2015), Published Dec. 2014, pp. 447-467, Elsevier Inc. [cited by applicant]
Lieto, et al., “Conceptual Spaces for Cognitive Architectures: A Lingua Franca for Different Levels of Representation”, Biologically Inspired Cognitive Architectures 19, May 2017, 17 pages, Cognitive Robotics and Social… [cited by applicant]
Ma, et al., “Granular computing and Dual Galois Connection”, Information Sciences, 177(23), 2007, pp. 5365-5377, Elsevier Inc. [cited by applicant]
Macaulay, Thomas, “Facebook's chief AI scientist says GPT-3 is ‘not a very good’ Q&A system”, online: https://thenextweb.com/news/facebooks-yann-lecun-says-gpt-3-is-not-very-good-as-a-qa-or-dialog-system, Oct. 28, 2020,… [cited by applicant]
Murahari, et la., “Improving Generative Visual Dialog by Answering Diverse Questions”, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on… [cited by applicant]
Patel, et al., “Video Representation and Suspicious Event Detection Using Semantic Technologies”, online: http://semantic-web-journal.net/system/files/swj2427.pdf, Semantic Web 0, Sep. 10, 2020, accessed Aug. 9, 2021, 2… [cited by applicant]
Pauli, Wolfgang, “Part I. General: (A) theory. Some relations between electrochemical behaviour and the structure of colloids”, Jan. 1935, pp. 11-27, Transactions of the Faraday Society, vol. 1. [cited by applicant]
Scarselli, et al., “The Graph Neural Network Model”, IEEE Transactions on Neural Networks (vol. 20, Issue: 1, Jan. 2009), pp. 61-80, IEEE. [cited by applicant]
Speer, et al., “ConceptNet 5.5: An Open Multilingual Graph of General Knowledge”, online: https://arxiv.org/pdf/1612.03975.pdf, 2017, 9 pagers, Association for the Advancement of Artificial Intelligence. [cited by applicant]
Swanson, Bret, “The Exponential Internet”, online: https://www.uschamberfoundation.org/bhq/exponential-internet, accessed Nov. 19, 2021, 8 pages, The U.S. Chamber of Commerce Foundation. [cited by applicant]
Tan, et al., “EfficientDet: Scalable and Efficient Object Detection”, online: https://arxiv.org/pdf/1911.09070.pdf, Jul. 2020, 10 pages. [cited by applicant]
Taylor, J. G., “Codam: A neural network model of consciousness”, Neural Networks 20 (2007), pp. 983-992, Elsevier Ltd. [cited by applicant]
Thomopoulous, Stelios C.A., “Risk Assessment and Automated Anomaly Detection Using a Deep Learning Architecture”, online: https://www.intechopen.com/chapters/75329, accessed Dec. 14, 2021, 30 pages, IntechOpen. [cited by applicant]
Thórisson, et al., “Cumulative Learning”, Artificial General Intelligence—12th International Conference, AGI 2019, Proceedings, pp. 198-208, Springer. [cited by applicant]
Thórisson, Kristinn R., “A New Constructivist AI: From Manual Methods to Self-Constructive Systems”, Chapter 9, Apr. 2012, pp. 147-174, Atlantis Press Book. [cited by applicant]
Thórisson, Kristinn R., “Integrated AI Systems”, Minds & Machines 17, Mar. 2007, pp. 11-25. [cited by applicant]
Tonioni, et al., “Product recognition in store shelves as a sub-graph isomorphism problem”, online: https://arxiv.org/abs/1707.08378, Sep. 2017, 14 pages. [cited by applicant]
Unger, et al., “The Singular Universe and the Reality of Time: A Proposal in Natural Philosophy”, 2015, 558 pages, Cambridge University Press. [cited by applicant]
Unger, R. M. 2014. “Roberto Unger: Free Classical Social Theory from Illusions of False Necessity”, Online Lecture. 45 pages Retrieved on Nov. 22, 2021 from https://www.youtube.com/watch?v=yYOOwNRFTcY. [cited by applicant]
Wang, et al., “Concept Analysis via Rough Set and AFS Algebra”, Information Sciences 178 (2008), pp. 4125-4137, Elsevier Inc. [cited by applicant]
Wang, Pei, “Experience-grounded semantics: a theory for intelligent systems”, Aug. 2004, 33 pages, Elsevier Science. [cited by applicant]
Wang, Pei, “Insufficient Knowledge and Resources—A Biological Constraint and Its Functional Implications”, Biologically Inspired Cognitive Architectures II: Papers from the AAAI Fall Symposium (FS-09-01), 2009, pp. 188-… [cited by applicant]
Wang, Pei, “Non-axiomatic logic (nal) specification”, online: http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.167.2069&rep=rep1&type=pdf, Oct. 2009, 88 pages. [cited by applicant]
Wang, Pei, “On Defining Artificial Intelligence”, Journal of Artificial General Intelligence 10(2) 2019, pp. 1-37, Sciendo. [cited by applicant]
Wang, et al. “Self in NARS, an AGI System”, vol. 5, Article 20, Mar. 2018, 15 pages, Frontiers in Robotics and AI. [cited by applicant]
Wang, et al., “SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems”, online: https://arxiv.org/pdf/1905.00537.pdf, 2019, 29 pages, 33rd Conference on Neural Information Processing Systems … [cited by applicant]
Wikipedia, “Wheat and chessboard problem”, online: https://en.wikipedia.org/wiki/Wheat_and_chessboard_problem, Oct. 2021, 5 pages, Wikimedia Foundation, Inc. [cited by applicant]
Wille, Rudolf, “Restructuring Lattice Theory: An Approach Based on Heirarchies of Concepts”, I. Rival (Ed.), Ordered Sets, 1982, pp. 314-339. [cited by applicant]
Yao, et al., “A Granular Computing Paradigm for Concept Learning”, Emerging Paradigms in Machine Learning, Springer, London, pp. 307-326, 2012. [cited by applicant]
Yao, Y. Y., “Information Granulation and Rough Set Approximation”, International Journal of Intelligent Systems, vol. 16, No. 1, 87-104, 2001. [cited by applicant]
Yao, Y. Y., “Integrative levels of granularity”, Human-Centric Information Processing Through Granular Modelling, 2009, 20 pages, Studies in Computational Intelligence, vol. 182. Springer, Berlin, Heidelberg. [cited by applicant]
Ying, et al., “Graph convolutional neural networks for web-scale recommender systems”, online: https://arxiv.org/pdf/1806.01973.pdf, In KDD '18: The 24th ACM SIGKDD International Conference on Knowledge Discovery & Data… [cited by applicant]
Zhou, et al., “Graph neural networks: A review of methods and applications”, AI Open, 2020, pp. 57-81, Elsevier B.V. [cited by applicant]
Zhu, et al., “Describing Unseen Videos via Multi-modal Cooperative Dialog Agents” Computer Vision—ECCV 2020, 17 pages, Lecture Notes in Computer Science, vol. 12368. Springer. [cited by applicant]