IP Library Granted Patent US 12,670,211
Granted Patent B2
US 12,670,211 · App. 19/039,607 · Granted Jun 30, 2026

Determining query complexity in video question answering

Inventors: Cristobal Eyzaguirre (Stanford, CA); Igor Vasiljevic (Chicago, IL); Achal Dave (San Francisco, CA); Jiajun Wu (Stanford, CA); Thomas Kollar (San Jose, CA); Juan Carlos Niebles (Mountain View, CA); Pavel Tokmakov (West Hollywood, CA)
Assignees: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSHA KABUSHIKI KAISHA; THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERSITY
G06F16/7343G06F16/735
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,211
App. No.
19/039,607
Filed
Jan 28, 2025
Granted
Jun 30, 2026
Kind
B2
Art Unit
2161
USPC
707/769
Abstract

A method for determining a complexity of a natural language query includes converting a first natural language query into executable program code, the first natural language query being a query for a first video to be answered by one or more video question answering (VideoQA) models. The method also includes generating, via a complexity model, an abstract syntax tree (AST) based on the executable program code. The method further includes determining, via the complexity model, a complexity of the first natural language query based on quantity of subtrees, from a group of subtrees, that are present in the AST.

Claims (40)

1 . A method for determining a complexity of a natural language query, comprising:

converting a first natural language query into executable program code, the first natural language query being a query for a first video to be answered by one or more video question answering (VideoQA) models;

generating, via a complexity model, an abstract syntax tree (AST) based on the executable program code;

determining, via the complexity model, a complexity of the first natural language query based on a quantity of subtrees, from a group of subtrees, that are present in the AST;

generating, via a large language model, a group of queries associated with a second video based on a prompt and a natural language summary of the second video; and

selecting, via the complexity model, one or more queries of the group of queries, a respective complexity of each of the one or more queries being greater than or equal to a complexity threshold.

2 . The method of claim 1 , wherein the group of subtrees are subtrees that are common among a group of natural language queries.

3 . The method of claim 1 , wherein each subtree in the group of subtrees decreases a probability of the one or more VideoQA models correctly answering the first natural language query.

4 . The method of claim 1 , wherein:

determining the complexity further comprises encoding, via one-hot encoding, the AST into a vector based on the quantity of subtrees, from a group of subtrees, that are present in the AST;

the complexity model determines the complexity based on the vector.

5 . The method of claim 1 , further comprising generating the natural language summary via an image captioning model.

6 . The method of claim 1 , wherein the complexity is based on a likelihood of each of the one or more VideoQA models failing to answer the query.

7 . An apparatus for determining a complexity of a natural language query, comprising:

one or more processors; and

one or more memories coupled with the one or more processors and storing processor-executable code that, when executed by the one or more processors, is configured to cause the apparatus to:

convert a first natural language query into executable program code, the first natural language query being a query for a first video to be answered by one or more video question answering (VideoQA) models;

generate, via a complexity model, an abstract syntax tree (AST) based on the executable program code; and

determine, via the complexity model, a complexity of the first natural language query based on a quantity of subtrees, from a group of subtrees, that are present in the AST;

generate, via a large language model, a group of queries associated with a second video based on a prompt and a natural language summary of the second video; and

select, via the complexity model, one or more queries of the group of queries, a respective complexity of each of the one or more queries being greater than or equal to a complexity threshold.

8 . The apparatus of claim 7 , wherein the group of subtrees are subtrees that are common among a group of natural language queries.

9 . The apparatus of claim 7 , wherein each subtree in the group of subtrees decreases a probability of the one or more VideoQA models correctly answering the first natural language query.

10 . The apparatus of claim 7 , wherein:

execution of the processor-executable code to determine the complexity further causes the apparatus to encode, via one-hot encoding, the AST into a vector based on the quantity of subtrees, from a group of subtrees, that are present in the AST;

the complexity model determines the complexity based on the vector.

11 . The apparatus of claim 7 , wherein execution of the processor-executable code further causes the apparatus to generate the natural language summary via an image captioning model.

12 . The apparatus of claim 7 , wherein the complexity is based on a likelihood of each of the one or more VideoQA models failing to answer the query.

13 . A non-transitory computer-readable medium having program code recorded thereon for determining a complexity of a natural language query, the program code executed by one or more processors and comprising:

program code to convert a first natural language query into executable program code, the first natural language query being a query for a first video to be answered by one or more video question answering (VideoQA) models;

program code to generate, via a complexity model, an abstract syntax tree (AST) based on the executable program code; and

program code to determine, via the complexity model, a complexity of the first natural language query based on a quantity of subtrees, from a group of subtrees, that are present in the AST;

program code to generate, via a large language model, a group of queries associated with a second video based on a prompt and a natural language summary of the second video; and

program code to select, via the complexity model, one or more queries of the group of queries, a respective complexity of each of the one or more queries being greater than or equal to a complexity threshold.

14 . The non-transitory computer-readable medium of claim 13 , wherein the group of subtrees are subtrees that are common among a group of natural language queries.

15 . The non-transitory computer-readable medium of claim 13 , wherein each subtree in the group of subtrees decreases a probability of the one or more VideoQA models correctly answering the first natural language query.

16 . The non-transitory computer-readable medium of claim 13 , wherein:

the program code to determine the complexity further comprises program code to encode, via one-hot encoding, the AST into a vector based on the quantity of subtrees, from a group of subtrees, that are present in the AST;

the complexity model determines the complexity based on the vector.

17 . The non-transitory computer-readable medium of claim 13 , wherein the program code further comprises program code to generate the natural language summary via an image captioning model.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 11, 2026
From: TOYOTA RESEARCH INSTITUTE, INC.
To: TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 075616/0089 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 28, 2025
From: EYZAGUIRRE, CRISTOBAL; WU, JIAJUN; NIEBLES, JUAN CARLOS
To: THE BOARD OF TRUSTEES OF THE LELAND STANFORD
Reel/Frame 070038/0279 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 28, 2025
From: VASILJEVIC, IGOR; DAVE, ACHAL; KOLLAR, THOMAS; TOKMAKOV, PAVEL
To: TOYOTA RESEARCH INSTITUTE, INC.; TOYOTA JIDOSHA KABUSHIKI KAISHA
Reel/Frame 070038/0358 →
Continuity (2)
Provisional Application 63649876 · May 20, 2024
Related Publication 20250355934A1 · Nov 20, 2025
References Cited (25)
US 11157554B2 · Gan et al. · 2021 [cited by applicant]
US 12019679B2 · Kim · 2024 [cited by examiner]
US 12124440B1 · Romero Calvo · 2024 [cited by examiner]
US 12399892B2 · Cao · 2025 [cited by examiner]
US 20150261744A1 · Suenbuel · 2015 [cited by examiner]
US 20180144065A1 · Yellai · 2018 [cited by examiner]
US 20210191938A1 · Galitsky · 2021 [cited by examiner]
US 20220180056A1 · Hong · 2022 [cited by examiner]
US 20220343903A1 · Mostafazadeh · 2022 [cited by examiner]
US 20230325154A1 · Arcadinho · 2023 [cited by examiner]
US 20240004623A1 · Groenewegen · 2024 [cited by examiner]
US 20240061997A1 · Kobayashi · 2024 [cited by examiner]
US 20250217266A1 · Hicks · 2025 [cited by examiner]
CN 115391602A · 2022 [cited by applicant]
Yang Liu et al., “A Robust Multivariate Reranking Algorithm for Question Answering Enrichment”, IEEE, pp. 1917-1920 (Year: 2012). [cited by examiner]
Khushboo Khurana et al., “Video Question-Answering Techniques, Benchmark Datasets and Evaluation Metrics Leveraging Video Captioning: A Comprehensive Survey”, Feb. 9, 2021, vol. 9, pp. 43799-43823 (Year: 2021). [cited by examiner]
Tat-Seng Chua, “Question Answering on Large News Video Archive”, pp. 289-294 (Year: 2003). [cited by examiner]
Vivek Gupta et al., “VQuAD: Video Question Answering Diagnostic Dataset”, IEEE, pp. 282-291 (Year: 2022). [cited by examiner]
Suris, Didac et al., “ViperGPT: Visual Inference via Python Execution for Reasoning,” https://arxiv.org/abs/2303.08128, Mar. 14, 2023. [cited by applicant]
Stankov, Emil et al., “A New Tool for Calculation of a New Source Code Metric,” 12th International Conference on Informatics and Information Technology (CiiT 2015), pp. 25-30. [cited by applicant]
Xiao, Junbin et al., “NEXT-QA: Next Phase of Question-Answering to Explaining Temporal Actions,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 9772-9781. [cited by applicant]
Su, Hung-Ting et al., “End-to-End Video Question-Answer Generation with Generator-Pretester Network,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, 2021. [cited by applicant]
Feng, Yunhe et al., “Investigating Code Generation Performance of ChatGPT with Crowdsourcing Social Data,” IEEE 47th Annual Computers, Software, and Applications Conference (COMPSAC), 2023, pp. 876-885, 2023. [cited by applicant]
Dang, Long Hoang, et al., “Hierarchical Object-Oriented Spatio-Temporal Reasoning for Video Question Answering,” Proceedings of the Thirteenth International Joint Conference on Artificial Intelligence (IJCAI-21), pp. 63… [cited by applicant]
Feitelson, Dror G., “From Code Complexity Metrics to Program Comprehension,” Communications of the ACM, vol. 66, No. 5, pp. 52-61, May 2023. [cited by applicant]