IP Library › Granted Patent US 12,311,268
Granted Patent B2
US 12,311,268 · App. 17/983,046 · Granted May 27, 2025

Automated validation of video game environments

Inventors: Alessandro Sestini (Arezzo, IT); Linus Gisslén (Stockholm, SE); Joakim Bergdahl (Stockholm, SE)
Assignee: ELECTRONIC ARTS INC.
A63F13/67A63F13/69G06N3/045G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,311,268
App. No.
17/983,046
Granted
May 27, 2025
Kind
B2
Abstract

This specification provides a computer-implemented method comprising providing, by a user, one or more demonstrations of a video game entity interacting with a video game environment to achieve a goal. The method further comprises generating, from the one or more demonstrations, one or more training examples. The method further comprises training an agent to control the video game entity using a neural network. The training comprises generating one or more predicted actions for each training example by processing, using the neural network, input data derived from the training example, and updating parameters of the neural network based on a comparison between the one or more predicted actions of the training examples and the one or more corresponding target actions of the training examples. The method further comprises performing validation of the video game environment, comprising controlling the video game entity in the video game environment using the trained agent.

Claims (70)

1. A computer-implemented method comprising:

providing, by a user, one or more demonstrations of a video game entity interacting with a video game environment to achieve a goal, each of the demonstrations specifying one or more actions performed by the video game entity in the video game environment at each of one or more time steps;

generating, from the one or more demonstrations, one or more training examples, each training example associated with a time step of a demonstration and comprising: (i) a segmented map of an area of the video game environment comprising the video game entity, (ii) entity data for the video game entity, wherein the entity data comprises a representation of a state of the video game entity with respect to the goal, (iii) entity data for each of one or more other video game entities in the video game environment, and (iv) one or more target actions performed by the video game entity for the time step of the training example;

training an agent to control the video game entity using a neural network, comprising:

generating, for each training example, one or more predicted actions for the training example by processing, using the neural network, input data derived from the training example, wherein the segmented map of each training example comprises an integer-coded representation of the area of the video game environment comprising the video game entity, and the neural network comprises a segmented map processing portion comprising one or more convolutional neural network layers; and

updating parameters of the neural network based on a comparison between the one or more predicted actions of the training examples and the one or more corresponding target actions of the training examples; and

performing validation of the video game environment, comprising controlling the video game entity in the video game environment using the trained agent.

2. The method of claim 1 , wherein performing validation of the video game environment comprises:

providing, by the user, one or more corrected demonstrations in response to the agent controlling the video game entity in the video game environment using the updated neural network; and

further training the agent to control the video game entity using the one or more training examples and one or more corrected training examples generated from the one or more corrected demonstrations, comprising further updating parameters of the neural network.

3. The method of claim 1 , wherein performing validation of the video game environment comprises:

providing, by the user, a modification to the video game environment;

providing, by the user, one or more corrected demonstrations in response to the agent controlling the video game entity in the modified video game environment using the updated neural network; and

further training the agent to control the video game entity using the one or more training examples and one or more corrected training examples generated from the one or more corrected demonstrations, comprising further updating parameters of the neural network.

4. The method of claim 1 , wherein the input data derived from a training example comprises: (i) the segmented map of an area of the video game environment comprising the video game entity, (ii) the entity data for the video game entity, and (iii) the entity data for each of one or more other video game entities in the video game environment.

5. The method of claim 4 , wherein the neural network comprises a segmented map processing portion, an entity data processing portion, and a further entity data processing portion, and wherein processing, using the neural network, input data derived from the training example comprises:

processing the segmented map using the segmented map processing portion to generate a segmented map embedding;

processing the entity data for the video game entity using the entity data processing portion to generate an entity embedding;

processing, for each of the one or more other video game entities, the entity data for the other video game entity to generate another entity embedding;

combining, for each other entity embedding, the other entity embedding with the entity embedding to generate a combined entity embedding; and

processing the one or more combined entity embeddings and the segmented map embedding to generate an output representing one or more predicted actions for the training example.

6. The method of claim 5 , wherein processing the one or more combined entity embeddings and the segmented map embedding to generate the output representing one or more predicted actions for the training example comprises:

processing the one or more combined entity embeddings using an attention mechanism of the neural network to generate a single combined embedding; and

processing a combination of the single combined embedding and the segmented map embedding using an output portion of the neural network to generate the output.

7. The method of claim 6 , wherein processing the one or more combined entity embeddings using the attention mechanism of the neural network to generate a single combined embedding comprises processing the one or more combined entity embeddings using one or more transformer encoder layers of the neural network.

8. The method of claim 1 , wherein the integer-coded representation of the area of the video game environment comprising the video game entity is a two-dimensional integer coded representation or a three-dimensional integer-coded representation, and the one or more convolutional neural network layers are one or more two-dimensional convolutional neural network layers or one or more three-dimensional convolutional neural network layers, respectively.

9. The method of claim 1 , wherein the entity data comprises relative co-ordinate data based on the position of the video game entity in the video game environment and the position of a location corresponding to the goal.

10. The method of claim 1 , wherein the entity data further comprises attribute information for the video game entity.

11. A computing system comprising one or more computing devices configured to:

receive, from a user, one or more demonstrations of a video game entity interacting with a video game environment to achieve a goal, each of the demonstrations specifying one or more actions performed by the video game entity in the video game environment at each of one or more time steps;

generate, from the one or more demonstrations, one or more training examples, each training example associated with a time step of a demonstration and comprising: (i) a segmented map of an area of the video game environment comprising the video game entity, (ii) entity data for the video game entity, wherein the entity data comprises a representation of a state of the video game entity with respect to the goal, (iii) entity data for each of one or more other video game entities in the video game environment, and (iv) one or more target actions performed by the video game entity for the time step of the training example;

train an agent to control the video game entity using a neural network, comprising:

generating, for each training example, one or more predicted actions for the training example by processing, using the neural network, input data derived from the training example, wherein the segmented map of each training example comprises an integer-coded representation of the area of the video game environment comprising the video game entity, and the neural network comprises a segmented map processing portion comprising one or more convolutional neural network layers; and

updating parameters of the neural network based on a comparison between the one or more predicted actions of the training examples and the one or more corresponding target actions of the training examples; and

perform validation of the video game environment, comprising controlling the video game entity in the video game environment using the trained agent.

12. The computing system of claim 11 , wherein performing validation of the video game environment comprises:

providing, by the user, one or more corrected demonstrations in response to the agent controlling the video game entity in the video game environment using the updated neural network; and

further training the agent to control the video game entity using the one or more training examples and one or more corrected training examples generated from the one or more corrected demonstrations, comprising further updating parameters of the neural network.

13. The computing system of claim 11 , wherein performing validation of the video game environment comprises:

providing, by the user, a modification to the video game environment;

providing, by the user, one or more corrected demonstrations in response to the agent controlling the video game entity in the modified video game environment using the updated neural network; and

further training the agent to control the video game entity using the one or more training examples and one or more corrected training examples generated from the one or more corrected demonstrations, comprising further updating parameters of the neural network.

14. The computing system of claim 11 , wherein the input data derived from a training example comprises: (i) the segmented map of an area of the video game environment comprising the video game entity, (ii) the entity data for the video game entity, and (iii) the entity data for each of one or more other video game entities in the video game environment.

15. The computing system of claim 14 , wherein the neural network comprises a segmented map processing portion, an entity data processing portion, and a further entity data processing portion, and wherein processing, using the neural network, input data derived from the training example comprises:

processing the segmented map using the segmented map processing portion to generate a segmented map embedding;

processing the entity data for the video game entity using the entity data processing portion to generate an entity embedding;

processing, for each of the one or more other video game entities, the entity data for the other video game entity to generate another entity embedding;

combining, for each other entity embedding, the other entity embedding with the entity embedding to generate a combined entity embedding; and

processing the one or more combined entity embeddings and the segmented map embedding to generate an output representing one or more predicted actions for the training example.

16. A non-transitory computer-readable medium, which when executed by a processor, cause the processor to:

receive, from a user, one or more demonstrations of a video game entity interacting with a video game environment to achieve a goal, each of the demonstrations specifying one or more actions performed by the video game entity in the video game environment at each of one or more time steps;

generate, from the one or more demonstrations, one or more training examples, each training example associated with a time step of a demonstration and comprising: (i) a segmented map of an area of the video game environment comprising the video game entity, (ii) entity data for the video game entity, wherein the entity data comprises a representation of a state of the video game entity with respect to the goal, (iii) entity data for each of one or more other video game entities in the video game environment, and (iv) one or more target actions performed by the video game entity for the time step of the training example;

train an agent to control the video game entity using a neural network, comprising:

generating, for each training example, one or more predicted actions for the training example by processing, using the neural network, input data derived from the training example, wherein the segmented map of each training example comprises an integer-coded representation of the area of the video game environment comprising the video game entity, and the neural network comprises a segmented map processing portion comprising one or more convolutional neural network layers; and

updating parameters of the neural network based on a comparison between the one or more predicted actions of the training examples and the one or more corresponding target actions of the training examples; and

perform validation of the video game environment, comprising controlling the video game entity in the video game environment using the trained agent.

17. The non-transitory computer-readable medium of claim 16 , wherein performing validation of the video game environment comprises:

providing, by the user, one or more corrected demonstrations in response to the agent controlling the video game entity in the video game environment using the updated neural network; and

further training the agent to control the video game entity using the one or more training examples and one or more corrected training examples generated from the one or more corrected demonstrations, comprising further updating parameters of the neural network.

18. The non-transitory computer-readable medium of claim 16 , wherein performing validation of the video game environment comprises:

providing, by the user, a modification to the video game environment;

providing, by the user, one or more corrected demonstrations in response to the agent controlling the video game entity in the modified video game environment using the updated neural network; and

further training the agent to control the video game entity using the one or more training examples and one or more corrected training examples generated from the one or more corrected demonstrations, comprising further updating parameters of the neural network.

19. The non-transitory computer-readable medium of claim 16 , wherein the input data derived from a training example comprises: (i) the segmented map of an area of the video game environment comprising the video game entity, (ii) the entity data for the video game entity, and (iii) the entity data for each of one or more other video game entities in the video game environment.

20. The non-transitory computer-readable medium of claim 16 , wherein the neural network comprises a segmented map processing portion, an entity data processing portion, and a further entity data processing portion, and wherein processing, using the neural network, input data derived from the training example comprises:

processing the segmented map using the segmented map processing portion to generate a segmented map embedding;

processing the entity data for the video game entity using the entity data processing portion to generate an entity embedding;

processing, for each of the one or more other video game entities, the entity data for the other video game entity to generate another entity embedding;

combining, for each other entity embedding, the other entity embedding with the entity embedding to generate a combined entity embedding; and

processing the one or more combined entity embeddings and the segmented map embedding to generate an output representing one or more predicted actions for the training example.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 23, 2022
From: BERGDAHL, JOAKIM
To: ELECTRONIC ARTS INC.
Reel/Frame 062195/0087 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2022
From: SESTINI, ALESSANDRO; GISSLÉN, LINUS
To: ELECTRONIC ARTS INC.
Reel/Frame 061704/0363 →
Continuity (2)
Provisional Application 63397253 · Aug 11, 2022
Related Publication 20240050858A1 · Feb 15, 2024
References Cited (50)
US 20200269136A1 · Gurumurthy · 2020 [cited by examiner]
US 20210001229A1 · Somers · 2021 [cited by examiner]
US 20210346798A1 · Borovikov · 2021 [cited by examiner]
Stahlke, Samantha et al., Artificial Players in the Design Process: Developing an Automated Testing Tool for Game Level and World Design, CHI Play, ACM, 14 pages, dated Nov. 2, 2020. [cited by applicant]
Gordillo, Camilo et al., Improving Playtesting Coverage via Curiosity Driven Reinforcement Learning Agents, arXiv:2103.13798v2, 8 pages, dated Jun. 23, 2021. [cited by applicant]
Sestini, Alessandro et al., CCPT: Automatic Gameplay Testing and Validation with Curiosity-Conditioned Proximal Trajectories, arXiv:2202.10057v1, 11 pages, dated Feb. 21, 2022. [cited by applicant]
Gisslen, Linus et al., Adversarial Reinforcement Learning for Procedural Content Generation, arXiv:2103.04847v2, 8 pages, dated Jun. 10, 2021. [cited by applicant]
Ross, Stephane et al., A Reduction of Limitation Learning and Structured Prediction to No-Regret Online Learning, 14th International Conference on AISTATS, 9 pages, dated 2011. [cited by applicant]
Holmgard, Christoffer et al., Automated Playtesting with Procedural Personas through MCTS with Evolved Heuristics, arXiv:1802.06881v1, 10 pages, dated Feb. 19, 2018. [cited by applicant]
Chang, Kenneth et al., Reveal-More: Amplifying Human Effort in Quality Assurance Testing Using Automated Exploration, University of California, 8 pages, dated 2019. [cited by applicant]
Mugrai, Luvneesh et al., Automated Playtesting of Matching Tile Games, arXiv:1907.06570v1, 7 pages, dated Jul. 15, 2019. [cited by applicant]
Iftikhar, Sidra et al., An Automated Model Based Testing Approach for Platform Games, retrieved from: https://www.researchgate.net/profile/Muhammad-Uzair-Khan-2/publication/292150724_An_Automated_Model_Based_Testing_App… [cited by applicant]
Lovreto, Gabriel et al., Automated Tests for Mobile Games: an Experience Report, SBC Proceedings of SBGames, 9 pages, dated Nov. 1, 2018. [cited by applicant]
Cho, Chang-Sik et al., Online Game Testing Using Scenario-based Control of Massive Virtual Users, ICACT, 5 pages, dated Feb. 7, 2010. [cited by applicant]
Xiao, Gang et al., Software Testing by Active Learning for Commerical Games, AAAI, 6 pages, dated 2005. [cited by applicant]
Bergdahl, Joakim et al., Augmenting Automated Game Testing with Deep Reinforcement Learning, arXiv:2103.1519v1, 4 pages, dated Mar. 29, 2021. [cited by applicant]
Alonso, Eloi et al., Deep Reinforcement Learning for Navigation in AAA Video Games, arXiv:2011.04764v2, 13 pages, dated Nov. 17, 2020. [cited by applicant]
Devlin, Sam et al., Navigation Turing Test (NTT): Learning to Evaluate Human-Like Navigation, 38th International Conference on Machine Learning, 10 pages, dated 2021. [cited by applicant]
Zheng, Yan et al., Wuji: Automatic Online Combat Game Testing Using Evolutionary Deep Reinforcement Learning, Institutional Knowledge at Singapore Management University, 34th IEEE Conference on Automated Software Engine… [cited by applicant]
Agarwal, Shivam et al., Visualizing AI Playtesting Data of 2D Side-Scrolling Games, IEEE Conference on Games, 7 pages, dated Aug. 24, 2020. [cited by applicant]
Bain, Michael et al., A Framework for Behavioural Cloning, University of South Wales, 37 pages, dated Jul. 30, 2001. [cited by applicant]
Knox, W. Bradley et al., Interactively Shaping Agents via Human Reinforcement, K-CAP, 8 pages, dated 2009. [cited by applicant]
Ho, Jonathan et al., Generative Adversarial Imitation Learning, 30th Conference on Neural Information Processing Systems, 9 pages, dated 2016. [cited by applicant]
Finn, Chelsea et al., A Connection Between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models, arXiv:1611.03852v3, 10 pages, dated Nov. 25, 2016. [cited by applicant]
Peng, Xue Bin et al., AMP: Adversarial Motion Priors for Stylized Physics-Based Character Control, ACM Trans. Graph., vol. 40, No. 4, Article 144, 20 pages, dated Aug. 2021. [cited by applicant]
Zhao, Yunqi et al., Winning Isn't Everything: Enhancing Game Development with Intelligent Agents, arXiv:1903.10545v5, 14 pages, dated Apr. 28, 2020. [cited by applicant]
Harmer, Jack et al., Imitation Learning with Concurrent Actions in 3D Games, arXiv:1803.05402v5, 8 pages, dated Sep. 6, 2018. [cited by applicant]
Tucker, Aaron et al., Inverse Reinforcement Learning for Video Games, arXiv1810.10593v1, 10 pages, dated Oct. 24, 2018. [cited by applicant]
Bellemare, Marc G., The Arcade Learning Environment: An Evaluation Platform for General Agents, Journal of Artificial Intelligence, 27 pages, date Jun. 2013. [cited by applicant]
Schulman, John et al., Proximal Policy Optimization Algorithms, arXiv:1707.06347v2, 12 pages, dated Aug. 28, 2017. [cited by applicant]
Huang, Sandy, et al., Adversarial Attacks on Neural Network Policies, arXiv:1702.02284v1, 10 pages, dated Feb. 8, 2017. [cited by applicant]
Fu, Justin et al., Learning Robust Rewards with Adversarial Inverse Reinforcement Learning, arX1701.11248v2, 15 pages, dated Aug. 13, 2018. [cited by applicant]
Torabi, Faraz et al., Behavioral Cloning from Observation, Proceedings of the 27th International Joint Conference on Artificial Intelligence, arXiv:1805.01954v2, 8 pages, dated May 11, 2018. [cited by applicant]
Monteiro, Juarez et al., Augmented Behavioral Cloning from Observation, IEEE, arXiv:2004.13529v1, 8 pages, dated Apr. 28, 2020. [cited by applicant]
Xu, Tian et al., On Generalization of Adversarial Imitation Learning and Beyond, arXiv:2106.10424v3, 80 pages, dated Feb. 11, 2022. [cited by applicant]
Sinha, Samarth et al., S4RL: Surprisingly Simple Self-Supervision for Offline Reinforcement Learning in Robotics, 5th Conference on Robot Learning, 11 pages, dated 2021. [cited by applicant]
Fujimoto, Scott et al., A Minimalist Approach to Offline Reinforcement Learning, 35th Conference on Neural Information Processing Systems, 14 pages, dated 2021. [cited by applicant]
Woillemont, Pierre et al., Configurable Agent with Reward As Input: A Play-Style Continuum Generation, IEEE, arXiv:2211.16221v1, 8 pages, dated Nov. 29, 2022. [cited by applicant]
Roy, Julien et al., Direct Behavior Specification via Constrained Reinforcement Learning, 39th International Conference on Machine Learning, arXiv:2112.12228v6, 16 pages, dated Jun. 18, 2022. [cited by applicant]
Sestini, Alessandro et al., Policy Fusion for Adaptive and Customizable Reinforcement Learning Agents, arXiv:2104.10610v1, 8 pages, dated Apr. 21, 2021. [cited by applicant]
Burda, Yuri et al., Exploration by Random Network Distillation, arXiv:1810.12894v1, 17 pages, dated Oct. 30, 2018. [cited by applicant]
Druce, Jeff et al., Explainable Artificial Intelligence for Increasing User Trust in Deep Reinforcement Learning Driven Autonomous Systems, 33rd Conference on Neural Information Processing Systems, arXiv:2106.03775v1, 9… [cited by applicant]
Duan, Yan et al., One-Shot Imitation Learning, 31st Conference on Neural Information Processing Systems, 12 pages, dated 2017. [cited by applicant]
Hakhamaneshi, Kourosh et al., Hierarchical Few-Shot Imitation with Skill Transition Models, ICLR, arXiv:2107.08981v2, 19 pages, dated Mar. 10, 2022. [cited by applicant]
Finn, Chelsea et al., Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks, 34th International Conference on Machine Learning, 10 pages, dated 2017. [cited by applicant]
Kumar, Aviral et al., Conservative Q-Learning for Offline Reinforcement Learning, 34th Conference on Neural Information Processing Systems, 13 pages, dated 2020. [cited by applicant]
Chen, Lili et al., Decision Transformer: Reinforcement Learning via Sequence Modeling, 35th Conference on Neural Information Processing Systems, 14 pages, dated 2021. [cited by applicant]
Le, Hoang M. et al., Coordinated Multi-Agent Imitation Learning, 34th International Conference on Machine Learning, 9 pages, dated 2017. [cited by applicant]
Stahlke, Samantha et al., Artificial Playfulness: A Tool for Automated Agent-Based Playtesting, CHI, 7 pages, dated May 4, 2019. [cited by applicant]
Schaefer, Christopher et al., Crushinator: A Framework towards Game-Independent Testing, IEEE, 4 pages, dated 2013. [cited by applicant]