IP Library Granted Patent US 12,327,421
Granted Patent B2
US 12,327,421 · App. 18/160,860 · Granted Jun 10, 2025

Using automatically uncovered failure cases to improve the performance of neural networks

Inventors: Ethan Josean Perez (Allen, TX); Saffron Shan Huang (London, GB); Nathaniel John McAleese-Park (London, GB); Geoffrey Irving (London, GB)
Assignee: DeepMind Technologies Limited
G06V30/1916G06N3/092
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,327,421
App. No.
18/160,860
Granted
Jun 10, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for adjusting a target neural network using automatically generated test cases before deployment of the target neural network in a deployment environment. One of the methods may include generating a plurality of test inputs by using a test case generation neural network; processing the plurality of test inputs using a target neural network to generate one or more test outputs for each test input; and identifying, from the one or more test outputs generated by the target neural network for each test input, failing test inputs that result in generation of test outputs by the target neural network that fail one or more criteria.

Claims (43)

1. A method performed by one or more computers comprising:

generating a plurality of test inputs by using a test case generation neural network having a plurality of test case network parameters;

processing the plurality of test inputs using a target neural network having a plurality of target network parameters to generate one or more test outputs for each test input, wherein the target neural network is configured to process each test input in accordance with the plurality of target network parameters to generate a target network output that specifies a corresponding test output for the test input; and

identifying, from the one or more test outputs generated by the target neural network for each test input, failing test inputs that result in generation of test outputs by the target neural network that fail one or more criteria.

2. The method of claim 1 , wherein:

the test inputs and the test outputs each comprise text;

the test case generation neural network is a generative neural network pre-trained on a first natural language modeling task using first unlabeled text training data and configured to generate a test case network output by generating, at each of multiple test case network output time steps, a respective test case network score for each candidate token in a vocabulary of candidate tokens; and

the target neural network is another generative neural network pre-trained on a second natural language modeling task using second unlabeled text training data and configured to generate the target network output by generating, at each of multiple target network output time steps, a respective target network score for each candidate token in the vocabulary of candidate tokens.

3. The method of claim 2 , wherein:

the test case generation neural network and the target neural network have the same network architecture,

the first and second natural language modeling tasks are the same natural language modeling task, and

the first and second unlabeled text training data are the same unlabeled text training data.

4. The method of claim 1 , wherein generating each test input by using the test case generation neural network comprises:

generating, as a test case network input, a natural language prompt comprising text, numbers, punctuations, or a combination thereof;

processing, using the test case generation neural network and in accordance with the plurality of test case network parameters, the test case network input to generate the test case network output that includes the respective test case network score for each candidate token in the vocabulary of candidate tokens at each of the multiple test case network output time steps; and

generating the test input based on sampling tokens in accordance with the test case network scores included in the test case network output.

5. The method of claim 4 , wherein generating the test input based on sampling tokens in accordance with the test case network scores included in the test case network output comprises, at each of some of the multiple test case network output time steps:

sampling, as a token to be included in the test input, a sampled token from a subset of candidate tokens for which a test case network score that is greater than a certain threshold has been generated by the test case generation neural network.

6. The method of claim 4 , wherein generating the test input based on sampling tokens in accordance with the test case network scores included in the test case network output comprises:

repeatedly sampling different tokens to be included in different test inputs until a predetermined number of different test inputs that each satisfy a validity criterion have been generated.

7. The method of claim 1 , further comprising:

generating each of a plurality of additional test inputs by:

sampling one or more test inputs from the plurality of test inputs that have been generated by using the test case generation neural network,

generating another test case network input that comprises the one or more sampled test inputs,

processing, using the test case generation neural network and in accordance with the plurality of test case network parameters, the other test case network input to generate another test case network output, and

generating the additional test input based on sampling tokens in accordance with test case network scores included in the other test case network output; and

processing the plurality of additional test inputs using the target neural network to identify further failing test inputs.

8. The method of claim 7 , wherein sampling the one or more test inputs from the plurality of test inputs comprises:

sampling the one or more test inputs from a distribution that assigns a greater likelihood of sampling the failing test inputs than of sampling other test inputs from the plurality of test inputs that are not failing test inputs.

9. The method of claim 7 , further comprising, prior to generating each of the plurality of additional test inputs:

using a supervised learning technique to fine-tune pre-trained values of the plurality of test case network parameters to encourage the test case generation neural network to generate test case network outputs from which the failing test inputs are more likely to be generated.

10. The method of claim 9 , further comprising, prior to generating each of the plurality of additional test inputs further comprises:

using a reinforcement learning technique to (i) further adjust the fine-tuned values or (ii) fine-tune pre-trained values of the plurality of test case network parameters to encourage the test case generation neural network to generate test case network outputs from which the failing test inputs are more likely to be generated.

11. The method of claim 10 , wherein the reinforcement learning technique comprises an actor-critic technique that optimizes a reinforcement learning loss comprising a loss term dependent on a diversity of the test case network outputs.

12. The method of claim 1 , wherein the one or more criteria specify that the test outputs should not include one or more of: offensive content, misinformation, confidential or private information, text from any private training data that is used to pre-train the target neural network, or text from any protected training data that is used to pre-train the target neural network.

13. The method of claim 1 , wherein identifying the failing test inputs comprises:

processing each test output generated by the target neural network using a text classifier neural network to generate a predicted likelihood that the test output will fail the one or more criteria.

14. The method of claim 1 , wherein identifying the failing test inputs comprises:

processing each test output generated by the target neural network using a text-based classifier to determine whether the test output fails the one or more criteria.

15. The method of claim 14 , wherein the text-based classifier comprises a black-box text classifier or a deterministic text-based classification algorithm.

16. The method of claim 1 , further comprising using the identified failing test inputs to adjust the target neural network to encourage the target neural network to generate test outputs that are less likely to fail the one or more criteria.

17. The method of claim 16 , wherein using the identified failing test inputs to adjust the target neural network comprises using natural language processing techniques including text clustering techniques to analyze the identified failing test inputs to determine text segments that, when included in the test inputs, have highest probabilities in resulting in the generation of the test outputs that fail the one or more criteria by the target neural network.

18. The method of claim 16 , wherein using the identified failing test inputs to adjust the target neural network comprises one or more of: removing a particular training example from the unlabeled text training data, or adjusting a target network input for the target neural network before processing the target network input using the target neural network, including removing a first text segment from or adding a second text segment to the target network input.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2024
From: PEREZ, ETHAN JOSEAN; HUANG, SAFFRON SHAN; MCALEESE-PARK, NATHANIEL JOHN; IRVING, GEOFFREY
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 068287/0246 →
Continuity (2)
Provisional Application 63303958 · Jan 27, 2022
Related Publication 20230237826A1 · Jul 27, 2023
References Cited (11)
US 20200073788A1 · Saha · 2020 [cited by examiner]
US 20200097742A1 · Ratnesh Kumar · 2020 [cited by examiner]
US 20200226421A1 · Almazan · 2020 [cited by examiner]
US 20220129760A1 · Ravikumar · 2022 [cited by examiner]
JP 2018063504 · 2018 [cited by applicant]
JP 2019125014 · 2019 [cited by applicant]
JP 2020112915 · 2020 [cited by applicant]
Xu, “Bot adversarial dialogue for safe conversational agents”, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 2950-296… [cited by examiner]
Rae, “Scaling Language Models: Methods, Analysis & Insights from Training Gopher,” arXiv:2112.11446, Jan. 21, 2022. (Year: 2022). [cited by examiner]
Office Action in Japanese Appln. No. 202310303, dated Mar. 11, 2024, 7 pages (with English translation). [cited by applicant]
Decision to Grant Patent in Japanese Appln. No. 202310303, dated Sep. 9, 2024, 5 pages (with English translation). [cited by applicant]