IP Library › Granted Patent US 12,265,890
Granted Patent B2
US 12,265,890 · App. 17/115,941 · Granted Apr 1, 2025

Extracted model adversaries for improved black box attacks

Inventors: Naveen Jafer Nizar (Chennai, IN); Ariel Gedaliah Kobren (Cambridge, MA)
Assignee: Oracle International Corporation
G06N20/00G06F18/2113G06F18/214G06F18/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,265,890
App. No.
17/115,941
Granted
Apr 1, 2025
Kind
B2
Abstract

Techniques are described for identifying successful adversarial attacks for a black box reading comprehension model using an extracted white box reading comprehension model. The system trains a white box reading comprehension model that behaves similar to the black box reading comprehension model using the set of queries and corresponding responses from the black box reading comprehension model as training data. The system tests adversarial attacks, involving modified informational content for execution of queries, against the trained white box reading comprehension model. Queries used for successful attacks on the white box model may be applied to the black box model itself as part of a black box improvement process.

Claims (74)

1. One or more non-transitory machine-readable media storing instructions which, when executed by one or more processors, cause:

executing a first plurality of queries on a first model to obtain a first set of results corresponding to the first plurality of queries;

generating training data comprising the first plurality of queries and the first set of results corresponding to the first plurality of queries;

applying the training data to train a second model to generate a second set of results in response to a second plurality of queries, the second set of results meeting one or more similarity criteria to a third set of results for the second plurality of queries generated by the first model;

modifying informational content for query execution to include a first set of one or more adversarial perturbations;

executing a query on the second model to generate a fourth set of results based on the modified informational content comprising the first set of one or more adversarial perturbations;

determining that k highest ranked results in the fourth set of results are incorrect; and

responsive to determining that the k highest ranked results in the fourth set of results are incorrect, identifying the query for analysis or improvement of the first model.

2. The media of claim 1 , further comprising:

generating a plurality of confidence scores corresponding to the fourth set of results; and

using the confidence scores to rank the corresponding results from a highest confidence score to a lowest confidence score.

3. The media of claim 1 , wherein:

the first model comprises a black box model displaying a query interface and a response interface configured to display the responses to each of the queries of the first plurality of queries; and

the second model comprises a white box model that is a simulation of the black box model, the white box model displaying a second query interface configured to display the fourth set of results in response to the query and also display a first set of confidence intervals corresponding to the fourth set of results.

4. The media of claim 1 , wherein determining the k highest ranked results in the first set of results are incorrect comprises one or both of:

generating an F1 accuracy score indicating that each of the k highest ranked results is incorrect; and

responsive to comparing the plurality of k highest ranked results in the fourth set of results with the information content, determining an absence of exact matches there between.

5. The media of claim 1 , wherein the adversarial perturbations comprise a plurality of terms selected from the informational content or terms having a similarity score above a threshold with the plurality of terms selected from the informational content.

6. The media of claim 5 , wherein:

the plurality of terms are selected randomly from the informational content; and

appended to one of a beginning or an ending of the informational content.

7. The media of claim 1 , wherein the first model is a target black box model and the second model is an extracted white box model.

8. The media of claim 1 , wherein:

the first model comprises a black box model;

the second model comprises a white box model;

applying the training data to train the second model comprises training the white box model to simulate the black box model; and

the fourth set of results comprises a plurality of results and a plurality of confidence levels corresponding to the plurality of results.

9. The media of claim 1 , wherein:

the fourth set of results comprises a plurality of individual results and a plurality of confidence levels corresponding to the plurality of individual results; and

determining that the k highest ranked results in the fourth set of results are incorrect comprises:

ranking the plurality of individual results based on the corresponding confidence levels of the plurality of confidence levels.

10. One or more non-transitory machine-readable media storing instructions which, when executed by one or more processors, cause:

executing a first plurality of queries on a first model to obtain a first set of results corresponding to the first plurality of queries;

generating training data comprising the first plurality of queries and the first set of results corresponding to the first plurality of queries;

applying the training data to train a second model to generate a second set of results in response to a second plurality of queries, the second set of results meeting one or more similarity criteria to a third set of results for the second plurality of queries generated by the first model;

modifying informational content for query execution to include one or more adversarial perturbations;

executing a query on the second model to generate a fourth set of results based on the modified informational content comprising the one or more adversarial perturbations;

wherein the fourth set of results comprises a correct answer to the query; and

responsive to determining one of (1) a confidence level of the correct answer is below a threshold confidence value or (2) a ranking associated with the correct answer is below a threshold rank value:

identifying the query for analysis or improvement of the first model.

11. The media of claim 10 , wherein:

the first model comprises a black box model displaying a query interface and a response interface configured to display the responses to each of the queries of the first plurality of queries; and

the second model comprises a white box model that is a simulation of the black box model, the white box model displaying a second query interface configured to display the fourth set of results in response to the query and also display a first set of confidence intervals corresponding to the first set of results.

12. The media of claim 10 , wherein determining the fourth set of results comprises a correct answer comprises one or both of:

generating an F1 accuracy score indicating a correct result; and

responsive to comparing the correct answer with the information content, determining corresponding exact matches there between.

13. The media of claim 10 , wherein the adversarial perturbations comprise a plurality of terms selected from the informational content or terms having a similarity score above a threshold with the plurality of terms selected from the informational content.

14. The media of claim 13 , wherein:

the plurality of terms are selected randomly from the informational content; and

appended to one of a beginning or an ending of the informational content.

15. The media of claim 10 , wherein the first model is a target black box model and the second model is an extracted white box model.

16. One or more non-transitory machine-readable media storing instructions which, when executed by one or more processors, cause:

executing a first plurality of queries on a first model to obtain a first set of results corresponding to the first plurality of queries;

generating training data comprising the first plurality of queries and the first set of results corresponding to the first plurality of queries;

applying the training data to train a second model to generate a second set of results in response to a second plurality of queries, the second set of results meeting one or more similarity criteria to a third set of results for the second plurality of queries generated by the first model;

modifying informational content for query execution to include a first set of one or more adversarial perturbations;

executing a query on the second model to generate a fourth set of results based on the modified informational content comprising the first set of one or more adversarial perturbations;

determining none of a plurality of k highest ranked results in the fourth set of results are incorrect;

responsive to determining that none of the k highest ranked results in the fourth set of results are incorrect, modifying the first set of one or more adversarial perturbations to a second set of one or more adversarial perturbations to generate a revised modified information content; and

executing the query on the second model to generate a fifth set of results based on the revised modified informational content comprising the second set of one or more adversarial perturbations.

17. The media of claim 16 , further comprising:

generating a plurality of confidence scores corresponding to the fourth set of results; and

using the confidence scores to rank the corresponding results from a highest confidence score to a lowest confidence score.

18. The media of claim 16 , wherein:

the first model comprises a black box model displaying a query interface and a response interface configured to display the responses to each of the queries of the first plurality of queries; and

the second model comprises a white box model that is a simulation of the black box model, the white box model displaying a second query interface configured to display the fourth set of results in response to the query and also display a first set of confidence intervals corresponding to the fourth set of results.

19. The media of claim 16 , wherein determining the plurality of k highest ranked results in the fourth set of results are not incorrect comprises one or both of:

generating an F1 accuracy score indicating correct results; and

responsive to comparing the plurality of k highest ranked results in the fourth set of results with the information content, determining corresponding exact matches there between.

20. The media of claim 16 , wherein the adversarial perturbations comprise a plurality of terms selected from the informational content or terms having a similarity score above a threshold with the plurality of terms selected from the informational content.

21. The media of claim 20 , wherein:

the plurality of terms are selected randomly from the informational content; and

appended to one of a beginning or an ending of the informational content.

22. The media of claim 16 , wherein the first model is a target black box model and the second model is an extracted white box model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2020
From: NIZAR, NAVEEN JAFER; KOBREN, ARIEL GEDALIAH
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 054594/0843 →
Continuity (1)
Related Publication 20220051134A1 · Feb 17, 2022
References Cited (37)
Krishna, Kalpesh, et al. “Thieves on sesame street! model extraction of bert-based apis.” arXiv preprint arXiv:1910.12366 (2019). (Year: 2019). [cited by examiner]
Papernot, Nicolas, et al. “Practical black-box attacks against machine learning.” Proceedings of the 2017 ACM on Asia conference on computer and communications security. 2017. (Year: 2017). [cited by examiner]
Jia, Robin, and Percy Liang. “Adversarial examples for evaluating reading comprehension systems.” arXiv preprint arXiv: 1707.07328 (2017). (Year: 2017). [cited by examiner]
Cai, Jinghui, et al. “Accelerate black-box attack with white-box prior knowledge.” Intelligence Science and Big Data Engineering. Big Data and Machine Learning: 9th International Conference, IScIDE 2019, Nanjing, China,… [cited by examiner]
Alzantot et al., “Generating Natural Language Adversarial Examples”, CoRR, 2018, arXiv:1804.07998. [cited by applicant]
Belinkov et al., “Synthetic and Natural Noise Both Break Neural Machine Translation”, 2017, arXiv:1711.02173. [cited by applicant]
Botta, “Getting to know a black-box model”, 12 pages, 2018, https://towardsdatascience.com/getting-to-know-a-black-box-model-374e180589ce. [cited by applicant]
Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, 2019, arXiv:1810.04805. [cited by applicant]
Du et al., “A Hybrid Adversarial Attack for Different Application Scenarios†”, Applied Sciences, vol. 10, No. 10, https://doi.org/10.3390/app10103559. [cited by applicant]
Feng et al., “Pathologies of Neural Models Make Interpretations Difficult”, In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2018, pp. 3719-3728. [cited by applicant]
Francis et al., “Brown corpus manual”, Technical report, Department of Linguistics, Brown University, Providence, Rhode Island, US, 1979. http://korpus.uib.no/icame/manuals/BROWN/INDEX.HTM. [cited by applicant]
Hosseini et al., “Deceiving Google's Perspective API Built for Detecting Toxic Comments”, CoRR, 2017, arXiv:1702.08138. [cited by applicant]
Jansen et al., “Discourse Complements Lexical Semantics for Non-factoid Answer Reranking”, In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics, 2014, pp. 977-986. [cited by applicant]
Jia et al., “Adversarial Examples for Evaluating Reading Comprehension Systems”, CoRR, 2017, arXiv:1707.07328. [cited by applicant]
Jin et al., “Is BERT really robust? natural language attack on text classification and entailment”, CoRR, 2019, abs/1907.11932. [cited by applicant]
Krishna et al., “How to steal modern NLP systems with gibberish?”, 2020. http://www.cleverhans.io/2020/04/06/stealing-bert.html. [cited by applicant]
Krishna et al., “Thieves of sesame street: Model extraction on bert-based apis”, ICLR, 2020. [cited by applicant]
Li et al., “BERT-ATTACK: Adversarial Attack Against BERT Using BERT”, 2020, arXiv:2004.09984. [cited by applicant]
Li et al., “TextBugger: Generating Adversarial Text Against Real-world Applications”, CoRR, 2018, arXiv:1812.05271. [cited by applicant]
Manning et al., “The Stanford CoreNLP Natural Language Processing Toolkit”, Proceedings of 52nd Annual Meeting of the Association for Computational Linguistics: System Demonstrations, 2014, pp. 55-60. [cited by applicant]
Merity et al., “Pointer Sentinel Mixture Models”, CoRR, 2016, arXiv:1609.07843. [cited by applicant]
Miller, “WordNet: a lexical database for English”, Commun. ACM., vol. 38, No. 11, 1995, 39-41. [cited by applicant]
Papernot et al., “Practical Black-Box Attacks against Machine Learning”, In Proceedings of the 2017 ACM on Asia Conference on Computer and Communications Security, ASIA CCS '17, 2017, pp. 506-519. [cited by applicant]
Pennington et al., “GloVe: Global Vectors for Word Representation”, Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2014, pp. 1532-1543. [cited by applicant]
Radford et al., “Language Models are Unsupervised Multitask Learners”, 2019. [cited by applicant]
Rajpurkar et al., “SQuAD: 100,000+ Questions for Machine Comprehension of Text”, Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2016, pp. 2383-2392. [cited by applicant]
Ribeiro et al., “Semantically Equivalent Adversarial Rules for Debugging NLP Models”, In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (vol. 1: Long Papers), 2018, pp. 856-865. [cited by applicant]
Seo et al., “Bidirectional Attention Flow for Machine Comprehension”, ICLR, 2017, arXiv:1611.01603. [cited by applicant]
Shi et al., “Next Sentence Prediction helps Implicit Discourse Relation Classification within and across Domains”, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th Inter… [cited by applicant]
Szegedy et al., “Intriguing properties of neural networks”, 2014, arXiv:1312.6199. [cited by applicant]
Szyller et al., “PRADA: Protecting Against DNN Model Stealing Attacks”, IEEE European Symposium on Security and Privacy (EuroS&P), 2019. [cited by applicant]
Vaswani et al., “Attention is all you need,” In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS'17, 2017, pp. 6000-6010. [cited by applicant]
Vaz, “Adversarial Attack Strategies on Machine Reading Comprehension Models”, 2019. [cited by applicant]
Vijayaraghavan et al., “Generating Black-Box Adversarial Examples for Text Classifiers Using a Deep Reinforced Model”, Joint European Conference on Machine Learning and Knowledge Discovery in Databases, 2019, pp. 711-72… [cited by applicant]
Wallace et al., “Universal Adversarial Triggers for Attacking and Analyzing NLP”, In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on N… [cited by applicant]
Xie et al., “Unsupervised Data Augmentation for Consistency Training”, CoRR, 2019, arXiv:1904.12848. [cited by applicant]
Zhang et al., “Adversarial Attacks on Deep learning Models in Natural Language Processing: A Survey”, ACM Transactions on Intelligent Systems and Technology, 2019. [cited by applicant]