IP Library › Granted Patent US 12,353,981
Granted Patent B2
US 12,353,981 · App. 18/661,499 · Granted Jul 8, 2025

Training of large neural networks

Inventors: Slav Petrov (New York, NY); Yonghui Wu (Fremont, CA); Andrew M. Dai (San Francisco, CA); David Richard So (Brooklyn, NY); Dmitry Lepikhin (Menlo Park, CA); Erica Ann Moreira (Fremont, CA); Gaurav Mishra (Sunnyvale, CA); Jonathan Hudson Clark (Seattle, WA); Maxim Krikun (Castro Valley, CA); Melvin Jose Johnson Premkumar (Sunnyvale, CA); Nan Du (San Jose, CA); Orhan Firat (Mountain View, CA); Rohan Anil (San Francisco, CA); Siamak Shakeri (New York, NY); Xavier Garcia (New York, NY); Yanping Huang (Mountain View, CA); Yong Cheng (Mountain View, CA); Yuanzhong Xu (Mountain View, CA); Yujing Zhang (Sunnyvale, CA); Zachary Alexander Nado (Brookline, MA); Eric Jun Jie Ni (Mountain View, CA); Kefan Xiao (Sunnyvale, CA); Vladimir Feinberg (San Francisco, CA); Jin Young Sohn (Jersey City, NJ); Aurko Roy (San Francisco, CA)
Assignee: Google LLC
G06N3/0475G06F40/284G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,353,981
App. No.
18/661,499
Granted
Jul 8, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a neural network to perform any one or more of a variety of machine learning tasks. For example, the neural network can be configured as a generative neural network, e.g., an autoregressive generative neural network.

Claims (80)

1. A method performed by one or more computers, wherein the method comprises:

obtaining a plurality of unlabeled text sequences, wherein each unlabeled text sequence comprises a plurality of text tokens;

training an autoregressive generative neural network comprising one or more self-attention layers based on optimizing multiple different pre-training objective functions that comprise (i) a causal language modeling objective function and (ii) a prefix language modeling objective function, wherein training the autoregressive generative neural network based on optimizing the multiple different pre-training objective functions comprises:

obtaining data specifying a respective weight assigned to each of the multiple different pre-training objective functions; and

repeatedly (a) selecting, based on the respective weights, a pre-training objective function from the multiple different pre-training objective functions and (b) training the autoregressive generative neural network on the selected pre-training objective function,

wherein training the autoregressive generative neural network based on optimizing the causal language modeling objective function comprises:

generating, from the plurality of unlabeled text sequences, a plurality of causal language modeling text sequences, wherein generating each causal language modeling text sequence comprises using a corresponding unlabeled text sequence as the causal language modeling text sequence without further processing the corresponding unlabeled text sequence to add to the corresponding unlabeled text sequence any additional tokens that were not included in the corresponding unlabeled text sequence;

processing, using the autoregressive generative neural network, each causal language modeling text sequence to generate, for each token in the causal language modeling text sequence, a causal prediction of a text token that should occupy a particular position of the text token in the causal language modeling text sequence conditioned on text tokens at preceding positions in the causal language modeling text sequence, wherein the one or more self-attention layers within the autoregressive generative neural network apply a masked self-attention mechanism over the preceding positions in the causal language modeling text sequence; and

determining, based on a quality of the causal predictions, an update to parameter values of the autoregressive generative neural network, and

wherein training the autoregressive generative neural network based on optimizing the prefix language modeling objective function comprises:

generating, from the plurality of unlabeled text sequences, a plurality of prefix language modeling text sequences, wherein generating each prefix language modeling text sequence comprises further processing a corresponding unlabeled text sequence to divide the corresponding unlabeled text sequence into a prefix text sequence and a suffix text sequence that follows the prefix text sequence;

processing, using the autoregressive generative neural network, each prefix language modeling text sequence to generate, for each token in the suffix text sequence, a prefix prediction of a text token that should occupy a particular position of the token in the suffix text sequence conditioned on tokens in the prefix text sequence and tokens at any preceding positions in the suffix text sequence, wherein the one or more self-attention layers within the autoregressive generative neural network applies a bidirectional, unmasked attention mechanism over the positions in the prefix text sequence and applies a masked self-attention mechanism over positions in the suffix text sequence so that each position in the suffix text sequence attend over the positions in the prefix text sequence and any preceding positions in the suffix text sequence; and

determining, based on a quality of the prefix predictions, an update to the parameter values of the autoregressive generative neural network.

2. The method of claim 1 , wherein the multiple different pre-training objective functions comprise (iii) a span corruption objective function, and wherein training the autoregressive generative neural network based on optimizing the span corruption objective function comprises:

generating, from the plurality of unlabeled text sequences, a plurality of span masked text sequences, wherein each span masked text sequence comprises a plurality of text tokens separated by one or more mask tokens, and wherein generating each span masked text sequence comprises further processing a corresponding unlabeled text sequence to replace one or more text tokens that were included in the corresponding unlabeled text sequence with the one or more mask tokens that were not included in the corresponding unlabeled text sequence;

processing, using the autoregressive generative neural network, each span masked text sequence to generate a span prediction of the one or more text tokens that should occupy respective positions of the one or more mask tokens in the span masked text sequence; and

determining, based on a quality of the span prediction, an update to the parameter values of the autoregressive generative neural network.

3. The method of claim 1 , further comprising:

sampling a batch of unlabeled text sequences; and

processing the batch of unlabeled text sequences according to the selected pre-training objective function.

4. The method of claim 1 , wherein repeatedly selecting the pre-training objective function from the multiple different pre-training objective functions comprises:

sampling a batch of unlabeled text sequences;

selecting, based on the respective weights, a first pre-training objective function from the multiple different pre-training objective functions;

processing a first subset of the batch of unlabeled text sequences according to the first selected pre-training objective function;

selecting, based on the respective weights, a second pre-training objective function from the multiple different pre-training objective functions; and

processing a second subset of the batch of unlabeled text sequences according to the second selected pre-training objective function.

5. The method of claim 1 , wherein the autoregressive generative neural network is a decoder-only attention neural network.

6. The method of claim 1 , wherein the autoregressive generative neural network is an encoder-decoder attention neural network.

7. The method of claim 1 , further comprising:

after the training, adapting the autoregressive generative neural network to perform one or more downstream tasks.

8. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one more computers to perform operations comprising:

obtaining a plurality of unlabeled text sequences, wherein each unlabeled text sequence comprises a plurality of text tokens;

training an autoregressive generative neural network comprising one or more self-attention layers based on optimizing multiple different pre-training objective functions that comprise (i) a causal language modeling objective function and (ii) a prefix language modeling objective function, wherein training the autoregressive generative neural network based on optimizing the multiple different pre-training objective functions comprises:

obtaining data specifying a respective weight assigned to each of the multiple different pre-training objective functions; and

repeatedly (a) selecting, based on the respective weights, a pre-training objective function from the multiple different pre-training objective functions and (b) training the autoregressive generative neural network on the selected pre-training objective function,

wherein training the autoregressive generative neural network based on optimizing the causal language modeling objective function comprises:

generating, from the plurality of unlabeled text sequences, a plurality of causal language modeling text sequences, wherein generating each causal language modeling text sequence comprises using a corresponding unlabeled text sequence as the causal language modeling text sequence without further processing the corresponding unlabeled text sequence to add to the corresponding unlabeled text sequence any additional tokens that were not included in the corresponding unlabeled text sequence;

processing, using the autoregressive generative neural network, each causal language modeling text sequence to generate, for each token in the causal language modeling text sequence, a causal prediction of a text token that should occupy a particular position of the text token in the causal language modeling text sequence conditioned on text tokens at preceding positions in the causal language modeling text sequence, wherein the one or more self-attention layers within the autoregressive generative neural network apply a masked self-attention mechanism over the preceding positions in the causal language modeling text sequence; and

determining, based on a quality of the causal predictions, an update to parameter values of the autoregressive generative neural network, and

wherein training the autoregressive generative neural network based on optimizing the prefix language modeling objective function comprises:

generating, from the plurality of unlabeled text sequences, a plurality of prefix language modeling text sequences, wherein generating each prefix language modeling text sequence comprises further processing a corresponding unlabeled text sequence to divide the corresponding unlabeled text sequence into a prefix text sequence and a suffix text sequence that follows the prefix text sequence;

processing, using the autoregressive generative neural network, each prefix language modeling text sequence to generate, for each token in the suffix text sequence, a prefix prediction of a text token that should occupy a particular position of the token in the suffix text sequence conditioned on tokens in the prefix text sequence and tokens at any preceding positions in the suffix text sequence, wherein the one or more self-attention layers within the autoregressive generative neural network applies a bidirectional, unmasked attention mechanism over the positions in the prefix text sequence and applies a masked self-attention mechanism over positions in the suffix text sequence so that each position in the suffix text sequence attend over the positions in the prefix text sequence and any preceding positions in the suffix text sequence; and

determining, based on a quality of the prefix predictions, an update to the parameter values of the autoregressive generative neural network.

9. The system of claim 8 , wherein the multiple different pre-training objective functions comprise (iii) a span corruption objective function, and wherein training the autoregressive generative neural network based on optimizing the span corruption objective function comprises:

generating, from the plurality of unlabeled text sequences, a plurality of span masked text sequences, wherein each span masked text sequence comprises a plurality of text tokens separated by one or more mask tokens, and wherein generating each span masked text sequence comprises further processing a corresponding unlabeled text sequence to replace one or more text tokens that were included in the corresponding unlabeled text sequence with the one or more mask tokens that were not included in the corresponding unlabeled text sequence;

processing, using the autoregressive generative neural network, each span masked text sequence to generate a span prediction of the one or more text tokens that should occupy respective positions of the one or more mask tokens in the span masked text sequence; and

determining, based on a quality of the span prediction, an update to the parameter values of the autoregressive generative neural network.

10. The system of claim 8 , wherein the operations further comprise:

sampling a batch of unlabeled text sequences; and

processing the batch of unlabeled text sequences according to the selected pre-training objective function.

11. The system of claim 8 , wherein repeatedly selecting the pre-training objective function from the multiple different pre-training objective functions comprises:

sampling a batch of unlabeled text sequences;

selecting, based on the respective weights, a first pre-training objective function from the multiple different pre-training objective functions;

processing a first subset of the batch of unlabeled text sequences according to the first selected pre-training objective function;

selecting, based on the respective weights, a second pre-training objective function from the multiple different pre-training objective functions; and

processing a second subset of the batch of unlabeled text sequences according to the second selected pre-training objective function.

12. The system of claim 8 , wherein the autoregressive generative neural network is a decoder-only attention neural network.

13. The system of claim 8 , wherein the autoregressive generative neural network is an encoder-decoder attention neural network.

14. The system of claim 8 , the operations further comprising:

after the training, adapting the autoregressive generative neural network to perform one or more downstream tasks.

15. One or more computer storage media storing instructions that when executed by one or more computers cause the one more computers to perform operations comprising:

obtaining a plurality of unlabeled text sequences, wherein each unlabeled text sequence comprises a plurality of text tokens;

training an autoregressive generative neural network comprising one or more self-attention layers based on optimizing multiple different pre-training objective functions that comprise (i) a causal language modeling objective function and (ii) a prefix language modeling objective function, wherein training the autoregressive generative neural network based on optimizing the multiple different pre-training objective functions comprises:

obtaining data specifying a respective weight assigned to each of the multiple different pre-training objective functions; and

repeatedly (a) selecting, based on the respective weights, a pre-training objective function from the multiple different pre-training objective functions and (b) training the autoregressive generative neural network on the selected pre-training objective function,

wherein training the autoregressive generative neural network based on optimizing the causal language modeling objective function comprises:

generating, from the plurality of unlabeled text sequences, a plurality of causal language modeling text sequences, wherein generating each causal language modeling text sequence comprises using a corresponding unlabeled text sequence as the causal language modeling text sequence without further processing the corresponding unlabeled text sequence to add to the corresponding unlabeled text sequence any additional tokens that were not included in the corresponding unlabeled text sequence;

processing, using the autoregressive generative neural network, each causal language modeling text sequence to generate, for each token in the causal language modeling text sequence, a causal prediction of a text token that should occupy a particular position of the text token in the causal language modeling text sequence conditioned on text tokens at preceding positions in the causal language modeling text sequence, wherein the one or more self-attention layers within the autoregressive generative neural network apply a masked self-attention mechanism over the preceding positions in the causal language modeling text sequence; and

determining, based on a quality of the causal predictions, an update to parameter values of the autoregressive generative neural network, and

wherein training the autoregressive generative neural network based on optimizing the prefix language modeling objective function comprises:

generating, from the plurality of unlabeled text sequences, a plurality of prefix language modeling text sequences, wherein generating each prefix language modeling text sequence comprises further processing a corresponding unlabeled text sequence to divide the corresponding unlabeled text sequence into a prefix text sequence and a suffix text sequence that follows the prefix text sequence;

processing, using the autoregressive generative neural network, each prefix language modeling text sequence to generate, for each token in the suffix text sequence, a prefix prediction of a text token that should occupy a particular position of the token in the suffix text sequence conditioned on tokens in the prefix text sequence and tokens at any preceding positions in the suffix text sequence, wherein the one or more self-attention layers within the autoregressive generative neural network applies a bidirectional, unmasked attention mechanism over the positions in the prefix text sequence and applies a masked self-attention mechanism over positions in the suffix text sequence so that each position in the suffix text sequence attend over the positions in the prefix text sequence and any preceding positions in the suffix text sequence; and

determining, based on a quality of the prefix predictions, an update to the parameter values of the autoregressive generative neural network.

16. The computer storage media of claim 15 , wherein the multiple different pre-training objective functions comprise (iii) a span corruption objective function, and wherein training the autoregressive generative neural network based on optimizing the span corruption objective function comprises:

generating, from the plurality of unlabeled text sequences, a plurality of span masked text sequences, wherein each span masked text sequence comprises a plurality of text tokens separated by one or more mask tokens, and wherein generating each span masked text sequence comprises further processing a corresponding unlabeled text sequence to replace one or more text tokens that were included in the corresponding unlabeled text sequence with the one or more mask tokens that were not included in the corresponding unlabeled text sequence;

processing, using the autoregressive generative neural network, each span masked text sequence to generate a span prediction of the one or more text tokens that should occupy respective positions of the one or more mask tokens in the span masked text sequence; and

determining, based on a quality of the span prediction, an update to the parameter values of the autoregressive generative neural network.

17. The computer storage media of claim 15 , wherein the operations further comprise:

sampling a batch of unlabeled text sequences; and

processing the batch of unlabeled text sequences according to the selected pre-training objective function.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 19, 2024
From: PETROV, SLAV; WU, YONGHUI; DAI, ANDREW M.; SO, DAVID RICHARD; LEPIKHIN, DMITRY; MOREIRA, ERICA ANN; MISHRA, GAURAV; CLARK, JONATHAN HUDSON; KRIKUN, MAXIM; JOHNSON PREMKUMAR, MELVIN JOSE; DU, NAN; FIRAT, ORHAN; ANIL, ROHAN; SHAKERI, SIAMAK; GARCIA, XAVIER; HUANG, YANPING; CHENG, YONG; XU, YUANZHONG; ZHANG, YUJING; NADO, ZACHARY ALEXANDER; NI, ERIC JUN JIE; XIAO, KEFAN; FEINBERG, VLADIMIR; SOHN, JIN YOUNG; ROY, AURKO
To: GOOGLE LLC
Reel/Frame 068330/0851 →
Continuity (2)
Provisional Application 63465487 · May 10, 2023
Related Publication 20240378427A1 · Nov 14, 2024
References Cited (196)
US 10885436B1 · Saleh · 2021 [cited by examiner]
US 11055639B1 · Cay · 2021 [cited by examiner]
US 11941356B2 · Liu · 2024 [cited by examiner]
US 20210174215A1 · Chan · 2021 [cited by examiner]
US 20210232773A1 · Wang · 2021 [cited by examiner]
US 20220092416A1 · Houlsby · 2022 [cited by examiner]
US 20220164626A1 · Bird · 2022 [cited by examiner]
US 20220382527A1 · Wang · 2022 [cited by examiner]
US 20230034401A1 · Weston · 2023 [cited by examiner]
US 20230090148A1 · Gamble, IV · 2023 [cited by examiner]
US 20230124177A1 · Jayakumar · 2023 [cited by examiner]
US 20230153546A1 · Peleg · 2023 [cited by examiner]
US 20230195066A1 · Ramanasankaran · 2023 [cited by examiner]
US 20230237773A1 · Li · 2023 [cited by examiner]
US 20230244938A1 · Wei · 2023 [cited by examiner]
US 20230281400A1 · Wang · 2023 [cited by examiner]
US 20230316001A1 · Araki · 2023 [cited by examiner]
US 20230325725A1 · Lester · 2023 [cited by examiner]
US 20230376676A1 · Mallinson · 2023 [cited by examiner]
US 20230376841A1 · Le · 2023 [cited by examiner]
US 20240061835A1 · Subramanian · 2024 [cited by examiner]
US 20240095275A1 · Tambi · 2024 [cited by examiner]
Devlin et al., BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, 2019, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: H… [cited by examiner]
Tay et al., Unifying Language Learning Paradigms, 2022, arXiv:2205.05131 [cs.CL], pp. 1-32 (Year: 2022). [cited by examiner]
Dong et al., Unified Language Model Pre-training for Natural Language Understanding and Generation, 2019, arXiv:1905.03197 [cs.CL], pp. 1-14 (Year: 2019). [cited by examiner]
Liu et al., FCM: Forgetful Causal Masking Makes Causal Language Models Better Zero-Shot Learners, 2022, arXiv:2210.13432, pp. 1-16 (Year: 2022). [cited by examiner]
Abid et al., “Persistent anti-muslim bias in large language models,” arXiv:2101.05783, Jan. 2021, 17 pages. [cited by applicant]
Adiwardana et al., “Towards a human-like open-domain chatbot,” arXiv:2001.09977, Jan. 2020, updated Feb. 2020, 38 pages. [cited by applicant]
Ai.google [online], “Our Principles,” 2018, retrieved on Jun. 10, 2024, retrieved from URL <https://ai.google/responsibility/principles/>, 6 pages. [cited by applicant]
Ai.google.dev [online], “PaLM API and MakerSuite Additional Terms of Service,” last modified Aug. 28, 2023, retrieved on Jun. 10, 2024, retrieved from URL <https://ai.google.dev/gemini-api/terms-archive/terms_08_28_23>,… [cited by applicant]
Akhbardeh et al., “Findings of the 2021 conference on machine translation (WMT21),” Proceedings of the Sixth Conference on Machine Translation, Jan. 2021, 88 pages. [cited by applicant]
Alayrac et al., “Flamingo: a Visual Language Model for Few-Shot Learning,” Advances in Neural Information Processing Systems, Oct. 2022, 35:23716-23736, 21 pages. [cited by applicant]
Austin et al., “Program synthesis with large language models,” arXiv: 2108.07732, Aug. 2021, 34 pages. [cited by applicant]
Bai et al., “Training a helpful and harmless assistant with reinforcement learning from human feedback,” arXiv: 2204.05862, Apr. 2022, 74 pages. [cited by applicant]
Bapna et al., “Building machine translation systems for the next thousand languages,” arXiv: 2205.03983, May 2022, 77 pages. [cited by applicant]
Barham et al., “Pathways: Asynchronous distributed dataflow for ml,” Proceedings of Machine Learning and Systems, Mar. 2022, 430-449. [cited by applicant]
Barocas et al., “Designing disaggregated evaluations of AI systems: Choices, considerations, and tradeoffs,” Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, Jul. 2021, 368-378. [cited by applicant]
Barocas et al., “Limitations and opportunities,” Fairness and Machine Learning, 2017, 294 pages. [cited by applicant]
Bender et al., “Data statements for natural language processing: Toward mitigating system bias and enabling better science,” Transactions of the Association for Computational Linguistics, Dec. 2018, 6: 587-604. [cited by applicant]
Berant et al., “Semantic parsing on Freebase from question-answer pairs,” Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, Oct. 2013, 1533-1544. [cited by applicant]
Bhatt et al., “Re-contextualizing fairness in NLP: The case of India,” arXiv:2209.12226, Sep. 2022, updated Nov. 2022, 14 pages. [cited by applicant]
Bisk et al., “PIQA: Reasoning about Physical Commonsense in Natural Language,” Proceedings of the AAAI Conference on Artificial Intelligence, Apr. 2020, 7432-7439. [cited by applicant]
Blodgett et al., “Language (Technology) is Power: A Critical Survey of “Bias” in NLP,”arXiv: 2005.14050, May 2020, 23 pages. [cited by applicant]
Blodgett et al., “Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets,” Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Internation… [cited by applicant]
Borkan et al., “Nuanced metrics for measuring unintended bias with real data for text classification,” arXiv: 1903.04561, Mar. 2019, updated May 2019, 10 pages. [cited by applicant]
Bowman et al., “What Will it Take to Fix Benchmarking in Natural Language Understanding?,” Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Languag… [cited by applicant]
Brown et al., “Language models are few-shot learners,” Advances in Neural Information Processing Systems, Jul. 2020, 1877-1901. [cited by applicant]
Carlini et al., “Extracting training data from large language models,” USENIX Security Symposium, Aug. 11-13, 2021, 19 pages. [cited by applicant]
Carlini et al., “Quantifying Memorization Across Neural Language Models,” arXiv:2202.07646, Feb. 2022, updated Mar. 2023, 19 pages. [cited by applicant]
Carlini et al., “The secret sharer: Evaluating and testing unintended memorization in neural networks,” USENIX Security Symposium, Aug. 14-16, 2019, 19 pages. [cited by applicant]
Casad et al., “Stereotype Threat Among Girls: Differences by Gender Identity and Math Education Context,” Psychology of Women Quarterly, Dec. 2017, 41(4):513-529, 18 pages. [cited by applicant]
Chen et al., “Evaluating large language models trained on code,” arXiv: 2107.03374, Jul. 2021, 35 pages. [cited by applicant]
Chen et al., “Microsoft COCO Captions: Data Collection and Evaluation Server,” arXiv:1504.00325, Apr. 2015, 1-7. [cited by applicant]
Chen et al., “PaLI: A Jointly-Scaled Multilingual Language-Image Model,” arXiv:2209.06794, Sep. 2022, updated Jun. 2023, 33 pages. [cited by applicant]
Chen et al., “Question Directed Graph Attention Network for Numerical Reasoning over Text,” Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Nov. 2020, 6759-6768. [cited by applicant]
Chowdhery et al., “PaLM: Scaling language modeling with Pathways,” arXiv:2204.02311, Apr. 2022, updated Oct. 2022, 87 pages. [cited by applicant]
Chung et al., “Scaling instruction-finetuned language models,” arXiv:2210.11416, Dec. 2022, 54 pages. [cited by applicant]
Clark et al., “Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge,” arXiv: 1803.05457, Mar. 2018, 24 pages. [cited by applicant]
Clark et al., “TyDiQA: A benchmark for information-seeking question answering in typologically diverse languages,” TACL, Jul. 2020, 17 pages. [cited by applicant]
Cobbe et al., “Training Verifiers to Solve Math Word Problems,” arXiv:2110.14168, Oct. 2021, updated Nov. 2021, 22 pages. [cited by applicant]
Crenshaw, “Demarginalizing the intersection of race and sex: A black feminist critique of antidiscrimination doctrine, feminist theory and antiracist politics,” Columbia Law School Scholarship Archive, 1989, 30 pages. [cited by applicant]
Dai et al., “Semi-supervised sequence learning,” Advances in Neural Information Processing Systems, Nov. 2015, 10 pages. [cited by applicant]
Denton et al., “Bringing the People Back In: Contesting Benchmark Machine Learning Datasets,” arXiv: 2007.07399, Jul. 2020, 6 pages. [cited by applicant]
Dev et al., “Harms of Gender Exclusivity and Challenges in Non-Binary Representation in Language Technologies,” arXiv: 2108.12084, Aug. 2021, 26 pages. [cited by applicant]
Dev et al., “On Measures of Biases and Harms in NLP,” arXiv: 2108.03362, Aug. 2021, 16 pages. [cited by applicant]
Dev et al., “On measures of biases and harms in NLP,” arXiv: 2108.03362, Aug. 2021, updated Oct. 2022, 23 pages. [cited by applicant]
Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human… [cited by applicant]
Diaz et al., “CrowdWorkSheets: Accounting for Individual and Collective Identities Underlying Crowdsourced Dataset Annotation,” 2022 ACM Conference on Fairness, Accountability, and Transparency, retrieved from URL <http… [cited by applicant]
Dinan et al., “Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack,” Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Jo… [cited by applicant]
Dodge et al., “Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus,” arXiv: 2104.08758, Aug. 2021, updated Sep. 2021, 20 pages. [cited by applicant]
Du et al., “GLaM: Efficient Scaling of Language Models with Mixture-of-Experts,” arXiv: 2112.06905, Dec. 2021, updated Aug. 2022, 23 pages. [cited by applicant]
Dua et al., “DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs,” Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Huma… [cited by applicant]
Feldman, “Does Learning Require Memorization? A Short Tale about a Long Tail,” arXiv:1906.05271, Jun. 2019, updated Jan. 2021, 38 pages. [cited by applicant]
Freitag et al., “Experts, Errors, and Context: A Large-Scale Study of Human Evaluation for Machine Translation,” Transactions of the Association for Computational Linguistics, Dec. 2021, 9:1460-1474. [cited by applicant]
Freitag et al., “Results of WMT22 Metrics Shared Task: Stop Using BLEU—Neural Metrics Are Better and More Robust,” Proceedings of the Seventh Conference on Machine Translation (WMT), Dec. 2022, 46-68. [cited by applicant]
Ganguli et al., “Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned,” arXiv:2209.07858, Aug. 2022, updated Nov. 2022, 30 pages. [cited by applicant]
Garg et al., “Handling Bias in Toxic Speech Detection: A Survey,” arXiv:2202.00126, Jan. 2022, updated Jan. 2023, 30 pages. [cited by applicant]
Garg et al., “Word embeddings quantify 100 years of gender and ethnic stereotypes,” Proceedings of the National Academy of Sciences, Apr. 2018, 115(16): E3635-E3644. [cited by applicant]
Gebru et al., “Datasheets for Datasets,” Communications of the ACM, 2021, 64(12): 86-92. [cited by applicant]
Gehman et al., “RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models,” Findings of the Association for Computational Linguistics: EMNLP 2020, Nov. 2020, 3356-3369. [cited by applicant]
Geva et al., “Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies,” Transactions of the Association for Computational Linguistics, Apr. 2021, 346-361. [cited by applicant]
github.com [online], “Pax,” 2022, retrieved on Jun. 7, 2024, retrieved from URL <URL https://github.com/google/paxml>, 12 pages. [cited by applicant]
github.com [online], “Sax,” 2022, retrieved on Jun. 10, 2024, retrieved from URL <https://github.com/google/saxml>, 7 pages. [cited by applicant]
Glaese et al., “Improving alignment of dialogue agents via targeted human judgements,” arXiv: 2209.14375, Sep. 2022, 77 pages. [cited by applicant]
Goldfarb-Tarrant et al., “Intrinsic Bias Metrics Do Not Correlate with Application Bias,” Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conferen… [cited by applicant]
Goyal et al., “Is Your Toxicity My Toxicity? Exploring the Impact of Rater Identity on Toxicity Annotation,” arXiv: 2205.00501, May 2022, 17 pages. [cited by applicant]
Goyal et al., “Making the v in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 2017, 6904-6… [cited by applicant]
Graves, “Generating Sequences with Recurrent Neural Networks,” arXiv:1308.0850, Aug. 2013, updated Jun. 2014, 1-43. [cited by applicant]
Hanna et al., “Towards a critical race methodology in algorithmic fairness,” Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, Jan. 2020, 501-512. [cited by applicant]
Hasan et al., “XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages,” Findings of the Association for Computational Linguistics: ACL-IJCNLP, Aug. 2021, 4693-4703. [cited by applicant]
Hendricks et al., “Women Also Snowboard: Overcoming Bias in Captioning Models,” Proceedings of the European Conference on Computer Vision (ECCV), 2018, 771-787, 17 pages. [cited by applicant]
Hendrycks et al., “Measuring Mathematical Problem Solving with the MATH Dataset,” arXiv: 2103.03874, Mar. 2021, updated Nov. 2021, 22 pages. [cited by applicant]
Hochreiter et al., “Long Short-Term Memory,” Neural Computation, Nov. 1997, 9(8): 1735-1780, 32 pages. [cited by applicant]
Hoffmann et al., “Training Compute-Optimal Large Language Models,” arXiv: 2203.15556, Mar. 2022, 36 pages. [cited by applicant]
Howard et al., “Universal Language Model Fine-tuning for Text Classification,” Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Jul. 2018, 328-339. [cited by applicant]
Hsiao et al., “Try Bard and share your feedback,” Mar. 21, 2023, retrieved on Jun. 6, 2024, retrieved from URL <https://blog.google/technology/ai/try-bard/>, 7 pages. [cited by applicant]
Hu et al., “LoRA: Low-Rank Adaptation of Large Language Models,” arXiv:2106.09685, Jun. 2021, updated Oct. 2021, 1-26. [cited by applicant]
Ippolito et al., “Preventing Verbatim Memorization in Language Models Gives a False Sense of Privacy,” arXiv: 2210.17546, Oct. 2022, updated Sep. 2023, 26 pages. [cited by applicant]
Jacobs et al., “Measurement and Fairness,” Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, Mar. 2021, 375-385. [cited by applicant]
Jagielski et al., “Measuring Forgetting of Memorized Training Examples,” arXiv:2207.00099, Jun. 2022, updated May 2023, 1-22. [cited by applicant]
Ji et al., “Survey of Hallucination in Natural Language Generation,” arXiv:2202.03629, Feb. 2022, last updated Feb. 2024, 60 pages. [cited by applicant]
Jigsaw, “Exploring the Role of Human Raters in Creating NLP Datasets,” Nov. 19, 2019, retrieved on Jun. 10, 2024, retrieved from URL <https://medium.com/jigsaw/creating-labeled-datasets-and-exploring-the-role-of- human-… [cited by applicant]
Joshi et al., “TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension,” Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Jul. 2017, 1601-1611. [cited by applicant]
Jouppi et al., “A domain-specific supercomputer for training deep neural networks,” Communications of the ACM, Jun. 2020, 63(7): 67-78. [cited by applicant]
kaggle.com [online], “Jigsaw Multilingual Toxic Comment Classification,” 2019, retrieved on Jun. 10, 2024, retrieved from URL <https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification>, 6 pages. [cited by applicant]
kaggle.com [online], “Toxic Comment Classification Challenge,” 2018, retrieved on Jun. 10, 2024, retrieved from URL <https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge>, 3 pages. [cited by applicant]
Kaplan et al., “Scaling Laws for Neural Language Models,” arXiv: 2001.08361, Jan. 2020, 1-30. [cited by applicant]
Keyes, “The Misgendering Machines: Trans/HCI Implications of Automatic Gender Recognition,” Proceedings of the ACM on Human-Computer Interaction, Nov. 2018, 2(CSCW): 22 pages. [cited by applicant]
Kneser et al., “Improved backing-off for M-gram language modeling,” International Conference on Acoustics, Speech, and Signal Processing, May 1995, 181-184. [cited by applicant]
Korbak et al., “Pretraining Language Models with Human Preferences,” arXiv:2302.08582, Feb. 2023, updated Jun. 2023, 1-28. [cited by applicant]
Kreutzer et al., “Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets,” Transactions of the Association for Computational Linguistics, Jan. 2022, 10: 50-72. [cited by applicant]
Kwiatkowski et al., “Natural Questions: A Benchmark for Question Answering Research,” Transactions of the Association for Computational Linguistics, Aug. 2019, 7: 453-466. [cited by applicant]
Ladhak et al., “WikiLingua: A New Benchmark Dataset for Cross-Lingual Abstractive Summarization,” Findings of the Association for Computational Linguistics: EMNLP, Nov. 2020, 4034-4048. [cited by applicant]
Lai et al., “RACE: Large-scale Reading Comprehension Dataset From Examinations, ” Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, Sep. 2017, 785-794. [cited by applicant]
Lee et al., “Deduplicating Training Data Makes Language Models Better,” arXiv:2107.06499, Jul. 2021, updated Mac. 2022, 22 pages. [cited by applicant]
Lee, “Welcome, singular ”they“,” Oct. 31, 2019, retrieved on Jun. 6, 2024, retrieved from URL <https://apastyle.apa.org/blog/singular-they>, 8 pages. [cited by applicant]
Lester et al., “The Power of Scale for Parameter-Efficient Prompt Tuning,” Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Nov. 2021, 3045-3059. [cited by applicant]
Levesque et al., “The Winograd Schema Challenge,” Thirteenth international conference on the principles of knowledge representation and reasoning, Jun. 2012, 552-561. [cited by applicant]
Lewkowycz et al., “Solving Quantitative Reasoning Problems with Language Models,” arXiv: 2206.14858, Jun. 2022, updated Jul. 2022, 1-54. [cited by applicant]
Liang et al., “Holistic Evaluation of Language Models,” arXiv: 2211.09110, Nov. 2022, updated Oct. 2023, 1-162. [cited by applicant]
Longpre et al., “The Flan Collection: Designing Data and Methods for Effective Instruction Tuning,” arXiv:2301.13688, Jan. 2023, updated Feb. 2023, 1-22. [cited by applicant]
Luccioni et al., “What's in the Box? A Preliminary Analysis of Undesirable Content in the Common Crawl Corpus,” arXiv:2105.02732, May 2021, 8 pages. [cited by applicant]
Marino et al., “OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2019, 3195-3204. [cited by applicant]
Meurer et al., “SymPy: symbolic computing in Python,” PeerJ Computer Science, Jan. 2017, 27 pages. [cited by applicant]
Mihaylov et al., “Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering,” Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Oct.-Nov. 2018, 2381-23… [cited by applicant]
Mitchell et al., “Model Cards for Model Reporting, ” Proceedings of the Conference on Fairness, Accountability, and Transparency, Jan. 2019, 220-229. [cited by applicant]
Mostafazadeh et al., “A corpus and cloze evaluation for deeper understanding of commonsense stories,” Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Hu… [cited by applicant]
Movva et al., “Coarse race data conceals disparities in clinical risk score performance,” Proceedings of the 8th Machine Learning for Healthcare Conference, PMLR, 2023, 219: 443-472, 47 pages. [cited by applicant]
Mozes et al., “Towards Agile Text Classifiers for Everyone,” arXiv: 2302.06541, Feb. 2023, updated Oct. 2023, 15 pages. [cited by applicant]
Narayan et al., “Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization,” Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,… [cited by applicant]
Nie et al., “Adversarial NLI: A New Benchmark for Natural Language Understanding,” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 2020, 4885-4901. [cited by applicant]
OpenAI, “GPT-4 Technical Report,” arXiv:2303.08774, Mar. 2023, updated Mar. 2024, 1-100. [cited by applicant]
openai.com [online], “ChatGPT plugins,” Mar. 23, 2023, retrieved on Jun. 7, 2024, retrieved from URL <https://openai.com/blog/chatgpt-plugins>, 13 pages. [cited by applicant]
openai.com [online], “Introducing ChatGPT,” Nov. 30, 2022, retrieved on Jun. 7, 2024, retrieved from URL <https://openai.com/blog/chatgpt>, 8 pages. [cited by applicant]
Orlanski et al., “Measuring the Impact of Programming Language Distribution,” Proceedings of the 40th International Conference on Machine Learning, PMLR, Feb. 2023, 202: 26619-26645, 27 pages. [cited by applicant]
Ouyang et al., “Training language models to follow instructions with human feedback,” Advances in Neural Information Processing Systems, Oct. 2022, 15 pages. [cited by applicant]
Paperno et al., “The LAMBADA dataset: Word prediction requiring a broad discourse context,” Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, Aug. 2016, 1525-1534. [cited by applicant]
Papineni et al., “Bleu: a method for automatic evaluation of machine translation,” Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Jul. 2002, 311-318. [cited by applicant]
Parrish et al., “BBQ: A Hand-Built Bias Benchmark for Question Answering,” arXiv:2110.08193, Oct. 2021, updated Mar. 2022, 20 pages. [cited by applicant]
Paullada et al., “Data and its (dis)contents: A survey of dataset development and use in machine learning research,” Patterns, Nov. 2021, 2(11):100336, 1-14. [cited by applicant]
policies.google.com [online], “Generative AI Prohibited Use Policy,” last modified Mar. 14, 2023, retrieved on Jun. 10, 2024, retrieved from URL <https://policies.google.com/terms/generative-ai/use-policy>, 2 pages. [cited by applicant]
Ponti et al., “XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning,” Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Nov. 2020, 2362-2376. [cited by applicant]
Pozzobon et al., “On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research,” arXiv:2304.12397, Apr. 2023, 1-19. [cited by applicant]
Prabhakaran et al., “Cultural Incongruencies in Artificial Intelligence,” arXiv:211.13069, Nov. 2022, 5 pages. [cited by applicant]
Prabhu et al., “Large image datasets: A pyrrhic win for computer vision?,” arXiv:2006.16923, Jun. 2020, updated Jul. 2020. [cited by applicant]
Pushkarna et al., “Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI,” Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, Jun. 2022, 1776-1826. [cited by applicant]
Rae et al., “Scaling Language Models: Methods, Analysis & Insights from Training Gopher,” arXiv: 2112.11446, Dec. 2021, 1-118. [cited by applicant]
Rae et al., “Scaling Language Models: Methods, Analysis & Insights from Training Gopher,” arXiv: 2112.11446, Jan. 2022, 120 pages. [cited by applicant]
Raffel et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” arXiv:1910.10683, Oct. 2019, updated Sep. 2023, 67 pages. [cited by applicant]
Raji et al., “AI and the Everything in the Whole Wide World Benchmark,” arXiv:2111.16366, Nov. 2021, 20 pages. [cited by applicant]
Rajpurkar et al., “Know What You Don't Know: Unanswerable Questions for SQuAD,” Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Jul. 2018, 2:784-789. [cited by applicant]
Ramesh et al., “How Platform-User Power Relations Shape Algorithmic Accountability: A Case Study of Instant Loan Platforms and Financially Stressed Users in India,” FAccT '22: 2022 ACM Conference on Fairness, Accountabi… [cited by applicant]
Rauh et al., “Characteristics of Harmful Text: Towards Rigorous Benchmarking of Language Models,” NeurIPS 2022 Datasets and Benchmarks, 2022, 1-20. [cited by applicant]
Riley et al., “FRMT: A Benchmark for Few-Shot Region-Aware Machine Translation,” Transactions of the Association for Computational Linguistics, Jun. 2023, 11: 671-685. [cited by applicant]
Roberts et al., “How Much Knowledge Can You Pack Into the Parameters of a Language Model?,” Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Nov. 2020, 5418-5426. [cited by applicant]
Rodriguez et al., “Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?,” Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Interna… [cited by applicant]
Ruder et al., “Square One Bias in NLP: Towards a Multi-Dimensional Exploration of the Research Manifold,” Findings of the Association for Computational Linguistics: ACL 2022, May 2022, 2340-2354. [cited by applicant]
Sakaguchi, “WinoGrande: an adversarial winograd schema challenge at scale,” Communications of the ACM, Aug. 2021, 64(9): 99-106. [cited by applicant]
Sambasivan et al., “Re-imagining Algorithmic Fairness in India and Beyond, ” FAccT '21: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, Mar. 2021, 315-328. [cited by applicant]
Sap et al., “Annotators with Attitudes: How Annotator Beliefs and Identities Bias Toxic Language Detection,” arXiv: 2111.07997, Nov. 2021, updated May 2022, 23 pages. [cited by applicant]
Sap et al., “Social Bias Frames: Reasoning about Social and Power Implications of Language,” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 2020, 5477-5490. [cited by applicant]
Schick et al., “Self-Diagnosis and Self-Debiasing: A Proposal for Reducing Corpus- Based Bias in NLP,” arXiv: 2103.00453, Feb. 2021, updated Sep. 2021, 16 pages. [cited by applicant]
Schlangen, “Targeting the Benchmark: On Methodology in Current Natural Language Processing Research,” arXiv:2007.04792, Jul. 2020, 5 pages. [cited by applicant]
Selbst et al., “Fairness and Abstraction in Sociotechnical Systems,” Proceedings of the Conference on Fairness, Accountability, and Transparency, Jan. 2019, 59-68. [cited by applicant]
Sellam et al., “BLEURT: Learning Robust Metrics for Text Generation,” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 2020, pp. 7881-7892. [cited by applicant]
Shannon, “Prediction and Entropy of Printed English,” The Bell System Technical Journal, Jan. 1951, 30(1): 50-64. [cited by applicant]
Shelby et al., “Identifying Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction,” arXiv: 2210.05791, Oct. 2022, updated Feb. 2023, 1-33. [cited by applicant]
Shi et al., “Language Models are Multilingual Chain-of-Thought Reasoners,” arXiv:2210.03057, Oct. 2022, 1-20. [cited by applicant]
Singh et al., “Towards VQA Models That Can Read,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2019, 8317-8326. [cited by applicant]
Smith et al., ““I'm sorry to hear that”: Finding New Biases in Language Models with a Holistic Descriptor Dataset,” arXiv: 2205.09209, May 2022, updated Oct. 2022, 32 pages. [cited by applicant]
Srivastava et al., “Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models,” arXiv: 2206.04615, Jun. 2022, updated Jun. 2023, 1-95. [cited by applicant]
success.appen.com [online], “Guide to: Fair Pay,” 2023, retrieved on Jun. 10, 2024, retrieved from URL <https://success.appen.com/hc/en-US/articles/9557008940941-Guide-to-Fair-Pay>, 5 pages. [cited by applicant]
Suzgan et al., “Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them,” arXiv: 2210.09261, Oct. 2022, 1-49. [cited by applicant]
Tabachnyk et al., “ML-Enhanced Code Completion Improves Developer Productivity,” Jul. 26, 2022, retrieved on Jun. 7, 2024, retrieved from URL <https://research.google/blog/ml-enhanced-code-completion-improves-developer-… [cited by applicant]
Talmor et al., “CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge,” Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human La… [cited by applicant]
Tay et al., “UL2: Unifying Language Learning Paradigms,” The Eleventh International Conference on Learning Representations, 2023, 1-33. [cited by applicant]
The Replit Team, “Meet Replit Ghostwriter, your partner in code,” Oct. 31, 2022, retrieved on Jun. 7, 2024, retrieved from URL <https://blog.replit.com/ghostwriter>, 6 pages. [cited by applicant]
Thoppilan et al., “LaMDA: Language Models for Dialog Applications,” arXiv:2201.08239, Jan. 2022, updated Feb. 2022, 1-47. [cited by applicant]
Tomasev et al., “Fairness for Unobserved Characteristics: Insights from Technological Impacts on Queer Communities,” Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, Jul. 2021, 254-265. [cited by applicant]
Van Esch et al., “Writing System and Speaker Metadata for 2,800+ Language Varieties,” Proceedings of the Thirteenth Language Resources and Evaluation Conference, Jun. 2022, 5035-5046. [cited by applicant]
Vaswani et al., “Attention is all you need,” 31st Conference on Neural Information Processing Systems (NIPS 2017), Jun. 2017, 1-11. [cited by applicant]
Vilar et al., “Prompting PaLM for Translation: Assessing Strategies and Performance,” arXiv: 2211.09102, Nov. 2022, updated Jun. 2023, 20 pages. [cited by applicant]
Wang et al., “Self-Consistency Improves Chain of Thought Reasoning in Language Models,” arXiv:2203.11171, Mar. 2022, updated Mar. 2023, 1-24. [cited by applicant]
Wang et al., “SuperGlue: A stickier benchmark for general-purpose language understanding systems,” 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), 2019, 1-15. [cited by applicant]
Wei et al., “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,” arXiv:2201.11903, Jan. 2022, updated Jan. 2023, 1-43. [cited by applicant]
Weidinger et al., “Ethical and social risks of harm from Language Models,” arXiv:2112.04359, Dec. 2021, 1-64. [cited by applicant]
Welty et al., “Metrology for AI: From Benchmarks to Instruments,” arXiv:1911.01875, Nov. 2019, 11 pages. [cited by applicant]
Xu et al., “Detoxifying Language Models Risks Marginalizing Minority Voices,” arXiv: 2104.06390, Apr. 2021, 8 pages. [cited by applicant]
Xu et al., “GSPMD: General and Scalable Parallelization for ML Computation Graphs,” arXiv: 2105.04663, May 2021, updated Dec. 2021, 1-16. [cited by applicant]
Xu et al., “Human Parity on CommonsenseQA: Augmenting Self-Attention with External Attention,” arXiv:2112.03254, Dec. 2021, updated May 2022, 8 pages. [cited by applicant]
Xue et al., “mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer,” Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technol… [cited by applicant]
Yin et al., “Natural Language to Code Generation in Interactive Data Science Notebooks,” arXiv:2212.09248, Dec. 2022, 46 pages. [cited by applicant]
Yu et al., “CoCa: Contrastive Captioners are Image-Text Foundation Models,” arXiv:2205.01917, May 2022, updated Jun. 2022, 1-19. [cited by applicant]
Zellers et al., “HellaSwag: Can a Machine Really Finish Your Sentence?,” Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Jul. 2019, 4791-4800. [cited by applicant]
Invitation to Pay Additional Fees in International Appln. No. PCT/IB2024/000231, mailed on Oct. 2, 2024, 16 pages. [cited by applicant]
Liu et al., “Mitigating Unintended Memorization in Language Models via Alternating Teaching,” Presented at International Conference on Acoustics, Speech and Signal Processing, Rhodes Island, Greece, Jun. 4-10, 2023, 5 p… [cited by applicant]