---
title: Item writing flaws contaminated one of the most widely used AI datasets
description: AI benchmarks can be flawed by item writing errors, impacting their reliability and the AI's learning process.
image: https://questionwell.org/hubfs/image-png-Jan-22-2026-09-19-32-8694-PM.png
---

[Skip to content](https://questionwell.org/blog/item-writing-flaws-contaminated-one-of-the-most-widely-used-ai-datasets#main-content)

[![QuestionWell](https://questionwell.org/hs-fs/hubfs/QW%20LOGO%20(2).png?width=180&height=45&name=QW%20LOGO%20(2).png)](https://questionwell.org/)

- [Home](https://questionwell.org/)
- [Product](https://questionwell.org/product)
- [Partners](https://questionwell.org/partnerships)
- [Research](https://questionwell.org/research)
- [FAQs](https://questionwell.org/faqs)
- [Pricing](http://app.questionwell.org/pricing)
- [Privacy](https://questionwell.org/privacy-policy)
- [Contact](https://questionwell.org/contact)

Open main navigation

Close main navigation

- [Home](https://questionwell.org/)
- [Product](https://questionwell.org/product)
- [Partners](https://questionwell.org/partnerships)
- [Research](https://questionwell.org/research)
- [FAQs](https://questionwell.org/faqs)
- [Pricing](http://app.questionwell.org/pricing)
- [Privacy](https://questionwell.org/privacy-policy)
- [Contact](https://questionwell.org/contact)
- [Login](https://app.questionwell.org/)

[Login](https://app.questionwell.org/)

[← Back to all posts](https://questionwell.org/blog)  1/22/26, 4:41 PM

# Item writing flaws contaminated one of the most widely used AI datasets

[Will Cummings](https://questionwell.org/blog/author/will-cummings)

Share: [facebook-f icon](http://www.facebook.com/share.php?u=https://questionwell.org/blog/item-writing-flaws-contaminated-one-of-the-most-widely-used-ai-datasets) [linkedin-in icon](http://www.linkedin.com/shareArticle?mini=true&url=https://questionwell.org/blog/item-writing-flaws-contaminated-one-of-the-most-widely-used-ai-datasets) [Twitter icon](https://twitter.com/intent/tweet?url=https://questionwell.org/blog/item-writing-flaws-contaminated-one-of-the-most-widely-used-ai-datasets) [pinterest-p icon](http://pinterest.com/pin/create/link/?url=https://questionwell.org/blog/item-writing-flaws-contaminated-one-of-the-most-widely-used-ai-datasets) [envelope icon](mailto:?body=https://questionwell.org/blog/item-writing-flaws-contaminated-one-of-the-most-widely-used-ai-datasets)

Much like how teachers evaluate human students, AI researchers put models through their paces using a suite of tests, called benchmarks. Benchmarks can be complex, like [SWE-bench](https://www.swebench.com/) which grades models on how well they fix bugs in a large piece of software, or simple question/answer pairs. There are benchmarks for [math](https://epoch.ai/frontiermath), science, [abstract reasoning,](https://arcprize.org/arc-agi/2/) [playing chess](https://maxim-saplin.github.io/llm_chess/) and all sorts of other things. 

As discussed in my previous blog post (if you want more background on item writing flaws, read this), [We Built a Model Thats Writes Better Multiple Choice Questions:](https://questionwell.org/blog/we-built-a-model-that-writes-better-multiple-choice-questions.-heres-the-evidence)

> *We noticed that the LLMs we were using were making many of the same mistakes that humans made while writing multiple choice questions based on a text.*

As an example of how AI datasets can be poisoned by item writing flaws, I present [MMLU-Pro](https://huggingface.co/spaces/TIGER-Lab/MMLU-Pro), a public benchmark which presents the LLM with multiple choice questions across a number of domains. [Eric Tramel](https://x.com/fujikanaeda), a research scientist at NVidia, recently discovered this dataset was poisoned by an obvious cueing error: correct answers are consistently proceeded by a space!

![](https://questionwell.org/hs-fs/hubfs/image-png-Jan-22-2026-09-19-32-8694-PM.png?width=453&height=468&name=image-png-Jan-22-2026-09-19-32-8694-PM.png)

We can see selecting the answer proceeded by a space doubles the number of correct answers compared to random selection in Math, Physics and Chemistry. Another user,[Peter Barnett](https://x.com/peterbarnett_/status/2011958022592180639?s=46), a MIRI (Machine Intelligence Research Institute) researcher, points out another cueing problem that may be more familiar to educators: the longest answer is correct.

![](https://questionwell.org/hs-fs/hubfs/image-png-Jan-22-2026-09-12-08-1982-PM.png?width=502&height=314&name=image-png-Jan-22-2026-09-12-08-1982-PM.png)

Naively choosing the longest answer every time provides a similar boost in performance, this time across all domains in the benchmark. If these types of errors persist, even in popular and well regarded public benchmarks, it's likely they're also pervasive in the training data. Research has shown that LLMs learn shortcuts to cheat on this benchmark, such as Changing Answer Order Can Decrease MMLU Accuracy ([arXiv:2406.19470v2](https://arxiv.org/html/2406.19470v2)). The benchmark will not be able to differentiate between a model which truly knows the correct answer to an important biology question, and which model has merely learned to select the conspicuously long answer.

Sound familiar? This is the same problem teachers face when evaluating students with multiple choice questions.

Don't worry, we're on top of it. [Learn how we fix cueing errors like "longest answer" in our models.](https://questionwell.org/blog/we-built-a-model-that-writes-better-multiple-choice-questions.-heres-the-evidence)

[Engineering](https://questionwell.org/blog/tag/engineering), [Multiple Choice Questions](https://questionwell.org/blog/tag/multiple-choice-questions), [Assessment](https://questionwell.org/blog/tag/assessment), [MMLU-Pro](https://questionwell.org/blog/tag/mmlu-pro), [Benchmarks](https://questionwell.org/blog/tag/benchmarks)

## Related posts

[![multiple choice revamped](https://miro.medium.com/v2/resize:fit:1400/0*4MDyZDzIfCNTbdd3.png)](https://questionwell.org/blog/the-case-for-multiple-choice)

## [The Case for Multiple Choice](https://questionwell.org/blog/the-case-for-multiple-choice)

![Picture of Maya Bialik](https://app.hubspot.com/settings/avatar/bb19a2b2a05c8222df4ad9de14d49662) [Maya Bialik](https://questionwell.org/blog/author/maya-bialik) 

 8/27/24, 12:03 PM

Multiple choice questions have become almost synonymous with high stakes testing, and thus with...

[Read more](https://questionwell.org/blog/the-case-for-multiple-choice)

[![](https://questionwell.org/hs-fs/hubfs/Screenshot%202026-01-03%20at%205.07.18%20PM.png?height=200&name=Screenshot%202026-01-03%20at%205.07.18%20PM.png)](https://questionwell.org/blog/we-trained-an-ai-model-that-writes-better-multiple-choice-questions-heres-the-evidence)

[Engineering](https://questionwell.org/blog/tag/engineering), [Research](https://questionwell.org/blog/tag/research), [Multiple Choice Questions](https://questionwell.org/blog/tag/multiple-choice-questions), [Assessment](https://questionwell.org/blog/tag/assessment)

## [We Trained an AI Model That Writes Better Multiple Choice Questions. Here’s the Evidence.](https://questionwell.org/blog/we-trained-an-ai-model-that-writes-better-multiple-choice-questions-heres-the-evidence)

[Will Cummings](https://questionwell.org/blog/author/will-cummings) 

 12/28/25, 6:29 AM

Writing good multiple-choice questions (MCQs) is well studied and well known to be difficult, even...

[Read more](https://questionwell.org/blog/we-trained-an-ai-model-that-writes-better-multiple-choice-questions-heres-the-evidence)

[![](https://questionwell.org/hs-fs/hubfs/tectonics.png?height=200&name=tectonics.png)](https://questionwell.org/blog/your-ai-tool-promised-grade-level-reading-did-you-check)

[Research](https://questionwell.org/blog/tag/research), [Benchmarks](https://questionwell.org/blog/tag/benchmarks), [EdTech](https://questionwell.org/blog/tag/edtech), [Reading Level](https://questionwell.org/blog/tag/reading-level)

## [Your AI Tool Promised Grade-Level Reading. Did You Check?](https://questionwell.org/blog/your-ai-tool-promised-grade-level-reading-did-you-check)

[Will Cummings](https://questionwell.org/blog/author/will-cummings) 

 3/31/26, 4:59 PM

EdTech AI tools make bold promises. They'll differentiate for you. They'll match your standards....

[Read more](https://questionwell.org/blog/your-ai-tool-promised-grade-level-reading-did-you-check)

![QW LOGO (2)](https://questionwell.org/hs-fs/hubfs/QW%20LOGO%20(2).png?width=1000&height=248&name=QW%20LOGO%20(2).png "QW LOGO (2)")

[Terms of Service](https://questionwell.org/terms-of-service)

[Privacy Policy](https://questionwell.org/privacy-policy)

[FAQs](https://questionwell.org/faqs)

[Contact Us](https://questionwell.org/contact)

<https://www.linkedin.com/company/questionwell-ai/> <https://x.com/questionwellai> <https://www.instagram.com/questionwellai/> <https://www.facebook.com/questionwellai> <https://www.tiktok.com/@questionwellai>

```json
{
  "@context" : "https://schema.org",
  "@type" : "BlogPosting",
  "author" : {
    "@type" : "Person",
    "name" : "Will Cummings",
    "url" : "https://questionwell.org/blog/author/will-cummings"
  },
  "dateModified" : "2026-05-01T14:38:53.141Z",
  "datePublished" : "2026-01-22T21:41:52.000Z",
  "headline" : "Item writing flaws contaminated one of the most widely used AI datasets",
  "image" : [ "https://questionwell.org/hubfs/image-png-Jan-22-2026-09-19-32-8694-PM.png" ],
  "mainEntityOfPage" : {
    "@id" : "https://questionwell.org/blog/item-writing-flaws-contaminated-one-of-the-most-widely-used-ai-datasets",
    "@type" : "WebPage"
  },
  "publisher" : {
    "@type" : "Organization",
    "logo" : {
      "@type" : "ImageObject",
      "url" : "https://questionwell.org/hubfs/QW%20LOGO%20(1)-1.png"
    },
    "name" : "QuestionWell AI"
  }
}
```