How accurate are AI-written exam questions? What real jobs show
Every question this site writes is checked before a student sees it, and the checks keep a count. This is what the 10 days since of real jobs on medical textbooks, lecture decks and notes show: how many questions a model writes that cannot be traced to the page, that a second model answers differently, or that repeat one already written.
The numbers, counted from the call log
| What the checks counted | Share of questions written |
|---|---|
| Questions the models wrote | 6,071 in 2,403 calls |
| Kept, every check passed | 44.2% |
| Thrown out, all reasons | 55.8% |
| Quote not found on the cited page | 6.0% |
| Second model answered differently | 5.3% |
| Repeat of a question already written | 11.3% |
| Structural flaw (all of the above, a vague stem, an option that gives itself away) | 31.1% |
Counted from the call log over September 14, 2026 to September 24, 2026, refreshed twice an hour. The same report, split by model family and with a CSV, is on the rejection rates page.
What each rejection means
Quote not found on the cited page. Every question has to carry one or two sentences copied out of the document, with the page they sit on. The server looks the sentence up in the stored text of that page. If it is not there, the model wrote a question about something the material does not say, or misremembered the wording, and the question is dropped. This is the check that catches an invented fact.
Second model answered differently. A model from a different family is given the question and the quotes, and nothing else: no answer key, no document. If it cannot arrive at the keyed answer from the quotes alone, the quotes do not prove the answer, and the question is dropped. This is the check that catches a right-sounding question with the wrong key, or a key the quote does not support.
Repeat of a question already written. Each new question is compared with everything already generated from the document. A near repeat is dropped, so a long run gives different questions rather than the same facts restated.
Structural flaw. "All of the above", "none of the above", a stem that asks which statement is true, an option conspicuously longer than the others, a question that refers to "the text". These are the tells examiners are trained to remove, and they are removed here by rule before the other checks run. The MCQ flaw checker runs the same rules on any question you paste.
Why a generator that shows no quote cannot tell you its accuracy
A tool that hands you a quiz without the sentence each question came from has no way to count how many of its questions were unsupported, because nothing in its pipeline looked. Its accuracy is whatever you find when you check the questions yourself, which is the work you were trying to hand off. One competing tool's own help page asks its users to review a generated quiz before relying on it, which is the honest thing to say when nothing was checked. The numbers above exist because the checking happens before the student sees anything, and because what was thrown out was counted rather than quietly dropped.
How to check an AI question yourself in twenty seconds
- Find the sentence. If the question does not show where it came from, search the document for its key phrase. If the phrase is not there, the question is not from your material.
- Cover the key and answer from the quote. Read only the quoted sentence and pick an option. If the quote does not settle it, the question is testing something the quote does not say.
- Look at the options. One option much longer than the rest, an option that repeats the stem, "all of the above": each is a tell, and each is a reason to distrust the question, not to learn from it.
What it means for exam preparation
A set of questions that survived these checks is a set you can practice on without re-reading the chapter to make sure each one is real. The rejected share is not wasted effort; it is the part of any AI-generated question bank that a student using an unchecked tool would have studied from. The free plan on this site runs the same checks on every question it writes, and the sample page shows three that survived, with the page each one quotes.
Questions students ask
Does a high rejection rate mean a bad model?
Not on its own. A model that writes ambitious clinical vignettes loses more of them to the second model than one that writes recall questions, and a long run from a short chapter loses questions to the duplicate check because the chapter has only so many facts. The rate is a property of the model, the material and the run together, which is why the report is published with the number it was counted from.
Are the questions that pass guaranteed to be correct?
No. The checks catch a question whose quote is not on the page it cites, one a second model answers differently from the key, one that repeats an earlier question, and a set of structural flaws. They cannot judge the medicine. That is why every kept question is shown with its quote and page number, so the student can read the line the question came from.
Why does a repeated question count as a rejection?
Because a generator that is asked for forty questions from a twelve page handout will restate the same twelve facts in different words if nothing stops it. Each new question is compared with everything already written from the document, and a near repeat is thrown out rather than handed to the student as new practice.
Where is the raw data?
The rejection rates page publishes the report by model family and by lane with a CSV download, and the numbers on this page are read from the same report. The call log has been kept since September 14, 2026.
No account needed to try it. Registering keeps your library, grades and review queue across devices.
Published , updated .