MedUni Exam

How to turn your notes into practice questions worth answering

Getting usable practice questions out of one lecture costs about fourteen minutes before you answer anything, and most of that is you writing three by hand to set a standard. Generating takes two minutes, and so does the checking that decides whether the set is worth answering. The worked run at the end of this article takes fifty minutes end to end, because answering the questions cold and reading back the misses is the part that teaches you something. Skip the checking and you study whatever the model guessed. A systematic review published in May 2026 put the error rate across 71 studies at anywhere from under 1 percent to 45 percent.

Write three by hand first, then never again

Open the deck. Find three lines that could decide an answer on their own: a mechanism, a threshold, a first line drug, a characteristic finding. Build one question from each, five options, one right. Ten minutes, and you are done with hand writing for good.

Those ten minutes buy you a standard. Having written three, you know that a usable stem asks something you could answer with the options covered up, that five options must be the same kind of thing in the same grammatical form, and that exactly one sentence in your source has to settle the answer. You also learn how slow it is, which is why nobody writes forty.

Keep the three. They become your reference for everything a model hands you later, and when a generated question looks wrong and you cannot say why, setting it beside one of yours usually locates the problem in the options.

What to ask for, in the words that change the output

Quiz me on this gets you trivia about the document instead of questions about the subject. Six instructions change what comes back. Paste them above your text, changing the count and the subject.

Use only the text below. Write 12 single best answer questions on second year cardiovascular physiology, each with five options and exactly one correct answer. After each question, print the verbatim sentence from the text that makes the answer correct, and the page or slide number it sits on. Do not write a question unless that one sentence settles the answer by itself. No all of the above, no none of the above, no which of the following is true stems. Keep the five options the same length and the same grammatical form.

  • Use only the text below. Left out, the model answers from what it already knows, and you practice on material your examiner never taught.
  • Print the verbatim sentence and the page. That clause does the most work of the six. Quotes you can look up settle an argument with the answer key in ten seconds, and a model required to produce one writes fewer claims it cannot support.
  • One sentence has to settle it. Otherwise you get items needing the quoted line plus three facts from elsewhere, impossible to check and unfair to answer.
  • Name the item format. Say questions and you get true or false, fill in the blank and one line recall mixed together.
  • Ban the two stems. All of the above and which of the following is true are the flaws generators produce most, and both go away on instruction.
  • Same length, same form. Item writers elaborate the option they know is correct, and a model trained on their output copies the habit.

Ask for a count your material can carry. The post on daily question supply counts a line at a time and lands near the slide count, which is a count of markable lines rather than of questions you keep. In practice a 40 slide deck yields 15 to 20 you would keep, because asking for 50 mostly buys duplicates the checks then remove.

What comes back wrong, and how often: 71 studies, counted

Two published numbers set a fair expectation. A systematic review in the Postgraduate Medical Journal, published May 20, 2026, pooled 71 studies from 24 countries on AI written multiple choice questions in medical education. Error rates ran from under 1 percent to 45 percent. Median item difficulty was 0.67 and median discrimination 0.28, which describes an item that separates students who know the material from those who do not, but not sharply. Most of those studies judged the questions by expert review rather than by putting them in front of learners, and the review concludes that the evidence does not yet support "unsupervised use in summative assessment".

Older and blunter: Boston University's Chobanian and Avedisian School of Medicine reported in December 2023 that ChatGPT produced a correct question with a correct answer and explanation in 32 percent of cases when asked to write items for a medical school exam.

Third number, from this site, which counts what its checks throw out and publishes the count. Read on September 20, 2026: between September 14, when the call log starts, and that date the models wrote 3,564 questions across 1,398 calls. Of those, 43.1 percent survived everything. Structural rules removed 32.8 percent, the duplicate check 13.5 percent, a second model answering from the quote without the key 6.7 percent, and 2.2 percent either quoted a sentence the checker could not find on the page they cited or cited a page outside the section. The remaining 1.7 percent arrived after the job was already full or came back in a form the parser could not read.

Those shares move with every job, and the duplicate share moves most. The post published a day earlier counted it at 17.6 percent over a shorter window, September 14 to 19, 2026; this read counts 13.5 percent over September 14 to 20, 2026. Both are correct reads of the same rolling report, which is why the rejection rates page carries the live figure and everything above is one snapshot of it.

Put together, those three do not say avoid generating questions. They say close to half of what comes back belongs in the bin, and your only decision is who does the binning.

When to stop checking, and what a failed question is worth

The three checks that catch most of it are already written out on the page behind those counts: find the quoted sentence in your document, cover the key and answer from the quote alone, then read the five options without the stem. Run them on the first five questions of a set, a minute in total, and the article on accuracy works through what each one catches and why.

Those checks are not mine. They come from the item writing taxonomy those two sources set out: the 31 rules Haladyna, Downing and Rodriguez set out in Applied Measurement in Education in 2002, and the working version item writers are usually given, the NBME Item-Writing Guide, now in its sixth edition and posted in February 2021.

What that page does not tell you is when to stop. Here is the number: check the first five questions and no more. At two failures out of five, stop checking and go back to the prompt. A set that fails twice that early almost always fails for one reason, and one reason is usually one line of the prompt, so tightening it and generating again costs two minutes. Checking all fifteen costs twenty and leaves the same defect in the next batch.

A question that fails goes in the bin, not into repair. Rewriting a stem whose quoted sentence does not settle the answer means deciding the answer yourself, at which point you are the item writer and the model wasted your evening. Under one failure in five, bin that one and answer the rest.

The binned question still earns something. You opened the slide to check it, which means you read the line the question was built on, in the source's words. Write that line down where your misses go. A question that failed a check and a question you answered wrong end up in the same place, which is the start of tomorrow.

The nursing version, and what a lecture slide does not give you

Nursing students usually arrive here with a narrower job: getting NCLEX-style items out of a class handout. Situation is the missing piece. Your pathophysiology slide hands you a mechanism, while an NCLEX-style item needs a client, a setting, a few findings and a lead in asking for a priority or an action, with four options that are all defensible and one that is best. None of that is on the slide, so the prompt has to put it there.

Add a line: write each question as a short client scenario in a hospital setting, ending with which action the nurse should take first, and make all four options safe nursing actions. Add a fourth check for these: does the stem carry enough information to choose between the four? Where any of them could reasonably be first, the item is not difficult, it is unanswerable, and it teaches you to guess.

Out of reach from a handout, completely: the Next Generation formats. Case study, matrix, bow tie, highlight and drop down items are built around unfolding patient data with scoring rules of their own, and nothing generated from a lecture deck produces them. Anyone selling you that from a PDF is overclaiming. The NCLEX-RN page here says the same. This is not an NCLEX question bank and has no connection to the NCSBN; items are written in that style from the file you upload.

Doing it at volume without drowning in near duplicates

Request 60 questions from a 12 page handout and you receive the same 12 facts in 60 costumes. Nothing in a chat window prevents it: by question 47 the model has lost question 14, and spotting a paraphrase of something you answered twenty minutes ago is hard.

Three habits keep a long run honest. Generate one chapter at a time rather than a whole module, because duplicates cluster inside a source. Change the item type on a second pass, so single best answer becomes a clinical vignette and the same fact is tested as reasoning. Skim the stems before answering any of them, since a repeat caught early costs nothing and one caught on question 30 costs you the session.

Machines are better at this part than you are. A generator that compares every new question against everything already written from the same document drops a near repeat before you ever see it, and that is what the duplicate rejections counted above are: questions that never reached a reader.

A worked run from one 40 slide deck

Tuesday evening, a 40 slide cardiology lecture, one sitting. Below are the times it takes, not the ones a plan would like.

  1. First ten minutes. Three questions by hand, from slides where you remember the lecturer slowing down.
  2. Two minutes. Paste the prompt above the deck text and ask for 15 questions, not 50. Fifteen is roughly what 40 slides support once the repeats come out.
  3. Another two. Run the checks on the first five items. Expect one or two to fail. At two failures, fix the prompt and generate again, which costs two minutes and saves twenty.
  4. Twenty three minutes. Answer all 15 cold, about 90 seconds each, options covered until you have committed.
  5. Thirteen minutes. For every miss, open the slide the question cites and read the line it came from. Copy it down in the source's words, not your summary of it.
  6. Tomorrow. Those misses open the next session, before anything new. Where in the following weeks they come back is the subject of the spaced repetition schedule.

Fifty minutes, 15 questions you can defend, and a handwritten list of the facts that beat you.

Questions people ask

What prompt makes ChatGPT write good exam questions from my notes?

Four clauses carry almost all of the value. Tell it to use only the text you pasted, name the format as single best answer with five options and one correct answer, require the verbatim sentence and page number that make the answer correct, and forbid all of the above and which of the following is true stems. Adding that the quoted sentence must settle the answer on its own removes most of what is left.

How accurate are AI generated practice questions?

Published studies do not give one number. A systematic review of 71 studies in the Postgraduate Medical Journal, published May 20, 2026, found error rates from under 1 percent to 45 percent, and Boston University reported in December 2023 that ChatGPT produced a correct question with a correct answer and explanation in 32 percent of cases. What decides your own set is whether anything checked it. This site publishes what its checks throw out: read on September 20, 2026, 43.1 percent of 3,564 generated questions survived every check, with the duplicate check accounting for 13.5 percent of the rest. Those shares move with the window, so the rejection rates page carries the live figure.

How many practice questions should I make from one lecture?

Counting markable lines, a 40 slide deck lands near one per slide, since that is how many lines could decide an answer on their own. Counting questions you would keep after checking, it is nearer 15 to 20, because not every markable line survives. Asking for 50 does not create more facts, it restates the same ones in new wording and the duplicate check removes them. Generate one lecture at a time and stop when the new questions stop surprising you.

When should I stop checking a generated set and start over?

Check the first five questions and stop there. At two failures out of five, go back to the prompt instead of repairing the set, because a set that fails twice that early usually fails for one reason and one reason is usually one line of the prompt. Regenerating costs two minutes; repairing fifteen questions one at a time costs the evening.

Can AI write NCLEX style questions from my class notes?

It can write scenario items in that style if you tell it to build a client situation with a priority or action lead in and four safe options, since a lecture slide supplies the pathophysiology but never the setting. The Next Generation formats, meaning case study, matrix, bow tie, highlight and drop down, cannot be produced from a document. This is not an NCLEX question bank and has no connection to the NCSBN.

Published .