How to design an OSCE case that shows what a learner can do
Start with the decision the station should support, define observable behaviors, give learners a fair chance to show them, train the SP, and pilot the scoring before the station counts.
ClinicalSim Team
ClinicalSim
Start with the decision the station should support. Then define the observable behaviors, write a case that gives learners a fair chance to show them, train the standardized patient, and pilot the scoring before the station counts.
Start with the decision
A communication station and a diagnostic reasoning station may use the same encounter, but they should not use the same rubric. Write down what the score will mean before you write the patient. If the program cannot say what a passing performance allows it to conclude, the case is not ready.
Decide whether the station is for practice, coaching, progression review, or a high-stakes assessment. That choice determines the evidence required, the people who need to review it, and how much standardization the station needs.
Define behaviors a reviewer can see
Write objectives as actions. "Acknowledges the patient's concern before giving more information" is easier to observe than "shows empathy." A reviewer should be able to point to what the learner said or did.
Keep the rubric focused on the decision. A checklist can work for required actions, while a rating scale may fit communication quality or clinical reasoning. Do not add items simply because they are easy to count.
Give the learner a fair chance to show the skill
The case must create an opening for every scored behavior. If the rubric asks whether a learner responds to emotion, the patient needs a clear emotional cue. If the station does not give that opportunity, mark the item not assessable instead of treating silence as failure.
Keep the patient history internally consistent. Separate facts the patient volunteers from facts revealed only after an appropriate question. Tell the standardized patient how to respond to common approaches without scripting every sentence.
Make "not assessable" a real outcome, not a zero
The item that no station design survives contact with is the one the encounter never gave the learner a chance to demonstrate. Scoring it as a failure is the single most common way a well-built rubric produces a misleading total, and it is worth building the escape hatch on purpose rather than leaving raters to improvise one.
The encounters we publish show what that looks like in practice, because a voice-only format makes the problem unavoidable. A leukemia disclosure scored against SPIKES totals 25 out of 30, and the first step, setting up the interview, comes back at 3 out of 5 with the reason stated: privacy, seating, and interruption management cannot be assessed in a voice-only encounter, so the score rests only on the elements that could be. A vaccine hesitancy encounter marks privacy not assessable in a spoken format for the same reason.
Two things follow for a station design. Say which items the station's own modality cannot reach, in the rubric, before anyone runs it. And report the count of excluded items alongside the total, so a reviewer reading 25 out of 30 knows whether the denominator was 30 or something less.
Match the case to the learner
An early medical student may need to gather a focused history and explain the next step. A resident may need to manage uncertainty, competing priorities, or a family meeting. The clinical task, case complexity, and framework should change with the learner's stage.
Avoid testing obscure recall when the station is meant to assess performance. If a written question can measure the objective more directly, the OSCE may be the wrong format.
Train the standardized patient and the rater together
The patient portrayal and the scoring rules are part of the same instrument. Use sample learner responses to calibrate when the patient offers information, how emotion changes, and what evidence meets each rating.
Review disagreements before launch. If experienced raters interpret an item differently, rewrite the item or provide an anchor. Do not ask the live administration to resolve an ambiguity the design team left behind.
This is also the question to put to any vendor scoring your station, including us. Consistency between runs of a model is not the same thing as agreement with your expert human raters, and we do not publish a faculty-rater agreement figure because that work has not been done on a customer's own rubric. Ask for the number. If a vendor has it, they will give it to you, and if the answer is that they have measured their own repeatability instead, that is a different claim and worth naming as one.
Pilot before the score counts
Run the complete station with learners who resemble the intended group. Watch for unclear instructions, cues that appear too early or too late, equipment problems, and items that almost everyone passes or misses.
Use the pilot to revise the case, rubric, and rater guidance. A polished script is not enough. The station is ready when the encounter gives learners a fair opportunity and reviewers can explain what the score means.
Where AI patients fit
AI patients can add repeatable spoken practice before an OSCE, scored against the rubric written for the practice case, or the station's own rubric where the program supplies it, with the learner's own words quoted under every score. Learners see what worked and what to practice next, while faculty get specific evidence for coaching. Standardized patients and faculty should keep the live assessment and human judgment that high-stakes decisions require.
Review unedited AI patient simulations.
References
- Cook DA, Brydges R, Ginsburg S, Hatala R. A contemporary approach to validity arguments: a practical guide to Kane's framework. Medical Education. 2015. doi:10.1111/medu.12678
- Regehr G, MacRae H, Reznick RK, Szalay D. Comparing the psychometric properties of checklists and global rating scales for assessing performance on an OSCE-format examination. Academic Medicine. 1998. doi:10.1097/00001888-199809000-00020
- Boursicot K, Kemp S, Wilkinson T, et al.. Performance assessment: consensus statement and recommendations from the 2020 Ottawa Conference. Medical Teacher. 2020. doi:10.1080/0142159X.2020.1830052
- Lewis KL, Bohnert CA, Gammon WL, et al.. The Association of Standardized Patient Educators (ASPE) Standards of Best Practice (SOBP). Advances in Simulation. 2017. doi:10.1186/s41077-017-0043-4
- Pell G, Fuller R, Homer M, Roberts T. How to measure the quality of the OSCE: a review of metrics. AMEE guide no. 49. Medical Teacher. 2010. doi:10.3109/0142159X.2010.507716