What 12 medical students want from AI patient simulation
Researchers interviewed 12 clinical-year medical students and ran three codesign workshops. The students put feedback, case quality, and faculty involvement ahead of novelty.
ClinicalSim Team
ClinicalSim
Researchers from CUHK Shenzhen, Nankai University, and the MIT Media Lab interviewed 12 clinical-year medical students and ran three codesign workshops. The students named six requirements for an AI patient they would actually use, and none of the six is realism.
The sample is small and qualitative, so it should guide design questions rather than stand in for the preferences of all medical students. The authors state the headline finding plainly:
"Our findings position AI-SPs as tools for deliberate practice and show that instructional usability, rather than conversational realism alone, drives learner trust, engagement, and educational value."
Gao Z, Zhu G, Luo H, et al., Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems, 2026
That sentence is worth sitting with, because most of the attention in this category has gone the other way. The race has been toward patients that sound more human. These students, who are the people expected to use the thing, asked instead for cases that state their purpose, reveal information for clinical reasons, and produce a review they can act on.
What the six requirements amount to
Students wanted the mode to match the goal. Exam preparation, open practice, and scaffolded skill development are three different jobs, and a single encounter design cannot serve each one equally well. A learner should be told whether they are exploring, practicing against a standard, or being assessed, because difficulty and feedback should follow from that.
They wanted the information rules to be legible. Visible findings should appear without a special prompt, other information should follow an appropriate clinical question, and the rule should hold across encounters. That is a case design problem more than a language model problem. The patient reveals something because of the clinical interaction, not because the learner guessed the right phrase.
They wanted control over difficulty and over how much help arrives mid-encounter, with the structured review saved for afterward. They wanted patient variation in age, emotional response, cultural background, and health beliefs, but tied to a learning objective: variation earns its place when it changes the behavior the learner has to demonstrate, and random novelty just makes cases less comparable.
And they wanted faculty in the loop. The researchers position AI patients as a complement to human standardized patients, who remain necessary for embodied interaction, live coaching, and high-stakes assessment.
What "instructional usability" looks like in a real report
The abstract phrase is easier to judge against something concrete, so here is one of the encounters we publish in full.
A pediatric critical care fellow takes a mother through consent for a central venous catheter in her five-year-old son. It runs fifteen minutes across 44 conversational turns, the longest encounter in the set. Scored against the six elements of an AMA-aligned consent discussion, it totals 22 out of 30.
The fellow scored 5 out of 5 on disclosing risks, benefits, and outcomes, and 5 out of 5 on confirming understanding. On assessing decision-making capacity, the first element, they scored 2 out of 5. The report is specific about why: this is a surrogate consent conversation, and the fellow never established the mother's role as decision maker, never checked what she already knew, and never asked how much detail she wanted. They answered her questions well as they came. They never set up the conversation.
That 2 is the whole value of the report. A learner who reads "22 out of 30" learns nothing. A learner who reads that they are strong at explaining and weak at opening has one thing to practice, and can go do it again that afternoon. This is what the students were asking for, and it has nothing to do with how convincing the voice on the other end sounds.
One requirement we do not meet
The students asked for a text fallback, keyword shortcuts, and hints, on the reasonable grounds that voice recognition fails and a technical error should never become evidence about clinical performance.
ClinicalSim is audio-only, deliberately. Voice practice surfaces pacing, silence, word choice, and how someone responds to emotion in real time, and a typed version of the same encounter does not test the skill we are trying to measure. That is a defensible trade, but it is a trade, and this study is a fair place to say so rather than quietly skipping the requirement we do not satisfy.
The test worth holding a case to
Every design decision above collapses into three questions a program can ask about any AI patient, ours included. Does the case state its objective? Is the conversation scored against a named framework or the program's own rubric? Is the learner's own language quoted under every score, so a faculty member can check the rating instead of trusting it?
A case that fails any of the three is entertainment, however good the voice is.
References
- Gao Z, Zhu G, Luo H, Pan DP, Tang H, Zhang B, Pei J, Li J, Wang B. "It Talks Like a Patient, But Feels Different": Co-Designing AI Standardized Patients with Medical Learners. Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems. 2026. doi:10.1145/3772363.3798336