Methodology: case creation, standards alignment, and feedback generation
Jacqueline Ponczek, MD, MS, FAAP
VP of MedEd: Quality & Standards, ClinicalSim
This page explains how ClinicalSim builds cases, aligns them to communication and governing-body frameworks, scores them, and turns each encounter into high-quality, actionable feedback. One engine, rubric, and dashboard serve learners across the medical education continuum. Every session produces timestamped, competency-based documentation for learners, faculty, and program leadership. For the institutional view of the same loop, see what clinical communication intelligence means.
Key takeaway: every ClinicalSim case is anchored to a specific, published competency or communication standard. Every score traces to a verbatim excerpt from the encounter transcript, never to an unexplained rating.
1. Purpose and scope
Every case is anchored to the relevant governing body’s framework for the learner’s level. That holds whatever primary measure a program chooses. A program may make an internal or externally validated tool its primary focus, or fold ClinicalSim cases into a broader curriculum. Either way, each case’s scoring and feedback rest on a specific, published standard.
Three commitments
We share this methodology to keep our work transparent, so the people who rely on it can trust it. Three commitments anchor it:
- Quality, because every case is built from primary sources.
- Consistency, because the same scoring logic applies to every case.
- Alignment, because every score traces to a published competency or a validated communication framework.
Quality
Every case is built from primary sources, written to a defined purpose, and reviewed by practicing physicians with strong academic backgrounds before release.
Consistency
The same scoring logic applies to every case, regardless of specialty, learner level, or which communication frameworks are applied alongside it.
Alignment
Every score traces to a program-approved competency standard or a published communication framework, never to an unexplained rating.
2. How the methodology works
The method below applies to every case at every learner level. Where something does vary by level, the subsections under Evidence and scoring describe it.
2.1 Building a case
A defined purpose
Every case begins with a defined purpose. The purpose names the communication and clinical skills the case should exercise and the competencies it should assess. The content is written to that purpose, with explicit learning objectives and a clinical evidence base drawn from foundational and other applicable literature.
Physician review
Physicians then review each case for accuracy, content, alignment, and fit to its objectives. The reviewers are practicing physicians with strong academic backgrounds and decades of collective experience. They include program directors, simulation facilitators, and UME and GME educators. Faculty development cases are also reviewed by someone with faculty development or clinical teaching expertise.
Testing before release
Before release, each case is run repeatedly to confirm three things:
- The AI character convincingly plays the role the case requires.
- Scoring and feedback perform as intended.
- What the case asks can be assessed within the limits of voice-based simulation.
Refinements are made in coordination with ClinicalSim’s clinical and technical leadership.
2.2 Competency alignment and communication frameworks
Three terms recur here. The competency framework anchors the competency assessment, and the program must supply or approve it. Communication frameworks are then applied to characterize how the learner communicated.
The two are distinct. The competency score reflects the learner’s developmental level. The communication frameworks capture the specific skills behind communication technique.
Competency framework
The standard a program approves for the case, including any rating scale and the behaviors the conversation can assess.
Communication framework
A validated, published model of communication behavior, such as SPIKES or Calgary-Cambridge, applied to characterize how the learner communicated.
Rubric
The scored instrument that turns a framework into rated items, including a program's own internal or externally validated tools.
How frameworks are chosen
Each communication framework comes from a cited, published source, and the frameworks are a floor, not a ceiling. One or more may be applied to a case, each is scored independently, and programs may add their own internal or externally validated rubrics.
These frameworks and rubrics work at different scopes, from whole-encounter structures to task-specific routines to discrete micro-skills. So ClinicalSim selects the ones best suited to each case’s communication task.
2.3 Evidence and scoring
Evidence from the transcript
Each encounter is a voice conversation between the learner and an AI role designed for the case, captured as a timestamped transcript. For every scored competency and framework step, the platform pulls one or two verbatim excerpts that show the behavior, or notes that it was absent. Each score traces to the moment that supports it, so the output holds up to review instead of standing as an unexplained rating.
What gets scored
Scoring follows the competency framework the case is built on. The unit of assessment is the individual competency the case exercises. Each communication framework or program rubric applied to the case is scored independently of the competency.
Rating scales
When an instrument publishes its own rating scale, we use it. Most communication frameworks do not. For those, we apply a ClinicalSim scale to the framework’s own steps and say so in the case, so a score is never read as though the framework’s validation stood behind it.
Reading scores by stage of training
These frameworks are developmental. The same result means different things at different stages of training, and it is always read that way. All scores are shown together with their verbatim evidence, so the learner or reviewer sees the complete picture.
What varies by learner level is the competency framework a case is anchored to and how the competency itself is scored. The sections below describe each level.
Graduate medical education
Residency and fellowship cases use the competency framework and rating scale that the program approves for that case. Each case targets a high-stakes conversation the specialty needs to rehearse. It scores only the behaviors the conversation can show.
The report names the standard behind each score and cites the learner’s words. Faculty and the Clinical Competency Committee can review it alongside direct observation and the program’s other evidence. ClinicalSim does not replace their judgment.
Undergraduate medical education
Cases align to two standards:
- The Foundational Competencies for Undergraduate Medical Education (AAMC, AACOM, and ACGME)
- The AAMC Core Entrustable Professional Activities (EPAs) for Entering Residency
The Core EPAs were originally mapped to the Physician Competency Reference Set (PCRS, 2013). The 2024 Foundational Competencies now supersede the PCRS. An updated set of EPAs aligned to the Foundational Competencies is anticipated but not yet published.
Until it is, ClinicalSim maps UME cases to the EPAs and to the Foundational Competencies independently. It does not assert a fixed crosswalk between them.
For UME, ClinicalSim records each competency on three points:
- Demonstrated
- Partially demonstrated
- Not demonstrated
Performance is scored through the applied communication or skill rubric. Entrustment, the pre-entrustable to entrustable judgment, remains a program decision that this evidence informs.
Development focuses on foundational encounters that mature alongside clinical knowledge, from history-taking to delivering a diagnosis. The aim is to prepare students for the transition to residency.
Faculty development
Faculty cases assess a faculty member or other teaching clinician, and the assessment is formative.
Faculty cases use the framework or rubric that the program approves for the teaching task. The report is formative evidence that informs judgment, not a grade.
A case scores only the subcompetencies a voice conversation can actually show, usually one or two. That keeps a case from taking on more than it can assess, and keeps every score to something the platform can evidence.
Cases are built in four types, according to who the faculty member is talking to:
- A student, resident, or fellow
- Another faculty member
- A patient or caregiver
- Other healthcare staff
2.4 Feedback
Each encounter produces a single feedback report. It includes:
- Verbatim evidence inside the grading rubrics, which justifies the level a learner reached or the specific step assessed
- An overall impression covering strengths, priority gaps, and top action items
- Targeted recommendations
Depending on the case, the report shows where a learner sits developmentally. It also gives reviewers transcript-grounded evidence for decisions about:
- Progression
- Remediation
- Readiness for practice
- Readiness to perform a particular task
- Familiarity with a given subject area
2.5 What a score claims, and what it does not
The frameworks ClinicalSim applies were built for trained human raters observing real encounters. That is how their published reliability was established. Scoring them with AI in a simulated encounter goes beyond that context, so a framework’s reliability does not carry over to a ClinicalSim score. Each score is a formative signal backed by verbatim transcript evidence.
How we are testing it
We are testing that rather than asserting it. In our current pilot, program directors review ClinicalSim output alongside their own assessment of the same encounters. That comparison is how we find out where the platform holds up against the standard, and where it complements faculty judgment rather than substituting for it.
3. Commitment to accuracy
Read every result as evidence, not a verdict. We are committed to accuracy and to fidelity to the source documents behind every case. Each result is a transparent statement of the evidence in the encounter. It informs the learner and the reviewer, and it never replaces final human judgment.
4. Common questions
How does ClinicalSim's AI scoring work?
Each ClinicalSim encounter is a voice conversation between the learner and an AI patient built for that case, captured as a timestamped transcript. For every scored competency and framework step, the platform pulls one or two verbatim excerpts from the transcript that show the behavior, or notes that it was absent.
Scoring follows the competency framework the case is anchored to, and the unit of assessment is the individual competency the case exercises. Any communication framework or program rubric applied alongside it is scored separately, so the two never collapse into one number.
Which competency frameworks does a ClinicalSim score map to?
ClinicalSim can map a case to the competency framework a program supplies or approves. The case uses only the behaviors the conversation can show, and the report names the standard behind each score.
Communication frameworks and program rubrics are scored separately.
Can faculty see the evidence behind a ClinicalSim score?
Every score in a ClinicalSim report carries the verbatim transcript excerpt that produced it. A reviewer reads the moment in the conversation instead of taking the rating on trust.
The report shows all scores together with their evidence and adds an overall impression covering strengths, priority gaps, and top action items. That gives faculty transcript-grounded evidence for decisions about progression, remediation, or readiness.
What does a ClinicalSim score claim, and what does it not claim?
The communication frameworks ClinicalSim applies were built for trained human raters observing real encounters. Their published reliability was established in that context.
Scoring those frameworks with AI in a simulated encounter goes beyond it, so a framework's published reliability does not transfer to a ClinicalSim score. Each score is a formative signal backed by verbatim transcript evidence, which is why this methodology asks readers to treat every result as evidence, not a verdict.
Should AI-generated scores be used for promotion or remediation decisions?
ClinicalSim scores are formative and do not stand alone behind a decision about promotion or remediation. A competency committee weighs the report alongside direct observation and faculty judgment, and people make the final decision.
In the current pilot, program directors assess the same encounters and compare their own read with the platform's output.
Who writes and reviews ClinicalSim cases?
Every ClinicalSim case starts from a defined purpose: the communication and clinical skills it should exercise and the competencies it should assess. It is written to that purpose, with explicit learning objectives and a clinical evidence base drawn from the literature.
Practicing physicians, including program directors, simulation facilitators, and educators from undergraduate and graduate medical education, then review it for accuracy, content, alignment, and fit to its objectives. Faculty development cases get an extra review from someone with faculty development or clinical teaching expertise, and every case is run repeatedly before release.
5. References
Undergraduate medical education
- AAMC, AACOM, and ACGME. Foundational Competencies for Undergraduate Medical Education. 2024.
- AAMC. The Core Entrustable Professional Activities (EPAs) for Entering Residency. 2014. aamc.org
Communication frameworks (representative; full citations in the ClinicalSim Frameworks Bibliography)
- SPIKES. Baile WF, et al. The Oncologist. 2000;5(4):302-311.
- KEECC-A (Kalamazoo). Makoul G. Acad Med. 2001;76(4):390-393.
- SEGUE. Makoul G. Patient Educ Couns. 2001;45(1):23-34.
- NURSE. Back AL, et al. CA Cancer J Clin. 2005;55(3):164-177.
- REMAP. Childers JW, et al. J Oncol Pract. 2017;13(10):e844-e850.
- SBAR. Haig KM, et al. Jt Comm J Qual Patient Saf. 2006;32(3):167-175.
- I-PASS. Starmer AJ, et al. Pediatrics. 2012;129(2):201-204.
- TeamSTEPPS. King HB, et al. AHRQ; 2008. CANDOR. AHRQ; updated 2023.
- Calgary-Cambridge. Silverman J, Kurtz S, Draper J. Skills for Communicating with Patients. 3rd ed. Radcliffe Publishing; 2013. Companion volume: Kurtz S, Silverman J, Draper J. Teaching and Learning Communication Skills in Medicine. 2nd ed. Radcliffe; 2005.
- R2C2. Sargeant J, Lockyer J, Mann K, et al. Acad Med. 2015;90(12):1698-1706.