ClinicalSim

Methodology: case creation, standards alignment, and feedback generation

JP

Jacqueline Ponczek, MD, MS, FAAP

VP of MedEd: Quality & Standards, ClinicalSim

Last updated: August 2026

This page outlines ClinicalSim case development, communication and governing-body framework alignment, scoring, and how each encounter generates high-quality, actionable feedback. A single engine, rubric, and dashboard serve learners across the medical-education continuum, and every session produces timestamped, competency-based documentation for learners, faculty, and program leadership.

Key takeaway: every ClinicalSim case is anchored to a specific, published competency or communication standard, and every score traces to a verbatim excerpt from the encounter transcript, never to an unexplained rating.

1. Purpose and scope

Every case is anchored to the relevant governing body’s framework for the learner’s level. This anchoring holds regardless of a program’s chosen primary measure. A program may adopt an internal or externally validated tool as its primary focus, or incorporate ClinicalSim cases into a broader curriculum. Each case’s scoring and feedback are always grounded in a specific, published standard.

Sharing our methodology keeps our work transparent, so those who rely on it can trust it. Three commitments anchor it: quality, because every case is built from primary sources; consistency, because the same scoring logic applies to every case; and alignment, because every score traces to a published competency or a validated communication framework.

Quality

Every case is built from primary sources, written to a defined purpose, and reviewed by practicing physicians with strong academic backgrounds before release.

Consistency

The same scoring logic applies to every case, regardless of specialty, learner level, or which communication frameworks are applied alongside it.

Alignment

Every score traces to a published competency or a validated communication framework, never to an unexplained rating.

2. How the methodology works

The method below applies to every case, regardless of learner level. Within Evidence and scoring, the level subsections describe what varies by learner level.

2.1 Building a case

Every case begins with a defined purpose: the communication and clinical skills it should exercise and the competencies it should assess. Content is written to that purpose, with explicit learning objectives and a clinical evidence base drawn from foundational and other applicable literature.

Physicians then review each case for accuracy, content, alignment, and fit to its objectives; reviewers are practicing physicians with strong academic backgrounds and decades of collective experience, including program directors, simulation facilitators, and UME and GME educators. Faculty development cases are also reviewed by someone with faculty development or clinical teaching expertise.

Before release, each case is run repeatedly to confirm three things: that the AI character convincingly plays the role the case requires; that scoring and feedback perform as intended; and that what the case asks can be assessed within the limits of voice-based simulation. Refinements are made in coordination with ClinicalSim’s clinical and technical leadership.

2.2 Competency alignment and communication frameworks

Three terms recur here. The competency framework is the anchor for the competency assessment; where it defines level descriptors, as the ACGME milestones do, those are quoted verbatim from the primary source. The communication frameworks are then applied to characterize how the learner communicated. The two are distinct: the competency score reflects the learner’s developmental level, while the communication frameworks capture the specific skills underlying communication technique.

Competency framework

The governing-body standard a case is assessed against: the ACGME Milestones 2.0 in graduate medical education, the Foundational Competencies in undergraduate medical education, and the ACGME Clinician Educator Milestones in faculty development.

Communication framework

A validated, published model of communication behavior, such as SPIKES or Calgary-Cambridge, applied to characterize how the learner communicated.

Rubric

The scored instrument that turns a framework into rated items, including a program's own internal or externally validated tools.

Each communication framework comes from a cited, published source and is a floor, not a ceiling: one or more may be applied to a case, each scored independently, and programs may add their own internal or externally validated rubrics. Because these frameworks and rubrics operate at different scopes, from whole-encounter structures to task-specific routines to discrete micro-skills, ClinicalSim selects those best suited to each case’s communication task.

2.3 Evidence and scoring

Each encounter is a voice conversation between the learner and an AI role designed for the case, captured as a timestamped transcript. For every scored competency and framework step, the platform draws one or two verbatim excerpts that demonstrate the behavior, or documents its absence. Because each score is traceable to the moment that supports it, the output withstands review rather than serving as an unexplained rating.

Scoring follows the competency framework on which a case is built, and the unit of assessment is the individual competency the case exercises. Each applied communication framework or program rubric is scored independently of the competency.

Where an instrument publishes its own rating scale, we use it. Most communication frameworks do not, and there we apply a ClinicalSim scale to the framework’s own steps and say so in the case, so a score is never read as though the framework’s validation stood behind it.

Because these frameworks are developmental, a given result carries different meaning at different stages of training and is always interpreted accordingly. All scores are presented together, with their verbatim evidence, so the learner or reviewer sees a complete picture. What varies is the competency framework a case is anchored to and how the competency itself is scored, described by learner level below.

Graduate medical education

Cases align to the specialty-specific ACGME Milestones 2.0, with milestone text quoted verbatim from each specialty’s own document, and target the high-stakes conversations a specialty most needs to rehearse. The Milestones 2.0 describe six core competencies across five developmental levels; several were harmonized across specialties in 2017 and then adapted by each specialty, which is why text is drawn from the specialty’s own version.

Scoring reflects whichever subcompetencies the scenario exercises, most often interpersonal and communication skills and professionalism, and systems-based practice or other domains where the encounter warrants. Each is scored on the Dreyfus scale (1 to 5), read against the milestone’s verbatim level descriptors:

  • Level 1, Novice
  • Level 2, Advanced Beginner
  • Level 3, Competent
  • Level 4, Proficient (readiness for unsupervised practice)
  • Level 5, Expert (aspirational)

The result is milestone-placed and ready for Clinical Competency Committee review. Because the milestones are formative and were not designed for high-stakes external decisions, ClinicalSim treats milestone-aligned output accordingly, as evidence that informs program judgment. With that in mind, we do not provide a milestone score if the case cannot achieve the level of complexity required to evaluate through level 4.

Undergraduate medical education

Cases align to the Foundational Competencies for Undergraduate Medical Education (AAMC, AACOM, and ACGME) and the AAMC Core Entrustable Professional Activities (EPAs) for Entering Residency. The Core EPAs were originally mapped to the Physician Competency Reference Set (PCRS, 2013), which the 2024 Foundational Competencies now supersede; an updated set of EPAs aligned to the Foundational Competencies is anticipated but not yet published. Until it is, ClinicalSim maps UME cases to the EPAs and to the Foundational Competencies independently, without asserting a fixed crosswalk between them.

For UME, ClinicalSim records each competency on three points, demonstrated, partially demonstrated, or not demonstrated, and scores performance through the applied communication or skill rubric. Entrustment, the pre-entrustable to entrustable judgment, remains a program decision that this evidence informs.

Development emphasizes foundational encounters that mature alongside clinical knowledge, from history-taking to delivering a diagnosis, preparing students for the transition to residency.

Faculty development

Faculty cases assess a faculty member or other teaching clinician, and the assessment is formative.

Cases align to the ACGME Clinician Educator Milestones, a 2022 joint initiative of the ACGME, the ACCME, the AAMC, and the AACOM. Its 19 subcompetencies carry the same five developmental levels as the ACGME Milestones used in graduate medical education, and level descriptors are quoted verbatim from the source.

The Clinician Educator Milestones are guidance rather than an accreditation requirement, in the ACGME’s own words. ClinicalSim treats faculty output the way it treats milestone-aligned output in residency: evidence that informs judgment, not a grade.

A case scores only the subcompetencies a voice conversation can actually show, usually one or two. That keeps a case from taking on more than it can assess, and keeps every score to something the platform can evidence.

Cases are built in four types, according to who the faculty member is talking to:

  • A student, resident, or fellow
  • Another faculty member
  • A patient or caregiver
  • Other healthcare staff

2.4 Feedback

Each encounter produces a single feedback report.

Verbatim evidence is incorporated into the grading rubrics, justifying the level a learner reached or the specific step assessed. The report then offers an overall impression (strengths, priority gaps, and top action items) and targeted recommendations. Depending on the case, it indicates where a learner sits developmentally and provides reviewers with transcript-grounded evidence for decisions about progression, remediation, readiness for practice, readiness to perform a particular task, or familiarity with a given subject area.

2.5 What a score claims, and what it does not

The frameworks ClinicalSim applies were built for trained human raters observing real encounters, and that is how their published reliability was established. Scoring them with AI in a simulated encounter is an extension beyond that context, so a framework’s reliability does not carry over to a ClinicalSim score. Each score is a formative signal backed by verbatim transcript evidence.

We are testing that rather than asserting it. In our current pilot, program directors review ClinicalSim output alongside their own assessment of the same encounters, which is how we find out where the platform holds up against the standard and where it complements faculty judgment rather than substituting for it.

3. Commitment to accuracy

Read every result as evidence, not a verdict. We are committed to accuracy and to fidelity to the source documents behind every case. Each result is a transparent statement of the evidence in the encounter: it informs the learner and the reviewer, and it never replaces final human judgment.

4. Common questions

How does ClinicalSim's AI scoring work?

Each ClinicalSim encounter is a voice conversation between the learner and an AI patient built for that case, captured as a timestamped transcript. For every scored competency and framework step, the platform pulls one or two verbatim excerpts from that transcript showing the behavior, or documents that it was absent. Scoring follows the competency framework the case is anchored to, and the unit of assessment is the individual competency the case exercises. Any communication framework or program rubric applied alongside it is scored separately, so the two are never collapsed into one number.

Which competency frameworks does a ClinicalSim score map to?

ClinicalSim anchors each case to a published competency framework and quotes that framework's level descriptors verbatim from the primary source. Residency and fellowship cases use the specialty-specific ACGME Milestones 2.0, scored on the Dreyfus scale from 1 to 5. Medical school cases use the Foundational Competencies for Undergraduate Medical Education (AAMC, AACOM, and ACGME) and the AAMC Core Entrustable Professional Activities, recorded on three points as demonstrated, partially demonstrated, or not demonstrated. Faculty cases use the ACGME Clinician Educator Milestones. Published communication frameworks are then applied on top of the competency score, each scored independently, and a program can add its own internal or externally validated rubrics.

Can faculty see the evidence behind a ClinicalSim score?

Every score in a ClinicalSim report carries the verbatim transcript excerpt that produced it, so a reviewer reads the moment in the conversation rather than taking the rating on trust. The report presents all scores together with their evidence, adds an overall impression covering strengths, priority gaps, and top action items, and gives faculty transcript-grounded evidence for decisions about progression, remediation, or readiness.

What does a ClinicalSim score claim, and what does it not claim?

The communication frameworks ClinicalSim applies were built for trained human raters observing real encounters, and that is the context in which their published reliability was established. Scoring those frameworks with AI in a simulated encounter goes beyond that context, so a framework's published reliability does not transfer to a ClinicalSim score. Each score is a formative signal backed by verbatim transcript evidence, which is why this methodology asks a reader to treat every result as evidence rather than a verdict.

Should AI-generated scores be used for promotion or remediation decisions?

ClinicalSim scores are formative and are not built to stand alone behind a decision about promotion or remediation. The ACGME milestones themselves were designed as formative tools rather than instruments for high-stakes external decisions, and ClinicalSim treats milestone-aligned output the same way, as evidence that informs program judgment. A competency committee weighs it alongside direct observation and faculty judgment, and the final judgment stays with people. ClinicalSim is testing that alignment rather than asserting it: in the current pilot, program directors assess the same encounters themselves and compare their own read against the platform's output.

Who writes and reviews ClinicalSim cases?

Every ClinicalSim case starts from a defined purpose, meaning the communication and clinical skills it should exercise and the competencies it should assess, and it is written to that purpose with explicit learning objectives and a clinical evidence base drawn from the literature. Practicing physicians then review it for accuracy, content, alignment, and fit to its objectives, among them program directors, simulation facilitators, and educators from both undergraduate and graduate medical education. Faculty development cases carry an additional review by someone with faculty development or clinical teaching expertise, and every case is run repeatedly before release.

5. References

Graduate medical education

  1. Edgar L, Roberts S, Holmboe E. Milestones 2.0: A Step Forward. J Grad Med Educ. 2018;10(3):367-369.
  2. Morrison LJ, Joyce BL, Meyer LE, et al. Strengthening Interpersonal and Communication Skills Assessment Through Harmonized Milestones. J Grad Med Educ.
  3. ACGME. Use of Individual Milestones Data by External Entities for High-Stakes Decisions: A Function for Which They Are Not Designed or Intended. October 2022.
  4. Specialty-specific ACGME Milestones and Supplemental Guides, sourced per specialty.

Undergraduate medical education

  1. AAMC, AACOM, and ACGME. Foundational Competencies for Undergraduate Medical Education. 2024.
  2. AAMC. The Core Entrustable Professional Activities (EPAs) for Entering Residency. 2014. aamc.org

Faculty development

  1. The Clinician Educator Milestone Project. Version 1.1, August 2022. A joint initiative of the Accreditation Council for Graduate Medical Education, the Accreditation Council for Continuing Medical Education, the Association of American Medical Colleges, and the American Association of Colleges of Osteopathic Medicine.
  2. The Clinician Educator Supplemental Guide. Version 1.1, October 2025.
  3. ACGME. Use of Individual Milestones Data by External Entities for High-Stakes Decisions. October 2022.

Communication frameworks (representative; full citations in the ClinicalSim Frameworks Bibliography)

  1. SPIKES. Baile WF, et al. The Oncologist. 2000;5(4):302-311.
  2. KEECC-A (Kalamazoo). Makoul G. Acad Med. 2001;76(4):390-393.
  3. SEGUE. Makoul G. Patient Educ Couns. 2001;45(1):23-34.
  4. NURSE. Back AL, et al. CA Cancer J Clin. 2005;55(3):164-177.
  5. REMAP. Childers JW, et al. J Oncol Pract. 2017;13(10):e844-e850.
  6. SBAR. Haig KM, et al. Jt Comm J Qual Patient Saf. 2006;32(3):167-175.
  7. I-PASS. Starmer AJ, et al. Pediatrics. 2012;129(2):201-204.
  8. TeamSTEPPS. King HB, et al. AHRQ; 2008. CANDOR. AHRQ; updated 2023.
  9. Calgary-Cambridge. Silverman J, Kurtz S, Draper J. Skills for Communicating with Patients. 3rd ed. Radcliffe Publishing; 2013. Companion volume: Kurtz S, Silverman J, Draper J. Teaching and Learning Communication Skills in Medicine. 2nd ed. Radcliffe; 2005.
  10. R2C2. Sargeant J, Lockyer J, Mann K, et al. Acad Med. 2015;90(12):1698-1706.

Questions about how this works?

Read the wider FAQ for questions about pricing, rollout, and program fit, or talk to us about piloting ClinicalSim at your program.