My colleague recently used an AI-powered clinical reference tool to look up the symptoms of psoriasis. The answer she got was thorough, confident, and wrong in one important way: the diagnostic description focused entirely on how the condition presents on lighter skin. Only when she explicitly typed "darker skin" did the tool adjust and give her a fuller picture. This wasn't a glitch – it was reflecting decades of medical literature that treats lighter skin as the clinical default. This kind of result is why experts have urged developers to measure equity in addition to fairness within clinical AI models.
Built in, not bolted on
Health equity means giving every patient a fair and just opportunity to reach their best possible health outcome. At Doximity, advancing health equity isn't a side initiative — it’s core to how we build tools clinicians can trust.
As we've expanded into clinical AI with Doximity Ask, our approach has evolved into a health equity model designed to keep bias from hiding within confident-sounding answers. This means actively testing across topics like race, skin tone, language and dialect, socioeconomic status, immigration status, and disability — not waiting for a user to notice something's missing, as my colleague did.
The problem is bigger than one tool
The example about psoriasis is not unique. We've seen versions of this problem before: pulse oximeters that read less accurately on darker skin and a widely used hospital algorithm that used healthcare cost as a proxy for illness — a choice that made Black patients half as likely to be flagged for extra care. Digitizing these clinical processes without an equity lens didn't remove the underlying bias. It scaled it, quietly and at speed.
When clinical and research leaders choose to confront these biases directly, the results have been positive. After removing race from kidney function equations, transplant programs revisited prior evaluations and it resulted in 5.3 more transplants for every thousand Black patients on the waitlist. Equity-minded redesign isn't a compliance exercise. In clinical contexts, it can be the difference between someone getting care and dying on the waitlist.
Why "human-in-the-loop" isn't the whole answer
"Human-in-the-loop" has become shorthand for AI accountability, but in clinical settings it's often underspecified. A more precise framework, borrowed from a recent NEJM piece, splits oversight into three loops: clinical (a physician catches a bad answer at the point of care), governance (how tools get approved and deployed), and learning (how models are monitored and updated after release). Most conversations about AI and equity stop at the clinical loop — useful, but insufficient, since catching an error once doesn't stop it from recurring at scale. And we know this isn't hypothetical. A recent Stanford-Harvard evaluation of clinical AI found the most serious safety failures came from omitted recommendations, not incorrect ones — the kind of gap a single point-of-care review is least likely to catch.
How we govern it: Detect, Measure, Report, Mitigate
Over the past year, our team has built a model that captures those three loops: Detect, Measure, Report, Mitigate.
Detect defines what bias looks like at the point of use — in Doximity Ask's case, auditing outputs across topics like racial bias, skin tone representation, and the safety of immigrant and LGBTQ+ populations.
Measure turns that into a repeatable benchmark. We built more than 30 specialized prompts and 150 judge criteria, developed with physician colleagues, to generate a daily Health Equity Score — allowing us to catch directional shifts in model behavior over time, not just at a single audit point.
Report converts findings into institutional memory through ongoing case reporting and PeerCheck expert feedback, giving our developers actionable, contextualized insight instead of just a flagged error.
Mitigate is where that insight becomes engineering work — fine-tuning and retraining cycles that move health equity from a stated value to a technical input.
This is the model that helped close the gap for psoriasis and other skin conditions. Reporting the issue in social and clinical context led to a systemic solution — ensuring dermatologic descriptions are grounded across the full Fitzpatrick skin tone scale, not just on lighter skin.

This is a practice, not a project
Developing health equity strategies in clinical AI isn't a milestone you hit once — it's a system that keeps observing and learning. By understanding an equitable answer is simply a more accurate one, we can continue to prioritize closing the gap between a tool's output and the clinical reality each patient deserves.
