There is a comforting story about weight stigma in medicine, and it goes like this: doctors were never taught better, so teach them better. Add a module. Bring in a fat patient to speak. Show the students the evidence on why bodies differ. The next generation will be different.
Someone has now added up thirty years of attempts to do exactly that. The answer is not that it fails. The answer is more uncomfortable than failure: it works on the part that is easy to measure and easy to fake, and it does not detectably move anything else.
What was actually counted
Ravisha S. Jayawickrama and colleagues at Curtin University in Perth, with co-authors at Monash University, the University of Leeds and the Swedish School of Sport and Health Sciences, screened 3,463 journal articles and dissertations and found 67 studies that tested an intervention meant to reduce weight bias in healthcare students. Thirty-five of them reported enough data to be pooled statistically; the rest were described narratively. Together the studies cover 7,528 participants, mostly women (62 percent), average age 22.75 years, from studies published between 1991 and August 2023.
This is not news, and we are not presenting it as news. The paper went online on 8 October 2024 and appeared in the February 2025 issue of Obesity Reviews (26(2):e13847). It is simply the most complete answer currently available to a question that gets asked every time a hospital announces a training day.
Two numbers carry the paper.
Explicit bias — what students report on a questionnaire — improved a little. The pooled effect was g = −0.31, with a 95 percent confidence interval of −0.43 to −0.19. In plain terms: a small but statistically reliable shift in the intended direction.
Implicit bias — the automatic association measured by a reaction-time test — did not move. g = −0.12, confidence interval −0.26 to +0.02, p = 0.105, pooled from ten studies. The interval crosses zero, which means the data are compatible with a small improvement, with nothing at all, and with a slight worsening.
The authors grade the certainty of both results, using the standard GRADE system, as “very low”. They also record that the risk of bias in most of the individual studies was high. That is their assessment of their own evidence base, not our gloss on it.
The sentence that should be quoted more than the effect size
Buried in the results is a prediction interval for the explicit-bias finding: −0.93 to +0.31.
A confidence interval tells you how precisely the average across these studies was estimated. A prediction interval tells you what to expect from the next one. Here it says that if you run a weight-bias intervention with a new group of healthcare students, the true effect could plausibly be a large reduction, or nothing, or a modest increase in weight bias. That is what heterogeneity of 74 percent looks like when it is written out honestly.
So the pooled figure is real, and it is also nearly useless as a prediction for any specific course. Which brings us to the finding that should worry course designers most.
Nothing about the course design explained the difference
The team ran eight subgroup comparisons on the explicit-bias data: healthcare discipline, one-off versus multi-session, in person versus online, active versus passive learning, single-strategy versus multifaceted, underlying theory, which outcome measure was used, and whether students had contact with an actual fat person during the intervention — active contact, passive contact, or none.
Not one comparison came out significant. A meta-regression on the proportion of male participants found nothing either.
The honest reading of that is not “these things don’t matter”. With this few studies per cell, the analysis could not have detected moderate differences even if they existed, and the authors say as much: eyeballing the estimates, multi-session, multifaceted and active-learning formats look better, but the comparison groups are too small to trust. The reading that does hold is narrower and still awkward: after thirty years of varied designs, there is no evidence base telling anyone which design to buy. A hospital choosing between a two-hour e-learning module and a semester-long curriculum with patient speakers is choosing on intuition.
And it does not stay
Only 10 of the 67 studies followed students beyond the end of the intervention at all.
The example the authors highlight is a study by Kushner and colleagues with 127 medical students. Empathy and confidence in clinical interaction were still improved a year later. The negative stereotypes about fat patients had returned to baseline.
That is a pattern worth naming precisely, because it is easy to get backwards: the part that felt good to the clinician persisted. The part that describes the patient did not.
What the questionnaire is actually measuring
Explicit weight bias, in these studies, means a score on instruments like the Beliefs About Obese Persons Scale, the Antifat Attitudes Test, or the Fat Phobia Scale — a student ticking boxes about whether fat people are lazy, or whether obesity is within a person’s control.
Only four of the 35 pooled studies checked whether students were simply giving the answer they knew was wanted. The findings from those four were mixed, and in one of them, controlling for socially desirable responding erased the difference between the intervention group and the control group entirely. The authors are careful not to over-read four studies, and so are we. But it means the most obvious alternative explanation for a small post-course improvement in questionnaire scores — that students learned which box to tick — has barely been tested.
The implicit measure has its own problem, and the paper does not hide it. Twelve of the thirteen implicit studies used the Implicit Association Test, and the authors cite the standing criticism of it, including Ulrich Schimmack’s argument that the construct validity evidence does not support the claim that the IAT measures implicit bias at all.
Put those two limitations together and you get the real state of knowledge. The first instrument measures what a student is willing to say. The second measures a reaction time whose meaning is contested. Almost nothing measures what a student later does to a patient. The authors call for exactly that: studies that assess stigmatising behaviour, not just attitudes.
Who is speaking, and why that is worth knowing
Obesity Reviews is the journal of the World Obesity Federation, and the paper is written from inside that frame. It opens with the Federation’s projection that around 51 percent of people aged five and over will be living with overweight or obesity by 2035, uses “people living with overweight or obesity” throughout, and argues for reducing weight bias partly on the grounds that stigma worsens the conditions it is attached to.
We link it anyway, because the analysis is careful and the numbers are the numbers. But readers should know the frame, and they should know the declared interests: co-author Stuart W. Flint reports research grants and meeting support from, among others, Novo Nordisk, the Novo Nordisk Foundation and Johnson & Johnson, all declared as unrelated to this manuscript, and co-author Erik Hemmingsson reports royalties from a book on weight stigma. The first author and five others declare no competing interests.
None of that invalidates a meta-analysis. It does mean that when a paper written from the disease frame reports that anti-bias training barely works, that finding is not the one its institutional context would have preferred.
Germany just made this a “should”
In October 2024, the German Adipositas Society and its partner societies published version 5.0 of the S3 guideline on the prevention and treatment of obesity (AWMF register 050/001). For the first time it contains a dedicated chapter on stigma, and inside that chapter, recommendation 2.4:
Training curricula for the health professions involved in prevention and treatment should not only inform about the aetiology, mechanisms, prevention and treatment of obesity, but should also educate about weight-related stigmatisation and self-stigmatisation and their clinical implications, and should teach practical skills for non-stigmatising interaction with people with obesity.
It carries recommendation grade A — the strongest “should” the system has — and was adopted with 94 percent consensus. The evidence quality printed directly underneath it is “very low”, based on indirect evidence.
The guideline is not hiding this. Its own implications section states that promising results exist for at least small, short-term improvements in weight stigma among health professions, and that all meta-analyses on weight-related stigmatisation showed low methodological quality.
So Germany has now committed its health professions curricula, at the highest recommendation grade, to an intervention its own guideline describes as producing small, short-lived, poorly evidenced change. That is not a scandal. Given what stigma costs patients, acting on weak evidence is defensible, and the alternative — teaching nothing — has been tested for a century and produced the situation the guideline is responding to.
It does mean something specific for patients, though: the training is a promise made about your doctor, upstream of you, and nobody checks whether it took.
What is actually checkable
The same chapter of the same guideline contains something a patient can verify from the waiting room. Recommendation 2.2, adopted with 88 percent consensus, says facilities should provide adequate equipment — it names heavy-duty chairs and scales with a sufficient weighing range — or maintain a referral network for onward care.
That is the difference between the two kinds of measure. Whether a practice bought a chair that holds you is visible in ten seconds. Whether the doctor sitting in it completed an anti-bias module in 2019 is not visible at all, and on this evidence would not tell you much if it were.
Which is why the practical advice on this site does not run through hoping your doctor was trained well:
- When your symptom is being answered with your weight, the documentation and escalation steps are in What to Do When a Doctor Blames Everything on Your Weight.
- When you are looking for a different practice, the search criteria that actually predict weight-neutral care are in How to Find a Weight-Neutral (HAES-Aligned) Doctor.
- When the problem is discrimination rather than bad care, the legal picture is in Weight Discrimination: What It Is and What You Can Do.
The point that survives all the caveats
Thirty years of teaching produced a small, real improvement in what healthcare students are willing to say about fat patients, no measurable change in their automatic associations, no durable change in stereotypes a year out, and almost no data on behaviour.
Read carelessly, that is an argument for giving up on training. It is not. A small genuine reduction in explicit bias, spread across everyone who will treat patients for the next forty years, is worth having, and the alternative has no evidence behind it either.
Read carefully, it is an argument about where to put the weight of expectation. A course is not a safeguard. It produces an effect that cannot be predicted for any individual cohort, cannot be attributed to any particular design, and fades on the measure that matters most for how you get spoken to. If the goal is that fat patients are treated properly, the mechanisms that do that job are the ones that stay in place when everyone has forgotten the workshop: equipment that fits, referral paths, complaint routes, documentation, and rules about what may be said and done in a treatment room.
Prejudice, it turns out, is not primarily an information deficit. You cannot explain it away, and the people who study it hardest say so in their own conclusion: a real shift in implicit bias may only come from a large societal change in beliefs and attitudes towards people in larger bodies. That is a longer job than a module, and it is the one this magazine exists for.
Sources: Jayawickrama RS, Hill B, O’Connor M, Flint SW, Hemmingsson E, Ellis LR, Du Y, Lawrence BJ. “Efficacy of interventions aimed at reducing explicit and implicit weight bias in healthcare students: A systematic review and meta-analysis.” Obesity Reviews 26(2):e13847, published online 8 October 2024, DOI 10.1111/obr.13847, open access: https://pmc.ncbi.nlm.nih.gov/articles/PMC11711078/ · PROSPERO registration CRD42020209407 · Deutsche Adipositas-Gesellschaft et al., “Interdisziplinäre Leitlinie der Qualität S3 zur Prävention und Therapie der Adipositas”, AWMF register 050/001, version 5.0, October 2024, chapter 2 “Stigmatisierung”, recommendations 2.1 to 2.4 and section 2.4 “Implikationen”: https://register.awmf.org/assets/guidelines/050-001l_S3_Praevention-Therapie-Adipositas_2024-10.pdf










