Why Nutrition Advice Keeps Changing (and Why Both Sides Have a Study)
Your doctor and your trainer give you opposite nutrition advice and both are quoting research. Here is the structural reason the literature contradicts itself, and how to judge a claim.
You were told eggs were dangerous. Then you were told they were fine.
If you feel like nutrition advice reverses itself every few years, you are not imagining it. The reversals are public record.
The cap on dietary cholesterol was dropped from the US dietary guidelines in 2015 after decades on the books. A generation was told to switch from butter to margarine, and the partially hydrogenated oils that made margarine work were later pulled from the FDA's generally-recognized-as-safe list. The low-fat era reshaped a grocery aisle, and much of what replaced the fat was sugar. Salt, coffee, saturated fat, and eggs have each been villain and hero inside one lifetime.
I am Ravi Dewangan, CFL3, MS in Strength and Conditioning, and CrossFit Seminar Staff. I coach at Persistence Athletics in Belltown with my wife Jacque, and I wrote the MetFix course. Before I teach anyone anything about metabolism, I teach this. If you do not understand why the literature contradicts itself, every nutrition claim feels like a coin flip. It is not. There is a structure underneath, and once you see it you can sort claims yourself. Published August 2026.
Table of Contents
- Why your doctor and your trainer both have a study
- The question a p-value answers is not your question
- Why a field can be honest and still be mostly wrong
- The incentives, and what nutrition adds on top
- Five questions I ask before I believe a claim
- How we handle this at Persistence Athletics
- Frequently Asked Questions
Why your doctor and your trainer both have a study
Your physician says one thing about fat, or protein, or meal frequency. A coach says close to the opposite. Both are sincere, and both can point at published research. They are reading different literatures aimed at different questions.
Clinical guidance is built on long-horizon, population-level endpoints: cardiovascular events, all-cause mortality, disease incidence, tracked across decades in large mixed populations, most of whom do not train. Performance guidance comes from shorter trials, weeks to months, in people who already lift or run, measuring lean mass, strength output, or time to exhaustion.
Different questions do not have to give the same answer. Advice tuned to the average risk profile of a sedentary population is not automatically right for a 32-year-old training five days a week, and the reverse is also true. Guidelines also move slowly on purpose: acting on thin evidence at population scale is how you get the margarine episode.
So the disagreement is rarely proof that either is a fraud. More often it is proof that neither is looking at your situation, which is why anything specific to your health belongs with your physician.
The question a p-value answers is not your question
Here is the part almost nobody is taught.
When a study reports a statistically significant result, the p-value answers this question:
If there were truly no effect, how often would data at least this extreme show up by chance?
That is a statement about the data, assuming nothing is going on. The question in your head reading the headline is different:
Given this data, how likely is it that this effect is real?
The two are not interchangeable, and you cannot get from the first to the second without one more thing: how plausible the hypothesis was before the study ran.
This is the same structure as a medical screening test, the analogy I use in the course. A test with a 5 percent false positive rate, applied to a genuinely rare condition, produces more false positives than true ones. The test is not broken. That is what the math does when you screen for something uncommon, and a positive tells you far less than it feels like it does.
Published research works the same way. The 5 percent threshold is the false positive rate. The rarity of the condition is the fraction of tested hypotheses that happen to be true.
Why a field can be honest and still be mostly wrong
The base rates below are illustrative; the arithmetic is not. Imagine a field runs 1,000 studies. Power is the chance of detecting a real effect when one exists.
| Share of hypotheses that are true | Power | True positives found | False positives | Share of "significant" results that are wrong |
|---|---|---|---|---|
| 10 percent | 80 percent | 80 | 45 | about 1 in 3 |
| 10 percent | 30 percent | 30 | 45 | about 3 in 5 |
| 2 percent | 80 percent | 16 | 49 | about 3 in 4 |
Nothing dishonest happened in any row. Everyone followed the rules and the false positive rate stayed at 5 percent. The share of published claims that are wrong still moved from one third to three quarters, driven entirely by how speculative the hypotheses were and how underpowered the studies were.
Two things follow.
First, the more surprising a finding is, the more likely it is to be a false positive. Surprising means the prior plausibility was low. The findings most likely to earn a press release are, structurally, the ones most likely to be noise.
Second, small studies do not just miss real effects, they exaggerate the ones they catch. An underpowered study only crosses the line when the effect lands unusually large, so its published effect sizes skew high. A larger study later finds a smaller effect, and the public reads that as science reversing itself.
This is not a fringe concern. John Ioannidis published a paper in 2005 titled "Why Most Published Research Findings Are False," laying out this logic. Organized replication efforts since then have found a substantial share of published findings do not hold when independent teams rerun them. Fewer than half replicated in the large psychology effort. In industry attempts to reproduce landmark preclinical cancer papers, only a small minority did.
The incentives, and what nutrition adds on top
The statistics explain why noise is possible. The incentives explain why it piles up.
Publication bias. A study finding "this supplement did nothing" is far less likely to reach print than one finding an effect. The published literature is a filtered sample, not a complete record. You read the winners of a competition whose losers you never see.
Researcher degrees of freedom. One dataset can be analyzed dozens of defensible ways: which covariates to adjust for, which outcome to feature, which subgroup to report, where to cut a continuous variable. Every fork raises the chance something crosses the line on luck alone. No intent to deceive required.
Careers reward novelty. Grants, citations, and coverage flow to surprising results. Careful replication of a boring finding is among the least rewarded work in science and the most useful.
Nothing labels which is which. Durable findings and noise sit side by side in the same databases, formatted identically. There is no badge on the abstract.
Nutrition inherits all of that and adds three difficulties of its own.
You cannot randomize a lifetime. Nobody will assign 5,000 people to a controlled diet for 30 years while a control group does something else. So most long-horizon diet evidence is observational, which shows associations but cannot isolate causes.
Healthy-user confounding is enormous. Someone eating whole grains in 1985 also, on average, smoked less, exercised more, and saw a doctor more often. Adjustment removes the confounders you measured, not the ones you did not.
The measurement is soft. Most large diet studies rely on people recalling what they ate, and validation work using objective energy expenditure methods has consistently shown self-reported intake is unreliable and skews low. Fine-grained statistics, rough inputs.
Put those together and you get what researchers demonstrated by pulling common cooking ingredients off a list and searching the literature: published cancer-risk associations for the large majority of them, often pointing opposite directions depending on the paper.
Five questions I ask before I believe a claim
When a nutrition claim reaches me, it goes through this before it changes anything I do or coach.
Compared to what, in whom, over how long, measured how? Most claims fall apart right here, because "improves metabolic health" turns out to mean a small marker change over six weeks in 22 people unlike you.
Does it predict something I can check? A claim that cannot be wrong is not useful. One that says "if this is true, over eight weeks I should see X" is, because you can go look.
Is there a mechanism I can follow? Mechanism is not proof; medicine has a long list of interventions that made perfect mechanistic sense and failed in trials. But a claim with no mechanism and no replication is a correlation with good public relations.
Has anyone independent reproduced it? When observational data, randomized trials, and mechanistic work point the same direction, you are probably looking at something durable. When one design carries the whole claim, wait.
What does being wrong cost me? People skip this and it may be the most valuable. Prefer changes with a wide safety margin and a short feedback loop, so being wrong is cheap and you find out fast. Protein at each meal and more walking are cheap bets. A restrictive protocol built on one paper is not.
Boring survives. The parts of nutrition science that have held for forty years are unexciting on purpose: enough protein, mostly whole foods, enough fiber, enough energy for your training, consistency measured in years. The churn is at the edges, which is why our posts on macros for CrossFit and HYROX and supplements come from that center.
How we handle this at Persistence Athletics
When a member walks in with a screenshot and asks whether they should be doing the thing in it, we do not argue about the study. We turn it into a test.
Pick one change. Write down what you expect to see and by when. Pick two or three things you will measure: training performance, bodyweight trend across weeks rather than days, energy at 3 PM, sleep. Run a fixed window, then look honestly at whether the prediction held.
That is slower than a headline and far more reliable, because it swaps a population average you may not belong to for data on the one person you care about. It is the backbone of how Jacque runs our nutrition coaching program, and she runs most of our in-person MetFix assessments too.
None of this is medical advice. If you have a diagnosed condition, take medication, or have bloodwork that looks off, that conversation belongs with a doctor. We help with training, food habits, and the testing loop around them. My background is on my coach page.
Frequently Asked Questions
Why does nutrition advice keep changing?
Three problems stack on top of each other. Most long-horizon diet research is observational, so it can show associations but cannot isolate causes. Journals publish surprising results far more readily than boring or null ones, which makes the published literature a filtered sample rather than a complete record. And the standard statistical threshold answers a different question than the one readers think it answers. The result is a literature that holds both durable findings and noise, with nothing on the surface telling you which is which.
Why do my doctor and my trainer give me opposite nutrition advice?
Usually because they are reading different literatures aimed at different questions. Clinical guidance is built on long-horizon, population-level outcomes such as cardiovascular events and all-cause mortality, mostly in general populations. Performance and body composition advice comes from shorter trials in people who already train. Both bodies of evidence can be real and still point in different directions, because they are answering different questions about different people. Neither person is necessarily lying to you.
What does a p-value actually mean?
A p-value answers this question: if there were truly no effect, how often would data at least this extreme turn up by chance? It does not answer the question most readers care about, which is: given this data, how likely is it that the effect is real. Getting from the first to the second requires knowing how plausible the hypothesis was before the study ran. When a field tests mostly long-shot hypotheses, a large share of its statistically significant results will be false positives even when everyone runs the analysis honestly.
How do I know if a nutrition study is trustworthy?
Ask five things: compared to what, in whom, over how long, measured how, and has anyone independent reproduced it. Prefer findings that hold across different study designs, that come with a mechanism you can follow and check, and that are large enough to notice in real life. Be skeptical of any single study, especially a surprising one, because surprising findings are exactly the ones that started with a low prior probability of being true.
Should I ignore nutrition science entirely?
No. The durable core has held for decades: adequate protein, mostly whole foods, enough fiber, enough total energy to support your training, and consistency over time. That part is boring and heavily replicated. The churn happens at the edges, where single studies about single nutrients get press releases. Anchor on the boring core and treat the edges as hypotheses to test on yourself rather than instructions to follow.
Where to go next
This article is the foundation of the MetFix series. Everything after it, on metabolic flexibility, mitochondria, insulin, and fuel use, assumes you will check claims rather than collect them.
I wrote the whole course to be free. Eight short modules, no pitch inside them, enough mechanism to evaluate the next headline yourself. Start the free course whenever you want.
If you would rather see where you stand first, the 2-minute metabolic health score is a starting point for a conversation, not a diagnosis.
Want to take this further?
Talk to a coach about metfix programming at Persistence Athletics.
