This is a different failure from the ones I've written about before on this site. It isn't about a fabricated citation, or a source that doesn't actually support the claim attached to it, or a claim that used to be true and no longer is. It's about what happens when several sources are genuine, current, and individually credible, and still don't agree. Synthesis has to do something with that. When it goes wrong, it can make the disagreement disappear.

Diagnose before you resolve

Not every apparent disagreement is a real one, and treating all of them the same way is itself a mistake. A useful sequence has three stages, and they build on each other rather than sorting a disagreement into one exclusive bucket.

The first stage is scoping: what is each claim actually about, population, outcome, measure, definition, time, and the conditions under which it was found. A study on patients over seventy and a study on patients under forty are answering questions about two different populations. Relative risk and absolute risk are different numbers answering different questions, even when they're about the same intervention. A related but distinct case belongs here too: a claim that was accurate when a source was published but has since been overtaken by events, the specific failure I wrote about in the previous article on this site, where citation validity and temporal validity turned out to be separate properties. Once timestamped, that kind of claim is perfectly comparable to a newer one, it's just answering an older version of the question, and needs different handling than a genuine disagreement does. Scoping doesn't settle the disagreement. It establishes the proposition that is actually being compared.

The second stage asks what explains or changes the evidential weight of a divergence once the scope is clear. Making the claims comparable doesn't automatically settle the disagreement, sometimes it explains it instead, and there are two different ways that can happen. One is conditioning: two findings may become compatible once the conclusion is conditioned on context, the evidence may support A under condition X and B under condition Y. That population difference from stage one might turn out to belong here too, if the evidence suggests the intervention genuinely behaves differently by age rather than merely being described by two non-comparable studies. The other is evidential weighting: two findings address the same question under the same conditions, but one source of evidence gives stronger grounds for belief than the other. A randomized trial and an observational study reaching different conclusions on the same question isn't necessarily revealing that the effect is conditional, the gap could instead reflect confounding, differences in precision, or other features of study design, which may justify giving one result more weight than the other. But weighting one finding more heavily doesn't make the other disappear. It shifts belief. It doesn't settle the question.

Scoping and weighting can do legitimate work, and material disagreement can still remain after both. Only once a claim has been properly scoped, and any conditioning or evidential weighting that's actually warranted has been applied, does the third stage matter, whatever disagreement is still standing after that. That's residual disagreement, and it's the case a system actually has to preserve rather than resolve, not because nothing could be said about it, but because what can be said doesn't make it go away.

The three-stage diagnostic above is a proposed way to think about the problem, not something the research below set out to test. What that research tests is narrower: whether a model can represent and explain conflicting evidence it has already been handed.

What models currently lose

ConfRAG assembled 1,814 questions from three public datasets and a generated, manually reviewed set spanning six topical domains, then paired them with retrieved documents from heterogeneous web sources, evaluating whether models could organize conflicting material into distinct positions, cover all of them, and preserve the reasoning behind each one. One of the strongest models evaluated, GPT-4.1, scored 0.157 on the benchmark's reason-coverage metric, which is itself gated on correctly matching the associated answer, well below the human panel's 0.275 on the same task. The benchmark also supplies the true number of answer clusters in advance, to rule out degenerate but structurally valid partitions. Real deployment does not get that help.

WikiContradict tests something ConfRAG doesn't: it gives models two passages from Wikipedia that contain conflicting answers, and the benchmark treats them as equally trustworthy by construction, so source authority can't be used to pick a winner. Models struggle especially with implicit conflicts, ones that require actually reasoning about what the two passages say, not just noticing contradictory wording. Between the two benchmarks, the pattern holds: ConfRAG removes the uncertainty of not knowing how many positions exist, WikiContradict removes source reputation as a ranking shortcut, and in both settings, models still struggle to organize the competing positions and explain the disagreement between them.

What preserving disagreement requires

At minimum, a synthesis needs to keep distinct positions separate rather than blending them, preserve the source and reasoning attached to each, identify the conditions to which each applies where those are known, and mark what remains unresolved.

There's recent work suggesting the reasoning-attached-to-position part is achievable, at least in a narrower setting. A 2026 framework for automated fact-checking found that grounding a model's uncertainty explanation in the specific evidence spans that conflict or agree, rather than a generic hedge like "sources disagree," tracked model uncertainty more faithfully and was judged by human raters as more helpful and informative. That doesn't establish that claim-clustering and provenance-tracking solve general research synthesis, only that grounding a disagreement explanation in the specific evidence actually in conflict is feasible and useful, at least in the fact-checking settings tested.

When resolution is justified

Not every disagreement should survive into the final output unresolved. But resolving one needs a reason, and that reason should be visible, not just implied by how confident the answer sounds. A scope mismatch is handled by separating or conditioning the claims, not averaging them. An evidence-quality difference gets a stated, justified weighting, applied honestly, which can shift where the balance of belief sits without necessarily eliminating the minority position. Whatever disagreement is still standing after both of those, the residual kind, doesn't get resolved at all. It gets preserved.

The short version: resolution needs a warrant. Fluency is not a warrant.

When unresolved evidence can still support a decision

A system can still be useful even when sources disagree. Disagreement doesn't necessarily need to block a decision, and treating every one of them as a reason to stall is its own kind of failure.

The test is whether the decision changes across the credible interpretations that remain, and that test only means something relative to an explicit decision rule. Suppose credible studies disagree about exactly how large an effect is, but under every credible interpretation that remains, the effect is still above the threshold at which one intervention is preferable to the alternative. The scientific disagreement is real, and worth preserving for anyone who needs the full picture. But the immediate decision is robust to it, nothing about choosing between the interventions actually depends on resolving which estimate is closest to correct. Now suppose the credible interpretations fall on opposite sides of that threshold. The disagreement hasn't gotten any bigger, but it's become decision-material, and a system that quietly averages those estimates into one confident number has destroyed exactly the piece of information the decision actually turned on.

Disagreement capable of changing the decision needs to remain visible, not smoothed into an answer that sounds more settled than the evidence actually is.