Meridian — Evidence, augmented
Capabilities / Contribution analysis and process tracing

Contribution analysis and process tracing, grounded in theory-based evaluation and in the strongest constraints on AI use in the practice.

Contribution analysis and process tracing are the two theory-based causal inference methods that answer the question every serious evaluation eventually reaches. Did the programme contribute to the observed change, and if so, how much and through what mechanism? Both are used where counterfactual designs are not available or not credible, and where the honest answer to the causal question is not attribution but contribution.

The methods are related but distinct. Contribution analysis is the broader six-step framework: build the theory of change, assess evidence link by link along the results chain, address external factors and other interventions, assemble the contribution narrative, iterate. Process tracing is a more specific evidentiary method that uses four defined empirical tests, straw in the wind, hoop, smoking gun, and doubly decisive, to evaluate the strength of causal inferences at specific points in a mechanism. In Meridian's practice both are used, sometimes together, sometimes separately, depending on what the evaluation question requires and what evidence is available.

Applying AI to this work carries risks that are different in kind from the other capabilities Meridian delivers. Elsewhere, the failure modes have visible signatures. A fabricated outcome has no source. An invented quote does not match the transcript. Here, the failure mode has no signature. A contribution narrative written by a model reads exactly like one grounded in rigorous link-by-link assessment. A process tracing write-up with test types assigned after the evidence was read reads exactly like one with predictions locked in advance. The output is fluent, correctly formatted, and methodologically worthless.

01 / What can go wrong

Failures with no visible signature.

The most dangerous failure in contribution analysis is coherent narrative produced independently of the evidence. A model asked to write a contribution story will produce one that reads as credible, because narrative coherence is what it optimises. Coherence is generated regardless of whether the underlying evidence supports the links being asserted. The result passes the method's own quality test on inspection. A reader, including the evaluator who commissioned it, cannot distinguish a well-evidenced contribution narrative from a well-written one by reading it. The failure has no signature. The output looks exactly like success.

Related to this is smoothing. Given a results chain with three strong links and two weak ones, a model will write five links that read equally solid, because connective tissue is generated at uniform confidence. The distribution of evidence strength across the chain, which is the most important thing a reader of a contribution analysis needs to understand, disappears into prose that flows too well.

The distinctive failure in process tracing is the sequencing problem. Process tracing works because predictions come before evidence. The evaluator states what would be expected if a hypothesis were true, and what would be expected if it were false, and only then goes looking. The formal tests are meaningful only against predictions specified in advance. A language model given a document pack and a hypothesis reads the material first and generates test labels fitted to the evidence in front of it. It is not lying. It is pattern-completing. Post-hoc test labelling is exactly the pattern that a corpus of process tracing write-ups teaches. The output is fluent, correctly formatted, and methodologically worthless.

Both methods share a rival hypothesis failure. A model asked to generate alternative explanations will produce plausible ones, and will tend to produce ones that are easy to defeat, because the surrounding context is a case being built for the primary hypothesis. If the same process generates the rivals and evaluates them, it is marking its own homework, and the resulting analysis shows the primary hypothesis surviving a field of straw men.

The fifth failure is rhetorical dismissal. In contribution analysis, the fourth of the standard conditions requires that the contribution of external factors be dismissed or demonstrated. A paragraph that acknowledges alternatives and moves on has not dismissed anything. In process tracing, "we did not find evidence that X" is a statement about the search, not about the world. Eliminating a hypothesis requires positive evidence of absence, not the absence of positive evidence. Both failures produce write-ups that look like they have addressed the alternatives when they have only mentioned them.

02 / How the workflow answers

Five structural controls, each answering a specific failure mode.

Register before narrative answers coherent-narrative-without-evidence. The link-by-link evidence assessment exists as a structured register before any prose is drafted. Every link in the results chain carries its evidence, its sources, and a strength rating assigned by the evaluator. The contribution narrative is then a rendering of the register, not a document written and sourced afterwards. What is checkable is the register. Prose written first and sourced afterwards is how an unsupported contribution claim acquires a bibliography, and the workflow makes that sequence structurally impossible.

Visible strength shading answers smoothing. The register export produces a client-facing workbook that colours every link cell by its evidence strength. A reader sees the distribution across the chain at a glance. Where the evidence is solid, where it thins, and where it runs out. A narrative can absorb a weak link. A table shaded by evidence strength cannot. An automated check compares the draft narrative against the ratings and flags weak or absent links written in confirming language.

Predictions locked before evidence answers the sequencing problem. Predictions of what would be expected if a hypothesis were true, and what would be expected if it were false, are written and locked before any evidence is retrieved. The lock is a hash, so it can be shown afterwards that the predictions were not adjusted to fit what was found. Evidence is then retrieved against the predictions, not the other way round. The formal test types are assigned to predictions before the evidence pass begins, with the necessity-and-sufficiency reasoning recorded, not fitted to results after the fact.

Inference bounded by test type answers post-hoc labelling and rhetorical elimination. The workflow enforces the logic of each test mechanically. A passed straw-in-the-wind supports weakly and does not confirm. A failed smoking gun proves nothing. A passed hoop keeps a hypothesis alive but does not confirm it. A hoop can only be failed on positive evidence of absence, or on a search whose completeness the evaluator has verified, not on a nil return. Doubly decisive tests, which require genuinely necessary-and-sufficient evidence, are flagged for review by default because they are rare in development work and their misuse is a reliable signal that classifications have drifted.

Separated rival generation and evaluation answers the marking-its-own-homework problem. AI generates candidate rival hypotheses generously and argues for each as strongly as it can. The evaluator extends the field from theory and from stakeholder interviews, records where each rival came from, and assesses whether each has been credibly addressed. A rival that only ever existed in the register, that no stakeholder ever proposed, is weak evidence that alternatives were seriously considered. The origin of every rival is recorded and reported.

03 / Anchors

Theory-based evaluation, and the current research on AI in interpretive causal analysis.

Contribution analysis and process tracing have developed over the last two decades, in evaluation practice and, for process tracing, in political science and historical methods. Structured results-chain assessment through defined conditions. The four empirical tests: straw in the wind, hoop, smoking gun, and doubly decisive. The theory-based logic both methods share, the recognition that development change usually has many contributors, and the rejection of attribution language in favour of contribution language. Meridian works within these traditions; it did not build them.

The AI-use controls are grounded in the current research on generative AI in interpretive analysis and on the specific failure modes of large language models in causal reasoning tasks. Three strands matter for this capability. Methodological work on when AI use is congruent with the analytic tradition being applied and when it is not, which distinguishes technological incongruence, asking a model to do what it structurally cannot, from methodological incongruence, reproducing the surface procedure of a method while discarding what makes it valid. Empirical work on narrative coherence as an optimisation target for language models, which documents how models produce plausible causal explanations independently of evidence. And research on the specific pattern-completion behaviour of models given hypothesis-and-evidence pairs, which shows how post-hoc test labelling arises as a default behaviour and how sequencing controls can prevent it.

Meridian's workflow applies findings from all three strands. The register-before-narrative rule and the visible strength shading come from the narrative coherence research. The predictions-locked-before-evidence rule and the inference lattice come from the pattern-completion research. The separation of AI-generated rivals from AI-evaluated rivals comes from both the congruence framework and the pattern-completion work. Every control has a specific research anchor, and every anchor names a specific failure mode the control is designed to prevent.

04 / Outputs

A register, a locked prediction set, a narrative rendered from ratings, a factors sheet, a disclosure.

Every Meridian contribution analysis or process tracing assignment produces four or five artefacts depending on which method is applied. For contribution analysis: the links register, showing every link in the results chain with its evidence, sources, and strength rating; the exported client-facing workbook, with strength shading visible across the chain and every link readable in the context of the whole; the external factors sheet, showing each factor, its evidence, its disposition, and the basis for it; and the contribution narrative, rendered from the register with language calibrated to the ratings. For process tracing: the hypothesis and rival register with origins recorded; the locked prediction set for each test, with the hash preserved as evidence of the sequencing; the test results with every verdict carrying a verbatim excerpt and location; and the write-up with inferences that match what each test permits. Both methods produce the disclosure text, a short methodology annex drafted for inclusion in the evaluation report.

The artefacts together are what make the causal reasoning inspectable. For contribution analysis, the shaded workbook is the most useful thing to put in front of a client who is sceptical that a plausible story is an evidenced one. The register in both cases is what preserves the audit trail from causal claim through to reported finding.

The locked prediction set is delivered to the client as the audit trail for the sequencing, and is the single most persuasive artefact in the practice for clients who are sceptical of qualitative causal claims. It shows what the evaluation expected to find before it looked. Clients who would dismiss any qualitative causal write-up as after-the-fact rationalisation cannot dismiss predictions locked with a hash before the evidence pass began. No amount of write-up achieves the same effect.

Where the evidence is genuinely weak, this is reported rather than smoothed. A contribution narrative that opens by naming its own soft links is read as more credible than one that has to be interrogated for them, which is also the honest ordering. Reporting that the evidence could not sustain a stronger claim is a legitimate output of both methods, and one the workflow preserves rather than writing over.

On method

The differentiation is the discipline.

The methods are contribution analysis and process tracing as they have been practised for over two decades. The failure modes are named in the research on interpretive causal analysis and language model behaviour. The controls are structural rather than prompted, and stronger here than anywhere else in the practice. This is a workflow designed by someone who has read the research on what goes wrong when AI is used carelessly for causal reasoning, and has built the workflow around the answer.