Meridian — Evidence, augmented
Capabilities / Theory of Change reconstruction

Theory of Change reconstruction, grounded in theory-based evaluation and in the current research on AI-assisted causal reasoning.

Theory of Change reconstruction is one of the standard early tasks in an evaluation. When a programme has been running for years without a coherent results architecture, or has run through multiple design iterations, or has never had its causal logic written down in a form that can be tested, the reconstruction produces the framework the rest of the evaluation works against. Activities. Outputs. Intermediate outcomes. Impact. The assumptions that link each level to the next. The distinction between what the design said and what the programme actually did.

The distinctive risk here is that the reconstruction sits upstream of every subsequent stage. Errors made now propagate through every downstream method, and each method's rigour then makes the underlying flaw harder to see. Meridian's workflow uses provenance classes and versioning to keep silent gap-filling structurally impossible.

01 / What can go wrong

Where the fiction enters, and why it propagates.

The most common failure when AI is used for Theory of Change reconstruction is causal fabrication. The model produces a pathway that reads coherently, links activities to outcomes through plausible mechanisms, and identifies assumptions that fit the logic. None of it is actually anchored in the programme documents. The links are inferred from the model's general knowledge of how programmes of this type typically work, not surfaced from the specific evidence base. When the reconstruction is presented to the programme team, the response is often that it looks right in the abstract but doesn't describe what this programme actually does.

The second failure is over-linearisation. A theory-based evaluation approaches causality as a set of interacting pathways, feedback loops, and context-dependent mechanisms. Large language models trained on standard results-chain notation tend to produce clean linear diagrams even when the underlying programme logic is genuinely non-linear. Complex causality gets collapsed into a sequence of arrows. Feedback and mutual influence disappear. The reconstruction misrepresents the programme by making it look tidier than it is.

The third failure is confusion between design-as-stated and design-as-evolved. Programme documents accumulate over time. The original proposal, the inception report, the mid-term redesign, the current logframe, the internal team's working understanding of what they now think the programme does. A model asked to reconstruct the Theory of Change tends to blend these into a single synthetic version that matches none of them. The distinction between what the programme was designed to do, what it actually did, and what the team now believes it should do gets lost.

Any of these failures matters more here than in other capabilities because the reconstruction is what the rest of the evaluation works against. Contribution analysis then assesses evidence link by link against invented links. Process tracing writes predictions against assumptions nobody ever held. Outcome harvesting maps outcomes onto a chain the programme never had. Each downstream method applies real rigour to a fictional object, and the rigour makes the fiction harder to see.

02 / How the workflow answers

Three structural controls, each answering a specific failure mode.

Three provenance classes answer causal fabrication. Every element of the reconstructed Theory of Change carries one of three markings on the artefact itself. Stated: written verbatim in a programme document, quotable with a reference. Synthesised: the evaluator's formulation, drawing on the documents but not found in them. Inferred: derived by reasoning where the documents are silent. Synthesised and inferred elements are never silently promoted to stated; they carry a validation route and a named person to validate with. The deck shows which is which. This is the discipline that makes silent gap-filling structurally impossible rather than only discouraged.

The complexity discipline answers over-linearisation. The reconstruction workflow produces two artefacts, not one. A linear results chain for donor-facing use, which serves the reporting purpose it exists for. And a more complex causal map alongside it, which shows the feedback loops, contextual dependencies, and non-linear mechanisms that a linear chain cannot represent. When a programme's logic is genuinely linear, the two artefacts converge. When it isn't, the difference between them is where the evaluation actually needs to look. Programme teams and evaluators are asked to review both, not the linear chain alone.

The design-as-stated / design-as-evolved pairing answers design drift, and is the single most valuable move in a reconstruction. Every reconstruction pairs the design as stated at inception with the design as it evolved in practice, with the design-as-currently-understood by the programme team held alongside where useful. Each version cites its source documents. The divergence between them is documented as an analytical finding rather than smoothed into a synthetic single-version output. A reconstruction that shows only the paper design describes a programme that may not exist. Pairing the two is what distinguishes a live reconstruction from a paper one, and it is the move most often skipped.

03 / Anchors

Theory-based evaluation, and the current research on AI in causal reasoning.

Theory-based evaluation has developed over three decades of practice: structured results chains, explicit assumptions registers at each causal link, the distinction between necessary and sufficient conditions, the recognition that programmes work through context-dependent mechanisms rather than universal ones. Meridian did not develop this tradition. It works within it.

The structural controls around AI use are grounded in the current research on generative AI in causal reasoning and in evaluation methodology specifically. Three strands of that research matter for Theory of Change work. Work on how large language models handle causal reasoning tasks, which documents the model tendency to generate plausible mechanisms that are not evidence-grounded, and the effectiveness of source-anchoring requirements in reducing this. Empirical work on AI use in evaluation inception, which surfaces the specific failure modes of applying models to unstructured programme documentation. And methodological work on the epistemological status of AI-produced analytic frameworks, which draws the line between AI as a structuring tool that produces frameworks for evaluator review and AI as a claimed source of findings about programme logic.

Meridian's workflow applies findings from all three strands. The provenance classes come from the causal reasoning research on source-anchoring. The complexity discipline and the two-artefact output come from the evaluation-specific research on how programme logic actually behaves in fragile and adaptive settings. The framing of reconstruction as a preliminary framework for validation rather than as a finding comes from the epistemological positioning work.

04 / Outputs

A validation deck carrying the chain, the provenance markings, the assumptions register, the version history, and the disclosure.

Every Meridian Theory of Change reconstruction is delivered as a validation deck. The deck carries the results chain, the complexity map, the provenance markings on every element, the assumptions register mapped to each causal link, the version history distinguishing design-as-stated, design-as-evolved, and design-as-currently-understood, and the disclosure text. The deck is a working analytical artefact, designed to be argued over in a validation meeting rather than filed. A reconstruction that comes back unchanged after validation usually means it was not read.

The reconstruction is delivered as a preliminary framework for stakeholder validation, not as a finding. Programme teams review it, refine it, contest it. The validated version is what the evaluation then works against. A Theory of Change reconstruction that arrives as a finding closes conversations that need to stay open.

The deck is what makes the reconstruction defensible under scrutiny. Every element of the ToC can be traced to its provenance class. Every distinction between what was designed and what evolved is documented.

On method

The differentiation is the discipline.

Theory of Change reconstruction is where AI failure compounds hardest, because every downstream method then works honestly against a fictional object. The provenance classes make silent gap-filling structurally impossible. The validation deck goes to a meeting to be argued over, not to file. A reconstruction that comes back unchanged after validation usually means it was not read.