Outcome harvesting, grounded in Wilson-Grau and in the current research on AI-assisted qualitative coding.
Outcome harvesting is one of the standard evaluation methodologies for capturing what actually changed as a result of a programme. Including outcomes the programme didn't set out to produce, and outcomes it produced without meaning to. The method was developed by Wilson-Grau and colleagues and remains the anchor for how the sector understands the approach.
The distinctive risk of AI in outcome harvesting is that a workbook can be filled with plausible-looking rows without the interpretive commitments that give the method its standing. Meridian's workflow keeps those commitments structural rather than optional.
The two ways an AI-assisted harvest goes wrong.
The failure modes here sit within two named categories. Technological incongruence is asking a model to do what it structurally cannot. Methodological incongruence is reproducing the surface procedure of the method (rows, classifications, a filled workbook) while stripping out the interpretive commitments that make the method valid.
Source fabrication is the technological incongruence. The model produces an outcome that reads plausibly, cites documents that exist, and references passages that were never actually written. It appears because the model is being asked to do something it structurally cannot do well: hold a long document pack in working memory and cite from it faithfully across many outputs.
Interpretive overreach is one of the methodological incongruences. The workbook a client receives contains six substantive columns, and only some of them are retrieval tasks. Significance, contribution plausibility, and the classification of an outcome as positive or negative are interpretive judgements, not extractions. When AI is asked to draft all six columns in one pass, the interpretive columns get produced without the reasoning behind them being human reasoning. The workbook then contains claims the evaluator cannot defend when challenged.
Epistemological drift is the other. Outcome harvesting sits within a non-positivist analytic tradition. It is a structured interpretive act by an evaluator working with the data, not a search for pre-existing findings in it. When a model surfaces themes from interview transcripts, it is producing a candidate reading, not making a discovery. Workflows that treat AI output as objective finding reproduce the surface procedure while incorporating a stance the method does not use.
Three structural controls, each answering a specific failure mode.
The candidate register answers source fabrication. Documents are read in chunks. Every candidate outcome the AI identifies is written to a structured register with a verbatim excerpt, the document it came from, and the location within the document. At the surfacing stage, a candidate is a question for the evaluator, not a claim; under-surfacing is the more expensive error than over-surfacing, because the evaluator can reject cheaply but cannot recover what was never shown. The workbook builder reads only from accepted register rows. The register also holds a rejection log, showing candidates that were surfaced and removed with the reasoning recorded, which is usually the strongest evidence of methodological rigour when a client asks whether alternative outcomes were considered.
The task split answers interpretive overreach. AI is assigned to the retrieval work. Reading document packs. Identifying passages that describe change. Extracting verbatim excerpts. Proposing classifications with reasoning. The interpretive judgements are reserved for the evaluator. Significance is drafted by the evaluator using AI-surfaced material as evidence, then confirmed with the client. Contribution plausibility is reasoned through by the evaluator, not generated. Classifications proposed by AI are confirmed rather than accepted. The line is drawn structurally in the tooling, not left to prompting discipline.
Declared epistemological positioning answers epistemological drift. Every Meridian harvest is accompanied by a short disclosure that names the analytic tradition the harvest sits within, the AI tooling used, the tasks it was assigned to, and the tasks it was reserved from. This is drafted for inclusion in the evaluation report's methods section. It exists so that a reader of the report can understand how the evidence was produced, and so that Meridian is held to a stated position rather than an implicit one.
The Wilson-Grau method, and the current literature on AI in qualitative analysis.
Outcome harvesting is not a Meridian invention. The method, its outcome and output tests, and its classification criteria across positive and negative and intended and unintended outcomes have been in use across the sector for over a decade. The intellectual credit sits with its originators. Meridian applies it in the form its originators built.
The practice is medium Q: non-positivist in its treatment of meaning and contribution, but structured by an a priori framework (the logframe organisation, the SMART criteria, the classification schema). Substantiation runs through client review as a form of member checking, not through inter-rater reliability or agreement scores. This matters because importing a post-positivist validation metric into an interpretive method is a specific form of methodological incongruence: adopting the appearance of rigour by borrowing a metric the method does not use.
The structural controls around AI use are grounded in the current research on generative AI in qualitative research. Three strands matter for outcome harvesting specifically. Work on methodological congruence, which sets out when AI use is honest to the analytic tradition being applied. Empirical documentation of AI-coding failure modes in real research settings. And work on the epistemological positioning of AI-assisted coding, which draws the line between what AI can do as a retrieval and pattern-surfacing tool and what it cannot do as an interpretive one.
Meridian's workflow applies findings from all three strands. The candidate register and verbatim-excerpt requirement come from the empirical failure-mode work. The task split between retrieval and interpretive judgement comes from the epistemological positioning work. The disclosure statement and the declared analytic tradition come from the methodological congruence work.
A workbook, a register, a rejection log, a disclosure.
Every Meridian outcome harvest produces four artefacts. The workbook, in the client's preferred format, with the six substantive columns populated and every source citation resolvable to a specific document and location. The candidate register, showing every candidate that was considered, its verbatim excerpt, and its disposition. The rejection log, showing candidates that were surfaced and removed with the reasoning recorded. The disclosure text, a short methodology annex drafted for inclusion in the evaluation report.
The disclosure text is not a legal disclaimer. It is a methodological statement, and its purpose is to let a reader of the report understand how the evidence was produced. Meridian believes this should be standard practice across the sector. It isn't yet.
The four artefacts together are what makes the harvest defensible under scrutiny. When a client presents the workbook to a board, a donor, or an external evaluator, the underlying evidence is inspectable at every level.
The differentiation is the discipline.
Outcome harvesting stops being outcome harvesting the moment its interpretive commitments are stripped away and only the workbook rows remain. The register and rejection log preserve the interpretive audit trail. Substantiation runs through client review, not through agreement scores that would import a validation metric the method does not use. Meridian's workflow is built to keep the method recognisable as the method its originators developed.