Augmentation, not replacement.
AI tooling earns its place where it adds genuine leverage: document review at scale, mechanical cleaning, first-pass coding, citation and consistency checks. Interpretation, contextual sensitivity, ethical judgement and stakeholder accountability stay with the evaluator. The line is drawn deliberately, and it's drawn in writing.
Six stages. One accountable evaluator.
Inception
Scoping conversations, evaluation questions, theory of change interrogation, evaluation matrix, sampling strategy. LLM-assisted document scans run in parallel to compress timelines without skipping ground.
Data collection
Survey, KII, FGD and observation work, usually with a partner field team. Instruments are tested in language, enumerator training is real, and pipelines are wired up before the first form is uploaded.
AI-assisted processing
Automated cleaning and validation on the quantitative side. LLM-assisted thematic coding on the qualitative side, with every coded segment linked back to its source. Nothing is synthesised away from evidence.
Human review & synthesis
The evaluator reads, recodes a sample, interrogates the model's blind spots, and writes the synthesis. This is where judgement lives, and where AI is deliberately kept out of the driver's seat.
Validation
Findings are tested with programme teams, partners and (where ethical and feasible) affected populations. Disagreement is documented, not smoothed over.
Reporting
Structured drafting, citation and consistency QA, plain-language summaries, translation review. Final outputs land on time and in a form decision-makers can actually use.
Six methods, in detail.
Each page sets out how the method is run, where AI use is constrained and why, and what gets delivered.
Tooling transparency
Pipelines run on auditable open-source components: Kobo and ODK CAPI for collection, reproducible Python notebooks for analysis, and custom retrieval pipelines over programme archives where document Q&A is in scope. Frontier models are used where they meaningfully outperform, and the specific models and versions are named explicitly in the report. No black boxes, no proprietary lock-in pitched as innovation. Methods, tools, and model documentation are designed to be shareable and reproducible.
The disciplines this workflow is held to
Quality assurance, AI disclosure, data protection, safeguarding, and how the practice contracts with the people who deliver the work are set out in full on the Practice page. They apply on every engagement, not selectively.
Read the practice standards →