Evaluation, rebuilt around the tools that work.
An independent monitoring, evaluation and learning practice working across humanitarian and development programmes. Mixed-methods evaluation, built on AI-augmented pipelines where they earn their place. Faster cycles. Tighter evidence. Analytical effort spent where it changes the answer, without giving up the rigour donors and programme teams need.
Start a conversation →The sector's analytical toolkit hasn't changed in a decade. The programmes have.
Larger, faster, more complex, more data-rich, and operating in environments where access is shrinking. Evaluation timelines routinely exceed programme decision cycles. Reports arrive after the window for course correction has closed. Recommendations sit on shelves because they were written for accountability, not utility.
Meridian exists to close the gap between what evaluation could deliver with the tools now available and what the sector is still settling for. Not by replacing evaluator judgement, but by removing the mechanical bottlenecks that consume most of the analytical cycle, so the evaluator can focus on the parts that require contextual understanding, ethical sensitivity, and interpretive skill.
A MEL practice that understands the technology — and its limitations.
Meridian is an evaluation practice first. The methods are the ones the sector already recognises — theory-based design, contribution analysis, process tracing, outcome harvesting — run by an evaluator who stays accountable for every claim in the report.
What is different is knowing precisely where the technology helps and where it fails. Each method has a specific way it breaks when a language model is applied to it carelessly, and the workflow is built around those failure points rather than around what the tools do well. That means transparent pipelines you can audit, a clear separation between machine output and human interpretation, and a refusal to treat language models as black boxes, or as a substitute for contextual understanding.
The analytical budget, spent on analysis.
The point is not that the evaluation finishes sooner. It is that the mechanical work stops consuming the part of the cycle where judgement, interrogation and interpretation actually determine whether the findings hold.
Reading a programme archive end to end, by hand, to find the handful of documents that matter.
Cleaning, structuring and EQ-mapping transcripts before any of them can be analysed.
Assembling the evaluation matrix, sampling strategy and instrument suite from scratch.
Reconciling survey quotas across instruments and sites, manually and after the fact.
Rebuilding the analysis into slides, and reformatting them again after every comment round.
Interrogating the theory of change against what the programme actually did, not what the design said.
Re-coding samples by hand, and working out where the model's reading of a transcript breaks down.
Following contradictions between stakeholders back to the respondents who raised them.
Generating rival explanations and testing them before a contribution claim is written.
Validation with programme teams and partners, with the disagreements documented rather than smoothed.
Theory of Change reconstruction
→Unstructured donor proposals, logframes, and programme narratives synthesised into a structured ToC with mapped assumptions and testable causal pathways, in a single session. These are preliminary frameworks to be tested and validated through fieldwork, not pre-cooked findings.
Evaluation matrix and tool design
→A fully mapped evaluation matrix (questions, sub-questions, data sources, analysis types) built directly from the ToR. Stakeholder-specific KII guides generated with every question traced to an evaluation sub-question.
Evidence synthesis and outcome harvesting
→Outcomes harvested systematically from programme archives, or structured syntheses drawn across multiple evaluations and country portfolios, identifying what works, what doesn't, and where the evidence gaps are. Source-traced, flagged as preliminary, and designed for stakeholder validation before any conclusions are drawn.
Transcript processing and cross-stakeholder analysis
→Raw interview recordings processed through an AI-assisted pipeline trained on programme documents: cleaned, structured, and mapped to evaluation sub-questions. Across a full transcript set, the model identifies convergence, divergence, and outlier perspectives with source attribution.
Branded knowledge products
Workshop presentations, dissemination briefs, and visual summaries designed in consistent brand systems, with every figure verifiable against the underlying analysis. Produced in hours, not days.
"AI earns its place in evaluation when it makes evidence more traceable, not less. Augmentation over replacement. The evaluator stays accountable for every claim."