Meridian — Evidence, augmented
Capabilities / Evaluation matrix and tool design

Evaluation matrix and tool design, grounded in evaluation practice and in the current research on AI-assisted structured generation.

An evaluation matrix is the operational backbone of an evaluation. It defines the questions being asked, the sub-questions that operationalise them, the data sources that will answer them, the methods of collection, and the analytical approaches that will be applied. From the matrix flows the instrument suite. The KII guides. The survey schedules. The FGD guides. The observation protocols. Every question asked in the field should route back to a specific sub-question in the matrix. Every sub-question in the matrix should be addressed by at least one instrument.

The distinctive risk here is that AI defaults to the standard frames it was trained on, whether the ToR calls for them or not. A matrix built on default framings sends the evaluation after the wrong questions, and no downstream discipline can recover from that. Meridian's workflow anchors generation to the ToR and traces every instrument question back to the sub-question it addresses.

01 / What can go wrong

Where a generated matrix quietly goes wrong.

The most common failure when AI is used to generate an evaluation matrix is standard-framing bias. The model defaults to the OECD-DAC criteria, relevance, coherence, effectiveness, efficiency, impact, sustainability, even when the ToR is asking a different set of questions. The result is a matrix that looks methodologically clean but does not actually address the evaluation's real terms of reference. The evaluation then proceeds to produce evidence on the wrong questions.

The second failure is sub-question sprawl. The model generates fifteen sub-questions for every high-level evaluation question, producing a matrix that is exhaustive on paper and unworkable in practice. Sample sizes get stretched thin across too many sub-questions. Data collection becomes unfocused. The evaluator ends up quietly ignoring half the matrix and the client discovers this only when the report cannot substantiate claims on the ignored sub-questions.

The third failure is instrument-matrix disconnection. The KII guide asks questions that do not route back cleanly to any sub-question in the matrix. Or asks the same question in three different guides without acknowledging the overlap. Or omits questions the matrix requires answers to. This failure mode is hard to detect by eye because each instrument reads as coherent on its own. It surfaces only during analysis, when data does not exist to answer sub-questions the matrix committed to.

The fourth failure is assumption inheritance. The model reads programme documents and inherits their assumptions into the matrix and the instruments. If the programme's theory of change is over-linearised, the matrix tests the linear pathway. If the programme assumes certain outcomes are self-evidently positive, the instruments do not test whether they were. The evaluator gets an instrument suite that reproduces the programme's own worldview rather than one that interrogates it.

02 / How the workflow answers

Four structural controls, each answering a specific failure mode.

ToR-anchored generation answers standard-framing bias. The AI is required to derive the question structure from the ToR and the programme documents, not from the model's training on standard evaluation frameworks. Where the ToR asks for something that does not map cleanly to OECD-DAC criteria, the matrix reflects that rather than smoothing it into a standard shape. Every high-level evaluation question in the matrix is required to cite the specific ToR paragraph it derives from. Standard framings are used where the ToR calls for them and rejected where it does not.

Bounded generation answers sub-question sprawl. The workflow sets defined sub-question limits per high-level question, tuned to the evaluation's budget, timeline, and sample sizes. The matrix is built for the evaluation the client is commissioning, not for a hypothetical unlimited-resource version. Where the sub-question count is constrained, the workflow requires the evaluator to make explicit choices about what is prioritised and what is left out, with the reasoning recorded rather than hidden.

Bidirectional traceability answers instrument-matrix disconnection. Every instrument question is required to link to a specific sub-question in the matrix. Every sub-question in the matrix is required to have at least one instrument routing to it. Both directions are checkable mechanically as part of the workflow. Overlap between instruments is surfaced and confirmed as intentional or removed. Gaps in instrument coverage are surfaced and closed before the instrument suite is finalised.

Assumption surfacing answers assumption inheritance. The matrix explicitly documents the assumptions it is inheriting from the programme documents. Assumptions the evaluation intends to test are flagged as such and reflected in the instrument design. Assumptions the evaluation is accepting as given are flagged as such and named in the disclosure. This is the equivalent of the temporal versioning move on the Theory of Change reconstruction workflow: making the workflow's own inheritance visible rather than hiding it.

03 / Anchors

Evaluation practice, and the current literature on AI-assisted structured generation.

The matrix-and-instruments approach has been standard evaluation practice for decades. Question-and-sub-question structures anchored to ToRs. Mixed-methods matrices mapping data sources and analysis types. Instrument suites that route back to the matrix. Pilot testing and refinement before deployment. The approach is the sector's, and the credit for it belongs to the practitioners who built it.

The structural controls around AI use are grounded in the current research on AI-assisted structured generation. Four strands of that research matter for evaluation matrix work specifically. Work on model bias toward training-distribution framings in structured generation tasks, which documents the model tendency to default to familiar frameworks even when the input asks for something else, and the effectiveness of anchored generation requirements in mitigating this. Empirical work on scope and length effects in AI-generated structured content, which shows that unbounded generation produces unusable outputs and that bounded generation with explicit trade-off decisions produces workable ones. Research on cross-artefact consistency in multi-document generation, which shows that generation across related artefacts produces predictable inconsistencies without traceability controls. And work on assumption inheritance in AI reading of source documents, which draws the line between the model surfacing what the documents say and the model unreflectively reproducing the documents' own analytical stance.

Meridian's workflow applies findings from all four strands. The ToR-anchored generation requirement comes from the framework-bias research. The bounded generation approach comes from the scope-and-length work. The bidirectional traceability comes from the cross-artefact consistency research. The assumption surfacing comes from the assumption-inheritance work.

04 / Outputs

An evaluation matrix, an instrument suite, a traceability report, an assumptions register, a disclosure.

Every Meridian matrix and instrument design assignment produces five artefacts. The evaluation matrix itself, in the client's preferred format, with every high-level question cited to the ToR paragraph it derives from and every sub-question mapped to data sources, methods, and analysis types. The instrument suite, KII guides, survey schedules, FGD guides, observation protocols as required, with every question tagged to the specific sub-question it addresses and the specific respondent group it is intended for. The traceability report, showing the mapping in both directions and confirming that every sub-question has instrument coverage and every instrument question has a matrix anchor. The assumptions register, distinguishing assumptions the evaluation intends to test from those it is accepting as given. And the disclosure text, a short methodology annex drafted for inclusion in the inception report.

The five artefacts together mean the matrix and instrument suite are inspectable. When a client asks whether the ToR is being addressed, the answer is a specific mapping. When a client asks whether the instrument suite is coherent, the answer is a bidirectional traceability check. When a client asks what assumptions the evaluation is inheriting from the programme documents, the answer is the assumptions register.

On method

The differentiation is the discipline.

Evaluation matrices and instruments sit upstream of every subsequent stage. A matrix defaulted to standard framings sends data collection after the wrong questions, and no downstream discipline can recover from that. ToR-anchoring and bidirectional traceability exist because the cost of getting this stage wrong is paid in the cost of the whole evaluation.