Inception report QA, grounded in evaluation practice and in the current research on AI-assisted error detection.
An inception report is the document a client reads to decide whether the evaluation is going to hold up. It sets out the theory of change interrogation, the evaluation matrix, the methods, the sampling, the workplan, and the risks. Errors at this stage propagate through every subsequent stage. A misspecified sub-question produces a misdirected data collection instrument. An unmapped assumption produces an analysis that misses what it should have tested. Inception QA is the mechanical check that catches those errors before they compound.
The distinctive risk here is that AI can produce QA output that looks thorough while catching less than a slower human check. Twenty flags of low value, no flags on the errors that matter, and a workflow that performs rigour while missing it. Meridian's workflow uses structured category-by-category checking and calibrated flagging so the check is inspectable rather than performative.
The failures that hide inside a clean-looking check.
The most consequential failure when AI is used for inception report QA is false negatives. The model reads the draft, cross-references against the evidence base, and reports that everything is consistent when it isn't. A factual error slips through because it appeared in a source the model gave less weight to. A logical inconsistency between two sections goes undetected because the sections were far apart in the document and the model's attention drifted between them. An outdated statistic gets passed because the model didn't recognise that a more recent version existed in the evidence base. These are the failures that undermine the entire purpose of the workflow.
The second failure is false positives at scale. The model flags twenty items per report where only three warrant evaluator attention. When the false positive rate is high, evaluator review of AI flags becomes a performance rather than a substantive check. Real errors get missed inside the noise. This is more damaging than it appears because it is hard to detect from the outside. A workflow producing many low-value flags looks thorough. It is often less thorough than a workflow producing fewer, better-calibrated flags.
The third failure is domain-convention blindness. The model flags things that experienced evaluators know are fine. Standard OECD-DAC phrasings. Conventional hedging language. Sector-specific formulations that are correct in context. Simultaneously, the model fails to flag things that are wrong. A value-for-money claim without underlying evidence. A conclusion that does not match the findings that support it. A recommendation that contradicts the theory of change the evaluation is working against. The AI has read the general research literature but has not internalised what makes an evaluation report actually work.
The fourth failure is source-absence confabulation. The model is asked whether the evidence base contains information on a specific topic. It confidently reports that it does when it does not. Or it reports that it does not when it does. This is a specific instance of the more general fabrication problem, but it arises differently in a QA context because the model is being asked to make presence-and-absence claims rather than to extract content. A false positive on presence produces a fabricated citation. A false negative on presence produces a missed source that should have been used.
Four structural controls, each answering a specific failure mode.
Structured checking against defined categories answers false negatives. The QA workflow does not ask the AI to identify errors in general. It asks the AI to check the draft against defined categories of finding, one category at a time. Factual claims against source. Internal consistency between sections. Alignment between findings and recommendations. Coverage of the evaluation matrix. Currency of data cited. Compliance with the ToR. Each category is a separate pass through the document with its own prompts and its own outputs. This structured approach catches more errors than open-ended QA prompts because it forces the model to look for specific problems rather than to look for problems in general.
Calibrated flagging answers false positives at scale. Every flag the AI raises is required to include a confidence level, a specific citation of the source the AI is referencing, and a suggested action. Flags below a defined confidence threshold are surfaced separately for evaluator awareness rather than mixed into the main review queue. This means the evaluator's substantive review is focused on the flags most likely to warrant attention, while lower-confidence flags remain visible without dominating the workflow. The threshold is tuned per assignment and documented in the disclosure.
Convention-aware prompting answers domain-convention blindness. The QA workflow is calibrated with sector-specific guidance on what constitutes a genuine error versus a conventional phrasing. Standard hedging language is treated as normal, not flagged. Standard OECD-DAC framings are recognised. Conversely, the workflow includes specific checks for the categories of error that experienced evaluators know are most consequential: value-for-money claims without underlying evidence, conclusions that do not match findings, recommendations that do not follow from analysis. These are treated as first-class check categories, not as things the model might notice.
Source-anchored presence claims answer source-absence confabulation. Every claim the AI makes about the presence or absence of information in the evidence base is required to include a specific citation. When the AI asserts that the evidence base contains information on X, the deliverable includes the document and location. When the AI asserts absence, the deliverable includes the list of documents searched and the specific queries run. This means the evaluator can verify presence and absence claims mechanically, rather than trusting the model's summary.
Evaluation practice, and the current literature on AI-assisted error detection.
Inception QA has been part of serious evaluation practice for as long as evaluation reports have existed. The check categories are familiar: cross-referencing draft findings against source documents, internal consistency across sections, alignment between recommendations, conclusions, and findings, coverage of the evaluation matrix and the ToR, currency and correctness of cited data. These are the practitioners' categories, not Meridian's.
The structural controls around AI use are grounded in the current research on AI-assisted error detection and document verification. Four strands of that research matter for inception QA specifically. Work on false-negative rates in AI-assisted document review, which documents how model attention allocation across long documents produces predictable patterns of missed errors, and the effectiveness of structured category-by-category checking in reducing this. Empirical work on false positive rates and reviewer fatigue in high-flag-volume workflows, which shows that calibrated confidence thresholds outperform unfiltered flagging. Research on domain-specific error detection, which shows that general-purpose language models require domain-convention calibration to distinguish genuine errors from conventional phrasings. And work on presence-and-absence claims in AI-assisted document verification, which documents the specific failure modes of the model reporting on what a source does and does not contain.
Meridian's workflow applies findings from all four strands. The structured category-by-category approach comes from the false-negative research. The calibrated flagging thresholds come from the reviewer fatigue research. The domain-convention prompting comes from the domain-specific error detection work. The source-anchored presence claims come from the presence-and-absence research.
A structured flag register, category summaries, evaluator sign-offs, a disclosure.
Every Meridian inception QA assignment produces four artefacts. The structured flag register, listing every flag the AI raised, with category, confidence level, source citation, and suggested action, together with the evaluator's disposition on each. The category summaries, showing what was checked in each of the QA categories and what the pass-or-flag distribution looked like. The evaluator sign-offs, category by category, with any residual concerns documented rather than dismissed. And the disclosure text, a short methodology annex drafted for inclusion in the inception report.
The four artefacts together mean the QA workflow is inspectable. When a client asks what was checked and what was found, the answer is a structured register. When a client asks about the AI's role, the answer is a category-by-category breakdown of what the AI proposed and what the evaluator confirmed. When a client asks about limitations, the answer is the residual concerns section of the sign-offs.
The differentiation is the discipline.
Inception QA is a check, not an analysis. The failure mode that matters most is thoroughness that performs itself: many low-value flags that look rigorous and hide the ones that count. Calibrated flagging and structured category-by-category checking exist because a report producing twenty flags is often catching fewer real errors than one producing three.