Meridian — Evidence, augmented
Capabilities / Transcript processing and cross-stakeholder analysis

Transcript processing, grounded in qualitative evaluation practice and in the current research on AI-assisted coding.

Transcript processing is one of the mechanically slowest steps in a qualitative evaluation. Recorded interviews and focus group discussions have to be transcribed, corrected for accent and terminology, structured against the evaluation questions, and analysed both individually and across the full transcript set. On a typical assignment with fifteen or twenty interviews this can take one to two weeks before any analytical work starts.

The distinctive risk here is that misrepresenting a respondent's testimony is not a methodological error in the abstract. It is an ethical one. The subjects are real people who spoke to an evaluator on record. Meridian's workflow is built so that verbatim source and named-source attribution stay preserved through to the reported theme.

01 / What can go wrong

Four ways a transcript gets misrepresented.

At the individual transcript level, the most common failure is invented content attributed to real respondents. A model asked to clean, structure, and summarise a transcript will smooth over unclear passages by producing text that reads naturally but was not what the respondent actually said. Half-finished sentences get completed. Ambiguous phrasing gets resolved into a specific meaning the respondent did not commit to. Quotes appear in the deliverable that the respondent never spoke. This is the failure mode that carries the greatest ethical weight.

The second individual-level failure is sentiment and meaning misreading, particularly in low-resource languages and in code-switching between languages. Standard transcription and coding tools trained predominantly on North American English routinely misread tone, sarcasm, and evaluative weight in Somali, Amharic, Arabic, or French-inflected regional variants. A statement that was clearly critical in the original language can appear neutral or even positive once processed. The evaluator working from the processed output has no way to see what was lost.

At the cross-stakeholder level, the most common failure is theme fabrication and false convergence. A model asked to identify themes across a transcript set will produce a set of themes that reads coherently and looks well-organised, but that includes patterns that are not actually present in the data. Related to this is false convergence: the model reports that a majority of respondents agree on a point when the actual pattern in the transcripts is more contested, or when the appearance of agreement is an artefact of how the questions were asked rather than a shared substantive view.

The fourth failure is first-transcript dominance. The reading the model develops from the first interview it processes tends to shape how it codes every subsequent interview. Themes that emerged strongly in interview one get over-detected in interviews two through twenty. Themes that only appear later in the corpus get under-weighted or missed. The result is a cross-stakeholder analysis that reflects the analytical trajectory of the model as much as the substance of the data.

02 / How the workflow answers

Four structural controls, each answering a specific failure mode.

Verbatim-anchored processing answers invented content. Every processed transcript is required to preserve the verbatim source alongside any cleaned or structured version. Where the AI has smoothed unclear passages, cleaned up disfluencies, or completed a truncated sentence, this is marked in the deliverable and the original passage is retained beneath. Quotes carried forward into the cross-stakeholder analysis are required to match the verbatim source, not the cleaned version. This means the evaluator, and any reader who later inspects the transcripts, can see exactly what was said and exactly what was changed.

Context-specific correction answers language and meaning misreading. Before processing, the AI is trained on programme-specific vocabulary from the evaluation documents: organisation names, place names, sector terminology, technical jargon. Where the assignment involves low-resource languages or code-switching, the model is used for structure and mapping but sentiment and evaluative weight are extracted directly from the source language passage rather than from a translation. A native-speaker evaluator or field partner reviews sentiment codes on a sample of transcripts as part of the standard workflow.

Structured cross-stakeholder analysis answers theme fabrication and false convergence. The cross-stakeholder analysis workflow is required to produce, for every theme reported, a named list of the specific respondents whose statements support it, and a verbatim excerpt from each. Convergence is reported as a count with sources: nine respondents raised X in these specific terms. Divergence is reported the same way: three respondents actively disagreed, with these specific statements. Outlier perspectives are surfaced with source attribution rather than smoothed into the majority reading. When a theme cannot be substantiated with named-source evidence in this format, it does not appear in the analysis.

Order-independent processing answers first-transcript dominance. Transcripts are processed independently against the evaluation questions before any cross-stakeholder analysis is attempted. The initial coding pass on each transcript uses only the evaluation matrix and programme documents as anchors, not the coding of other transcripts. Cross-stakeholder analysis is then a distinct workflow step run on the full set of independently coded transcripts, not a continuous process where later transcripts are read through the lens of earlier ones. This decouples the analytical trajectory of the AI from the order in which the data happened to arrive.

03 / Anchors

Qualitative evaluation practice, and the current literature on AI-assisted coding.

Qualitative evaluation as practised in the sector is the analytic method: semi-structured interviewing anchored to an evaluation matrix, thematic analysis against defined sub-questions, cross-stakeholder synthesis that distinguishes convergent from divergent views and preserves outlier perspectives, inter-coder reliability checks on a sample of coding decisions. The credit for these practices belongs to the researchers and evaluators who developed them, not to Meridian.

The structural controls around AI use are grounded in the current research on generative AI in qualitative coding. Four strands of that research matter for transcript work specifically. Empirical documentation of AI-coding failure modes in real research settings, including invented quotes, hallucinated themes, and confident summaries of material the model never actually processed. Work on the specific limitations of large language models in low-resource languages and in sentiment analysis of code-switched or non-Western-English speech. Research on order effects and first-example dominance in AI-assisted classification, and the effectiveness of order-independent processing designs in mitigating them. And methodological work on the epistemological status of AI-produced qualitative analysis, which draws the line between AI as a structuring and pattern-surfacing tool and AI as a claimed source of interpretive findings about what respondents meant.

Meridian's workflow applies findings from all four strands. The verbatim-preservation requirement and the named-source evidence requirement for every theme come from the empirical failure-mode research. The language-specific processing rules come from the research on model limitations in non-Western-English speech. The order-independent processing design comes from the research on order effects in AI-assisted classification. The distinction between structured pattern-surfacing and interpretive judgement runs through all four sections.

04 / Outputs

Processed transcripts, a cross-stakeholder analysis, sentiment review notes, a disclosure.

Every Meridian transcript-processing assignment produces four artefacts. The processed transcripts themselves, each preserving the verbatim source alongside the cleaned and structured version, with EQ-mapping marked against the evaluation matrix. The cross-stakeholder analysis, structured by evaluation sub-question, reporting convergence, divergence, and outlier perspectives with named-source attribution and verbatim excerpts for every theme. The sentiment review notes, where relevant, documenting the language-specific decisions taken during processing and any sentiment codes flagged for evaluator or field-partner review. And the disclosure text, a short methodology annex drafted for inclusion in the evaluation report.

The named-source attribution is what makes the cross-stakeholder analysis inspectable. When a client, donor, or external reviewer asks how a particular theme was substantiated, the answer is a named list of respondents and their verbatim statements. When a client asks whether there was disagreement, the answer is the specific respondents who disagreed and what they said. This is not the standard output of AI-assisted transcript work in the sector. It is what makes the analysis defensible.

The four artefacts together preserve the audit trail from raw recording through to reported theme. When a claim in the evaluation report is questioned six months later, the underlying evidence is still inspectable at every level.

On method

The differentiation is the discipline.

Transcript processing is the capability where AI failure carries the greatest ethical weight, because the subjects are real people who spoke on record and can be misrepresented. Verbatim preservation and named-source attribution keep the reader in contact with what was actually said. When a client asks who spoke to a theme and what they said, the answer is a name and a quote, not a summary.