Stage 6 · Response Coding

LLM-as-a-judge coding you can actually defend.

Interpretive outcomes — empathy, sycophancy, stance, quality — need a judge. AIScholar scores responses against versioned coding schemes with one or many judge models, measures their agreement, and keeps a human override as the final word.

From raw text to analyzable values

Structured outputs (labels, numbers, JSON) are extracted automatically at collection time. But many research questions hinge on judgments a regex cannot make: how empathetic is this response? Does it actually agree with the user's wrong claim? For those, AIScholar uses LLM-as-a-judge coding: a judge model reads each response and scores it against your coding scheme — a rubric with one or more variables, each with an explicit definition and value set.

Judge models are chosen freely — any model available through OpenRouter can serve as a judge — and judging runs on your own API key with the same up-front cost estimation as experiment execution. The judge is always separate from the subject: the model being studied never grades itself in the same breath.

Reliability is built in, not bolted on

  • Multi-judge coding

    Run several judge models — or several passes — over the same responses, then quantify their agreement with Cohen's kappa, Krippendorff's alpha, Fleiss' kappa, or percent agreement.

  • Human review & gold standard

    Review machine codes response by response. A human override becomes the authoritative value, and human-coded subsets can serve as a gold standard for judging the judges.

  • Versioned coding schemes

    Editing a scheme creates a new version, and every coded value records the scheme version that produced it — so re-coding after a rubric change is explicit, never silent.

  • Omission-aware templates

    Dual-axis scheme templates separate commission errors (the model said something wrong) from omission errors (it failed to say something required) — two failure modes that single-axis rubrics blur together.

  • Judge-prompt sensitivity check

    Samples your responses, generates variants of the judge's rubric wording, and reports how much agreement shifts — telling you whether your measure is robust or an artifact of one phrasing.

  • Coverage tracking

    Live progress shows exactly which responses are coded, pending, or failed, per scheme and per judge, so a coding pass never quietly ends at 93%.

A typical coding pass

  1. Define the coding scheme

    Name each variable, define its values, and write the judging instructions. Scheme templates — including the omission-aware dual-axis pattern — give you a rigorous starting point.

  2. Choose judges and estimate cost

    Pick one or more judge models and review the estimated cost of the pass before dispatching anything.

  3. Run and monitor

    Judges code responses concurrently with live progress. Parse failures are retried and surfaced, never swallowed.

  4. Check agreement, review, and override

    Inspect inter-judge reliability statistics, review disagreements by hand, and apply human overrides where the machines got it wrong. The final dataset records who — or what — decided every value.

Judges are instruments — treat them like it

A judge model is a measurement instrument, and AIScholar treats it with the same suspicion a psychometrician brings to a survey scale: agreement statistics, rubric sensitivity checks, and human gold standards are first-class features, not afterthoughts.

Coded data flows into analysis & visualization, where the statistics live.

Turn a question about LLMs into published research.

Start your first study today. No human participants required.