LLM-as-a-judge coding you can actually defend.
Interpretive outcomes — empathy, sycophancy, stance, quality — need a judge. AIScholar scores responses against versioned coding schemes with one or many judge models, measures their agreement, and keeps a human override as the final word.
From raw text to analyzable values
Structured outputs (labels, numbers, JSON) are extracted automatically at collection time. But many research questions hinge on judgments a regex cannot make: how empathetic is this response? Does it actually agree with the user's wrong claim? For those, AIScholar uses LLM-as-a-judge coding: a judge model reads each response and scores it against your coding scheme — a rubric with one or more variables, each with an explicit definition and value set.
Judge models are chosen freely — any model available through OpenRouter can serve as a judge — and judging runs on your own API key with the same up-front cost estimation as experiment execution. The judge is always separate from the subject: the model being studied never grades itself in the same breath.
Reliability is built in, not bolted on
Multi-judge coding
Run several judge models — or several passes — over the same responses, then quantify their agreement with Cohen's kappa, Krippendorff's alpha, Fleiss' kappa, or percent agreement.
Human review & gold standard
Review machine codes response by response. A human override becomes the authoritative value, and human-coded subsets can serve as a gold standard for judging the judges.
Versioned coding schemes
Editing a scheme creates a new version, and every coded value records the scheme version that produced it — so re-coding after a rubric change is explicit, never silent.
Omission-aware templates
Dual-axis scheme templates separate commission errors (the model said something wrong) from omission errors (it failed to say something required) — two failure modes that single-axis rubrics blur together.
Judge-prompt sensitivity check
Samples your responses, generates variants of the judge's rubric wording, and reports how much agreement shifts — telling you whether your measure is robust or an artifact of one phrasing.
Coverage tracking
Live progress shows exactly which responses are coded, pending, or failed, per scheme and per judge, so a coding pass never quietly ends at 93%.
A typical coding pass
Define the coding scheme
Name each variable, define its values, and write the judging instructions. Scheme templates — including the omission-aware dual-axis pattern — give you a rigorous starting point.
Choose judges and estimate cost
Pick one or more judge models and review the estimated cost of the pass before dispatching anything.
Run and monitor
Judges code responses concurrently with live progress. Parse failures are retried and surfaced, never swallowed.
Check agreement, review, and override
Inspect inter-judge reliability statistics, review disagreements by hand, and apply human overrides where the machines got it wrong. The final dataset records who — or what — decided every value.
Judges are instruments — treat them like it
A judge model is a measurement instrument, and AIScholar treats it with the same suspicion a psychometrician brings to a survey scale: agreement statistics, rubric sensitivity checks, and human gold standards are first-class features, not afterthoughts.
Coded data flows into analysis & visualization, where the statistics live.
Explore more of the pipeline
- Experiment executionDispatch trials to any OpenRouter model with concurrency, retries, pilot runs, prompt variants, and cost estimates.
- Analysis & visualizationRun mixed models, ANOVA, Bayesian and non-parametric tests, LLM-specific diagnostics, and talk through results with an AI Analyst.
- Hypotheses & experimental designFrame testable predictions, then build factorial, fractional, matched-pair, or custom designs with an AI design wizard.
Turn a question about LLMs into published research.
Start your first study today. No human participants required.