Statistics that respect how LLM data actually behaves.
Built-in mixed models, ANOVA, Bayesian and non-parametric tests run directly on your coded data — alongside LLM-specific diagnostics like variance decomposition, prompt sensitivity, and semantic entropy, and a conversational AI Analyst that always shows its work.
The standard toolkit, one click away
LLM experiment data has a particular shape: responses cluster within conditions, models, and prompt variants, distributions are often skewed or zero-inflated, and effect sizes matter more than p-values when replications are cheap. AIScholar's built-in analyses are chosen for exactly this terrain — mixed-effects models that treat prompt variant as a random effect, factorial ANOVA with effect sizes, Bayesian estimation with credible intervals, chi-square and non-parametric tests for categorical and ordinal outcomes, and descriptive statistics with cell-coverage tables so you always know how much data sits behind every mean.
Every result is rendered as both a table and a publication-ready chart, and every chart can be exported as an image. The full dashboard — all charts plus descriptive and coverage tables — exports as a paginated PDF or a ZIP bundle of PNGs and CSVs, in colour or greyscale, opening with an auto-generated study-context section.
Diagnostics designed for language models
Variance decomposition
Partitions total outcome variance into condition, model, prompt-variant, and residual components — telling you whether your effect is a real treatment effect or an artifact of wording.
Prompt sensitivity analysis
Analyzes performance spread across systematic prompt perturbations: how much results move, which perturbation dimensions drive it, and whether condition rankings stay stable.
Semantic uncertainty
Clusters free-text responses by meaning and computes Shannon entropy over clusters, flagging conditions where the model is roulette rather than measurement.
Psychometric validity
Alternate-forms reliability (Cronbach's α, split-half), an ICL/contamination assessment for published instruments, memorization probes, and an effect-size calibration table comparing human and LLM Cohen's d values.
Decoupling-gap analysis
For matched-pair designs: measures whether a model's stated stance and its justification move together across framings, or quietly decouple.
Judge agreement
Cohen's kappa, Krippendorff's alpha, Fleiss' kappa, and percent agreement across judges — reliability of the measurement itself, quantified.
The AI Analyst: a conversation with your data
Ask in plain language
"Is there an interaction between framing and model?" The Analyst sees your design, coded data, and prior analyses, and chooses appropriate statistics.
It runs real computations
The Analyst executes actual statistical analyses on your actual data — the same engines behind the built-in analyses — and can render charts inline. It does not eyeball numbers.
Every claim is inspectable
Results arrive with the test, the assumptions, and the numbers attached. Charts produced in the conversation can be promoted into your write-up as exhibits.
Interpretation discipline
AIScholar's framing is deliberately conservative: results describe the tested model versions under the tested conditions. The methodological warnings system flags overgeneralization risks — like treating three models as "LLMs in general" — before a reviewer does.
When the numbers are in, the write-up assistant turns them into manuscript prose — and the full dashboard export keeps a permanent visual record of exactly what you saw.
Explore more of the pipeline
- Response codingScore responses with LLM-as-a-judge, check agreement across judges, and keep humans in the loop with gold-standard review.
- Write-up assistanceDraft manuscript sections with AI, manage citations in five styles, verify references, and export to Word, Markdown, or LaTeX.
- Hypotheses & experimental designFrame testable predictions, then build factorial, fractional, matched-pair, or custom designs with an AI design wizard.
Turn a question about LLMs into published research.
Start your first study today. No human participants required.