Stage 8 · Analysis & Visualization

Statistics that respect how LLM data actually behaves.

Built-in mixed models, ANOVA, Bayesian and non-parametric tests run directly on your coded data — alongside LLM-specific diagnostics like variance decomposition, prompt sensitivity, and semantic entropy, and a conversational AI Analyst that always shows its work.

The standard toolkit, one click away

LLM experiment data has a particular shape: responses cluster within conditions, models, and prompt variants, distributions are often skewed or zero-inflated, and effect sizes matter more than p-values when replications are cheap. AIScholar's built-in analyses are chosen for exactly this terrain — mixed-effects models that treat prompt variant as a random effect, factorial ANOVA with effect sizes, Bayesian estimation with credible intervals, chi-square and non-parametric tests for categorical and ordinal outcomes, and descriptive statistics with cell-coverage tables so you always know how much data sits behind every mean.

Every result is rendered as both a table and a publication-ready chart, and every chart can be exported as an image. The full dashboard — all charts plus descriptive and coverage tables — exports as a paginated PDF or a ZIP bundle of PNGs and CSVs, in colour or greyscale, opening with an auto-generated study-context section.

Diagnostics designed for language models

  • Variance decomposition

    Partitions total outcome variance into condition, model, prompt-variant, and residual components — telling you whether your effect is a real treatment effect or an artifact of wording.

  • Prompt sensitivity analysis

    Analyzes performance spread across systematic prompt perturbations: how much results move, which perturbation dimensions drive it, and whether condition rankings stay stable.

  • Semantic uncertainty

    Clusters free-text responses by meaning and computes Shannon entropy over clusters, flagging conditions where the model is roulette rather than measurement.

  • Psychometric validity

    Alternate-forms reliability (Cronbach's α, split-half), an ICL/contamination assessment for published instruments, memorization probes, and an effect-size calibration table comparing human and LLM Cohen's d values.

  • Decoupling-gap analysis

    For matched-pair designs: measures whether a model's stated stance and its justification move together across framings, or quietly decouple.

  • Judge agreement

    Cohen's kappa, Krippendorff's alpha, Fleiss' kappa, and percent agreement across judges — reliability of the measurement itself, quantified.

The AI Analyst: a conversation with your data

  1. Ask in plain language

    "Is there an interaction between framing and model?" The Analyst sees your design, coded data, and prior analyses, and chooses appropriate statistics.

  2. It runs real computations

    The Analyst executes actual statistical analyses on your actual data — the same engines behind the built-in analyses — and can render charts inline. It does not eyeball numbers.

  3. Every claim is inspectable

    Results arrive with the test, the assumptions, and the numbers attached. Charts produced in the conversation can be promoted into your write-up as exhibits.

Interpretation discipline

AIScholar's framing is deliberately conservative: results describe the tested model versions under the tested conditions. The methodological warnings system flags overgeneralization risks — like treating three models as "LLMs in general" — before a reviewer does.

When the numbers are in, the write-up assistant turns them into manuscript prose — and the full dashboard export keeps a permanent visual record of exactly what you saw.

Turn a question about LLMs into published research.

Start your first study today. No human participants required.