AIScholar for AI-safety and behavioral researchers
Refusal rates, sycophancy, persona drift, jailbreak robustness — behavioral safety claims are experimental claims, and they deserve experimental discipline: controlled conditions, replications, agreement-checked measurement, and statistics that survive review. AIScholar turns ad-hoc eval scripts into studies.
From eval script to defensible study
Most behavioral evals are one prompt set, one pass, one number. The moment a claim matters — "model X refuses more than model Y", "sycophancy increases with user confidence" — the methodological questions arrive, and AIScholar has stock answers:
| Eval weakness | Platform fix |
|---|---|
| One phrasing per probe | Prompt-variant generation across six perturbation dimensions; wording treated as a random effect |
| Single-pass measurements | Replicated trials per condition at controlled temperature — behavior as a distribution, not an anecdote |
| Unvalidated auto-grading | Judge coding calibrated against human gold standards, with multi-judge agreement statistics |
| No uncertainty estimates | Mixed models, effect sizes, and semantic-entropy flags for unstable conditions |
| Unreproducible collections | Full parameter and model-version capture, exported as a reproducibility package |
Subject models run through OpenRouter on your own key, so the same design executes across many frontier and open models — with model identity as an explicit factor in the analysis, not a footnote.
Walkthroughs closest to safety work
- Prompt sensitivity & robustnessThe core robustness question: how much of your measurement is phrasing? Variance decomposition and ranking stability.
- Coding open-ended responsesThe measurement pipeline behind any judged eval — calibration, agreement gates, versioned rubrics.
- Framing effects across modelsMatched-pair designs and the decoupling gap: do stances flip while justifications stay put?
- Analyzing an existing datasetImport responses from your own eval harness and run the coding and analysis machinery on them.
Behavior change over time, on the record
Because designs, prompts, and parameters are versioned artifacts, a study is re-runnable when a model updates — same conditions, new snapshot, and a model-version comparison that is apples to apples. The reproducibility package (trial-level data, exact prompts, judge rubrics, Method Specification Prompt) is the format that lets another lab check your claim — increasingly the bar for safety findings that inform policy or deployment decisions.
- Experimental designProbes as factors and levels, with crossing strategies and replication planning.
- Prompt constructionVersioned snippets and systematic variants — perturbation studies without script sprawl.
- Response codingRefusal taxonomies and severity scales as versioned schemes with agreement checks.
- ReproducibilityEverything another lab needs to re-run your study verbatim, in one export.
Scope honestly
Single-turn text studies measure single-turn text behavior: no tool use, no multi-turn escalation, no deployment context. That covers a large and important slice of behavioral safety questions — and for that slice, the difference between an anecdote and a finding is exactly the machinery here. Start from the example studies or the feature tour.
AIScholar for other researchers
- Social & management scholarsBring experimental rigor to questions about framing, decision-making, and judgment in language models.
- Psychology researchersAdapt validated instruments, run psychometric checks, and study machine behavior with familiar methods.
- Educators & methods instructorsTeach the full arc of experimental research — design, execution, analysis, write-up — without an IRB or participant pool.
Turn a question about LLMs into published research.
Start your first study today. No human participants required.