Who it's for

AIScholar for AI-safety and behavioral researchers

Refusal rates, sycophancy, persona drift, jailbreak robustness — behavioral safety claims are experimental claims, and they deserve experimental discipline: controlled conditions, replications, agreement-checked measurement, and statistics that survive review. AIScholar turns ad-hoc eval scripts into studies.

From eval script to defensible study

Most behavioral evals are one prompt set, one pass, one number. The moment a claim matters — "model X refuses more than model Y", "sycophancy increases with user confidence" — the methodological questions arrive, and AIScholar has stock answers:

Common eval weaknesses and their experimental fixes
Eval weaknessPlatform fix
One phrasing per probePrompt-variant generation across six perturbation dimensions; wording treated as a random effect
Single-pass measurementsReplicated trials per condition at controlled temperature — behavior as a distribution, not an anecdote
Unvalidated auto-gradingJudge coding calibrated against human gold standards, with multi-judge agreement statistics
No uncertainty estimatesMixed models, effect sizes, and semantic-entropy flags for unstable conditions
Unreproducible collectionsFull parameter and model-version capture, exported as a reproducibility package

Subject models run through OpenRouter on your own key, so the same design executes across many frontier and open models — with model identity as an explicit factor in the analysis, not a footnote.

Behavior change over time, on the record

Because designs, prompts, and parameters are versioned artifacts, a study is re-runnable when a model updates — same conditions, new snapshot, and a model-version comparison that is apples to apples. The reproducibility package (trial-level data, exact prompts, judge rubrics, Method Specification Prompt) is the format that lets another lab check your claim — increasingly the bar for safety findings that inform policy or deployment decisions.

Scope honestly

Single-turn text studies measure single-turn text behavior: no tool use, no multi-turn escalation, no deployment context. That covers a large and important slice of behavioral safety questions — and for that slice, the difference between an anecdote and a finding is exactly the machinery here. Start from the example studies or the feature tour.

Turn a question about LLMs into published research.

Start your first study today. No human participants required.