Run hundreds of trials without babysitting a script.
AIScholar dispatches your full trial grid to any model on OpenRouter with orchestrated concurrency, automatic retries, pilot runs, and an up-front cost estimate — then collects every response with the metadata reproducibility demands.
One trial, many replications
The unit of work is the trial: one model × one fully-resolved prompt × one parameter set, run once. Your design multiplies conditions by models and replications into a structured grid of trials, and the execution engine works through that grid for you — no glue scripts, no notebooks, no half-finished CSVs.
Because a single model answer is just one draw from a probability distribution, AIScholar runs each condition many times. That is what turns "the model said X once" into an analyzable distribution with a central tendency, spread, and shape — the raw material of defensible inference.
What you control per run
Any OpenRouter model
Choose any subject model available on OpenRouter — and include several in one study, treating model identity as a factor for cross-model comparison.
Sampling parameters
Temperature, top-p, max tokens, and seed, per model and condition. AIScholar records whether the provider actually honored the seed you set.
Concurrency & retries
Parallel dispatch with automatic retries for transient API failures, so flaky endpoints don't leave holes in your grid.
Pilot mode
Run exactly one trial per cell as a cheap sanity check before committing to the full grid. Pilot trials are tracked separately and never contaminate your analysis.
Cost estimation
Before you launch, AIScholar estimates the OpenRouter cost of the full run — trials × models × token budgets — so there are no surprises on your bill.
Bring your own key
Subject-model calls run on your own OpenRouter API key, encrypted at rest with AES-256-GCM. You pay OpenRouter directly for exactly what you use; AIScholar never marks up model usage.
A run, start to finish
Review the resolved grid
See every condition, model, and fully-resolved prompt the run will dispatch, along with the trial count and estimated cost.
Pilot first (recommended)
A one-trial-per-cell pilot verifies that prompts resolve correctly, output formats parse, and models behave as expected — for a fraction of the cost.
Launch the full run
The orchestrator dispatches trials concurrently, retries transient failures, and streams live progress so you can watch cells fill in.
Collect with full metadata
Every response is stored with its exact prompt, model version string, parameters, latency, token usage, cost, and a validity classification — valid, recovered by rescue parse, invalid, or refused.
Refusals and failures are data, not noise
Each response is classified on extraction: parsed cleanly, recovered by an AI-assisted rescue parse, invalid, or detected as a refusal. Invalid rates are tracked per condition and flagged above a threshold — and refusal rate itself is a perfectly valid outcome to study.
Collected responses flow straight into response coding, where raw text becomes structured, analyzable data.
Explore more of the pipeline
- Prompt constructionTurn factor levels into concrete prompts with a versioned snippet library, a variation grid, and an AI prompt wizard.
- Response codingScore responses with LLM-as-a-judge, check agreement across judges, and keep humans in the loop with gold-standard review.
- Hypotheses & experimental designFrame testable predictions, then build factorial, fractional, matched-pair, or custom designs with an AI design wizard.
Turn a question about LLMs into published research.
Start your first study today. No human participants required.