Stages 4–5 · Execution & Collection

Run hundreds of trials without babysitting a script.

AIScholar dispatches your full trial grid to any model on OpenRouter with orchestrated concurrency, automatic retries, pilot runs, and an up-front cost estimate — then collects every response with the metadata reproducibility demands.

One trial, many replications

The unit of work is the trial: one model × one fully-resolved prompt × one parameter set, run once. Your design multiplies conditions by models and replications into a structured grid of trials, and the execution engine works through that grid for you — no glue scripts, no notebooks, no half-finished CSVs.

Because a single model answer is just one draw from a probability distribution, AIScholar runs each condition many times. That is what turns "the model said X once" into an analyzable distribution with a central tendency, spread, and shape — the raw material of defensible inference.

What you control per run

  • Any OpenRouter model

    Choose any subject model available on OpenRouter — and include several in one study, treating model identity as a factor for cross-model comparison.

  • Sampling parameters

    Temperature, top-p, max tokens, and seed, per model and condition. AIScholar records whether the provider actually honored the seed you set.

  • Concurrency & retries

    Parallel dispatch with automatic retries for transient API failures, so flaky endpoints don't leave holes in your grid.

  • Pilot mode

    Run exactly one trial per cell as a cheap sanity check before committing to the full grid. Pilot trials are tracked separately and never contaminate your analysis.

  • Cost estimation

    Before you launch, AIScholar estimates the OpenRouter cost of the full run — trials × models × token budgets — so there are no surprises on your bill.

  • Bring your own key

    Subject-model calls run on your own OpenRouter API key, encrypted at rest with AES-256-GCM. You pay OpenRouter directly for exactly what you use; AIScholar never marks up model usage.

A run, start to finish

  1. Review the resolved grid

    See every condition, model, and fully-resolved prompt the run will dispatch, along with the trial count and estimated cost.

  2. Pilot first (recommended)

    A one-trial-per-cell pilot verifies that prompts resolve correctly, output formats parse, and models behave as expected — for a fraction of the cost.

  3. Launch the full run

    The orchestrator dispatches trials concurrently, retries transient failures, and streams live progress so you can watch cells fill in.

  4. Collect with full metadata

    Every response is stored with its exact prompt, model version string, parameters, latency, token usage, cost, and a validity classification — valid, recovered by rescue parse, invalid, or refused.

Refusals and failures are data, not noise

Each response is classified on extraction: parsed cleanly, recovered by an AI-assisted rescue parse, invalid, or detected as a refusal. Invalid rates are tracked per condition and flagged above a threshold — and refusal rate itself is a perfectly valid outcome to study.

Collected responses flow straight into response coding, where raw text becomes structured, analyzable data.

Turn a question about LLMs into published research.

Start your first study today. No human participants required.