The LLM-as-research-subject platform

Run rigorous social-science experiments on LLMs, end to end.

Treat language models as research subjects and go from hypothesis to a publishable write-up, without recruiting a single human participant. AIScholar automates the entire experimental pipeline, with an AI co-researcher at every stage.

The only platform that turns LLM experimentation into publishable social science.

The research pipeline
Hypotheses01
Design02
Prompts03
Execute04
Collect05
Code06
Visualize07
Analyze08
Write-up09
Why researchers choose AIScholar

Everything a rigorous study needs, in one place.

Pioneer machine psychology

Investigate entirely new domains of LLM behavior. Treat models as research subjects and produce findings nobody has documented before.

An AI co-researcher at every step

Purpose-built AI wizards assist with hypotheses, experimental design, prompt construction, coding schemes, analysis conversations, and the academic write-up.

Methodological rigor built in

Contextual methodological warnings, psychometric validity checks, variance decomposition, semantic-uncertainty measurement, and prompt-sensitivity analysis.

Reproducible by default

Versioned prompt templates and coding schemes, full trial metadata, auto-generated method specifications, and a one-click reproducibility package.

Bring your own data

Already have model responses? Upload an existing dataset and jump straight to coding, visualization, and analysis with the same statistical toolkit.

Export everything you produce

Publication-ready charts, a drafted manuscript with a style-switchable bibliography, raw data as CSV or XLSX, and Word, Markdown, or LaTeX output.

How it works

One disciplined pipeline, from question to manuscript.

Each project walks through the same sequence of research stages. The platform keeps your work organized, versioned, and reproducible at every step.

STAGE 01

Hypotheses

Frame testable predictions about LLM behavior, with AI as a sounding board.

Learn more about Hypotheses

STAGE 02

Design

Build factorial, fractional, matched-pair, or custom designs with discrete factors and levels.

Learn more about Design

STAGE 03

Prompts

Construct prompt templates from a snippet library bound to your factor levels.

Learn more about Prompts

STAGE 04

Execute

Run trials at scale against any OpenRouter model, with retries, concurrency, and pilot runs.

Learn more about Execute

STAGE 05

Collect

Capture every response with full trial metadata for reproducibility.

Learn more about Collect

STAGE 06

Code

Score responses with LLM-as-a-judge, multi-judge agreement, and human review.

Learn more about Code

STAGE 07

Visualize

Explore results through charts, descriptive statistics, and cell coverage.

Learn more about Visualize

STAGE 08

Analyze

Run inferential models, effect sizes, and LLM-specific methodological checks.

Learn more about Analyze

STAGE 09

Write-up

Draft publishable prose with citations and export to Word, Markdown, or LaTeX.

Learn more about Write-up

Explore all features in depth →

What's inside

Explore the platform, capability by capability.

Every stage of the pipeline is backed by purpose-built tooling. Dig into any capability for a full walkthrough of how it works.

Hypotheses & experimental design

Full-factorial, fractional-factorial, matched-pair, and custom hand-picked designs, with an AI design wizard and contextual methodological warnings.

Learn more about Hypotheses & experimental design

Prompt construction

A Snippet Library binds versioned text to factor levels, a spreadsheet-style variation grid edits them in bulk, and a variation matrix previews every prompt your design will produce.

Learn more about Prompt construction

Experiment execution

Dispatch trials to any OpenRouter model with orchestrated concurrency, automatic retries, cheap pilot runs, prompt-variant generation, and up-front cost estimates.

Learn more about Experiment execution

Response coding

LLM-as-a-Judge scoring with any judge model, multi-judge agreement statistics, human review that overrides machine codes, and versioned coding schemes.

Learn more about Response coding

Analysis & visualization

Mixed models, ANOVA, GLMM, Bayesian tests, equivalence testing, and effect sizes — plus a conversational AI Analyst, variance decomposition, semantic uncertainty, prompt sensitivity, and psychometric validity checks.

Learn more about Analysis & visualization

Write-up assistance

AI-drafted sections grounded in your project data, traceable citations, switchable citation styles (APA, MLA, Harvard, Chicago, Vancouver), an editable BibTeX bibliography with reference verification, and Word, Markdown, or LaTeX export.

Learn more about Write-up assistance

Reproducibility

Every template and coding scheme is versioned, every trial carries its exact prompt, model, and parameters, and a one-click package bundles it all for replication.

Learn more about Reproducibility

R&D documentation

OECD Frascati-aligned uncertainty statements, a contemporaneous activity log, effort timesheets, and an export package built for R&D tax-incentive and grant reporting.

Learn more about R&D documentation

Methodology knowledge base

Reference curated methods entries in any AI prompt — the wizards read them as context.

Citation traceability

A curated citation library tracked per project, with inline references that survive into the write-up.

AI Analyst chat

Ask your coded data questions in plain English, on top of the formal statistical tests.

BYOD dataset mode

Analyze response datasets you collected elsewhere with the same coding and statistics stages.

Why AIScholar is different

No other tool runs the whole study.

Adjacent tools each cover one slice of the work. AIScholar is the only platform built to take an LLM experiment all the way to a defensible, publishable result.

vs. LLM eval & observability tools

Built for engineers shipping products. They tell a product team whether a prompt got better.

Built for researchers. Prove, with statistical and methodological rigor, why a model behaves the way it does, then publish it.

vs. generic design-of-experiments software

Classical DOE packages plan factorial designs but have no idea how to talk to a language model.

Factorial design and orchestrated LLM execution live in one pipeline, with no glue scripts between your stats tool and your model calls.

vs. academic research code

One-off scripts from a paper that rot the moment the authors move on.

A maintained platform that takes you from hypothesis to write-up, not a repository you have to resurrect.

vs. R&D tax & compliance tooling

Documents experimentation after the fact for reporting.

Runs the experiment and documents it, with an optional R&D module that captures the work as you do it.

AI wizards that know your study

Built-in AI wizards that understand both the app and your study.

AIScholar is packed with purpose-built AI wizards that know how each tool works and read the full context of your project: your hypotheses, design, prompts, and results. They assist across the whole pipeline, from drafting hypotheses and proposing designs to building prompts, coding responses, talking through your analysis, and writing up findings. (The model that assists you is always separate from the model you put under study.)

Hypotheses
Brainstorm and sharpen testable predictions.
Design
Propose factorial conditions and levels.
Prompts
Generate and validate snippet-bound prompt templates.
Coding
Draft coding schemes and judge prompts.
Analysis
Talk through results with the AI Analyst.
Write-up
Draft prose with citations and references.
Bring your own OpenRouter key
For running trials against subject models.
Wizards are included
AI assistance is part of your AIScholar plan — every account starts with 20 free wizard runs, and Pro makes them unlimited.
Subject models are BYOK
Studying LLM responses requires your own OpenRouter API key. Those OpenRouter usage costs are not included in AIScholar's pricing.
Estimated before you run
AIScholar estimates your OpenRouter cost up front, so you know what a run will cost before you start it.
Frequently asked questions

Honest answers about what the platform does.

What exactly can I study on AIScholar?

Anything you can express as a single-turn, text-only experiment on a language model: you manipulate prompt text and model parameters as discrete factors, send each condition to subject models many times, and measure outcomes readable from each text response — labels, numbers, refusal rates, response length, judge-rated scores, or semantic entropy. Because responses are distributions, every study is built on replications, not one-off answers.

What is out of scope?

AIScholar is deliberately honest about its boundaries: no multi-turn dialogue, no tool use or agentic loops, no image, audio, or video inputs, no fine-tuning or access to model internals, and no human participants. Subject models are limited to what OpenRouter serves.

Do I need my own API key?

The AI wizards that assist you are part of your AIScholar plan — no setup needed. Running trials against subject models uses your own OpenRouter API key, and those usage costs are billed by OpenRouter, not AIScholar. Every run shows a cost estimate before you start it.

Which models can be research subjects?

Any model available on OpenRouter — hundreds of models across providers. You can include several models in one study and treat model identity as an experimental factor for cross-model comparison.

What statistics are built in?

Linear mixed models, factorial and repeated-measures ANOVA, logistic GLMM, chi-square, Bayesian tests with Bayes factors, non-parametric tests, equivalence testing (TOST), multiple-comparison corrections, inter-rater reliability, and effect sizes with bootstrap confidence intervals — plus LLM-specific analyses like variance decomposition and prompt sensitivity.

What do I get on the Free plan?

The full manual pipeline through Visualization, up to 3 projects, data export, and 20 AI wizard runs to try the assistance. Pro unlocks unlimited AI wizards, the AI Analyst chat, and Write-Up assistance; Enterprise adds the compliance-grade R&D documentation module.

Read the full FAQ — models, costs, data ownership, reproducibility, and more.

Pricing

Start free. Upgrade when AI earns its place.

The Free plan runs the full research workflow through Visualization and includes 20 AI wizard runs to try the assistance. Pro unlocks unlimited AI wizards, the conversational AI Analyst, and Write-Up assistance; Enterprise adds compliance-grade R&D documentation.

Running trials against subject models uses your own OpenRouter API key. Those OpenRouter usage costs are billed by OpenRouter and are not included in any AIScholar plan — AIScholar estimates them for you before each run.

Free

Run the full research workflow and try AI on the house.

$0

Free forever

  • Up to 3 projects
  • 20 free AI wizard runs to try it out
  • Full pipeline through Visualization
  • Experiment execution & response collection
  • Built-in statistical analyses & charts
  • Data export (CSV / XLSX)
  • AI wizards capped — upgrade for unlimited
  • No AI Analyst chat
  • No Write-Up assistance
Most popular

Pro

Unlimited AI across every stage — go from idea to manuscript in a day.

$39.99/mo

  • Up to 100 projects
  • Everything in Free
  • Unlimited AI hypothesis, design, prompt & coding-scheme wizards
  • AI Analyst — ask your data questions in plain English
  • AI Write-Up — citation-backed draft from Intro to Discussion
  • Reproducibility package export

Enterprise

Everything in Pro, plus compliance-grade R&D documentation for grants & tax credits.

$99.99/mo

  • Unlimited projects
  • Everything in Pro
  • R&D documentation module (OECD Frascati-aligned)
  • Technological uncertainty tracking
  • Effort & expenditure timesheets
  • R&D tax-incentive export package
  • Documentation supports SR&ED tax credits in Canada

Have a discount code? Enter it at checkout to apply your savings.

Prices in USD. Cancel anytime from the billing portal. Need a custom plan? Contact us.

Turn a question about LLMs into published research.

Start your first study today. No human participants required.

You are responsible for your research

AIScholar's AI wizards are research aids, not a substitute for your judgment. AI-generated content can be incomplete or wrong, so review and verify everything before you rely on it. You remain responsible for the integrity of the research you produce and publish.