AIScholar for psychology researchers
Machine psychology applies the discipline's hard-won measurement standards to a new subject: language models that answer questionnaires, display response biases, and sometimes merely imitate the constructs we try to measure. AIScholar is built for researchers who refuse to skip the psychometrics.
Measurement discipline, not just prompting
The difference between "we gave GPT a personality test" and a publishable machine- psychology study is measurement validation. AIScholar makes the validation steps first-class citizens of the pipeline:
| Concern | Built-in support |
|---|---|
| Internal consistency | Cronbach's α and split-half (Spearman-Brown) reliability, computed per model over replicated item responses |
| Wording dependence | Alternate-forms designs with paraphrased item sets and alternate-forms correlations |
| Training-data contamination | Memorization probes, ICL detection, and generation of content-matched but surface-novel items |
| Construct checklist | A Psychometric Validity Checklist (after Löhn et al., 2024) filled from your actual results |
| Response distributions | Replications per item at controlled temperature — a model's answer is a distribution, and it's analyzed as one |
The psychometric-assessment walkthrough shows the full arc: a 10-item risk-attitude scale administered one item per trial, with reliability gates that decide whether the substantive comparison is even allowed to run.
Classic paradigms, re-run on models
Much of early machine psychology replays cognitive and social paradigms on models — and the platform's design vocabulary was shaped around exactly that work:
- Framing effects (Tversky & Kahneman)Matched-pair gain/loss framing with a decoupling-gap analysis of stance vs. justification.
- Psychometric assessmentValidated scales with reliability, alternate forms, and contamination probes.
- Coding open-ended responsesAnchored rubrics, human gold standards, Krippendorff's α — content analysis at LLM scale.
- Persona × task factorialRole-assignment effects with the interaction term your question actually needs.
From design to APA-ready draft
Designs are declared as factors and levels with explicit crossing strategies; item wording lives in a versioned snippet library so instruments are byte-reproducible; judged coding carries inter-rater agreement statistics; and the write-up assistant drafts Method and Results in familiar reporting conventions, with citation styles switchable per project (APA among them).
- Experimental designInstruments as designs: items × forms × models, with replication planning.
- Response codingVersioned schemes, judge calibration, human review, agreement statistics.
- Analysis & visualizationReliability coefficients, ANOVA, mixed models, effect-size calibration against human benchmarks.
- Contamination & validity checksMemorization probes, ICL detection, semantic-uncertainty flags for unstable conditions.
The epistemics stay honest
A reliable scale score still is not a mind. AIScholar's framing throughout — in warnings, in analysis interpretations, in drafted prose — is that you are measuring response distributions under an administration procedure, and claims about inner states need converging evidence the platform will not fabricate for you. If that is the intellectual standard you hold your own field to, you will feel at home. Browse the example studies to see it in practice.
AIScholar for other researchers
- Social & management scholarsBring experimental rigor to questions about framing, decision-making, and judgment in language models.
- AI-safety & behavioral researchersMeasure refusals, sycophancy, and robustness with reproducible, statistically defensible methods.
- Educators & methods instructorsTeach the full arc of experimental research — design, execution, analysis, write-up — without an IRB or participant pool.
Turn a question about LLMs into published research.
Start your first study today. No human participants required.