Methodology glossary
The vocabulary of rigorous LLM research.
Plain-language definitions of the methods behind defensible LLM studies — from LLM-as-a-Judge coding and semantic uncertainty to factorial designs and psychometric validity. Each term links to the AIScholar features that put it to work.
All terms
- Machine psychologyStudying LLMs as behavioral subjects with experimental methods from psychology.
- LLM-as-a-JudgeUsing a judge model to code free-text responses against a rubric, with human calibration.
- Semantic uncertaintyShannon entropy over meaning-clustered responses — response consistency as a measurable outcome.
- Variance decompositionPartitioning outcome variance into condition, model, prompt-variant, and residual components.
- Prompt sensitivityMeasuring how results change under semantically-equivalent rewordings of a prompt.
- Factorial & fractional factorial designFull vs. fractional crossing of factor levels, aliasing trade-offs, and when to use each.
- Matched-pair designPairing conditions on constant content for paired tests and decoupling-gap analysis.
- In-context-learning contaminationDetecting when instrument scores reflect memorization or in-prompt adaptation instead of the construct.
- Psychometric validity for LLMsRe-establishing reliability and validity when human instruments are administered to models.
- Method Specification Prompt (MSP)A self-contained method specification precise enough for independent re-implementation.
- BYOK (Bring Your Own Key)Trials run on the researcher's own OpenRouter key, billed by the provider at cost.
Wondering how these pieces fit together? The features overview walks the whole pipeline, and the example studies show each method applied to a real research question.
Put the vocabulary to work.
Every term in this glossary is a built-in capability. Start a study and use them for real.