CLI Reference¶
The package installs the rupsycho command; python -m rupsycho is equivalent. It covers the
everyday workflow without writing any Python: check a configuration, preview the prompts, run
the experiment, score the answers.
rupsycho --help # all commands
rupsycho run --help # options and examples of one command
rupsycho --version
| Command | Purpose |
|---|---|
run |
Run an experiment and collect the answers |
validate |
Check configuration files without loading any model |
prompt |
Print the fully assembled prompt of one item and persona |
examples |
List, show and copy the bundled example configurations |
postprocess |
Clean, validate and score the answers of a run |
configurator |
Launch the Streamlit configurator app |
Conventions¶
Data on stdout, diagnostics on stderr. Answer tables, prompts, validation verdicts and
example files go to stdout, so they can be piped or redirected. Progress bars, summaries,
warnings and errors go to stderr and always start with rupsycho:.
Exit codes are shell friendly:
| Code | Meaning |
|---|---|
0 |
Success |
1 |
Invalid configuration, missing or protected files, a model that cannot be loaded, any other error. The message is one line on stderr, without a traceback. |
2 |
Usage error (unknown option, missing argument, bad value); printed by the argument parser |
3 |
run finished, but some model calls failed and their answers are missing |
130 |
Interrupted with Ctrl-C (answers already streamed to a file are kept) |
Nothing is overwritten by accident. run, postprocess and examples copy refuse to
replace an existing file unless you pass -f / --force.
Environment variable. RUPSYCHO_DEBUG=1 prints the full traceback of an unexpected error
instead of the one-line message.
Models load lazily. Whatever parameters.lazy_load_models says in the file, the command
loads models one at a time when the run reaches them. validate, prompt and --dry-run
therefore never download or load a model.
run¶
Run an experiment: every model answers every item as every persona, once per seed.
rupsycho run config.json -o results.csv # stream the answers into a CSV file
rupsycho run config.json --seeds 1 2 3 -o runs.jsonl # three repetitions as JSON lines
rupsycho run config.json --model small --dry-run # preview the work, load no model
rupsycho run config.json # print the answers as a table
| Option | Description |
|---|---|
CONFIG |
Experiment configuration (JSON file, see the Configuration Reference) |
-o, --output FILE |
Write the answers to FILE. The format follows the extension: .csv, .jsonl or .json. Without it the answers are printed as a table. |
-f, --force |
Overwrite FILE if it exists |
--model ID [ID ...] |
Run only these model ids (the keys of models); may be repeated. Models that are not selected are not loaded. |
--seeds S [S ...] |
Use these seeds instead of parameters.seeds (non-negative integers) |
--cumulative |
Let each persona remember its earlier answers (needs a chat prompt) |
--max-concurrency N |
Calls to run at the same time. Only API models run in parallel; local Hugging Face models always run one call at a time and the command says so. |
--on-error {warn,raise,ignore} |
What to do when a model call fails: log it and go on (warn, default), stop at the first failure (raise) or go on silently (ignore) |
--dry-run |
Load no model and write nothing; print the number of calls and the first assembled prompt |
-q, --quiet |
Hide the progress bar and the final summary (failures are still reported) |
Output formats¶
| Extension | Content | Written |
|---|---|---|
.csv |
One row per answer: experiment_name, instruction_item_id, instruction_item, model_id, profile_id, random_seed, time, answer. This is the input of postprocess. |
while running |
.jsonl |
The same fields, one JSON object per line | while running |
.json |
The whole experiment including the answers, as written by experiment.export_to_file; load it again with rup.experiment_from_file |
at the end |
| none | A table with the columns Instruction ID, Instruction Question, Model ID, Persona ID, Run Seed, Answer on stdout |
at the end |
.csv and .jsonl files grow with the run, so an interrupted or crashed run keeps the
answers it has already collected. A failed call has an empty answer in the CSV file, null in
the JSON lines, and is left out of the table and of the .json export.
Reproducibility¶
A run is determined by the configuration and its seeds: running the same command twice
produces the same answers (the time column is the only difference). The seed of every
answer is recorded in the random_seed column. If the configuration defines no seeds, a
random one is drawn and the command says so; pass --seeds to make the run reproducible.
Failed calls¶
A failed call does not stop the run (unless --on-error raise). Its answer is missing, a
summary line names the first error, and the exit code is 3:
rupsycho: 8 model calls in 0.1s, 4 failed (first error: RuntimeError: cannot answer); answers in results.csv
The exit code is 3 for warn and ignore alike, so scripts never mistake an incomplete
result for a complete one.
rupsycho run study.json -o results.csv
case $? in
0) echo "complete" ;;
3) echo "finished, but some answers are missing" ;;
*) echo "aborted" ;;
esac
validate¶
Check one or more configuration files without loading any model. Prints one line per file and
exits with 1 if any file is invalid.
OK bfi.json: Generative Models for Big Five Inventory - 4 items x 2 personas x 1 models x 1 seeds = 8 calls
ERROR typo.json: unknown top-level key(s): 'modles' (did you mean 'models'?)
ERROR missing.json: no such file
A file is valid if it is a JSON object, passes the schema, has no unknown top-level keys (a
misspelled models would silently fall back to the default model) and can actually be run:
it needs a questionnaire with items, personas, at least one model, integer seeds and a prompt
template that loads, and every persona and every item must produce a prompt. run makes the
same checks before it loads the first model. The OK line shows the size of the experiment:
items x personas x models x seeds.
prompt¶
Print the fully assembled prompt of one item and persona, exactly as the model would receive it. The prompt is the only thing on stdout, so it can be redirected into a file.
rupsycho prompt config.json
rupsycho prompt config.json --item 3 --persona "Conservative Persona" > prompt.txt
| Option | Description |
|---|---|
CONFIG |
Experiment configuration (JSON file) |
--item I |
Index of the instruction item, counting from 0 (default 0) |
--persona P |
Index or id of the persona (default 0) |
An index outside the questionnaire is an error that names the valid range. Items without
answer options of their own use the questionnaire's default_answer_options. No model is
needed, so the configuration may have "models": {}.
examples¶
Work with the example configurations that ship with the package.
rupsycho examples list # names, one per line
rupsycho examples show bfi # print one as JSON
rupsycho examples copy bfi my-study.json # a starting point for your own experiment
rupsycho examples copy bfi studies/ # into a directory: studies/bfi.json
| Action | Description |
|---|---|
list |
Print the names of the bundled examples |
show NAME |
Print the example as JSON |
copy NAME DEST [-f] |
Write the example to DEST (a file or a directory). Existing files are only replaced with -f / --force. |
postprocess¶
Turn free-text answers into scores. The command reads the CSV files written by run -o
and applies the rule-based defaults of the postprocessing
pipeline: a cleaner, the combined validator (apologies,
"as an AI" and refusals) and the multiple-choice judge, which picks the answer option the
text matches best. Items without answer options of their own are judged against the
questionnaire's default options.
rupsycho postprocess config.json results.csv -o scored.csv
rupsycho postprocess config.json 'runs/*.csv' -o scored.csv --pattern 'answer:\s*"([^"]*)"'
| Option | Description |
|---|---|
CONFIG |
The configuration the results come from (it provides the answer options) |
RESULTS ... |
CSV files written by run -o; globs are allowed and may be quoted |
-o, --output OUT |
The processed CSV file (required) |
--cleaner {basic,regex} |
basic (default) removes line breaks and non-ASCII characters; regex extracts the answer with --pattern |
--pattern REGEX |
Regular expression with one capturing group that extracts the answer; implies --cleaner regex |
-f, --force |
Overwrite OUT if it exists |
-q, --quiet |
Hide the progress bars and the summary |
The output has all columns of the results plus four new ones:
| Column | Content |
|---|---|
cleaned_answer |
The answer after the cleaner |
validation_status |
The verdict of the validator, with details |
valid |
True if the validator found nothing wrong |
decision |
The answer option chosen by the judge; not present if the text matches no option, inconclusive if it matches several equally well |
postprocess never overwrites a results file, not even with --force. For other cleaners,
validators and judges (for example the model-based ones) use the
PostprocessingPipeline from Python.
configurator¶
Launch the Streamlit configurator app, a graphical editor for
experiment configurations. It needs the configurator extra:
pip install 'rupsycho[configurator]'
rupsycho configurator # the same as rup-configurator
rupsycho configurator -- --server.port 8502
Arguments after -- are passed on to streamlit run.
Worked example: from configuration to scores¶
The bundled Big Five example, from a fresh install to a scored table.
1. Get a configuration and check it. Copying gives you a file to edit; validate loads
no model.
rupsycho: wrote bfi.json
OK bfi.json: Generative Models for Big Five Inventory - 4 items x 2 personas x 1 models x 1 seeds = 8 calls
2. Preview the work. The dry run shows how many model calls three seeds mean and what the model will be asked. Nothing is loaded and nothing is written.
Dry run: no model was loaded and nothing was written.
experiment: Generative Models for Big Five Inventory
models: Qwen/Qwen2.5-0.5B-Instruct
seeds: 1, 2, 3
calls: 4 items x 2 personas x 1 models x 3 seeds = 24 calls
output: stdout (table)
Assembled prompt (item 0, persona 0 'Optimistic Persona'):
System: Objective: Act like you are Ms Muller is 18 years old and very open minded with a optimistic personality., a survey participant answering a questionnaire.
...
The solution must be provided in this format: {answer: "answer option"}
Human: Question: I see myself as someone who...
Answer Options: 1. Disagree strongly, 2. Disagree a little, 3. Neither agree nor disagree, 4. Agree a little, 5. Agree strongly
Answer:
3. Run it. The model is downloaded on first use. The answers stream into the CSV file while the run is going.
Generative Models for Big Five Inventory: 100%|██████████| 24/24 [00:41<00:00, 1.72s/ prompts]
rupsycho: 24 model calls in 41.3s; answers in results.csv
The instruction asks for the format {answer: "answer option"}, so each row of results.csv
holds the model's raw text:
experiment_name,instruction_item_id,instruction_item,model_id,profile_id,random_seed,time,answer
Generative Models for Big Five Inventory,0,I see myself as someone who...,Qwen/Qwen2.5-0.5B-Instruct,Optimistic Persona,1,1.912,"{answer: ""5. Agree strongly""}"
Generative Models for Big Five Inventory,0,I see myself as someone who...,Qwen/Qwen2.5-0.5B-Instruct,Conservative Persona,1,1.874,"{answer: ""3. Neither agree nor disagree""}"
...
4. Score it. The regular expression pulls the chosen option out of the braces; the judge maps it onto the answer options of the questionnaire.
5. Analyse it with any tool that reads CSV, for example pandas:
import pandas as pd
scores = pd.read_csv("scored.csv")
scores = scores[scores["valid"]]
print(scores.groupby("profile_id")["decision"].value_counts())
Everything the command does is available in Python as well:
| Command | Python |
|---|---|
run |
rup.experiment_from_file(...), then experiment.run(callbacks=[CSVCallback(...)]) |
validate |
rup.experiment_from_file(...) |
prompt |
experiment.print_assembled_prompt(item_idx, persona_idx) |
examples |
rup.list_examples(), rup.load_example_config(name) |
postprocess |
PostprocessingPipeline(...).run() |
configurator |
rup-configurator |
See Running Experiments and Callbacks for the Python side.