Examples¶
All examples live in the examples/ folder of the repository. Start with the quickstart notebook: it is executed, needs no model download and no API key, and walks through the whole workflow in about ten seconds.
Notebooks¶
| File | What it shows | Requirements |
|---|---|---|
quickstart.ipynb |
Executed notebook: bundled example, assembled prompt, adding a model, running with a CSV callback, RunSummary, on_error, max_concurrency, seeds and reproducibility, cleaning / validating / judging answers, scores from the answer option weights, the CLI |
none (offline, scripted stand-in model) |
Basics.ipynb |
Load a configuration, inspect it, run it with a real local model, collect and save the answers | huggingface extra, downloads a model |
Callbacks.ipynb |
Stream answers to the console, CSV and JSONL while the experiment runs; the file formats; a custom callback | huggingface extra, downloads a model |
Parsers.ipynb |
Every cleaner, validator and judge on its own, parsers as the output parser of an experiment | none; the two model-based parsers need the huggingface extra and a download |
Postprocessing.ipynb |
The PostprocessingPipeline on the sample results in examples/data/output/ |
none (offline) |
The notebooks are committed with their outputs, so you can read them on GitHub. Cells that load a real model are tagged skip-execution and have no stored output. Install the model back-ends with an extra, for example pip install "rupsycho[huggingface,notebook]".
Scripts¶
Run them from the repository root, for example python examples/reproducibility_check.py. All of them have --help.
| File | What it shows | Requirements |
|---|---|---|
run_experiment.py |
Run a configuration with a local Hugging Face model and stream the answers into a CSV file; --dry-run validates and previews without loading anything |
huggingface extra, downloads a model (except with --dry-run) |
reproducibility_check.py |
Same seeds give the same answers; what rupsycho.seeding does for models with and without seed support, including a tiny local Hugging Face model built on the fly |
none (offline); the local-model part needs the huggingface extra |
custom_callback.py |
A callback that stores every answer, failed calls included, in a SQLite database | none (offline) |
api_models.py |
Configuration snippets and environment variables for OpenAI, Ollama, Google Gemini and DeepSeek; --run tries one of them |
none by default; --run needs the extra, a key and network access |
The scripts that run offline use small scripted stand-in models that are defined in the script and marked as such. They are not language models; they only make the behaviour of the package observable without a download.
Minimal end-to-end script¶
import rupsycho as rup
from rupsycho.callbacks import CSVCallback
from rupsycho.parsers.cleaners import BasicCleaner
from rupsycho.parsers.judges import MultipleChoiceJudge
from rupsycho.parsers.validators import ValidatorParser
from rupsycho.postprocessing import PostprocessingPipeline
CONFIG = "examples/data/bfi_demo_config.json"
# 1. Load and validate the configuration: questionnaire, personas, models, prompt, seeds
experiment = rup.experiment_from_file(CONFIG)
# 2. Look at exactly what the model will see
experiment.print_assembled_prompt(item_idx=0, persona_idx=0)
# 3. Run every model x seed x persona x question; answers stream into a CSV file
# (callbacks append: delete results.csv before you run this a second time)
summary = experiment.run(callbacks=[CSVCallback("results.csv")], on_error="warn")
print(summary) # N model calls in T seconds
# 4. Clean, validate and judge the free-text answers
options = experiment.questionnaire.default_answer_options.get_options_as_list()
processed = PostprocessingPipeline(
CONFIG,
"results.csv",
cleaner=BasicCleaner(),
validator=ValidatorParser(),
judge=MultipleChoiceJudge(options),
output_path="processed.csv",
).run()
print(processed[["profile_id", "cleaned_answer", "valid", "decision"]])
The demo configuration loads Qwen/Qwen2.5-0.5B-Instruct (about 1 GB) on first use. To try the same code without a download, replace the models of the configuration and add a scripted one, as the quickstart notebook does:
config = rup.load_example_config("bfi")
config["models"] = {}
experiment = rup.experiment_from_dict(config)
experiment.add_model(my_model, identifier="my-model") # any LangChain model
Experiment configurations¶
examples/data/ holds ready-made configurations, all validated by the test suite. The demo has a twin inside the package: rup.load_example_experiment("bfi") loads the bundled copy. The others are the larger study configurations (their names read "Experiment 1" to "Experiment 5"); they have hundreds of personas and models up to 72B parameters, so look at them with rupsycho run --dry-run before you start one.
| File | Questionnaire | Questions | Personas | Models | Model calls |
|---|---|---|---|---|---|
bfi_demo_config.json |
Big Five Inventory, 5 questions | 5 | 2 | 1 | 10 |
bfi_small_and_mid.json |
Big Five Inventory, impact of model size | 44 | 250 | 5 | 55 000 |
bdi_qwen72.json |
Beck's Depression Inventory, prompt order | 21 | 250 | 1 | 5 250 |
trolley_qwen72.json |
Trolley dilemma, prior knowledge | 3 | 250 | 1 | 750 |
gsdb_new_qwen.json |
Gender/Sex Diversity Beliefs Scale, bias detection | 23 | 250 | 1 | 5 750 |
rfq_json_format_small.json, rfq_model_friendly_small.json, rfq_natural_language_small.json |
Regulatory Focus Questionnaire in three prompt formats | 11 | 250 | 1 | 2 750 each |
example_config_for_demographics_judge.json |
Demographic questions for the DemographicsJudge (3 seeds) |
2 | 1 | 1 | 6 |
examples/data/output/ contains sample result files in the format of CSVCallback and JSONLCallback and the output of the post-processing pipeline. They were created offline with a scripted model (examples/make_sample_outputs.py), not by a language model. The materials of the user study are in examples/user_testing/.
Command line¶
Everything above that does not need Python code is also available from the shell: rupsycho validate, rupsycho prompt, rupsycho run --dry-run, rupsycho run -o results.csv, rupsycho postprocess. See the CLI reference.
Benchmarks¶
The benchmarks/ folder contains scripts that need no model download. bench_run_loop.py (poe bench) times the run loop with scripted models: the framework overhead per call and what max_concurrency buys for slow API models. bench_import.py measures the cost of import rupsycho, and bench_legacy.py compares the current implementation with the one from before the restructuring. The docstring of each script lists its options; Running experiments explains the run options they exercise.