Skip to content

Experiment

rupsycho.experiment.ExperimentDocument

ExperimentDocument(**data: Any)

Class for storing an experiment and associated questionnaires along with metadata.

Extends LangChain's BaseMedia and Pydantic's BaseModel, integrating data validation and serialization capabilities with document handling features.

Example:

.. code-block:: python

    from my_module import ExperimentDocument

    experiment_doc = ExperimentDocument(
        name="Experiment 1",
        description="Test Experiment",
        demographic_profiles={...},
        models={...},
        questionnaire=...,
        metadata={"source": "Lab A"}
    )

Initialize ExperimentDocument with optional conversions for nested data.

name

name: str | None = Field(None, description='The name of the experiment', examples=['Generative Models for Big Five Inventory'])

description

description: str | None = Field(None, description='The description of the experiment', examples=['Testing Generative Models for BFI questionnaire using Rupsycho.'])

parameters

parameters: ExperimentParameters = Field(default_factory=ExperimentParameters, description='The parameters for the experiment and text generation')

prompt_template

prompt_template: NormalPromptTemplateConfig | ChatPromptTemplateConfig | LangchainPromptTemplateConfig = Field(None, description='The prompt template used by the model')

models

models: dict[str, LangChainModelConfig | LocalHuggingFaceModelConfig | RemoteHuggingFaceModelConfig | OllamaModelConfig | OpenAIModelConfig | GoogleModelConfig | DeepSeekModelConfig] = Field(default_factory=dict, description='The models in the experiment')

demographic_profiles

demographic_profiles: dict[str, DemographicProfile] = Field(default_factory=dict, description='The demographic profiles in the experiment')

questionnaire

questionnaire: Questionnaire | None = Field(None, description='The questionnaire in the experiment')

metadata

metadata: dict[str, Any] = Field(default_factory=dict, description='Additional metadata for the experiment document.')

set_questionnaire

set_questionnaire(questionnaire: Questionnaire) -> None

Sets the questionnaire for the experiment.

set_parser

set_parser(parser: Any) -> None

Adds a parser to the experiment.

rupsycho.experiment_collection.ExperimentCollection

ExperimentCollection(experiments: list[ExperimentDocument], **kwargs: Any)

Class for storing a collection of ExperimentDocument objects along with metadata.

Example:

.. code-block:: python

    from my_module import ExperimentCollection, ExperimentDocument

    experiment_doc1 = ExperimentDocument(
        name="Experiment 1",
        description="Test Experiment",
        demographic_profiles=[...],
        models=[...],
        questionnaire=...,
        metadata={"source": "Lab A"}
    )

    experiment_doc2 = ExperimentDocument(
        name="Experiment 2",
        description="Another Test",
        demographic_profiles=[...],
        models=[...],
        questionnaire=...,
        metadata={"source": "Lab B"}
    )

    experiment_collection = ExperimentCollection(
        experiments=[experiment_doc1, experiment_doc2],
        metadata={"source": "Lab Collection"}
    )

Initialize an ExperimentCollection with a list of ExperimentDocument objects.

experiments instance-attribute

experiments: list[ExperimentDocument]

List of ExperimentDocument objects representing the experiments.

is_lc_serializable classmethod

is_lc_serializable() -> bool

Return whether this class is serializable.

get_lc_namespace classmethod

get_lc_namespace() -> list[str]

Get the namespace of the LangChain object.

run_all

run_all(**run_kwargs: Any) -> list[RunSummary]

Run every experiment of the collection, one after another.

PARAMETER DESCRIPTION
**run_kwargs

Passed to run (callbacks, cumulative, max_concurrency, on_error, ...).

TYPE: Any DEFAULT: {}

RETURNS DESCRIPTION
list[RunSummary]

One RunSummary per experiment.

Mixins

rupsycho.mixins.experiment_processing.ExperimentProcessingMixin

Runs an experiment: models x seeds x items x personas.

The mixin relies on the attributes of ExperimentDocument (questionnaire, demographic_profiles, runnable_models, runnable_prompt, runnable_parser, parameters, models and name).

assemble_prompt

assemble_prompt(item_idx: int = 0, persona_idx: int = 0) -> str

Return the fully assembled prompt for one item and persona, exactly as sent.

PARAMETER DESCRIPTION
item_idx

Index of the instruction item.

TYPE: int DEFAULT: 0

persona_idx

Index of the demographic profile (in configuration order).

TYPE: int DEFAULT: 0

RETURNS DESCRIPTION
The prompt text; for chat prompts the messages are joined as ``"System

..."`` /

``"Human

..."`` blocks.

RAISES DESCRIPTION
IndexError

If item_idx or persona_idx is out of range.

ValueError

If the experiment has no questionnaire or prompt.

Note

Does not include the accumulated memory of cumulative runs.

Example
print(experiment.assemble_prompt(item_idx=0, persona_idx=1))

print_assembled_prompt

print_assembled_prompt(item_idx: int = 0, persona_idx: int = 0) -> None

Print the assembled prompt of one item and persona (see assemble_prompt).

Problems (for example an out-of-range index) are reported as a warning instead of raising, which is convenient in notebooks.

PARAMETER DESCRIPTION
item_idx

Index of the instruction item.

TYPE: int DEFAULT: 0

persona_idx

Index of the demographic profile.

TYPE: int DEFAULT: 0

process_single_experiment

process_single_experiment(cumulative: bool, pbar: Any = None, callbacks: Sequence = (), *, max_concurrency: int = 1, on_error: ErrorPolicy = 'warn') -> RunSummary

Process every model of the experiment and return a :class:RunSummary.

PARAMETER DESCRIPTION
cumulative

Let each persona remember its earlier answers.

TYPE: bool

pbar

Progress bar to update; one is created if omitted.

TYPE: Any DEFAULT: None

callbacks

Callbacks that receive every answer.

TYPE: Sequence DEFAULT: ()

max_concurrency

Calls to run at the same time (API models only; local Hugging Face models always run sequentially).

TYPE: int DEFAULT: 1

on_error

"warn" (default) logs failed calls and continues, "raise" stops at the first failure, "ignore" continues silently.

TYPE: ErrorPolicy DEFAULT: 'warn'

RETURNS DESCRIPTION
RunSummary

A summary of the calls made.

RAISES DESCRIPTION
ValueError

If on_error or max_concurrency is invalid or the experiment is incomplete.

run

run(callbacks: Sequence = (), cumulative: bool = False, *, max_concurrency: int = 1, on_error: ErrorPolicy = 'warn', show_progress: bool = True) -> RunSummary

Run the experiment.

Every model is asked every question as every persona, once per seed. Answers are stored on the questionnaire items (see get_answers / get_answers_as_dataframe) and passed to the callbacks as soon as they are generated.

PARAMETER DESCRIPTION
callbacks

Callbacks such as CSVCallback that receive every answer.

TYPE: Sequence DEFAULT: ()

cumulative

Let each persona remember its earlier answers ("response memory"). Needs a chat prompt whose user message is asked once per item.

TYPE: bool DEFAULT: False

max_concurrency

Number of calls to run in parallel. Useful for API models where the time is spent waiting; models running in this process (local Hugging Face) are always called sequentially. Results and callbacks keep their order.

TYPE: int DEFAULT: 1

on_error

What to do when a model call fails. "warn" (default) logs the failure, continues and warns once at the end; "raise" stops at the first failure; "ignore" continues silently. Failed calls have no answer.

TYPE: ErrorPolicy DEFAULT: 'warn'

show_progress

Show a progress bar.

TYPE: bool DEFAULT: True

RETURNS DESCRIPTION
RunSummary

A RunSummary with the number

RunSummary

of calls, failures and the elapsed time.

RAISES DESCRIPTION
ValueError

If the experiment is incomplete (no model, prompt, questionnaire).

Example
summary = experiment.run(callbacks=[CSVCallback("answers.csv")], max_concurrency=8)
print(summary)  # "160 model calls in 12.3s"

rupsycho.mixins.experiment_exporting.ExperimentExportMixin

Mixin providing methods to export the experiment results to a file or return the answers.

to_config

to_config(*, include_answers: bool = True) -> dict[str, Any]

Return the experiment as a plain, JSON-serialisable configuration dictionary.

Only the configuration is exported (name, description, parameters, models, prompt template, personas, questionnaire), never runtime objects. Secrets such as API keys are masked. The result can be passed to experiment_from_dict, so an exported experiment can be shared, versioned and re-loaded.

PARAMETER DESCRIPTION
include_answers

Keep the answers collected so far on the questionnaire items. Pass False to export only the experiment definition.

TYPE: bool DEFAULT: True

RETURNS DESCRIPTION
dict[str, Any]

The configuration dictionary.

Example
config = experiment.to_config(include_answers=False)

export_to_file

export_to_file(filename: str | PathLike[str], *, include_answers: bool = True) -> None

Write the experiment configuration (and answers) to a JSON file.

PARAMETER DESCRIPTION
filename

Target file; it is written as UTF-8 and overwritten if it exists.

TYPE: str | PathLike[str]

include_answers

Keep the answers collected so far; see to_config.

TYPE: bool DEFAULT: True

RAISES DESCRIPTION
OSError

If the file cannot be written.

get_answers

get_answers() -> list[dict[str, Any]]

Extracts and returns the answers from the experiment in a list holding the nested answer structure.

:return: A list of dictionaries representing the answers for each instruction item.

get_answers_as_dataframe

get_answers_as_dataframe() -> pd.DataFrame

Returns a flat pandas DataFrame with the experiment's answers.

Columns include: - "Instruction ID" - "Instruction Question" - "Model ID" - "Persona ID" - "Run Seed" - "Answer"

:return: A pandas DataFrame containing the flattened answers.

rupsycho.mixins.model_managing.ModelManagementMixin

Methods to manage the models of an experiment.

An experiment keeps two dictionaries under the same identifiers: models holds the configurations (what gets exported), runnable_models holds what is run - either the configuration itself (loaded lazily when the run reaches it) or a ready LangChain model added with add_model.

load_model

load_model(model_definition: dict[str, Any]) -> Any | None

Deserialize a model from its LangChain definition.

PARAMETER DESCRIPTION
model_definition

Serialized model as produced by langchain_core.load.dumpd.

TYPE: dict[str, Any]

RETURNS DESCRIPTION
Any | None

The model, or None (with a warning) if it cannot be deserialized.

add_model

add_model(model: Any, identifier: str | None = None) -> None

Add a ready LangChain model to the experiment.

The model is run as it is. Its serialized definition is stored in models so that the experiment can be exported; models that cannot be serialized (for example local pipelines) are exported as a placeholder and have to be added again after loading.

PARAMETER DESCRIPTION
model

Any LangChain runnable (chat model, LLM, ...).

TYPE: Any

identifier

Name of the model in the results. Defaults to its object id.

TYPE: str | None DEFAULT: None

Example
experiment.add_model(ChatOpenAI(model="gpt-4o-mini"), identifier="gpt-4o-mini")
Note

Adding a model under an identifier that exists replaces it and emits a warning.

set_runnable_models

set_runnable_models() -> None

Reset runnable_models to the configured models.

Every model is then loaded from its configuration when the run reaches it. Models that were added as live objects without a loadable configuration are no longer available afterwards.

get_model

get_model(identifier: str) -> Any | None

Return the runnable entry of a model.

PARAMETER DESCRIPTION
identifier

Identifier of the model.

TYPE: str

RETURNS DESCRIPTION
Any | None

The LangChain model, or - for models that are loaded lazily - its configuration;

Any | None

None if there is no such model.

remove_model

remove_model(identifier: str) -> None

Remove a model from the experiment.

PARAMETER DESCRIPTION
identifier

Identifier of the model. Unknown identifiers only emit a warning.

TYPE: str

list_models

list_models() -> list[str]

Return the identifiers of all models of the experiment.

has_model

has_model(identifier: str) -> bool

Return whether a model with this identifier exists.

replace_model

replace_model(identifier: str, new_model: Any) -> None

Replace an existing model by a ready LangChain model.

PARAMETER DESCRIPTION
identifier

Identifier of the model to replace. Unknown identifiers only emit a warning.

TYPE: str

new_model

The new model.

TYPE: Any

clear_models

clear_models() -> None

Remove all models from the experiment.

count_models

count_models() -> int

Return the number of models in the experiment.

get_all_runnable_models

get_all_runnable_models() -> dict[str, Any]

Return the runnable entries of all models, keyed by identifier.

rupsycho.mixins.persona_managing.PersonaManagementMixin

Methods to manage the personas of an experiment.

Personas are the demographic profiles the models answer as. They are stored in demographic_profiles under unique identifiers, which also label the rows of the results.

add_persona

add_persona(persona: DemographicProfile, identifier: str | None = None) -> None

Add a persona to the experiment.

PARAMETER DESCRIPTION
persona

The demographic profile to add.

TYPE: DemographicProfile

identifier

Name of the persona in the results. Defaults to the object id.

TYPE: str | None DEFAULT: None

Example
from rupsycho.models.questionnaire import DemographicProfile

experiment.add_persona(
    DemographicProfile(
        attributes={"name": "Alex", "age": 34},
        template="{name} is {age} years old.",
    ),
    identifier="Alex",
)
Note

Adding a persona under an identifier that exists replaces it and emits a warning.

get_persona

get_persona(identifier: str) -> DemographicProfile | None

Return the persona with this identifier, or None if there is none.

remove_persona

remove_persona(identifier: str) -> None

Remove a persona.

PARAMETER DESCRIPTION
identifier

Identifier of the persona. Unknown identifiers only emit a warning.

TYPE: str

list_personas

list_personas() -> list[str]

Return the identifiers of all personas.

clear_personas

clear_personas() -> None

Remove all personas from the experiment.

rupsycho.mixins.prompt_managing.PromptTemplateMixin

Mixin providing methods to manage the prompt template in the experiment.

This includes setting the prompt template, loading it, and converting it into a runnable form.

load_prompt

load_prompt(prompt_template: dict[str, Any]) -> Any | None

Load the prompt template from its serialized definition.

:param prompt_template: Serialized prompt template. :return: Loaded prompt, or None if an error occurs.

set_prompt

set_prompt(prompt: Any) -> None

Use a ready-made LangChain prompt for this experiment.

The prompt is kept as a serialized langchain prompt configuration, so it survives export_to_file and can be loaded again.

PARAMETER DESCRIPTION
prompt

A LangChain prompt template such as ChatPromptTemplate. Its input variables may be general_instruction, persona_description, question and answer_options.

TYPE: Any

Example
from langchain_core.prompts import ChatPromptTemplate

experiment.set_prompt(ChatPromptTemplate.from_messages([
    ("system", "Answer as {persona_description}."),
    ("user", "{question}\n{answer_options}"),
]))

get_prompt

get_prompt() -> Any | None

Return the current runnable prompt template.

RETURNS DESCRIPTION
Any | None

The LangChain prompt, or None if none is set.

get_prompt_config

get_prompt_config() -> Any | None

Return the prompt template configuration (normal, chat or langchain).

RETURNS DESCRIPTION
Any | None

The configuration object, or None if none is set.

set_prompt_config

set_prompt_config(prompt_template: dict[str, Any] | BaseModel) -> None

Set the prompt from a configuration, as it would appear in a JSON file.

PARAMETER DESCRIPTION
prompt_template

A configuration dictionary ({"type": "chat", "messages": [...]}, {"type": "normal", "template": "..."}) or LangChain's own serialization of a prompt (as produced by langchain_core.load.dumpd), or a config object.

TYPE: dict[str, Any] | BaseModel

RAISES DESCRIPTION
ValueError

If the configuration is not recognised.

reset_prompt

reset_prompt() -> None

Resets the current prompt template and runnable form to None.

has_prompt

has_prompt() -> bool

Check whether a runnable prompt is set.

:return: True if a runnable prompt is set, False otherwise.

Run summary

rupsycho.mixins.experiment_processing.RunSummary dataclass

RunSummary(n_calls: int = 0, n_failed: int = 0, elapsed: float = 0.0, errors: list[str] = list())

Outcome of :meth:ExperimentProcessingMixin.run.

ATTRIBUTE DESCRIPTION
n_calls

Number of model calls made.

TYPE: int

n_failed

Number of calls that raised; their answers are missing from the results.

TYPE: int

elapsed

Wall-clock seconds of the whole run.

TYPE: float

errors

The first distinct error messages (at most five).

TYPE: list[str]

Example
summary = experiment.run()
if summary.n_failed:
    print(summary.errors)

n_succeeded property

n_succeeded: int

Number of calls that returned an answer.