Skip to content

Postprocessing

rupsycho.postprocessing

Post-processing of experiment results: clean, validate and judge free-text answers.

Language models answer in free text ("I would say 4, because ..."). The PostprocessingPipeline reads the CSV files written by CSVCallback and adds, for every answer,

  • cleaned_answer - the answer after the cleaner,
  • validation_status - the verdict of the validator (a dictionary with the key "validation_status" that is "valid" or "invalid"), and valid - a boolean shortcut for it,
  • decision - the answer option chosen by the judge.

All three stages are LangChain output parsers from rupsycho.parsers.

REQUIRED_COLUMNS module-attribute

REQUIRED_COLUMNS = ('instruction_item_id', 'answer')

Columns that every results file must contain (CSVCallback writes them).

PostprocessingPipeline

PostprocessingPipeline(config_file_path: str | Path | Mapping[str, Any] | ExperimentDocument, results_file_patterns: str | PathLike[str] | Sequence[str | PathLike[str]], cleaner: Any, validator: Any, judge: Any, output_path: str | Path = 'processed_results.csv', *, errors: Literal['raise', 'coerce'] = 'raise', show_progress: bool = True)

Clean, validate and judge the answers stored in result CSV files.

The judge chooses among the answer options of the question; items that define no options of their own use the questionnaire's default_answer_options.

PARAMETER DESCRIPTION
config_file_path

The experiment the results belong to: the path of its JSON configuration, a configuration dictionary or an ExperimentDocument (only its questionnaire is used; no model is ever loaded).

TYPE: str | Path | Mapping[str, Any] | ExperimentDocument

results_file_patterns

A path / glob pattern or a list of them locating the result CSV files (as written by CSVCallback).

TYPE: str | PathLike[str] | Sequence[str | PathLike[str]]

cleaner

Output parser (or LangChain chain of parsers, such as BasicCleaner() | RegexExtractorCleaner(...)) applied to the raw answer.

TYPE: Any

validator

Output parser returning a dictionary with "validation_status", e.g. ValidatorParser().

TYPE: Any

judge

Output parser whose parse(text, possible_answers) returns the chosen answer option, e.g. MultipleChoiceJudge(...).

TYPE: Any

output_path

Where run writes the processed results.

TYPE: str | Path DEFAULT: 'processed_results.csv'

errors

"raise" (default) stops at the first row a parser cannot handle, "coerce" logs it and leaves the cell empty.

TYPE: Literal['raise', 'coerce'] DEFAULT: 'raise'

show_progress

Show progress bars.

TYPE: bool DEFAULT: True

Example
from rupsycho.parsers.cleaners import BasicCleaner
from rupsycho.parsers.judges import MultipleChoiceJudge
from rupsycho.parsers.validators import ValidatorParser
from rupsycho.postprocessing import PostprocessingPipeline

pipeline = PostprocessingPipeline(
    "config.json",
    "results.csv",
    cleaner=BasicCleaner(),
    validator=ValidatorParser(),
    judge=MultipleChoiceJudge(["1. Disagree", "2. Agree"]),
    output_path="processed.csv",
)
processed = pipeline.run()

run

run() -> pd.DataFrame

Load the results, process them and write output_path.

RETURNS DESCRIPTION
DataFrame

The processed results (also saved as CSV at output_path).

RAISES DESCRIPTION
FileNotFoundError

If no result file matches.

ValueError

If a result file is not a CSVCallback output.