Postprocessing¶
rupsycho.postprocessing
¶
Post-processing of experiment results: clean, validate and judge free-text answers.
Language models answer in free text ("I would say 4, because ..."). The
PostprocessingPipeline reads the CSV files
written by CSVCallback and adds, for every answer,
cleaned_answer- the answer after the cleaner,validation_status- the verdict of the validator (a dictionary with the key"validation_status"that is"valid"or"invalid"), andvalid- a boolean shortcut for it,decision- the answer option chosen by the judge.
All three stages are LangChain output parsers from rupsycho.parsers.
REQUIRED_COLUMNS
module-attribute
¶
Columns that every results file must contain (CSVCallback writes them).
PostprocessingPipeline
¶
PostprocessingPipeline(config_file_path: str | Path | Mapping[str, Any] | ExperimentDocument, results_file_patterns: str | PathLike[str] | Sequence[str | PathLike[str]], cleaner: Any, validator: Any, judge: Any, output_path: str | Path = 'processed_results.csv', *, errors: Literal['raise', 'coerce'] = 'raise', show_progress: bool = True)
Clean, validate and judge the answers stored in result CSV files.
The judge chooses among the answer options of the question; items that define no options
of their own use the questionnaire's default_answer_options.
| PARAMETER | DESCRIPTION |
|---|---|
config_file_path
|
The experiment the results belong to: the path of its JSON
configuration, a configuration dictionary or an
TYPE:
|
results_file_patterns
|
A path / glob pattern or a list of them locating the result
CSV files (as written by
TYPE:
|
cleaner
|
Output parser (or LangChain chain of parsers, such as
TYPE:
|
validator
|
Output parser returning a dictionary with
TYPE:
|
judge
|
Output parser whose
TYPE:
|
output_path
|
Where
TYPE:
|
errors
|
TYPE:
|
show_progress
|
Show progress bars.
TYPE:
|
Example
from rupsycho.parsers.cleaners import BasicCleaner
from rupsycho.parsers.judges import MultipleChoiceJudge
from rupsycho.parsers.validators import ValidatorParser
from rupsycho.postprocessing import PostprocessingPipeline
pipeline = PostprocessingPipeline(
"config.json",
"results.csv",
cleaner=BasicCleaner(),
validator=ValidatorParser(),
judge=MultipleChoiceJudge(["1. Disagree", "2. Agree"]),
output_path="processed.csv",
)
processed = pipeline.run()
run
¶
Load the results, process them and write output_path.
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
The processed results (also saved as CSV at |
| RAISES | DESCRIPTION |
|---|---|
FileNotFoundError
|
If no result file matches. |
ValueError
|
If a result file is not a |