Skip to content

Scoring

rupsycho.scoring

Scores from judged answers: item scores, scale scores and the rules behind them.

A questionnaire carries everything that is needed to score it: every answer option has a weight and can be ignored_for_scale, an item can be reversed (reverse-keyed) and belongs to a dimension (a scale such as "Extraversion"). This module applies that metadata to the output of the PostprocessingPipeline, whose decision column holds the text of the answer option the judge chose for every free-text answer.

  • score_answers gives every answer an item score and the dimension (scale) of its item.
  • scale_scores aggregates the item scores to one score per respondent (model, persona and seed) and scale.
  • score_experiment does both in one call.

The rules in short (the scoring tutorial explains them with tables):

  • The options of an item are its own answer_options or, if it has none, the questionnaire's default_answer_options. A decision is matched with an option text exactly or, failing that, ignoring case and white space.
  • The score of an option is its weight. On a reversed item it is min + max - weight, where min and max are taken over the options that are not ignored.
  • There is no score (NaN) for ignored options, for decisions that match no option (the judges' "not present" and "inconclusive", missing values, text that is not an option) and, by default, for answers that the validator marked as invalid.
  • A scale score aggregates the item scores that exist, by default their mean. The number of items without score is reported next to it, because a scale built from few answers is weak.
Example
import pandas as pd

from rupsycho.models.questionnaire import Questionnaire
from rupsycho.scoring import scale_scores, score_answers

questionnaire = Questionnaire(
    name="Mini inventory",
    general_instruction="Rate the statement.",
    attributes={"dimension": {"1": "Extraversion"}},
    default_answer_options={
        "1": {"text": "1. Disagree", "weight": 1},
        "2": {"text": "2. Neutral", "weight": 2},
        "3": {"text": "3. Agree", "weight": 3},
    },
    instruction_items=[
        {"question": "Is talkative", "attributes": {"dimension": "1"}},
        {"question": "Is reserved", "reversed": True, "attributes": {"dimension": "1"}},
    ],
)
processed = pd.DataFrame(
    {
        "instruction_item_id": [0, 1],
        "model_id": ["model", "model"],
        "decision": ["3. Agree", "1. Disagree"],
    }
)
scored = score_answers(processed, questionnaire)
print(scored["score"].tolist())
# [3.0, 3.0]
print(scale_scores(scored).to_string(index=False))
# model_id    dimension  score  n_items  n_missing
#    model Extraversion    3.0        2          0

Both answers are worth 3 points: "Agree" on the first item, and "Disagree" on the reverse-keyed second item.

TOTAL_SCALE module-attribute

TOTAL_SCALE = 'total'

Name of the scale that covers the items without a dimension (or, on request, all items).

DEFAULT_GROUPS module-attribute

DEFAULT_GROUPS: tuple[str, ...] = ('model_id', 'profile_id', 'random_seed')

Columns that identify one simulated respondent: the model, the persona and the seed.

JUDGE_SENTINELS module-attribute

JUDGE_SENTINELS: tuple[str, ...] = ('not present', 'inconclusive')

Decisions of the judges that mean "no answer option was chosen"; they never get a score.

score_answers

score_answers(processed: DataFrame, experiment: ExperimentDocument | Questionnaire, *, decision_column: str = 'decision', item_column: str = 'instruction_item_id', valid_column: str = 'valid', only_valid: bool = True) -> pd.DataFrame

Give every answer an item score and the scale (dimension) of its item.

The item of a row is found by its position in the questionnaire (item_column, the instruction_item_id that CSVCallback writes). The decision is matched with the texts of the item's answer options, and the weight of the matching option is the score.

  • Options. The options of an item are its own answer_options; an item without options of its own (or with an empty set of them) uses the questionnaire's default_answer_options.
  • Matching. A decision matches an option text exactly, or, if there is no exact match, when both are equal after collapsing white space and ignoring case. The judge's sentinels "not present" and "inconclusive" are only ever matched exactly, so that they cannot be mistaken for an option called "Not present". Numbers are compared as text (4 and 4.0 match the option "4"). Options with the same text that score differently cannot be told apart: such a text is not scored and a UserWarning names the items.
  • Reverse keying. The score of a reversed item is min + max - weight, with min and max taken over the options that are not ignored (a 1-5 scale turns 2 into 4; a 0-3 scale turns 0 into 3).
  • No score. The score is NaN for an ignored option (ignored_for_scale), for a decision that matches no option (sentinels, missing values, text that is not an option), for an item with no scored option, and, if only_valid is true and the frame has the valid_column, for rows with valid == False. A missing valid cell does not exclude a row.
  • Dimension. The "dimension" attribute of an item is looked up in the questionnaire's attributes["dimension"] mapping ({"1": "Extraversion"}) to get the name of the scale; an id that is not listed is used as it is (as text, so 1 and "1" are the same id), and an item without dimension gets None.

The scoring is vectorised: only the distinct (item, decision) pairs are looked up, so a frame with hundreds of thousands of rows is scored in a fraction of a second.

PARAMETER DESCRIPTION
processed

The post-processed answers, one row per answer, e.g. the result of PostprocessingPipeline.run (or the CSV it wrote, read with pandas.read_csv).

TYPE: DataFrame

experiment

The experiment the answers belong to, or just its questionnaire.

TYPE: ExperimentDocument | Questionnaire

decision_column

Column with the chosen answer option (its text).

TYPE: str DEFAULT: 'decision'

item_column

Column with the position of the item in the questionnaire (0-based).

TYPE: str DEFAULT: 'instruction_item_id'

valid_column

Column with the validator's verdict (True for a usable answer).

TYPE: str DEFAULT: 'valid'

only_valid

Give no score to rows with valid == False. Has no effect if the frame has no valid_column.

TYPE: bool DEFAULT: True

RETURNS DESCRIPTION
DataFrame

A copy of processed with two more columns, "score" (float, NaN where the

DataFrame

row cannot be scored) and "dimension" (the scale name, None for items without

DataFrame

dimension). Columns of these names are replaced. The order of the scales is recorded in

DataFrame

result.attrs so that scale_scores lists the

DataFrame

scales in the order of the questionnaire.

RAISES DESCRIPTION
TypeError

If processed is not a data frame, experiment is neither an experiment nor a questionnaire, or the decision or validity column holds lists or dictionaries (for instance the validation_status column instead of valid).

ValueError

If the item or decision column is missing, if an item id is not an integer position of the questionnaire (a hint that the experiment does not match the results), or if the experiment has no questionnaire.

WARNS DESCRIPTION
UserWarning

If options of an item share a text but not their score, or if no decision at all is the text of an option (a hint that the raw answers or the wrong experiment were passed).

Example
import pandas as pd

from rupsycho.models.questionnaire import Questionnaire
from rupsycho.scoring import score_answers

questionnaire = Questionnaire(
    name="Mini inventory",
    general_instruction="Rate the statement.",
    attributes={"dimension": {"1": "Extraversion"}},
    default_answer_options={
        "1": {"text": "1. Disagree", "weight": 1},
        "2": {"text": "2. Neutral", "weight": 2},
        "3": {"text": "3. Agree", "weight": 3},
        "4": {"text": "4. Don't know", "weight": 0, "ignored_for_scale": True},
    },
    instruction_items=[
        {"question": "Is talkative", "attributes": {"dimension": "1"}},
        {"question": "Is reserved", "reversed": True, "attributes": {"dimension": "1"}},
    ],
)
processed = pd.DataFrame(
    {
        "instruction_item_id": [0, 0, 1, 1, 1],
        "decision": [
            "3. Agree",
            " 2. neutral ",
            "1. Disagree",
            "4. Don't know",
            "not present",
        ],
        "valid": [True, True, True, True, False],
    }
)
scored = score_answers(processed, questionnaire)
print(scored[["decision", "score", "dimension"]].to_string())
#         decision  score     dimension
# 0       3. Agree    3.0  Extraversion
# 1    2. neutral     2.0  Extraversion
# 2    1. Disagree    3.0  Extraversion
# 3  4. Don't know    NaN  Extraversion
# 4    not present    NaN  Extraversion

Row 1 matches although its case and white space differ, row 2 is reverse-keyed (1 becomes 3), and rows 3 and 4 get no score: an ignored option and a judge's sentinel.

scale_scores

scale_scores(scored: DataFrame, *, by: str | Sequence[str] = DEFAULT_GROUPS, dimension_column: str = 'dimension', score_column: str = 'score', agg: str | Callable[[Series], Any] = 'mean', include_total: bool = False, dimensions: Sequence[str] | None = None) -> pd.DataFrame

Aggregate item scores to scale scores, one row per respondent and scale.

A respondent is one combination of the by columns (by default model, persona and seed), a scale is a dimension. The score of a scale is the aggregate (agg, by default the mean) of the item scores that exist; items without score (see score_answers) are not counted as zero but reported in n_missing. A scale without any item score gets the score NaN, whatever agg is.

  • Items without dimension form the scale "total". If include_total is true and there are items with a dimension, "total" is instead the aggregate over all item scores (not the mean of the scale scores).
  • Columns of by that the frame does not have are skipped, so results without a random_seed column can be aggregated as they are. The three standard columns are skipped silently, any other missing name triggers a UserWarning (a typo would otherwise merge respondents).
  • Rows with a missing value in a by column form a group of their own.
  • The result is sorted by the by columns, then by scale: in the order of the questionnaire (recorded by score_answers, or the order given in dimensions), scales that are not listed in the order of appearance, and "total" last.

Items of one scale should use the same weights; the mean of items scored 1-5 and items scored 0-3 is hard to interpret.

PARAMETER DESCRIPTION
scored

Item scores, usually the result of score_answers.

TYPE: DataFrame

by

Column(s) that identify a respondent. A single name is accepted.

TYPE: str | Sequence[str] DEFAULT: DEFAULT_GROUPS

dimension_column

Column with the scale of each row; missing values and blank text mean "no dimension". The result uses the same column name.

TYPE: str DEFAULT: 'dimension'

score_column

Column with the item scores. The result uses the same column name.

TYPE: str DEFAULT: 'score'

agg

How to aggregate the item scores: "mean", "sum", any other name of a pandas aggregation ("median", "std", "max", ...) or a function from a series to a number.

TYPE: str | Callable[[Series], Any] DEFAULT: 'mean'

include_total

Also report the overall "total" of questionnaires with dimensions.

TYPE: bool DEFAULT: False

dimensions

Order of the scales, if it should differ from the questionnaire's.

TYPE: Sequence[str] | None DEFAULT: None

RETURNS DESCRIPTION
DataFrame

A data frame with the by columns that exist, dimension_column (scale name),

DataFrame

score_column (the aggregate), n_items (how many item scores went into it) and

DataFrame

n_missing (how many rows had no score). It has a fresh index and no rows if

DataFrame

scored has none.

RAISES DESCRIPTION
TypeError

If scored is not a data frame or its dimension column holds lists or dictionaries.

ValueError

If the dimension or score column is missing, if the scores are not numbers, if by contains the dimension or score column, or if agg is not a valid aggregation.

WARNS DESCRIPTION
UserWarning

If by names a column other than the three standard ones that the frame does not have.

Example
import pandas as pd

from rupsycho.scoring import scale_scores

scored = pd.DataFrame(
    {
        "model_id": ["a", "a", "a", "a", "b", "b"],
        "dimension": ["Extraversion", "Extraversion", "Neuroticism", "Neuroticism"]
        + ["Extraversion", "Neuroticism"],
        "score": [5.0, 3.0, 2.0, float("nan"), 4.0, 1.0],
    }
)
print(scale_scores(scored, by="model_id").to_string(index=False))
# model_id    dimension  score  n_items  n_missing
#        a Extraversion    4.0        2          0
#        a  Neuroticism    2.0        1          1
#        b Extraversion    4.0        1          0
#        b  Neuroticism    1.0        1          0
totals = scale_scores(scored, by="model_id", agg="sum", include_total=True)
print(totals.to_string(index=False))
# model_id    dimension  score  n_items  n_missing
#        a Extraversion    8.0        2          0
#        a  Neuroticism    2.0        1          1
#        a        total   10.0        3          1
#        b Extraversion    4.0        1          0
#        b  Neuroticism    1.0        1          0
#        b        total    5.0        2          0

score_experiment

score_experiment(experiment: ExperimentDocument | Questionnaire, processed: DataFrame | None = None, *, decision_column: str = 'decision', item_column: str = 'instruction_item_id', valid_column: str = 'valid', only_valid: bool = True, by: str | Sequence[str] = DEFAULT_GROUPS, agg: str | Callable[[Series], Any] = 'mean', include_total: bool = False) -> pd.DataFrame

Score an experiment: item scores from score_answers, then scale_scores.

Pass the output of the PostprocessingPipeline as processed. Without it, the answers stored in the experiment (experiment.get_answers_as_dataframe()) are used, and the raw answer text is taken as the decision. That only works if the model answered with exactly the text of an answer option (and was not asked for anything else, such as a JSON object); real free-text answers need to go through the pipeline first (score_answers warns if not a single answer is an option). Calls that failed have no row in the stored answers, so they are not counted in n_missing.

PARAMETER DESCRIPTION
experiment

The experiment (or its questionnaire) with the scoring metadata. Without processed it must be an experiment that holds the answers.

TYPE: ExperimentDocument | Questionnaire

processed

The post-processed answers. Default: the answers of experiment (see above).

TYPE: DataFrame | None DEFAULT: None

decision_column

See score_answers; refers to processed (the frame built from the experiment uses the standard names).

TYPE: str DEFAULT: 'decision'

item_column

See score_answers; refers to processed.

TYPE: str DEFAULT: 'instruction_item_id'

valid_column

See score_answers; refers to processed.

TYPE: str DEFAULT: 'valid'

only_valid

See score_answers.

TYPE: bool DEFAULT: True

by

TYPE: str | Sequence[str] DEFAULT: DEFAULT_GROUPS

agg

See scale_scores.

TYPE: str | Callable[[Series], Any] DEFAULT: 'mean'

include_total

See scale_scores.

TYPE: bool DEFAULT: False

RETURNS DESCRIPTION
DataFrame

The scale scores, see scale_scores.

RAISES DESCRIPTION
ValueError

If processed is omitted and experiment holds no answers (it is a questionnaire), plus the errors of score_answers and scale_scores.

WARNS DESCRIPTION
UserWarning

The warnings of score_answers and scale_scores.

Example
import rupsycho as rup
from langchain_core.language_models.fake import FakeListLLM

from rupsycho.scoring import score_experiment

experiment = rup.ExperimentDocument(
    name="Mini inventory",
    models={},
    demographic_profiles={
        "Anna": {"attributes": {"name": "Anna", "age": 30}},
        "Ben": {"attributes": {"name": "Ben", "age": 45}},
    },
    questionnaire={
        "name": "Mini inventory",
        "general_instruction": "Rate the statement.",
        "attributes": {"dimension": {"1": "Extraversion"}},
        "default_answer_options": {
            "1": {"text": "1. Disagree", "weight": 1},
            "2": {"text": "2. Neutral", "weight": 2},
            "3": {"text": "3. Agree", "weight": 3},
        },
        "instruction_items": [
            {"question": "Is talkative", "attributes": {"dimension": "1"}},
            {"question": "Is reserved", "reversed": True, "attributes": {"dimension": "1"}},
        ],
    },
)
experiment.add_model(FakeListLLM(responses=["3. Agree"]), identifier="agreeable-model")
experiment.run(show_progress=False)
print(score_experiment(experiment, by="profile_id").to_string(index=False))
# profile_id    dimension  score  n_items  n_missing
#       Anna Extraversion    2.0        2          0
#        Ben Extraversion    2.0        2          0

A model that agrees with everything lands in the middle of the scale, because the second item is reverse-keyed.