Scoring¶
rupsycho.scoring
¶
Scores from judged answers: item scores, scale scores and the rules behind them.
A questionnaire carries everything that is needed to score it: every answer option has a
weight and can be ignored_for_scale, an item can be reversed (reverse-keyed) and
belongs to a dimension (a scale such as "Extraversion"). This module applies that metadata
to the output of the
PostprocessingPipeline, whose decision
column holds the text of the answer option the judge chose for every free-text answer.
score_answersgives every answer an item score and the dimension (scale) of its item.scale_scoresaggregates the item scores to one score per respondent (model, persona and seed) and scale.score_experimentdoes both in one call.
The rules in short (the scoring tutorial explains them with tables):
- The options of an item are its own
answer_optionsor, if it has none, the questionnaire'sdefault_answer_options. A decision is matched with an option text exactly or, failing that, ignoring case and white space. - The score of an option is its
weight. On a reversed item it ismin + max - weight, whereminandmaxare taken over the options that are not ignored. - There is no score (
NaN) for ignored options, for decisions that match no option (the judges'"not present"and"inconclusive", missing values, text that is not an option) and, by default, for answers that the validator marked as invalid. - A scale score aggregates the item scores that exist, by default their mean. The number of items without score is reported next to it, because a scale built from few answers is weak.
Example
import pandas as pd
from rupsycho.models.questionnaire import Questionnaire
from rupsycho.scoring import scale_scores, score_answers
questionnaire = Questionnaire(
name="Mini inventory",
general_instruction="Rate the statement.",
attributes={"dimension": {"1": "Extraversion"}},
default_answer_options={
"1": {"text": "1. Disagree", "weight": 1},
"2": {"text": "2. Neutral", "weight": 2},
"3": {"text": "3. Agree", "weight": 3},
},
instruction_items=[
{"question": "Is talkative", "attributes": {"dimension": "1"}},
{"question": "Is reserved", "reversed": True, "attributes": {"dimension": "1"}},
],
)
processed = pd.DataFrame(
{
"instruction_item_id": [0, 1],
"model_id": ["model", "model"],
"decision": ["3. Agree", "1. Disagree"],
}
)
scored = score_answers(processed, questionnaire)
print(scored["score"].tolist())
# [3.0, 3.0]
print(scale_scores(scored).to_string(index=False))
# model_id dimension score n_items n_missing
# model Extraversion 3.0 2 0
Both answers are worth 3 points: "Agree" on the first item, and "Disagree" on the reverse-keyed second item.
TOTAL_SCALE
module-attribute
¶
Name of the scale that covers the items without a dimension (or, on request, all items).
DEFAULT_GROUPS
module-attribute
¶
Columns that identify one simulated respondent: the model, the persona and the seed.
JUDGE_SENTINELS
module-attribute
¶
Decisions of the judges that mean "no answer option was chosen"; they never get a score.
score_answers
¶
score_answers(processed: DataFrame, experiment: ExperimentDocument | Questionnaire, *, decision_column: str = 'decision', item_column: str = 'instruction_item_id', valid_column: str = 'valid', only_valid: bool = True) -> pd.DataFrame
Give every answer an item score and the scale (dimension) of its item.
The item of a row is found by its position in the questionnaire (item_column, the
instruction_item_id that CSVCallback writes). The decision is matched with the
texts of the item's answer options, and the weight of the matching option is the score.
- Options. The options of an item are its own
answer_options; an item without options of its own (or with an empty set of them) uses the questionnaire'sdefault_answer_options. - Matching. A decision matches an option text exactly, or, if there is no exact match,
when both are equal after collapsing white space and ignoring case. The judge's
sentinels
"not present"and"inconclusive"are only ever matched exactly, so that they cannot be mistaken for an option called "Not present". Numbers are compared as text (4and4.0match the option"4"). Options with the same text that score differently cannot be told apart: such a text is not scored and aUserWarningnames the items. - Reverse keying. The score of a reversed item is
min + max - weight, withminandmaxtaken over the options that are not ignored (a 1-5 scale turns 2 into 4; a 0-3 scale turns 0 into 3). - No score. The score is
NaNfor an ignored option (ignored_for_scale), for a decision that matches no option (sentinels, missing values, text that is not an option), for an item with no scored option, and, ifonly_validis true and the frame has thevalid_column, for rows withvalid == False. A missingvalidcell does not exclude a row. - Dimension. The
"dimension"attribute of an item is looked up in the questionnaire'sattributes["dimension"]mapping ({"1": "Extraversion"}) to get the name of the scale; an id that is not listed is used as it is (as text, so1and"1"are the same id), and an item without dimension getsNone.
The scoring is vectorised: only the distinct (item, decision) pairs are looked up, so a frame with hundreds of thousands of rows is scored in a fraction of a second.
| PARAMETER | DESCRIPTION |
|---|---|
processed
|
The post-processed answers, one row per answer, e.g. the result of
TYPE:
|
experiment
|
The experiment the answers belong to, or just its questionnaire.
TYPE:
|
decision_column
|
Column with the chosen answer option (its text).
TYPE:
|
item_column
|
Column with the position of the item in the questionnaire (0-based).
TYPE:
|
valid_column
|
Column with the validator's verdict (
TYPE:
|
only_valid
|
Give no score to rows with
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
A copy of |
DataFrame
|
row cannot be scored) and |
DataFrame
|
dimension). Columns of these names are replaced. The order of the scales is recorded in |
DataFrame
|
|
DataFrame
|
scales in the order of the questionnaire. |
| RAISES | DESCRIPTION |
|---|---|
TypeError
|
If |
ValueError
|
If the item or decision column is missing, if an item id is not an integer position of the questionnaire (a hint that the experiment does not match the results), or if the experiment has no questionnaire. |
| WARNS | DESCRIPTION |
|---|---|
UserWarning
|
If options of an item share a text but not their score, or if no decision at all is the text of an option (a hint that the raw answers or the wrong experiment were passed). |
Example
import pandas as pd
from rupsycho.models.questionnaire import Questionnaire
from rupsycho.scoring import score_answers
questionnaire = Questionnaire(
name="Mini inventory",
general_instruction="Rate the statement.",
attributes={"dimension": {"1": "Extraversion"}},
default_answer_options={
"1": {"text": "1. Disagree", "weight": 1},
"2": {"text": "2. Neutral", "weight": 2},
"3": {"text": "3. Agree", "weight": 3},
"4": {"text": "4. Don't know", "weight": 0, "ignored_for_scale": True},
},
instruction_items=[
{"question": "Is talkative", "attributes": {"dimension": "1"}},
{"question": "Is reserved", "reversed": True, "attributes": {"dimension": "1"}},
],
)
processed = pd.DataFrame(
{
"instruction_item_id": [0, 0, 1, 1, 1],
"decision": [
"3. Agree",
" 2. neutral ",
"1. Disagree",
"4. Don't know",
"not present",
],
"valid": [True, True, True, True, False],
}
)
scored = score_answers(processed, questionnaire)
print(scored[["decision", "score", "dimension"]].to_string())
# decision score dimension
# 0 3. Agree 3.0 Extraversion
# 1 2. neutral 2.0 Extraversion
# 2 1. Disagree 3.0 Extraversion
# 3 4. Don't know NaN Extraversion
# 4 not present NaN Extraversion
Row 1 matches although its case and white space differ, row 2 is reverse-keyed (1 becomes 3), and rows 3 and 4 get no score: an ignored option and a judge's sentinel.
scale_scores
¶
scale_scores(scored: DataFrame, *, by: str | Sequence[str] = DEFAULT_GROUPS, dimension_column: str = 'dimension', score_column: str = 'score', agg: str | Callable[[Series], Any] = 'mean', include_total: bool = False, dimensions: Sequence[str] | None = None) -> pd.DataFrame
Aggregate item scores to scale scores, one row per respondent and scale.
A respondent is one combination of the by columns (by default model, persona and
seed), a scale is a dimension. The score of a scale is the aggregate (agg, by default
the mean) of the item scores that exist; items without score (see
score_answers) are not counted as zero but reported in
n_missing. A scale without any item score gets the score NaN, whatever agg is.
- Items without dimension form the scale
"total". Ifinclude_totalis true and there are items with a dimension,"total"is instead the aggregate over all item scores (not the mean of the scale scores). - Columns of
bythat the frame does not have are skipped, so results without arandom_seedcolumn can be aggregated as they are. The three standard columns are skipped silently, any other missing name triggers aUserWarning(a typo would otherwise merge respondents). - Rows with a missing value in a
bycolumn form a group of their own. - The result is sorted by the
bycolumns, then by scale: in the order of the questionnaire (recorded byscore_answers, or the order given indimensions), scales that are not listed in the order of appearance, and"total"last.
Items of one scale should use the same weights; the mean of items scored 1-5 and items scored 0-3 is hard to interpret.
| PARAMETER | DESCRIPTION |
|---|---|
scored
|
Item scores, usually the result of
TYPE:
|
by
|
Column(s) that identify a respondent. A single name is accepted.
TYPE:
|
dimension_column
|
Column with the scale of each row; missing values and blank text mean "no dimension". The result uses the same column name.
TYPE:
|
score_column
|
Column with the item scores. The result uses the same column name.
TYPE:
|
agg
|
How to aggregate the item scores:
TYPE:
|
include_total
|
Also report the overall
TYPE:
|
dimensions
|
Order of the scales, if it should differ from the questionnaire's.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
A data frame with the |
DataFrame
|
|
DataFrame
|
|
DataFrame
|
|
| RAISES | DESCRIPTION |
|---|---|
TypeError
|
If |
ValueError
|
If the dimension or score column is missing, if the scores are not
numbers, if |
| WARNS | DESCRIPTION |
|---|---|
UserWarning
|
If |
Example
import pandas as pd
from rupsycho.scoring import scale_scores
scored = pd.DataFrame(
{
"model_id": ["a", "a", "a", "a", "b", "b"],
"dimension": ["Extraversion", "Extraversion", "Neuroticism", "Neuroticism"]
+ ["Extraversion", "Neuroticism"],
"score": [5.0, 3.0, 2.0, float("nan"), 4.0, 1.0],
}
)
print(scale_scores(scored, by="model_id").to_string(index=False))
# model_id dimension score n_items n_missing
# a Extraversion 4.0 2 0
# a Neuroticism 2.0 1 1
# b Extraversion 4.0 1 0
# b Neuroticism 1.0 1 0
totals = scale_scores(scored, by="model_id", agg="sum", include_total=True)
print(totals.to_string(index=False))
# model_id dimension score n_items n_missing
# a Extraversion 8.0 2 0
# a Neuroticism 2.0 1 1
# a total 10.0 3 1
# b Extraversion 4.0 1 0
# b Neuroticism 1.0 1 0
# b total 5.0 2 0
score_experiment
¶
score_experiment(experiment: ExperimentDocument | Questionnaire, processed: DataFrame | None = None, *, decision_column: str = 'decision', item_column: str = 'instruction_item_id', valid_column: str = 'valid', only_valid: bool = True, by: str | Sequence[str] = DEFAULT_GROUPS, agg: str | Callable[[Series], Any] = 'mean', include_total: bool = False) -> pd.DataFrame
Score an experiment: item scores from score_answers, then scale_scores.
Pass the output of the
PostprocessingPipeline as
processed. Without it, the answers stored in the experiment
(experiment.get_answers_as_dataframe()) are used, and the raw answer text is taken as the
decision. That only works if the model answered with exactly the text of an answer option
(and was not asked for anything else, such as a JSON object); real free-text answers need to
go through the pipeline first (score_answers warns if not a single answer is an option).
Calls that failed have no row in the stored answers, so they are not counted in
n_missing.
| PARAMETER | DESCRIPTION |
|---|---|
experiment
|
The experiment (or its questionnaire) with the scoring metadata. Without
TYPE:
|
processed
|
The post-processed answers. Default: the answers of
TYPE:
|
decision_column
|
See
TYPE:
|
item_column
|
See
TYPE:
|
valid_column
|
See
TYPE:
|
only_valid
|
See
TYPE:
|
by
|
See
TYPE:
|
agg
|
See
TYPE:
|
include_total
|
See
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
DataFrame
|
The scale scores, see |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
If |
| WARNS | DESCRIPTION |
|---|---|
UserWarning
|
The warnings of |
Example
import rupsycho as rup
from langchain_core.language_models.fake import FakeListLLM
from rupsycho.scoring import score_experiment
experiment = rup.ExperimentDocument(
name="Mini inventory",
models={},
demographic_profiles={
"Anna": {"attributes": {"name": "Anna", "age": 30}},
"Ben": {"attributes": {"name": "Ben", "age": 45}},
},
questionnaire={
"name": "Mini inventory",
"general_instruction": "Rate the statement.",
"attributes": {"dimension": {"1": "Extraversion"}},
"default_answer_options": {
"1": {"text": "1. Disagree", "weight": 1},
"2": {"text": "2. Neutral", "weight": 2},
"3": {"text": "3. Agree", "weight": 3},
},
"instruction_items": [
{"question": "Is talkative", "attributes": {"dimension": "1"}},
{"question": "Is reserved", "reversed": True, "attributes": {"dimension": "1"}},
],
},
)
experiment.add_model(FakeListLLM(responses=["3. Agree"]), identifier="agreeable-model")
experiment.run(show_progress=False)
print(score_experiment(experiment, by="profile_id").to_string(index=False))
# profile_id dimension score n_items n_missing
# Anna Extraversion 2.0 2 0
# Ben Extraversion 2.0 2 0
A model that agrees with everything lands in the middle of the scale, because the second item is reverse-keyed.