Skip to content

Parsers

Cleaners

rupsycho.parsers.cleaners

BasicCleaner

A custom parser that processes and cleans text by removing line breaks and non-ASCII Unicode characters.

This parser is designed to work with LangChain and can be integrated into various chains or agents that require cleaned text output.

parse

parse(text: str) -> str

Parses the input text to remove line breaks and non-ASCII Unicode characters.

Parameters

text : str The input string to be cleaned.

Returns

str The cleaned string with line breaks and non-ASCII Unicode characters removed.

Raises

OutputParserException If an error occurs during parsing, an OutputParserException is raised with a descriptive error message.

PromptRemovalCleaner

PromptRemovalCleaner(prompt: str, similarity_threshold: float = 0.7, fast: bool = True)

A custom parser that removes prompt-related text from a completion. It uses the prompt_cleaner function to process and clean the input text.

Initializes the PromptRemovalCleanerParser with the specified prompt and parameters for cleaning the completion.

Parameters

prompt : str The prompt text that should be removed or considered for cleaning from the completion. similarity_threshold : float, optional The similarity threshold for the prompt_cleaner function (default is 0.7). fast : bool, optional A flag to indicate whether the cleaning process should be fast (default is True).

parse

parse(completion: str) -> str

Cleans the input completion text by removing prompt-related content.

Parameters

completion : str The input text (completion) that needs to be cleaned.

Returns

str The cleaned completion text.

Raises

OutputParserException If an error occurs during parsing, an OutputParserException is raised with a descriptive error message.

RegexExtractorCleaner

RegexExtractorCleaner(pattern: str)

A custom parser that extracts values from the input text based on a given regex pattern. If the extraction fails, it simply returns the original input text.

Initializes the RegexExtractorCleaner with the specified regex pattern.

Parameters

pattern : str The regex pattern used to extract values from the input text.

parse

parse(text: str) -> str

Parses the input text to extract values based on the regex pattern.

If the regex match fails, returns the original input text.

Parameters

text : str The input string to be processed.

Returns

str The extracted value based on the regex pattern, or the original input text if extraction fails.

Raises

OutputParserException If an error occurs during parsing, an OutputParserException is raised with a descriptive error message.

Validators

rupsycho.parsers.validators

normalize

normalize(text)

Removes leading noise (whitespaces, tabs, newlines, etc.) from the beginning of the text, maps typographic quotes to ASCII and transforms it to lowercase.

ApologiesValidatorParser

A parser that validates the input text for apologies-related content. Returns the original text along with a validation status.

BeingAiValidatorParser

A parser that validates the input text for being_ai-related content. Returns the original text along with a validation status.

RefusalValidatorParser

A parser that validates the input text for refusal-related content. Returns the original text along with a validation status.

ValidatorParser

A custom parser that combines the results of ApologiesValidatorParser, BeingAiValidatorParser, and RefusalValidatorParser to validate the input text against these checks and return the original text along with a combined validation status.

parse

parse(text: str) -> dict

Parses the input text by running it through all three validators and combining their results.

Parameters

text : str The input string to be validated.

Returns

dict A dictionary containing the original text, a combined validation status, and detailed results from each individual validator.

ModelBasedValidator

ModelBasedValidator(model_name: str = 'ProtectAI/distilroberta-base-rejection-v1', device: str | None = None)

A custom parser that uses a Hugging Face model to classify input text into two categories: 0 for normal output and 1 for rejection detected.

Initializes the parser with a Hugging Face model to classify rejection. See: https://huggingface.co/protectai/distilroberta-base-rejection-v1

Parameters

model_name : str The Hugging Face model name for the sequence classification model. device : str The device to run the model on (e.g., 'cuda' for GPU or 'cpu').

parse

parse(text: str) -> dict

Classifies the input text as normal or rejection detected. Returns a dictionary with the original text and the classification result.

Judges

rupsycho.parsers.judges

MultipleChoiceJudge

MultipleChoiceJudge(possible_answers: list[str], ignore_case: bool = True, **kwargs)

A custom parser that processes a text input and returns the most likely answer from a set of possible multiple-choice answers using the check_multiple_choice_answers function.

This parser is designed to work with LangChain and can be integrated into various chains or agents that require decision-making based on textual analysis.

parse

parse(text: str, possible_answers=None) -> str

Parses the input text to determine the most likely answer from the possible answers.

DemographicsJudge

Custom parser for demographic questionnaires that can either judge statements on age or on gender. Which of the two tasks is required has to be specified by a keyword ('gender' or 'age') in a single answer option for each item in the experiment config file.

ModelBasedAnswerJudge

A custom parser that processes a text input and returns the most likely answer from a set of possible answers using a Hugging Face model. If the entropy of the decision probabilities is greater than a threshold, it returns "inconclusive".

model_post_init

model_post_init(__context)

Set up LLM.

calculate_entropy

calculate_entropy(decision_list)

Calculates entropy from a list of decision probabilities.

predict_answer

predict_answer(answer_option: str, answer: str)

Predicts the probability and label for a given answer option.

predict_for_all_options

predict_for_all_options(answer: str, answer_options=None)

Iterates over all answer options and predicts for each.

parse

parse(text: str, possible_answers=None) -> str

Parses the input text to determine the most likely answer from the possible answers or returns "inconclusive" if the entropy is above the threshold.

Parser utilities

rupsycho.parsers.parser_utils

check_span

check_span(text: str, filter_dictionary: dict[str, list[str]], span: tuple[int, int] | None = None, ratio: float | None = None) -> dict[str, Any]

Checks the occurrence of specified answers in the text within a defined span or ratio of the text length.

PARAMETER DESCRIPTION
text

The input text to search.

TYPE: str

filter_dictionary

A dictionary where the keys are categories and the values are lists of answers to check for in the text.

TYPE: dict[str, list[str]]

span

A tuple defining the start and end positions within the text to search. Defaults to None.

TYPE: tuple[int, int] DEFAULT: None

ratio

A ratio (0.0 to 1.0) of the text length to define the portion of the text to search. Negative ratios start from the end of the text. Defaults to None.

TYPE: float DEFAULT: None

RETURNS DESCRIPTION
dict[str, bool]

A dictionary where keys are the categories and values are booleans indicating if any answers were found in the specified span or ratio of the text.

prompt_cleaner

prompt_cleaner(prompt, completion, fast: bool = True, similarity_threshold: float = 0.5, junk: str | None = None) -> dict[str, Any]

Cleans the model's output by trimming repeated sections of the input prompt.

This function identifies and removes any part of the model's output (`completion`) that closely matches a section of the input `prompt`. The goal is to clean the output by eliminating unnecessary repetitions.

Args:
    prompt (str): The original input prompt.
    completion (str): The output generated by the model.
    fast (bool): If True, completions that cannot reach the threshold are rejected with a cheap length-based bound (their reported score is then that upper bound, not the exact ratio). The decision to strip is exact either way. Defaults to True.
    similarity_threshold (float): The similarity threshold (between 0 and 1) above which a repeated section is considered for removal. Defaults to 0.5.
    junk (str): A string containing characters to be ignored when considering matches (e.g., whitespace, punctuation). Defaults to "

.,;:!?".

Returns:
    dict: A dictionary with two keys:
        - 'completion' (str): The cleaned output with any repeated prompt sections removed.
        - 'similarity_score' (float): The similarity score between the prompt and the completion.
        - 'similarity_threshold' (float): The similarity threshold that was set for the processing.

process_completion

process_completion(text: str, pattern_name: str | None = None, regex_dict_path: str | None = None, user_input_pattern: str | None = None) -> list

Process the text using a specified regex pattern.

PARAMETER DESCRIPTION
text

The input text to be processed.

TYPE: str

pattern_name

The name of the pattern to use from the regex dictionary ("numeric_alpha", "alpha_alpha", "numeric_numeric").

TYPE: str DEFAULT: None

regex_dict_path

The path to a JSON file containing regex patterns.

TYPE: str DEFAULT: None

user_input_pattern

A regex pattern provided by the user.

TYPE: str DEFAULT: None

RETURNS DESCRIPTION
list

A list of matched groups from the text based on the specified pattern.

RAISES DESCRIPTION
ValueError

If none of pattern_name, regex_dict_path, or user_input_pattern are provided.

ValueError

If the specified pattern_name does not exist in the default or provided regex dictionary.

check_multiple_choice_answers

check_multiple_choice_answers(text: str, possible_answers: list[str], ignore_case: bool = True) -> dict

Checks for the presence of possible multiple-choice answers in the given text. Separates the number from the rest of the answer option and removes punctuation.

PARAMETER DESCRIPTION
text

The input text to search within.

TYPE: str

possible_answers

A list of possible answer strings to check for.

TYPE: list[str]

RETURNS DESCRIPTION
dict

A dictionary where keys are the possible answers and values are the count of detections in the text.

number_to_words

number_to_words(n: int) -> str

Spell out an integer in British English, e.g. 42 -> "forty-two".

PARAMETER DESCRIPTION
n

A number from 0 to 999.

TYPE: int

RETURNS DESCRIPTION
str

The number in words ("one hundred and one" for 101).

RAISES DESCRIPTION
ValueError

If n is outside 0-999.

mk_age_keywords

mk_age_keywords(max_age: int = 100) -> list[list]

Return a list that contains (in order) for each number in the specified range a list of all common expectable representations of that number as stings.

For expample for the range 0-44: [['0', 'nil', 'nought', 'oh', 'zero'] ... ['44', 'forty four', 'forty-four', 'fortyfour']].

check_gender

check_gender(text: str, ignore_case: bool = True) -> str

Judge what gender a model claims to have in its answer.

Gender can be either 'male', 'female', or 'other'. The decision is made by counting the occurences of related words for each of the three options. The option that has the highest score wins.

check_age

check_age(text: str, max_age: int, ignore_case: bool = True) -> str

Judge what age a model claims to have in its answer.

RETURNS DESCRIPTION
str

string of the decided age in numbers (if judging was successful)

The decision is made based on the first occurence of a number (either written or in numbers) in the sentence. All possible subsequent numbers are ignored.

split_on_symbols

split_on_symbols(text: str) -> list[str]

Splits the input text based on numbers, special characters, and symbols.

PARAMETER DESCRIPTION
text

The input string to split.

TYPE: str

RETURNS DESCRIPTION
list[str]

A list of non-empty components split by numbers, special characters, and symbols.

json_saver

json_saver(data: dict, name: str = 'output', path: str = '') -> None

Save a dictionary as a JSON file at the specified path and file name. If no path is provided, saves to the current directory. Prints the full path upon successful save.

PARAMETER DESCRIPTION
data

Dictionary to save as JSON.

TYPE: dict

name

Name of the file to save. Defaults to 'output'.

TYPE: str DEFAULT: 'output'

path

Directory path where the file will be saved. Defaults to the current directory.

TYPE: str DEFAULT: ''

RETURNS DESCRIPTION
None

None

Legacy parser

BasicParser is the former name of BasicCleaner. It is also importable from the backwards-compatible module rupsycho.parser.

rupsycho.parsers.parser

BasicParser: the original name of :class:~rupsycho.parsers.cleaners.BasicCleaner.

BasicParser

Removes line breaks, unusual white space and non-ASCII characters from a text.

Identical to BasicCleaner; the name is kept for backwards compatibility.

Example
from rupsycho.parsers.parser import BasicParser

BasicParser().invoke("Hello,\nworld 😊")  # 'Hello, world'