Parsers¶
Cleaners¶
rupsycho.parsers.cleaners
¶
BasicCleaner
¶
A custom parser that processes and cleans text by removing line breaks and non-ASCII Unicode characters.
This parser is designed to work with LangChain and can be integrated into various chains or agents that require cleaned text output.
parse
¶
Parses the input text to remove line breaks and non-ASCII Unicode characters.
Parameters¶
text : str The input string to be cleaned.
Returns¶
str The cleaned string with line breaks and non-ASCII Unicode characters removed.
Raises¶
OutputParserException If an error occurs during parsing, an OutputParserException is raised with a descriptive error message.
PromptRemovalCleaner
¶
A custom parser that removes prompt-related text from a completion. It uses the prompt_cleaner function to process and clean the input text.
Initializes the PromptRemovalCleanerParser with the specified prompt and parameters for cleaning the completion.
Parameters¶
prompt : str The prompt text that should be removed or considered for cleaning from the completion. similarity_threshold : float, optional The similarity threshold for the prompt_cleaner function (default is 0.7). fast : bool, optional A flag to indicate whether the cleaning process should be fast (default is True).
parse
¶
Cleans the input completion text by removing prompt-related content.
Parameters¶
completion : str The input text (completion) that needs to be cleaned.
Returns¶
str The cleaned completion text.
Raises¶
OutputParserException If an error occurs during parsing, an OutputParserException is raised with a descriptive error message.
RegexExtractorCleaner
¶
A custom parser that extracts values from the input text based on a given regex pattern. If the extraction fails, it simply returns the original input text.
Initializes the RegexExtractorCleaner with the specified regex pattern.
Parameters¶
pattern : str The regex pattern used to extract values from the input text.
parse
¶
Parses the input text to extract values based on the regex pattern.
If the regex match fails, returns the original input text.
Parameters¶
text : str The input string to be processed.
Returns¶
str The extracted value based on the regex pattern, or the original input text if extraction fails.
Raises¶
OutputParserException If an error occurs during parsing, an OutputParserException is raised with a descriptive error message.
Validators¶
rupsycho.parsers.validators
¶
normalize
¶
Removes leading noise (whitespaces, tabs, newlines, etc.) from the beginning of the text, maps typographic quotes to ASCII and transforms it to lowercase.
ApologiesValidatorParser
¶
A parser that validates the input text for apologies-related content. Returns the original text along with a validation status.
BeingAiValidatorParser
¶
A parser that validates the input text for being_ai-related content. Returns the original text along with a validation status.
RefusalValidatorParser
¶
A parser that validates the input text for refusal-related content. Returns the original text along with a validation status.
ValidatorParser
¶
A custom parser that combines the results of ApologiesValidatorParser, BeingAiValidatorParser, and RefusalValidatorParser to validate the input text against these checks and return the original text along with a combined validation status.
parse
¶
ModelBasedValidator
¶
ModelBasedValidator(model_name: str = 'ProtectAI/distilroberta-base-rejection-v1', device: str | None = None)
A custom parser that uses a Hugging Face model to classify input text into two categories: 0 for normal output and 1 for rejection detected.
Initializes the parser with a Hugging Face model to classify rejection. See: https://huggingface.co/protectai/distilroberta-base-rejection-v1
Parameters¶
model_name : str The Hugging Face model name for the sequence classification model. device : str The device to run the model on (e.g., 'cuda' for GPU or 'cpu').
parse
¶
Classifies the input text as normal or rejection detected. Returns a dictionary with the original text and the classification result.
Judges¶
rupsycho.parsers.judges
¶
MultipleChoiceJudge
¶
A custom parser that processes a text input and returns the most likely answer from a set of possible multiple-choice answers using the check_multiple_choice_answers function.
This parser is designed to work with LangChain and can be integrated into various chains or agents that require decision-making based on textual analysis.
parse
¶
Parses the input text to determine the most likely answer from the possible answers.
DemographicsJudge
¶
Custom parser for demographic questionnaires that can either judge statements on age or on gender. Which of the two tasks is required has to be specified by a keyword ('gender' or 'age') in a single answer option for each item in the experiment config file.
ModelBasedAnswerJudge
¶
A custom parser that processes a text input and returns the most likely answer from a set of possible answers using a Hugging Face model. If the entropy of the decision probabilities is greater than a threshold, it returns "inconclusive".
calculate_entropy
¶
Calculates entropy from a list of decision probabilities.
predict_answer
¶
Predicts the probability and label for a given answer option.
predict_for_all_options
¶
Iterates over all answer options and predicts for each.
parse
¶
Parses the input text to determine the most likely answer from the possible answers or returns "inconclusive" if the entropy is above the threshold.
Parser utilities¶
rupsycho.parsers.parser_utils
¶
check_span
¶
check_span(text: str, filter_dictionary: dict[str, list[str]], span: tuple[int, int] | None = None, ratio: float | None = None) -> dict[str, Any]
Checks the occurrence of specified answers in the text within a defined span or ratio of the text length.
| PARAMETER | DESCRIPTION |
|---|---|
text
|
The input text to search.
TYPE:
|
filter_dictionary
|
A dictionary where the keys are categories and the values are lists of answers to check for in the text.
TYPE:
|
span
|
A tuple defining the start and end positions within the text to search. Defaults to None.
TYPE:
|
ratio
|
A ratio (0.0 to 1.0) of the text length to define the portion of the text to search. Negative ratios start from the end of the text. Defaults to None.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
dict[str, bool]
|
A dictionary where keys are the categories and values are booleans indicating if any answers were found in the specified span or ratio of the text. |
prompt_cleaner
¶
prompt_cleaner(prompt, completion, fast: bool = True, similarity_threshold: float = 0.5, junk: str | None = None) -> dict[str, Any]
Cleans the model's output by trimming repeated sections of the input prompt.
This function identifies and removes any part of the model's output (`completion`) that closely matches a section of the input `prompt`. The goal is to clean the output by eliminating unnecessary repetitions.
Args:
prompt (str): The original input prompt.
completion (str): The output generated by the model.
fast (bool): If True, completions that cannot reach the threshold are rejected with a cheap length-based bound (their reported score is then that upper bound, not the exact ratio). The decision to strip is exact either way. Defaults to True.
similarity_threshold (float): The similarity threshold (between 0 and 1) above which a repeated section is considered for removal. Defaults to 0.5.
junk (str): A string containing characters to be ignored when considering matches (e.g., whitespace, punctuation). Defaults to "
.,;:!?".
Returns:
dict: A dictionary with two keys:
- 'completion' (str): The cleaned output with any repeated prompt sections removed.
- 'similarity_score' (float): The similarity score between the prompt and the completion.
- 'similarity_threshold' (float): The similarity threshold that was set for the processing.
process_completion
¶
process_completion(text: str, pattern_name: str | None = None, regex_dict_path: str | None = None, user_input_pattern: str | None = None) -> list
Process the text using a specified regex pattern.
| PARAMETER | DESCRIPTION |
|---|---|
text
|
The input text to be processed.
TYPE:
|
pattern_name
|
The name of the pattern to use from the regex dictionary ("numeric_alpha", "alpha_alpha", "numeric_numeric").
TYPE:
|
regex_dict_path
|
The path to a JSON file containing regex patterns.
TYPE:
|
user_input_pattern
|
A regex pattern provided by the user.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
list
|
A list of matched groups from the text based on the specified pattern. |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
If none of pattern_name, regex_dict_path, or user_input_pattern are provided. |
ValueError
|
If the specified pattern_name does not exist in the default or provided regex dictionary. |
check_multiple_choice_answers
¶
check_multiple_choice_answers(text: str, possible_answers: list[str], ignore_case: bool = True) -> dict
Checks for the presence of possible multiple-choice answers in the given text. Separates the number from the rest of the answer option and removes punctuation.
| PARAMETER | DESCRIPTION |
|---|---|
text
|
The input text to search within.
TYPE:
|
possible_answers
|
A list of possible answer strings to check for.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
dict
|
A dictionary where keys are the possible answers and values are the count of detections in the text. |
number_to_words
¶
Spell out an integer in British English, e.g. 42 -> "forty-two".
| PARAMETER | DESCRIPTION |
|---|---|
n
|
A number from 0 to 999.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
str
|
The number in words ( |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
If |
mk_age_keywords
¶
Return a list that contains (in order) for each number in the specified range a list of all common expectable representations of that number as stings.
For expample for the range 0-44: [['0', 'nil', 'nought', 'oh', 'zero'] ... ['44', 'forty four', 'forty-four', 'fortyfour']].
check_gender
¶
Judge what gender a model claims to have in its answer.
Gender can be either 'male', 'female', or 'other'. The decision is made by counting the occurences of related words for each of the three options. The option that has the highest score wins.
check_age
¶
Judge what age a model claims to have in its answer.
| RETURNS | DESCRIPTION |
|---|---|
str
|
string of the decided age in numbers (if judging was successful) |
The decision is made based on the first occurence of a number (either written or in numbers) in the sentence. All possible subsequent numbers are ignored.
split_on_symbols
¶
Splits the input text based on numbers, special characters, and symbols.
| PARAMETER | DESCRIPTION |
|---|---|
text
|
The input string to split.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
list[str]
|
A list of non-empty components split by numbers, special characters, and symbols. |
json_saver
¶
Save a dictionary as a JSON file at the specified path and file name. If no path is provided, saves to the current directory. Prints the full path upon successful save.
| PARAMETER | DESCRIPTION |
|---|---|
data
|
Dictionary to save as JSON.
TYPE:
|
name
|
Name of the file to save. Defaults to 'output'.
TYPE:
|
path
|
Directory path where the file will be saved. Defaults to the current directory.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
None
|
None |
Legacy parser¶
BasicParser is the former name of BasicCleaner. It is also importable from the
backwards-compatible module rupsycho.parser.
rupsycho.parsers.parser
¶
BasicParser: the original name of :class:~rupsycho.parsers.cleaners.BasicCleaner.
BasicParser
¶
Removes line breaks, unusual white space and non-ASCII characters from a text.
Identical to BasicCleaner; the name is kept for
backwards compatibility.