Segmentation Module¶
Sentence splitting and tokenisation for span-based extractors.
BaseSegmenter¶
implicit_word_network.segmentation._base.BaseSegmenter
¶
Abstract base class for sentence splitting and tokenisation.
Segmenters provide the sentence and token structure that span-based
entity extractors (GLiNER, gazetteers, ...) lack. Subclasses must
implement segment; segment_many may be overridden for
batched processing.
Available implementations:
RegexSegmenter— dependency-free regular-expression segmenter.SpacySegmenter— spaCy pipeline (sentences, POS tags, lemmas, stop words).
segment
abstractmethod
¶
Split text into sentences and tokens.
| PARAMETER | DESCRIPTION |
|---|---|
text
|
Document text.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
Segmentation
|
A |
Segmentation
|
character offset. Every token belongs to exactly one sentence. |
segment_many
¶
Segment several texts, yielding one Segmentation per text.
| PARAMETER | DESCRIPTION |
|---|---|
texts
|
Document texts.
TYPE:
|
batch_size
|
Hint for implementations that process texts in batches.
TYPE:
|
RegexSegmenter¶
implicit_word_network.segmentation.regex.RegexSegmenter
¶
RegexSegmenter(*, stopwords: Collection[str] | None = None, sentence_boundary: str | Pattern[str] | None = None, token_pattern: str | Pattern[str] | None = None)
Sentence splitter and tokeniser based on regular expressions.
This segmenter has no external dependencies and is fast, but it knows
nothing about abbreviations, part-of-speech tags or lemmas. Use
SpacySegmenter when
linguistic quality matters.
| PARAMETER | DESCRIPTION |
|---|---|
stopwords
|
Stop word list (lowercase). Defaults to a compact English list; pass an empty collection to disable stop word flagging.
TYPE:
|
sentence_boundary
|
Regular expression matching sentence boundaries. When the first group matches closing quotes/brackets they are kept with the preceding sentence.
TYPE:
|
token_pattern
|
Regular expression matching one token.
TYPE:
|
Example
SpacySegmenter¶
implicit_word_network.segmentation.spacy.SpacySegmenter
¶
SpacySegmenter(model: str | Any = DEFAULT_SPACY_MODEL, *, disable: Sequence[str] = ('ner',), n_process: int = 1, max_length: int | None = None)
Sentence splitter and tokeniser backed by a spaCy pipeline.
Provides sentences, part-of-speech tags, lemmas and stop word flags, which
enable lemma-based term nodes and POS filtering in the network builder.
Requires the spacy extra and an installed pipeline.
| PARAMETER | DESCRIPTION |
|---|---|
model
|
Installed spaCy pipeline name, or an already loaded
TYPE:
|
disable
|
Components to disable when loading by name. NER is disabled by default because this class only segments.
TYPE:
|
n_process
|
Number of processes for
TYPE:
|
max_length
|
Raise spaCy's character limit for long documents.
TYPE:
|
Example
Helpers¶
implicit_word_network.segmentation.spacy.load_spacy_model
¶
Load a spaCy pipeline, making sure it can split sentences.
| PARAMETER | DESCRIPTION |
|---|---|
model
|
Name of an installed spaCy pipeline (e.g.
TYPE:
|
disable
|
Pipeline components to disable.
TYPE:
|
max_length
|
Raise spaCy's
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
Any
|
The loaded |
| RAISES | DESCRIPTION |
|---|---|
ImportError
|
If spaCy is not installed. |
OSError
|
If the pipeline is not installed (with download hint). |
implicit_word_network.segmentation.spacy.spacy_doc_to_segmentation
¶
Convert a spacy.tokens.Doc into a Segmentation.
Token positions are preserved, so spaCy token indices (e.g. ent.start)
index the returned token list directly.