Skip to content

Segmentation Module

Sentence splitting and tokenisation for span-based extractors.

BaseSegmenter

implicit_word_network.segmentation._base.BaseSegmenter

Abstract base class for sentence splitting and tokenisation.

Segmenters provide the sentence and token structure that span-based entity extractors (GLiNER, gazetteers, ...) lack. Subclasses must implement segment; segment_many may be overridden for batched processing.

Available implementations:

  • RegexSegmenter — dependency-free regular-expression segmenter.
  • SpacySegmenter — spaCy pipeline (sentences, POS tags, lemmas, stop words).

segment abstractmethod

segment(text: str) -> Segmentation

Split text into sentences and tokens.

PARAMETER DESCRIPTION
text

Document text.

TYPE: str

RETURNS DESCRIPTION
Segmentation

A Segmentation with sentences and tokens sorted by

Segmentation

character offset. Every token belongs to exactly one sentence.

segment_many

segment_many(texts: Iterable[str], *, batch_size: int = 64) -> Iterator[Segmentation]

Segment several texts, yielding one Segmentation per text.

PARAMETER DESCRIPTION
texts

Document texts.

TYPE: Iterable[str]

batch_size

Hint for implementations that process texts in batches.

TYPE: int DEFAULT: 64

RegexSegmenter

implicit_word_network.segmentation.regex.RegexSegmenter

RegexSegmenter(*, stopwords: Collection[str] | None = None, sentence_boundary: str | Pattern[str] | None = None, token_pattern: str | Pattern[str] | None = None)

Sentence splitter and tokeniser based on regular expressions.

This segmenter has no external dependencies and is fast, but it knows nothing about abbreviations, part-of-speech tags or lemmas. Use SpacySegmenter when linguistic quality matters.

PARAMETER DESCRIPTION
stopwords

Stop word list (lowercase). Defaults to a compact English list; pass an empty collection to disable stop word flagging.

TYPE: Collection[str] | None DEFAULT: None

sentence_boundary

Regular expression matching sentence boundaries. When the first group matches closing quotes/brackets they are kept with the preceding sentence.

TYPE: str | Pattern[str] | None DEFAULT: None

token_pattern

Regular expression matching one token.

TYPE: str | Pattern[str] | None DEFAULT: None

Example
from implicit_word_network import RegexSegmenter

segmenter = RegexSegmenter()
sentences, tokens = segmenter.segment("Hello world. Second sentence!")
print(len(sentences))  # 2
print(tokens[0].text)  # "Hello"

SpacySegmenter

implicit_word_network.segmentation.spacy.SpacySegmenter

SpacySegmenter(model: str | Any = DEFAULT_SPACY_MODEL, *, disable: Sequence[str] = ('ner',), n_process: int = 1, max_length: int | None = None)

Sentence splitter and tokeniser backed by a spaCy pipeline.

Provides sentences, part-of-speech tags, lemmas and stop word flags, which enable lemma-based term nodes and POS filtering in the network builder. Requires the spacy extra and an installed pipeline.

PARAMETER DESCRIPTION
model

Installed spaCy pipeline name, or an already loaded Language object.

TYPE: str | Any DEFAULT: DEFAULT_SPACY_MODEL

disable

Components to disable when loading by name. NER is disabled by default because this class only segments.

TYPE: Sequence[str] DEFAULT: ('ner',)

n_process

Number of processes for nlp.pipe.

TYPE: int DEFAULT: 1

max_length

Raise spaCy's character limit for long documents.

TYPE: int | None DEFAULT: None

Example
from implicit_word_network import SpacySegmenter

segmenter = SpacySegmenter("en_core_web_sm")
sentences, tokens = segmenter.segment("Feynman taught at Caltech.")
print(tokens[0].pos, tokens[0].lemma)  # "PROPN" "Feynman"

nlp property

nlp: Any

The underlying spaCy pipeline (loaded lazily).

Helpers

implicit_word_network.segmentation.spacy.load_spacy_model

load_spacy_model(model: str, *, disable: Sequence[str] = (), max_length: int | None = None) -> Any

Load a spaCy pipeline, making sure it can split sentences.

PARAMETER DESCRIPTION
model

Name of an installed spaCy pipeline (e.g. "en_core_web_sm").

TYPE: str

disable

Pipeline components to disable.

TYPE: Sequence[str] DEFAULT: ()

max_length

Raise spaCy's nlp.max_length (default 1,000,000 characters) to process longer documents.

TYPE: int | None DEFAULT: None

RETURNS DESCRIPTION
Any

The loaded spacy.language.Language object.

RAISES DESCRIPTION
ImportError

If spaCy is not installed.

OSError

If the pipeline is not installed (with download hint).

implicit_word_network.segmentation.spacy.spacy_doc_to_segmentation

spacy_doc_to_segmentation(doc: Any) -> Segmentation

Convert a spacy.tokens.Doc into a Segmentation.

Token positions are preserved, so spaCy token indices (e.g. ent.start) index the returned token list directly.