Skip to content

Extraction Module

Entity extractors turn documents into AnnotatedDocument objects. All extractors inherit from BaseEntityExtractor.

BaseEntityExtractor

implicit_word_network.extraction._base.BaseEntityExtractor

Abstract base class for entity extractors.

An entity extractor turns raw documents into AnnotatedDocument objects that carry sentences, tokens and entity mentions. Subclasses must implement annotate.

Available implementations:

  • SpacyEntityExtractor — spaCy NER with a fixed label set.
  • GLiNEREntityExtractor — zero-shot NER with GLiNER (any labels).
  • GazetteerEntityExtractor — dictionary matching, no models needed.

Extractors that only predict character spans should subclass SpanEntityExtractor instead, which handles segmentation and alignment.

annotate abstractmethod

annotate(documents: Iterable[Document | str], *, batch_size: int = 32, show_progress: bool = False) -> Iterator[AnnotatedDocument]

Annotate documents lazily, in corpus order.

PARAMETER DESCRIPTION
documents

Documents (or raw strings) to annotate.

TYPE: Iterable[Document | str]

batch_size

Number of documents processed per model call.

TYPE: int DEFAULT: 32

show_progress

Whether to display a progress bar.

TYPE: bool DEFAULT: False

YIELDS DESCRIPTION
AnnotatedDocument

One annotated document per input document.

annotate_text

annotate_text(text: str, *, doc_id: ID = 0) -> AnnotatedDocument

Annotate a single string.

PARAMETER DESCRIPTION
text

Document text.

TYPE: str

doc_id

Identifier of the resulting document.

TYPE: ID DEFAULT: 0

annotate_all

annotate_all(documents: Iterable[Document | str], *, batch_size: int = 32, show_progress: bool = False) -> list[AnnotatedDocument]

Eager variant of annotate returning a list.

SpanEntityExtractor

implicit_word_network.extraction._base.SpanEntityExtractor

SpanEntityExtractor(*, segmenter: BaseSegmenter | None = None)

Base class for extractors that predict character spans only.

Sentence splitting and tokenisation are delegated to a BaseSegmenter; predicted spans are aligned to the resulting structure with align_spans. Subclasses must implement extract_spans.

PARAMETER DESCRIPTION
segmenter

Segmenter providing sentences and tokens. Defaults to the dependency-free RegexSegmenter.

TYPE: BaseSegmenter | None DEFAULT: None

Example
from implicit_word_network import SpanEntityExtractor, EntitySpan

class UpperCaseExtractor(SpanEntityExtractor):
    def extract_spans(self, texts, sentences):
        import re
        return [
            [EntitySpan(m.start(), m.end(), "TERM") for m in re.finditer(r"\b[A-Z]{2,}\b", t)]
            for t in texts
        ]

doc = UpperCaseExtractor().annotate_text("NASA and ESA cooperate.")
print([m.text for m in doc.mentions])  # ["NASA", "ESA"]

extract_spans abstractmethod

extract_spans(texts: Sequence[str], sentences: Sequence[Sequence[Sentence]]) -> list[list[EntitySpan]]

Predict entity spans for a batch of texts.

PARAMETER DESCRIPTION
texts

Document texts of the batch.

TYPE: Sequence[str]

sentences

Sentence boundaries of every text (parallel to texts), useful for chunking long inputs.

TYPE: Sequence[Sequence[Sentence]]

RETURNS DESCRIPTION
list[list[EntitySpan]]

One list of EntitySpan

list[list[EntitySpan]]

per text.

SpacyEntityExtractor

implicit_word_network.extraction.spacy.SpacyEntityExtractor

SpacyEntityExtractor(model: str | Any = DEFAULT_SPACY_MODEL, *, labels: Sequence[str] | None = DEFAULT_SPACY_LABELS, n_process: int = 1, disable: Sequence[str] = (), max_length: int | None = None)

Extract entities with a spaCy NER pipeline.

Runs segmentation, tagging and NER in one batched nlp.pipe pass and keeps only entities of the requested types. Requires the spacy extra and an installed pipeline (python -m spacy download en_core_web_sm).

PARAMETER DESCRIPTION
model

Installed pipeline name, or an already loaded Language.

TYPE: str | Any DEFAULT: DEFAULT_SPACY_MODEL

labels

Entity types to keep, or None to keep every type.

TYPE: Sequence[str] | None DEFAULT: DEFAULT_SPACY_LABELS

n_process

Number of processes for nlp.pipe.

TYPE: int DEFAULT: 1

disable

Additional pipeline components to disable when loading.

TYPE: Sequence[str] DEFAULT: ()

max_length

Raise spaCy's character limit (default 1,000,000) for long documents.

TYPE: int | None DEFAULT: None

Example
from implicit_word_network import SpacyEntityExtractor

extractor = SpacyEntityExtractor("en_core_web_sm", labels=["PERSON", "ORG"])
doc = extractor.annotate_text("Feynman received the Nobel Prize at Caltech.")
print([(m.text, m.label) for m in doc.mentions])

nlp property

nlp: Any

The underlying spaCy pipeline (loaded lazily).

GLiNEREntityExtractor

implicit_word_network.extraction.gliner.GLiNEREntityExtractor

GLiNEREntityExtractor(model: str | Any = DEFAULT_GLINER_MODEL, *, labels: Sequence[str] = DEFAULT_GLINER_LABELS, threshold: float = 0.5, segmenter: BaseSegmenter | None = None, device: str = 'cpu', max_chunk_words: int = 200, batch_size: int = 8, flat_ner: bool = True, multi_label: bool = False, label_map: Mapping[str, str] | None = None, load_kwargs: Mapping[str, Any] | None = None)

Zero-shot entity extraction with a GLiNER model.

GLiNER predicts spans for arbitrary natural-language labels, so the entity types of the network can be chosen freely ("person", "disease", "ship", ...). Documents are split into chunks of whole sentences that fit the model's context window, and all chunks of a batch are scored in one batched call. Requires the gliner extra.

PARAMETER DESCRIPTION
model

Hugging Face model id / local path, or an already loaded gliner.GLiNER instance. Defaults to gliner-community/gliner_medium-v2.5.

TYPE: str | Any DEFAULT: DEFAULT_GLINER_MODEL

labels

Natural-language entity labels to extract.

TYPE: Sequence[str] DEFAULT: DEFAULT_GLINER_LABELS

threshold

Minimum confidence for a predicted span.

TYPE: float DEFAULT: 0.5

segmenter

Segmenter providing sentences and tokens (defaults to RegexSegmenter; use SpacySegmenter for lemmas and POS tags).

TYPE: BaseSegmenter | None DEFAULT: None

device

Device passed to GLiNER.from_pretrained ("cpu", "cuda", "mps").

TYPE: str DEFAULT: 'cpu'

max_chunk_words

Maximum number of whitespace-delimited words per chunk sent to the model.

TYPE: int DEFAULT: 200

batch_size

Number of chunks per forward pass.

TYPE: int DEFAULT: 8

flat_ner

Disallow overlapping spans.

TYPE: bool DEFAULT: True

multi_label

Allow several labels for the same span.

TYPE: bool DEFAULT: False

label_map

Optional mapping to rename predicted labels (e.g. {"person": "PERSON"}).

TYPE: Mapping[str, str] | None DEFAULT: None

load_kwargs

Extra keyword arguments for GLiNER.from_pretrained.

TYPE: Mapping[str, Any] | None DEFAULT: None

Example
from implicit_word_network import GLiNEREntityExtractor

extractor = GLiNEREntityExtractor(
    "gliner-community/gliner_medium-v2.5",
    labels=["person", "organization", "award"],
    threshold=0.4,
)
doc = extractor.annotate_text("Feynman received the Nobel Prize in 1965.")
print([(m.text, m.label, round(m.score, 2)) for m in doc.mentions])

model property

model: Any

The underlying GLiNER model (loaded lazily on first use).

GazetteerEntityExtractor

implicit_word_network.extraction.gazetteer.GazetteerEntityExtractor

GazetteerEntityExtractor(gazetteer: Mapping[str, Iterable[str]], *, segmenter: BaseSegmenter | None = None, case_sensitive: bool = False)

Extract entities by matching a dictionary of known surface forms.

Useful for closed vocabularies, for tests, and as a fully offline baseline. Longer surface forms win over shorter ones at the same position; matches never start or end inside a word.

PARAMETER DESCRIPTION
gazetteer

Mapping from entity label to surface forms, e.g. {"PERSON": ["Richard Feynman", "Feynman"], "ORG": ["Caltech"]}.

TYPE: Mapping[str, Iterable[str]]

segmenter

Segmenter providing sentences and tokens.

TYPE: BaseSegmenter | None DEFAULT: None

case_sensitive

Whether matching respects letter case.

TYPE: bool DEFAULT: False

Example
from implicit_word_network import GazetteerEntityExtractor

extractor = GazetteerEntityExtractor(
    {"PERSON": ["Feynman"], "ORG": ["Caltech", "Nobel Prize"]}
)
doc = extractor.annotate_text("Feynman worked at Caltech.")
print([(m.text, m.label) for m in doc.mentions])
# [("Feynman", "PERSON"), ("Caltech", "ORG")]

labels property

labels: frozenset[str]

Entity labels known to this extractor.