Extraction Module¶
Entity extractors turn documents into AnnotatedDocument objects. All
extractors inherit from BaseEntityExtractor.
BaseEntityExtractor¶
implicit_word_network.extraction._base.BaseEntityExtractor
¶
Abstract base class for entity extractors.
An entity extractor turns raw documents into
AnnotatedDocument objects that
carry sentences, tokens and entity mentions. Subclasses must implement
annotate.
Available implementations:
SpacyEntityExtractor— spaCy NER with a fixed label set.GLiNEREntityExtractor— zero-shot NER with GLiNER (any labels).GazetteerEntityExtractor— dictionary matching, no models needed.
Extractors that only predict character spans should subclass
SpanEntityExtractor instead, which handles segmentation and
alignment.
annotate
abstractmethod
¶
annotate(documents: Iterable[Document | str], *, batch_size: int = 32, show_progress: bool = False) -> Iterator[AnnotatedDocument]
Annotate documents lazily, in corpus order.
| PARAMETER | DESCRIPTION |
|---|---|
documents
|
Documents (or raw strings) to annotate.
TYPE:
|
batch_size
|
Number of documents processed per model call.
TYPE:
|
show_progress
|
Whether to display a progress bar.
TYPE:
|
| YIELDS | DESCRIPTION |
|---|---|
AnnotatedDocument
|
One annotated document per input document. |
annotate_text
¶
annotate_text(text: str, *, doc_id: ID = 0) -> AnnotatedDocument
Annotate a single string.
| PARAMETER | DESCRIPTION |
|---|---|
text
|
Document text.
TYPE:
|
doc_id
|
Identifier of the resulting document.
TYPE:
|
annotate_all
¶
annotate_all(documents: Iterable[Document | str], *, batch_size: int = 32, show_progress: bool = False) -> list[AnnotatedDocument]
Eager variant of annotate returning a list.
SpanEntityExtractor¶
implicit_word_network.extraction._base.SpanEntityExtractor
¶
SpanEntityExtractor(*, segmenter: BaseSegmenter | None = None)
Base class for extractors that predict character spans only.
Sentence splitting and tokenisation are delegated to a
BaseSegmenter; predicted
spans are aligned to the resulting structure with
align_spans. Subclasses must
implement extract_spans.
| PARAMETER | DESCRIPTION |
|---|---|
segmenter
|
Segmenter providing sentences and tokens. Defaults to the
dependency-free
TYPE:
|
Example
from implicit_word_network import SpanEntityExtractor, EntitySpan
class UpperCaseExtractor(SpanEntityExtractor):
def extract_spans(self, texts, sentences):
import re
return [
[EntitySpan(m.start(), m.end(), "TERM") for m in re.finditer(r"\b[A-Z]{2,}\b", t)]
for t in texts
]
doc = UpperCaseExtractor().annotate_text("NASA and ESA cooperate.")
print([m.text for m in doc.mentions]) # ["NASA", "ESA"]
extract_spans
abstractmethod
¶
extract_spans(texts: Sequence[str], sentences: Sequence[Sequence[Sentence]]) -> list[list[EntitySpan]]
Predict entity spans for a batch of texts.
| PARAMETER | DESCRIPTION |
|---|---|
texts
|
Document texts of the batch.
TYPE:
|
sentences
|
Sentence boundaries of every text (parallel to
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
list[list[EntitySpan]]
|
One list of |
list[list[EntitySpan]]
|
per text. |
SpacyEntityExtractor¶
implicit_word_network.extraction.spacy.SpacyEntityExtractor
¶
SpacyEntityExtractor(model: str | Any = DEFAULT_SPACY_MODEL, *, labels: Sequence[str] | None = DEFAULT_SPACY_LABELS, n_process: int = 1, disable: Sequence[str] = (), max_length: int | None = None)
Extract entities with a spaCy NER pipeline.
Runs segmentation, tagging and NER in one batched nlp.pipe pass and
keeps only entities of the requested types. Requires the spacy extra
and an installed pipeline (python -m spacy download en_core_web_sm).
| PARAMETER | DESCRIPTION |
|---|---|
model
|
Installed pipeline name, or an already loaded
TYPE:
|
labels
|
Entity types to keep, or
TYPE:
|
n_process
|
Number of processes for
TYPE:
|
disable
|
Additional pipeline components to disable when loading.
TYPE:
|
max_length
|
Raise spaCy's character limit (default 1,000,000) for long documents.
TYPE:
|
Example
GLiNEREntityExtractor¶
implicit_word_network.extraction.gliner.GLiNEREntityExtractor
¶
GLiNEREntityExtractor(model: str | Any = DEFAULT_GLINER_MODEL, *, labels: Sequence[str] = DEFAULT_GLINER_LABELS, threshold: float = 0.5, segmenter: BaseSegmenter | None = None, device: str = 'cpu', max_chunk_words: int = 200, batch_size: int = 8, flat_ner: bool = True, multi_label: bool = False, label_map: Mapping[str, str] | None = None, load_kwargs: Mapping[str, Any] | None = None)
Zero-shot entity extraction with a GLiNER model.
GLiNER predicts spans for arbitrary natural-language labels, so the
entity types of the network can be chosen freely ("person",
"disease", "ship", ...). Documents are split into chunks of whole
sentences that fit the model's context window, and all chunks of a batch
are scored in one batched call. Requires the gliner extra.
| PARAMETER | DESCRIPTION |
|---|---|
model
|
Hugging Face model id / local path, or an already loaded
TYPE:
|
labels
|
Natural-language entity labels to extract.
TYPE:
|
threshold
|
Minimum confidence for a predicted span.
TYPE:
|
segmenter
|
Segmenter providing sentences and tokens (defaults to
TYPE:
|
device
|
Device passed to
TYPE:
|
max_chunk_words
|
Maximum number of whitespace-delimited words per chunk sent to the model.
TYPE:
|
batch_size
|
Number of chunks per forward pass.
TYPE:
|
flat_ner
|
Disallow overlapping spans.
TYPE:
|
multi_label
|
Allow several labels for the same span.
TYPE:
|
label_map
|
Optional mapping to rename predicted labels
(e.g.
TYPE:
|
load_kwargs
|
Extra keyword arguments for
TYPE:
|
Example
from implicit_word_network import GLiNEREntityExtractor
extractor = GLiNEREntityExtractor(
"gliner-community/gliner_medium-v2.5",
labels=["person", "organization", "award"],
threshold=0.4,
)
doc = extractor.annotate_text("Feynman received the Nobel Prize in 1965.")
print([(m.text, m.label, round(m.score, 2)) for m in doc.mentions])
GazetteerEntityExtractor¶
implicit_word_network.extraction.gazetteer.GazetteerEntityExtractor
¶
GazetteerEntityExtractor(gazetteer: Mapping[str, Iterable[str]], *, segmenter: BaseSegmenter | None = None, case_sensitive: bool = False)
Extract entities by matching a dictionary of known surface forms.
Useful for closed vocabularies, for tests, and as a fully offline baseline. Longer surface forms win over shorter ones at the same position; matches never start or end inside a word.
| PARAMETER | DESCRIPTION |
|---|---|
gazetteer
|
Mapping from entity label to surface forms, e.g.
TYPE:
|
segmenter
|
Segmenter providing sentences and tokens.
TYPE:
|
case_sensitive
|
Whether matching respects letter case.
TYPE:
|
Example
from implicit_word_network import GazetteerEntityExtractor
extractor = GazetteerEntityExtractor(
{"PERSON": ["Feynman"], "ORG": ["Caltech", "Nobel Prize"]}
)
doc = extractor.annotate_text("Feynman worked at Caltech.")
print([(m.text, m.label) for m in doc.mentions])
# [("Feynman", "PERSON"), ("Caltech", "ORG")]