Skip to content

Pipeline Module

ImplicitNetworkPipeline

implicit_word_network.pipeline.ImplicitNetworkPipeline

ImplicitNetworkPipeline(extractor: BaseEntityExtractor | None = None, *, config: NetworkConfig | None = None, clusterer: ContextualEdgeClusterer | None = None, normalize_entity: Callable[[str], str] | None = None, **config_overrides: Any)

End-to-end construction of (contextual) implicit entity networks.

The pipeline composes three pluggable stages:

  1. an entity extractor (BaseEntityExtractor) that annotates raw documents,
  2. the network builder (ImplicitNetwork with a NetworkConfig) that aggregates cooccurrences into weighted edges, and
  3. an optional edge clusterer (ContextualEdgeClusterer) that splits edges by cooccurrence context.
PARAMETER DESCRIPTION
extractor

Entity extractor. Defaults to SpacyEntityExtractor (requires the spacy extra).

TYPE: BaseEntityExtractor | None DEFAULT: None

config

Network construction parameters. Mutually exclusive with keyword overrides such as window=3.

TYPE: NetworkConfig | None DEFAULT: None

clusterer

Optional contextual edge clusterer used by cluster.

TYPE: ContextualEdgeClusterer | None DEFAULT: None

normalize_entity

Custom entity-name normalisation.

TYPE: Callable[[str], str] | None DEFAULT: None

**config_overrides

Fields of NetworkConfig.

TYPE: Any DEFAULT: {}

Example
from implicit_word_network import Corpus, GLiNEREntityExtractor, ImplicitNetworkPipeline

corpus = Corpus.from_txt("news.txt")
pipeline = ImplicitNetworkPipeline(
    GLiNEREntityExtractor(labels=["person", "organization", "location"]),
    window=2,
)
network = pipeline.run(corpus, show_progress=True)
print(network.summary())
for edge in network.edges(top_k=5):
    print(edge.source.text, "—", edge.target.text, round(edge.weight, 2))

extractor property

The entity extractor (a spaCy extractor is created lazily by default).

annotate

annotate(documents: Iterable[Document | str], *, batch_size: int = 32, show_progress: bool = False) -> Iterator[AnnotatedDocument]

Run only the extraction stage.

run

run(documents: Iterable[Document | str], *, batch_size: int = 32, chunk_size: int = 256, show_progress: bool = False, network: ImplicitNetwork | None = None) -> ImplicitNetwork

Extract entities and build (or extend) an implicit network.

PARAMETER DESCRIPTION
documents

A Corpus, or any iterable of documents / strings.

TYPE: Iterable[Document | str]

batch_size

Documents per extractor call.

TYPE: int DEFAULT: 32

chunk_size

Documents per network update (bounds peak memory).

TYPE: int DEFAULT: 256

show_progress

Display a progress bar.

TYPE: bool DEFAULT: False

network

Existing network to extend instead of creating a new one.

TYPE: ImplicitNetwork | None DEFAULT: None

RETURNS DESCRIPTION
ImplicitNetwork

The resulting network.

cluster

cluster(network: ImplicitNetwork, **kwargs: Any) -> dict[tuple[int, int], list[EdgeContextCluster]]

Cluster edge contexts with the configured clusterer (see ContextualEdgeClusterer.cluster_edges).

build_network

implicit_word_network.pipeline.build_network

build_network(documents: Iterable[Document | str], *, extractor: BaseEntityExtractor | None = None, config: NetworkConfig | None = None, show_progress: bool = False, **config_overrides: Any) -> ImplicitNetwork

Convenience wrapper: build an implicit network in one call.

PARAMETER DESCRIPTION
documents

Corpus, documents or raw strings.

TYPE: Iterable[Document | str]

extractor

Entity extractor (defaults to spaCy).

TYPE: BaseEntityExtractor | None DEFAULT: None

config

Network parameters, or pass fields as keyword arguments.

TYPE: NetworkConfig | None DEFAULT: None

show_progress

Display a progress bar.

TYPE: bool DEFAULT: False

**config_overrides

Fields of NetworkConfig.

TYPE: Any DEFAULT: {}

Example
from implicit_word_network import build_network, GazetteerEntityExtractor

network = build_network(
    ["Feynman met Schwinger in Stockholm."],
    extractor=GazetteerEntityExtractor({"PERSON": ["Feynman", "Schwinger"], "LOC": ["Stockholm"]}),
    window=1,
)