Skip to content

Context Module

Contextual implicit entity networks: context embedding and edge clustering.

ContextualEdgeClusterer

implicit_word_network.context.clustering.ContextualEdgeClusterer

ContextualEdgeClusterer(embedder: BaseContextEmbedder | None = None, *, eps: float = 0.25, min_samples: int = 1, metric: Metric = 'cosine', batch_size: int = 256, keep_centroids: bool = True)

Cluster the cooccurrence contexts of entity pairs (CIEN edge splitting).

PARAMETER DESCRIPTION
embedder

Context embedder. Defaults to the dependency-free BagOfWordsEmbedder; use SentenceTransformerEmbedder for neural contexts as in ECCE.

TYPE: BaseContextEmbedder | None DEFAULT: None

eps

DBSCAN neighbourhood radius (0.25 cosine distance in ECCE).

TYPE: float DEFAULT: 0.25

min_samples

DBSCAN core-point threshold (1 = no noise).

TYPE: int DEFAULT: 1

metric

"cosine" or "euclidean".

TYPE: Metric DEFAULT: 'cosine'

batch_size

Number of contexts embedded per model call. Contexts of many edges are pooled into batches of this size.

TYPE: int DEFAULT: 256

keep_centroids

Store the mean embedding of every cluster.

TYPE: bool DEFAULT: True

Example
from implicit_word_network import ContextualEdgeClusterer

clusterer = ContextualEdgeClusterer(eps=0.4)
clusters = clusterer.cluster_edge(network, ("Feynman", "PERSON"), ("Caltech", "ORG"))
for cluster in clusters:
    print(cluster.size, round(cluster.weight, 2), cluster.contexts[0][:60])

cluster_edge

cluster_edge(network: ImplicitNetwork, a: EntityLike, b: EntityLike) -> list[EdgeContextCluster]

Cluster the cooccurrences of one entity pair on demand.

cluster_edges

cluster_edges(network: ImplicitNetwork, pairs: Iterable[tuple[EntityLike, EntityLike]] | None = None, *, min_weight: float = 0.0, top_k: int | None = None, show_progress: bool = False) -> dict[tuple[int, int], list[EdgeContextCluster]]

Cluster many edges, embedding their contexts in pooled batches.

PARAMETER DESCRIPTION
network

Source network (must store sentence texts).

TYPE: ImplicitNetwork

pairs

Entity pairs to cluster. Defaults to all edges passing min_weight / top_k.

TYPE: Iterable[tuple[EntityLike, EntityLike]] | None DEFAULT: None

min_weight

Edge weight threshold used when pairs is None.

TYPE: float DEFAULT: 0.0

top_k

Number of heaviest edges used when pairs is None.

TYPE: int | None DEFAULT: None

show_progress

Display a progress bar.

TYPE: bool DEFAULT: False

RETURNS DESCRIPTION
dict[tuple[int, int], list[EdgeContextCluster]]

Mapping (entity_id_a, entity_id_b) (a < b) to clusters.

EdgeContextCluster

implicit_word_network.context.clustering.EdgeContextCluster dataclass

EdgeContextCluster(label: int, cooccurrences: list[Cooccurrence], contexts: list[str], weight: float, centroid: NDArray[float32] | None = None)

A group of cooccurrences of two entities that share a similar context.

ATTRIBUTE DESCRIPTION
label

Cluster label (-1 for DBSCAN noise).

TYPE: int

cooccurrences

Member cooccurrences.

TYPE: list[Cooccurrence]

contexts

Context texts, parallel to cooccurrences.

TYPE: list[str]

weight

Sum of the members' decayed weights (the cluster's edge weight).

TYPE: float

centroid

Mean embedding of the members (None if not kept).

TYPE: NDArray[float32] | None

size property

size: int

Number of cooccurrences in the cluster.

Embedders

implicit_word_network.context.embedders.BaseContextEmbedder

Abstract base class for text embedders.

Embedders map cooccurrence contexts (the sentences spanned by two entity mentions) to dense vectors so that ContextualEdgeClusterer can group parallel edges by context. Subclasses must implement encode.

Available implementations:

  • BagOfWordsEmbedder — hashed, L2-normalised bag of words; no dependencies, deterministic, good for tests and lexical similarity.
  • SentenceTransformerEmbedder — neural sentence embeddings via sentence-transformers.

encode abstractmethod

encode(texts: Sequence[str], *, batch_size: int = 32) -> NDArray[float32]

Embed texts.

PARAMETER DESCRIPTION
texts

Input texts.

TYPE: Sequence[str]

batch_size

Batch size hint for model-based implementations.

TYPE: int DEFAULT: 32

RETURNS DESCRIPTION
NDArray[float32]

A (len(texts), dim) float32 array.

implicit_word_network.context.embedders.BagOfWordsEmbedder

BagOfWordsEmbedder(n_features: int = 2048, *, stopwords: Collection[str] | None = None)

Hashed bag-of-words embedder (dependency-free baseline).

Tokens are lowercased, stop words removed and hashed into n_features buckets; rows are L2-normalised so that cosine similarity equals the dot product.

PARAMETER DESCRIPTION
n_features

Dimensionality of the hashed space.

TYPE: int DEFAULT: 2048

stopwords

Words to ignore (defaults to English stop words).

TYPE: Collection[str] | None DEFAULT: None

implicit_word_network.context.embedders.SentenceTransformerEmbedder

SentenceTransformerEmbedder(model: str | Any = DEFAULT_SENTENCE_TRANSFORMER, *, device: str | None = None, normalize: bool = True)

Neural sentence embeddings via sentence-transformers.

Requires the embeddings extra. The model is loaded lazily on first use.

PARAMETER DESCRIPTION
model

Model name or path, or a loaded SentenceTransformer.

TYPE: str | Any DEFAULT: DEFAULT_SENTENCE_TRANSFORMER

device

Torch device (None lets the library choose).

TYPE: str | None DEFAULT: None

normalize

L2-normalise embeddings (recommended with cosine distance).

TYPE: bool DEFAULT: True

Example
from implicit_word_network import SentenceTransformerEmbedder

embedder = SentenceTransformerEmbedder("sentence-transformers/multi-qa-distilbert-cos-v1")
vectors = embedder.encode(["Feynman received the Nobel Prize."])
print(vectors.shape)  # (1, 768)

model property

model: Any

The underlying SentenceTransformer (loaded lazily).

Clustering primitives

implicit_word_network.context.clustering.dbscan

dbscan(distances: NDArray[floating], eps: float, min_samples: int = 1) -> NDArray[int64]

DBSCAN on a precomputed distance matrix.

A point is a core point when at least min_samples points (itself included) lie within eps. Clusters are the connected components of core points; non-core points within eps of a core point join that core point's cluster (border points); all others are noise (-1). Cluster labels are numbered 0, 1, ... in order of first appearance.

PARAMETER DESCRIPTION
distances

Square, symmetric distance matrix.

TYPE: NDArray[floating]

eps

Neighbourhood radius.

TYPE: float

min_samples

Minimum neighbourhood size of a core point.

TYPE: int DEFAULT: 1

RETURNS DESCRIPTION
NDArray[int64]

Integer cluster label per point (-1 = noise).

implicit_word_network.context.clustering.cosine_distances

cosine_distances(vectors: NDArray[floating]) -> NDArray[float64]

Pairwise cosine distances 1 - cos(x, y) (zero vectors get distance 1).

implicit_word_network.context.clustering.euclidean_distances

euclidean_distances(vectors: NDArray[floating]) -> NDArray[float64]

Pairwise Euclidean distances.