Context Module¶
Contextual implicit entity networks: context embedding and edge clustering.
ContextualEdgeClusterer¶
implicit_word_network.context.clustering.ContextualEdgeClusterer
¶
ContextualEdgeClusterer(embedder: BaseContextEmbedder | None = None, *, eps: float = 0.25, min_samples: int = 1, metric: Metric = 'cosine', batch_size: int = 256, keep_centroids: bool = True)
Cluster the cooccurrence contexts of entity pairs (CIEN edge splitting).
| PARAMETER | DESCRIPTION |
|---|---|
embedder
|
Context embedder. Defaults to the dependency-free
TYPE:
|
eps
|
DBSCAN neighbourhood radius (
TYPE:
|
min_samples
|
DBSCAN core-point threshold (
TYPE:
|
metric
|
TYPE:
|
batch_size
|
Number of contexts embedded per model call. Contexts of many edges are pooled into batches of this size.
TYPE:
|
keep_centroids
|
Store the mean embedding of every cluster.
TYPE:
|
Example
cluster_edge
¶
cluster_edge(network: ImplicitNetwork, a: EntityLike, b: EntityLike) -> list[EdgeContextCluster]
Cluster the cooccurrences of one entity pair on demand.
cluster_edges
¶
cluster_edges(network: ImplicitNetwork, pairs: Iterable[tuple[EntityLike, EntityLike]] | None = None, *, min_weight: float = 0.0, top_k: int | None = None, show_progress: bool = False) -> dict[tuple[int, int], list[EdgeContextCluster]]
Cluster many edges, embedding their contexts in pooled batches.
| PARAMETER | DESCRIPTION |
|---|---|
network
|
Source network (must store sentence texts).
TYPE:
|
pairs
|
Entity pairs to cluster. Defaults to all edges passing
TYPE:
|
min_weight
|
Edge weight threshold used when
TYPE:
|
top_k
|
Number of heaviest edges used when
TYPE:
|
show_progress
|
Display a progress bar.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
dict[tuple[int, int], list[EdgeContextCluster]]
|
Mapping |
EdgeContextCluster¶
implicit_word_network.context.clustering.EdgeContextCluster
dataclass
¶
EdgeContextCluster(label: int, cooccurrences: list[Cooccurrence], contexts: list[str], weight: float, centroid: NDArray[float32] | None = None)
A group of cooccurrences of two entities that share a similar context.
| ATTRIBUTE | DESCRIPTION |
|---|---|
label |
Cluster label (
TYPE:
|
cooccurrences |
Member cooccurrences.
TYPE:
|
contexts |
Context texts, parallel to
TYPE:
|
weight |
Sum of the members' decayed weights (the cluster's edge weight).
TYPE:
|
centroid |
Mean embedding of the members (
TYPE:
|
Embedders¶
implicit_word_network.context.embedders.BaseContextEmbedder
¶
Abstract base class for text embedders.
Embedders map cooccurrence contexts (the sentences spanned by two entity
mentions) to dense vectors so that
ContextualEdgeClusterer can group
parallel edges by context. Subclasses must implement encode.
Available implementations:
BagOfWordsEmbedder— hashed, L2-normalised bag of words; no dependencies, deterministic, good for tests and lexical similarity.SentenceTransformerEmbedder— neural sentence embeddings viasentence-transformers.
encode
abstractmethod
¶
Embed texts.
| PARAMETER | DESCRIPTION |
|---|---|
texts
|
Input texts.
TYPE:
|
batch_size
|
Batch size hint for model-based implementations.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
NDArray[float32]
|
A |
implicit_word_network.context.embedders.BagOfWordsEmbedder
¶
Hashed bag-of-words embedder (dependency-free baseline).
Tokens are lowercased, stop words removed and hashed into
n_features buckets; rows are L2-normalised so that cosine similarity
equals the dot product.
| PARAMETER | DESCRIPTION |
|---|---|
n_features
|
Dimensionality of the hashed space.
TYPE:
|
stopwords
|
Words to ignore (defaults to English stop words).
TYPE:
|
implicit_word_network.context.embedders.SentenceTransformerEmbedder
¶
SentenceTransformerEmbedder(model: str | Any = DEFAULT_SENTENCE_TRANSFORMER, *, device: str | None = None, normalize: bool = True)
Neural sentence embeddings via sentence-transformers.
Requires the embeddings extra. The model is loaded lazily on first use.
| PARAMETER | DESCRIPTION |
|---|---|
model
|
Model name or path, or a loaded
TYPE:
|
device
|
Torch device (
TYPE:
|
normalize
|
L2-normalise embeddings (recommended with cosine distance).
TYPE:
|
Example
Clustering primitives¶
implicit_word_network.context.clustering.dbscan
¶
DBSCAN on a precomputed distance matrix.
A point is a core point when at least min_samples points (itself
included) lie within eps. Clusters are the connected components of
core points; non-core points within eps of a core point join that
core point's cluster (border points); all others are noise (-1).
Cluster labels are numbered 0, 1, ... in order of first appearance.
| PARAMETER | DESCRIPTION |
|---|---|
distances
|
Square, symmetric distance matrix.
TYPE:
|
eps
|
Neighbourhood radius.
TYPE:
|
min_samples
|
Minimum neighbourhood size of a core point.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
NDArray[int64]
|
Integer cluster label per point ( |
implicit_word_network.context.clustering.cosine_distances
¶
Pairwise cosine distances 1 - cos(x, y) (zero vectors get distance 1).
implicit_word_network.context.clustering.euclidean_distances
¶
Pairwise Euclidean distances.