Contextual Edges (CIEN)¶
An implicit network aggregates all cooccurrences of two entities into one edge. When the same pair appears in different contexts, the aggregated edge hides that ambiguity. Contextual implicit entity networks (CIEN) split an edge into one sub-edge per context cluster. See Theory.
On-demand clustering of one edge¶
from implicit_word_network import ContextualEdgeClusterer
clusterer = ContextualEdgeClusterer(eps=0.25, min_samples=1, metric="cosine")
clusters = clusterer.cluster_edge(network, ("Feynman", "PERSON"), ("Nobel Prize", "ORG"))
for cluster in clusters:
print(f"cluster {cluster.label}: {cluster.size} cooccurrences, weight {cluster.weight:.2f}")
print(" ", cluster.contexts[0][:100])
Each EdgeContextCluster holds its member Cooccurrence objects, their
context texts, the summed weight (the sub-edge weight) and the mean embedding.
The cluster weights add up to the original edge weight.
Embedders¶
Contexts are embedded by a BaseContextEmbedder:
BagOfWordsEmbedder(default) — hashed bag of words, no dependencies. Groups contexts by lexical overlap; use a largereps(0.4–0.7).SentenceTransformerEmbedder— neural sentence embeddings as in ECCE (sentence-transformers/multi-qa-distilbert-cos-v1); requires theembeddingsextra. Works well with the ECCE defaults (eps=0.25).
from implicit_word_network import SentenceTransformerEmbedder
embedder = SentenceTransformerEmbedder("sentence-transformers/multi-qa-distilbert-cos-v1", device="cuda")
clusterer = ContextualEdgeClusterer(embedder, eps=0.25)
Custom embedders only need an encode(texts, batch_size=...) -> np.ndarray
method.
Precomputing many edges¶
cluster_edges pools the contexts of many edges and embeds them in batches of
batch_size contexts, which is far faster than one model call per edge:
results = clusterer.cluster_edges(network, min_weight=1.0, show_progress=True)
for (a, b), clusters in results.items():
source, target = network.entity_by_id(a), network.entity_by_id(b)
print(source.text, target.text, [c.size for c in clusters])
Pass explicit pairs=[...] to cluster a selection, or top_k= for the
heaviest edges only. Through the pipeline:
pipeline = ImplicitNetworkPipeline(extractor, clusterer=clusterer, window=2)
network = pipeline.run(corpus)
clusters = pipeline.cluster(network, top_k=100)
DBSCAN¶
The clustering step is a small NumPy DBSCAN operating on a precomputed
distance matrix (implicit_word_network.context.dbscan), so no scikit-learn
dependency is needed. min_samples=1 (the ECCE setting) means every context
belongs to a cluster; with larger values, outlying contexts are collected in a
noise cluster with label -1.