Performance & Scaling¶
The 0.0.x prototype stored every token as a Python dictionary, enumerated entity pairs with nested loops and quadratic list membership tests, and called the embedding model once per edge. Version 0.1 replaces this with column-oriented tables and sparse linear algebra.
Architecture for scale¶
| Concern | Design |
|---|---|
| Entity–entity weights | sparse triple product Ω = Sᵀ K S per added batch (sentence × entity incidence, banded sentence-distance kernel) |
| Cooccurrence counts, entity–term counts | the same products with a 0/1 kernel and S_eᵀ S_t |
| Growing corpora | append-only columns with geometric capacity growth (GrowableArray) and lazily compacted sparse accumulators, so 1 000 small updates cost about the same as one big build |
| Large entity vocabularies | integer ids everywhere; an entity → mention index built once per update makes per-entity queries independent of the corpus size |
| Long documents | instance pairs are enumerated with searchsorted + ragged range expansion in bounded chunks (iter_cooccurrences(chunk_size=...)), never per token |
| Query objects | node and sentence objects are cached; edge_table / cooccurrence_table return NumPy arrays for bulk work |
| Context embedding | contexts of many edges are pooled into model batches (ContextualEdgeClusterer.cluster_edges) |
| Extraction | spaCy nlp.pipe with n_process, GLiNER sentence-aligned chunks scored in batches; max_length for very long documents |
The remaining Python work per document is one pass over its tokens for vocabulary lookups; everything downstream is NumPy/SciPy.
Benchmarks¶
Synthetic corpora (benchmarks/synthetic.py, Zipf-distributed vocabularies,
≈1.5 mentions and 15 tokens per sentence), single core of an Apple M-series
laptop, network construction only (extraction is model-bound and excluded):
| Corpus | Tokens | Mentions | Entities | Edges | Build time | Throughput |
|---|---|---|---|---|---|---|
| 2 000 docs × 30 sentences | 1.0 M | 90 k | 12 k | 129 k | 0.7 s | 1.5 M tokens/s |
| same, added in 200 batches of 10 | 1.0 M | 90 k | 12 k | 129 k | 0.85 s | 1.2 M tokens/s |
| 10 000 docs × 30 sentences | 5.0 M | 450 k | 52 k | 558 k | 3.7 s | 1.4 M tokens/s |
| 200 docs × 2 000 sentences (long documents) | 6.8 M | 600 k | 5 k | 495 k | 5.1 s | 1.3 M tokens/s |
Against the original implementation on 300 documents (150 k tokens):
| Implementation | Time |
|---|---|
0.0.x buildGraph |
3.5 s |
0.1 ImplicitNetwork.from_documents |
0.14 s (26× faster) |
The gap widens with corpus size because the legacy pair enumeration is quadratic in the number of mentions per document.
Typical query latencies on the 5 M-token network: neighbors ≈ 3 ms,
edges(top_k=100) ≈ 70 ms (partial sort over 558 k edges),
cooccurrences on the busiest edge (45 k instances) ≈ 0.3 s, dominated by
building the Python result objects; use cooccurrence_table for arrays.
Reproduce with:
python benchmarks/bench_build.py --docs 2000 --queries
python benchmarks/bench_build.py --docs 10000 --chunk 500 --entities 100000 --terms 200000
python benchmarks/bench_build.py --docs 200 --sentences 2000 --entities 5000
python benchmarks/bench_legacy.py --docs 300
Memory¶
- Mentions and term occurrences are stored as NumPy int32/int64 columns, not per-token objects.
- Aggregated edges live in CSR matrices; instance-level cooccurrences are
not materialised by default.
cooccurrences(a, b)enumerates the pairs of one edge on demand,iter_cooccurrencesstreams all pairs in chunks andall_cooccurrencesmaterialises everything only when asked. NetworkConfig(store_text=False)drops sentence texts when contexts are not needed.ImplicitNetworkPipeline.run(chunk_size=...)adds documents to the network in chunks, so annotated documents of a large corpus never have to be held in memory all at once.
Throughput tips¶
- Extraction dominates end-to-end runtime. Use
batch_sizeand, for spaCy,n_process; for GLiNER, use a GPU (device="cuda") and a largerbatch_size. - Prefer
SpacyEntityExtractor(one pass) over GLiNER +SpacySegmenter(two passes) when spaCy's label set suffices. - Keep the window small (
window=2is the ECCE default); the exponential decay makes distant cooccurrences negligible while the number of pairs grows with the window. - Build once,
save()the network andload()it in analysis sessions.
Validation¶
See Validation: the engine is tested against a literal implementation of the published definitions and against the original 0.0.x code on random corpora.