API Reference¶
This section provides detailed documentation for the Python API, auto-generated from source code docstrings.
Core Modules¶
Document Module¶
The Document module provides corpus ingestion:
Document- A raw text with an identifier and metadataCorpus- Ordered collection of documents (from texts, TXT or CSV files)load_example_corpus- Bundled example data
Annotation Module¶
The Annotation module defines the typed output of extractors:
AnnotatedDocument- Sentences, tokens and mentions of one documentSentence,Token,EntitySpan,EntityMentionalign_spans- Map character spans to sentences and tokens
Segmentation Module¶
The Segmentation module splits text into sentences and tokens:
BaseSegmenter- Abstract interfaceRegexSegmenter- Dependency-free segmenterSpacySegmenter- spaCy-backed segmenter (POS tags, lemmas)
Extraction Module¶
The Extraction module provides entity extractors:
BaseEntityExtractor- Abstract interface (annotate())SpanEntityExtractor- Base class for span-only extractorsSpacyEntityExtractor- spaCy NERGLiNEREntityExtractor- Zero-shot NER with GLiNER v2.5GazetteerEntityExtractor- Dictionary matching
Network Module¶
The Network module builds and stores implicit networks:
ImplicitNetwork- Nodes, edges, cooccurrences, contexts, persistenceNetworkConfig- Window, decay and term filtersEntityNode,TermNode,SentenceRef,DocumentRef,Mention,Cooccurrence,EntityEdgerank_entities,rank_sentences,rank_documents,load_weight_matrix- LOAD/EVELIN queriesto_networkx,to_dict,to_json,to_edgelist,to_pandas- Exportsregister_decay- Custom decay functions
Context Module¶
The Context module implements contextual implicit entity networks:
ContextualEdgeClusterer- DBSCAN clustering of edge contextsEdgeContextCluster- One context cluster of an edgeBaseContextEmbedder,BagOfWordsEmbedder,SentenceTransformerEmbedderdbscan- DBSCAN on a distance matrix
Pipeline Module¶
The Pipeline module composes the stages:
ImplicitNetworkPipeline- Extractor + network (+ clusterer)build_network- One-call convenience wrapper
Visualization Module¶
The Visualization module provides plot_network.
Quick Reference¶
from implicit_word_network import Corpus, GLiNEREntityExtractor, ImplicitNetworkPipeline
corpus = Corpus.from_txt("documents.txt")
network = ImplicitNetworkPipeline(GLiNEREntityExtractor(), window=2).run(corpus)
network.edges(top_k=10)
network.neighbors(("Feynman", "person"))
network.contexts(("Feynman", "person"), ("Caltech", "organization"))
network.to_networkx()
network.save("network.npz")