Skip to content

Network Module

Construction, storage, querying and export of implicit networks.

ImplicitNetwork

implicit_word_network.network.graph.ImplicitNetwork

ImplicitNetwork(config: NetworkConfig | None = None, *, normalize_entity: Callable[[str], str] | None = None)

Implicit entity network of a document collection.

Networks are built incrementally from AnnotatedDocument objects (see :meth:add_documents) and expose entity, term, sentence and document nodes together with aggregated entity–entity edges, instance-level cooccurrences, LOAD importance weights and ranking queries.

PARAMETER DESCRIPTION
config

Construction parameters (window, decay, term filters).

TYPE: NetworkConfig | None DEFAULT: None

normalize_entity

Function mapping a mention surface form to the normalised name used for entity identity. Defaults to lowercasing and whitespace collapsing. Custom functions are not persisted by :meth:save.

TYPE: Callable[[str], str] | None DEFAULT: None

Example
from implicit_word_network import GazetteerEntityExtractor, ImplicitNetwork

extractor = GazetteerEntityExtractor({"PERSON": ["Feynman", "Schwinger"]})
docs = extractor.annotate_all(["Feynman shared the prize with Schwinger."])

network = ImplicitNetwork.from_documents(docs)
for edge in network.edges():
    print(edge.source.text, edge.target.text, edge.weight)

from_documents classmethod

from_documents(documents: Iterable[AnnotatedDocument], config: NetworkConfig | None = None, *, normalize_entity: Callable[[str], str] | None = None, show_progress: bool = False) -> ImplicitNetwork

Build a network from annotated documents.

PARAMETER DESCRIPTION
documents

Annotated documents.

TYPE: Iterable[AnnotatedDocument]

config

Construction parameters.

TYPE: NetworkConfig | None DEFAULT: None

normalize_entity

Entity name normalisation function.

TYPE: Callable[[str], str] | None DEFAULT: None

show_progress

Display a progress bar.

TYPE: bool DEFAULT: False

add_documents

add_documents(documents: Iterable[AnnotatedDocument], *, show_progress: bool = False) -> ImplicitNetwork

Add documents to the network, updating all nodes, tables and edges.

Cooccurrences never cross document boundaries, so adding documents is purely additive: the contribution of the new batch is computed with vectorised sparse operations and accumulated into the existing matrices. Appends are amortised O(1) per row, so networks can grow by many small batches.

PARAMETER DESCRIPTION
documents

Annotated documents that are not yet part of the network.

TYPE: Iterable[AnnotatedDocument]

show_progress

Display a progress bar.

TYPE: bool DEFAULT: False

RETURNS DESCRIPTION
ImplicitNetwork

self for chaining.

RAISES DESCRIPTION
ValueError

If a document id is already present or annotations are inconsistent.

n_documents property

n_documents: int

Number of document nodes.

n_sentences property

n_sentences: int

Number of sentence nodes.

n_entities property

n_entities: int

Number of entity nodes.

n_terms property

n_terms: int

Number of term nodes.

n_mentions property

n_mentions: int

Number of entity mentions.

n_edges property

n_edges: int

Number of (undirected) entity–entity edges.

window property

window: int

Context window in sentences.

entity_entity_matrix property

entity_entity_matrix: csr_matrix

Symmetric n_entities × n_entities matrix of aggregated edge weights.

entity_count_matrix property

entity_count_matrix: csr_matrix

Symmetric n_entities × n_entities matrix of cooccurrence counts.

entity_term_matrix property

entity_term_matrix: csr_matrix

n_entities × n_terms matrix of same-sentence cooccurrence counts.

sentence_entity_matrix property

sentence_entity_matrix: csr_matrix

n_sentences × n_entities matrix of mention counts.

sentence_term_matrix property

sentence_term_matrix: csr_matrix

n_sentences × n_terms matrix of term occurrence counts.

sentence_document

sentence_document() -> NDArray[int32]

Document index of every sentence (the document–sentence edges).

entity_counts

entity_counts() -> NDArray[int64]

Number of mentions per entity.

term_counts

term_counts() -> NDArray[int64]

Number of occurrences per term.

entity_label_ids

entity_label_ids() -> NDArray[int32]

Integer type id per entity (see :meth:entity_labels).

load_weight_matrix

load_weight_matrix() -> csr_matrix

Directed LOAD importance weights (see load_weight_matrix).

entity

entity(text: str, label: str) -> EntityNode | None

Look up an entity by (surface) name and type; None if absent.

entity_by_id

entity_by_id(entity_id: int) -> EntityNode

Return the entity node with integer id entity_id.

entities

entities(*, label: str | None = None) -> list[EntityNode]

All entity nodes (optionally restricted to one type), ordered by id.

iter_entities

iter_entities(*, label: str | None = None) -> Iterator[EntityNode]

Lazily iterate over entity nodes ordered by id.

entity_labels

entity_labels() -> list[str]

Distinct entity types present in the network.

top_entities

top_entities(k: int = 10, *, label: str | None = None) -> list[EntityNode]

The k most frequently mentioned entities.

term

term(text: str, pos: str = '') -> TermNode | None

Look up a term by normalised form and POS tag; None if absent.

term_by_id

term_by_id(term_id: int) -> TermNode

Return the term node with integer id term_id.

terms

terms() -> list[TermNode]

All term nodes ordered by id.

top_terms

top_terms(k: int = 10) -> list[TermNode]

The k most frequent terms.

document

document(index: int) -> DocumentRef

Return the document node at position index.

document_by_id

document_by_id(doc_id: ID) -> DocumentRef

Return the document node with identifier doc_id.

documents

documents() -> list[DocumentRef]

All document nodes in insertion order.

sentence

sentence(sentence_id: int) -> SentenceRef

Return the sentence node with global id sentence_id.

sentences_of_document

sentences_of_document(doc: int | ID) -> list[SentenceRef]

Sentence nodes of a document (by position or identifier).

mentions_of

mentions_of(entity: EntityLike) -> list[Mention]

All mentions of an entity in corpus order.

sentences_of

sentences_of(entity: EntityLike) -> list[SentenceRef]

Distinct sentences mentioning an entity.

mention

mention(row: int) -> Mention

Return the mention stored at row row of the mention table.

weight

weight(a: EntityLike, b: EntityLike) -> float

Aggregated edge weight between two entities (0.0 if not connected).

count

count(a: EntityLike, b: EntityLike) -> int

Number of cooccurrences between two entities.

load_weight

load_weight(x: EntityLike, y: EntityLike) -> float

Directed LOAD importance of y for x (see load_weight_matrix).

edge_table

edge_table(*, min_weight: float = 0.0, top_k: int | None = None, labels: Sequence[str] | None = None) -> tuple[NDArray[int64], NDArray[int64], NDArray[float64], NDArray[int64]]

Edges as arrays (source_ids, target_ids, weights, counts) with source < target.

Sorted by decreasing weight; top_k selects the heaviest edges with a partial sort. This is the allocation-free counterpart of :meth:edges for bulk analyses.

edges

edges(*, min_weight: float = 0.0, top_k: int | None = None, labels: Sequence[str] | None = None) -> list[EntityEdge]

Aggregated entity–entity edges sorted by decreasing weight.

PARAMETER DESCRIPTION
min_weight

Drop edges lighter than this.

TYPE: float DEFAULT: 0.0

top_k

Keep only the heaviest top_k edges.

TYPE: int | None DEFAULT: None

labels

Keep only edges whose endpoints both have one of these entity types.

TYPE: Sequence[str] | None DEFAULT: None

neighbors

neighbors(entity: EntityLike, *, k: int | None = None, min_weight: float = 0.0, weighting: Weighting = 'raw') -> list[tuple[EntityNode, float]]

Adjacent entities with edge weights, heaviest first.

PARAMETER DESCRIPTION
entity

The entity whose neighbourhood is returned.

TYPE: EntityLike

k

Number of neighbours (None for all).

TYPE: int | None DEFAULT: None

min_weight

Drop neighbours below this weight.

TYPE: float DEFAULT: 0.0

weighting

"raw" for ω = Σ exp(−δ) or "load" for the directed LOAD importance of the neighbour for entity.

TYPE: Weighting DEFAULT: 'raw'

cooccurrences

cooccurrences(a: EntityLike, b: EntityLike) -> list[Cooccurrence]

Instance-level cooccurrences of two entities (on-demand enumeration).

PARAMETER DESCRIPTION
a

First entity.

TYPE: EntityLike

b

Second entity (must differ from a).

TYPE: EntityLike

RETURNS DESCRIPTION
list[Cooccurrence]

One Cooccurrence per pair of mentions inside the context

list[Cooccurrence]

window, in corpus order.

cooccurrence_table

cooccurrence_table(a: EntityLike, b: EntityLike) -> CooccurrenceTable

Cooccurrences of two entities as column arrays (see :meth:all_cooccurrences).

all_cooccurrences

all_cooccurrences() -> CooccurrenceTable

All cooccurrence instances of the network as column arrays.

The rows reference the mention table (see :meth:mention). This is the eager counterpart of :meth:cooccurrences and is used for exports and validation; its size grows quadratically with the number of mentions per context window. Use :meth:iter_cooccurrences to stream the pairs in bounded chunks.

iter_cooccurrences

iter_cooccurrences(*, chunk_size: int = 1000000) -> Iterator[CooccurrenceTable]

Stream all cooccurrence instances in chunks of at most chunk_size pairs.

context_of

context_of(cooccurrence: Cooccurrence) -> str

Text of the sentences spanned by a cooccurrence (its context window).

context_of_span

context_of_span(first: int, last: int) -> str

Text of the global sentences first..last (inclusive).

contexts

contexts(a: EntityLike, b: EntityLike) -> list[str]

Context texts of all cooccurrences of two entities (parallel to :meth:cooccurrences).

entity_terms

entity_terms(entity: EntityLike, *, k: int | None = None) -> list[tuple[TermNode, int]]

Terms sharing sentences with an entity, with cooccurrence counts (most frequent first).

term_entities

term_entities(term: TermLike, *, k: int | None = None) -> list[tuple[EntityNode, int]]

Entities sharing sentences with a term, with cooccurrence counts.

rank_entities

rank_entities(query: EntityLike | Sequence[EntityLike], *, label: str | None = None, k: int | None = 10, weighting: Weighting = 'load') -> list[tuple[EntityNode, float]]

Rank entities related to one or more query entities (see rank_entities).

rank_sentences

rank_sentences(query: EntityLike | Sequence[EntityLike], *, k: int | None = 10, n_terms: int = 10) -> list[tuple[SentenceRef, float]]

Rank sentences describing the query entities (see rank_sentences).

rank_documents

rank_documents(query: EntityLike | Sequence[EntityLike], *, k: int | None = 10, n_terms: int = 10) -> list[tuple[DocumentRef, float]]

Rank documents for the query entities (see rank_documents).

to_networkx

to_networkx(**kwargs: Any) -> Graph

Convert to a NetworkX graph (see to_networkx).

to_dict

to_dict(**kwargs: Any) -> dict[str, Any]

JSON-serialisable node/edge lists (see to_dict).

save

save(path: str | Path) -> Path

Persist the network to a compressed .npz file.

PARAMETER DESCRIPTION
path

Target path (.npz is appended when missing).

TYPE: str | Path

RETURNS DESCRIPTION
Path

The written path.

load classmethod

load(path: str | Path, *, normalize_entity: Callable[[str], str] | None = None) -> ImplicitNetwork

Load a network written by :meth:save.

PARAMETER DESCRIPTION
path

Path of the .npz file.

TYPE: str | Path

normalize_entity

Entity normalisation function used when the network was built (needed for name lookups if it was custom).

TYPE: Callable[[str], str] | None DEFAULT: None

summary

summary() -> str

Human-readable summary of the network size and configuration.

NetworkConfig

implicit_word_network.network._types.NetworkConfig dataclass

NetworkConfig(window: int = 2, decay: str = 'exponential', include_stopwords: bool = False, include_punctuation: bool = False, term_pos: frozenset[str] | None = None, use_lemma: bool = True, lowercase_terms: bool = True, store_text: bool = True)

Parameters controlling how an implicit network is built.

ATTRIBUTE DESCRIPTION
window

Context window c in sentences. Two entity mentions cooccur when they appear in the same document at most window sentences apart (0 restricts cooccurrence to one sentence).

TYPE: int

decay

Name of the weighting function applied to the sentence distance δ of a cooccurrence. "exponential" (the default from Spitz & Gertz) uses exp(-δ); "constant" counts cooccurrences; "inverse" uses 1 / (1 + δ); "linear" uses 1 - δ / (window + 1). Custom functions can be added with register_decay.

TYPE: str

include_stopwords

Keep stop words as term nodes.

TYPE: bool

include_punctuation

Keep punctuation tokens as term nodes.

TYPE: bool

term_pos

Restrict term nodes to these coarse POS tags (e.g. {"NOUN", "PROPN", "VERB", "ADJ"}). None keeps every tag; note that the regex segmenter does not provide POS tags.

TYPE: frozenset[str] | None

use_lemma

Use the lemma (when available) instead of the surface form as term identity.

TYPE: bool

lowercase_terms

Lowercase term identities.

TYPE: bool

store_text

Keep sentence texts in the network. Required for cooccurrence contexts and contextual edge clustering.

TYPE: bool

Example
from implicit_word_network import NetworkConfig

config = NetworkConfig(window=3, decay="exponential", term_pos={"NOUN", "PROPN"})

to_dict

to_dict() -> dict[str, Any]

Return a JSON-serialisable representation.

from_dict classmethod

from_dict(data: dict[str, Any]) -> NetworkConfig

Rebuild a configuration from to_dict output.

Nodes and edges

implicit_word_network.network._types.EntityNode dataclass

EntityNode(id: int, text: str, norm: str, label: str, count: int)

An entity node, unique by normalised name and type.

ATTRIBUTE DESCRIPTION
id

Integer index of the node inside the network.

TYPE: int

text

Surface form of the first mention seen.

TYPE: str

norm

Normalised name used for identity.

TYPE: str

label

Entity type.

TYPE: str

count

Number of mentions in the corpus.

TYPE: int

key property

key: tuple[str, str]

(norm, label) identity tuple.

implicit_word_network.network._types.TermNode dataclass

TermNode(id: int, text: str, pos: str, count: int)

A term node (non-entity word), unique by normalised form and POS tag.

ATTRIBUTE DESCRIPTION
id

Integer index of the term inside the network.

TYPE: int

text

Normalised form (lemma or lowercased surface).

TYPE: str

pos

Coarse part-of-speech tag (empty when unknown).

TYPE: str

count

Number of occurrences in the corpus.

TYPE: int

implicit_word_network.network._types.DocumentRef dataclass

DocumentRef(index: int, id: ID, n_sentences: int, meta: dict[str, Any] = dict())

A document node.

ATTRIBUTE DESCRIPTION
index

Position of the document in the network.

TYPE: int

id

User-facing document identifier.

TYPE: ID

n_sentences

Number of sentences.

TYPE: int

meta

Document metadata.

TYPE: dict[str, Any]

implicit_word_network.network._types.SentenceRef dataclass

SentenceRef(id: int, document: ID, index: int, text: str)

A sentence node.

ATTRIBUTE DESCRIPTION
id

Global sentence index inside the network.

TYPE: int

document

Identifier of the containing document.

TYPE: ID

index

Position of the sentence inside its document.

TYPE: int

text

Sentence text (empty when texts are not stored).

TYPE: str

implicit_word_network.network._types.Mention dataclass

Mention(id: int, entity: EntityNode, sentence: SentenceRef, start: int, end: int, score: float)

One occurrence of an entity in a sentence.

ATTRIBUTE DESCRIPTION
id

Row index in the mention table.

TYPE: int

entity

The entity node.

TYPE: EntityNode

sentence

The containing sentence.

TYPE: SentenceRef

start

Character offset in the document text.

TYPE: int

end

Character offset one past the last character.

TYPE: int

score

Extractor confidence.

TYPE: float

implicit_word_network.network._types.Cooccurrence dataclass

Cooccurrence(source: Mention, target: Mention, delta: int, weight: float)

One cooccurrence of two entity mentions inside the context window.

ATTRIBUTE DESCRIPTION
source

First mention (earlier in the document).

TYPE: Mention

target

Second mention.

TYPE: Mention

delta

Sentence distance between the mentions.

TYPE: int

weight

Decayed weight contributed to the edge (e.g. exp(-delta)).

TYPE: float

document property

document: ID

Identifier of the document containing both mentions.

sentence_span property

sentence_span: tuple[int, int]

Global ids of the first and last sentence of the context.

implicit_word_network.network._types.EntityEdge dataclass

EntityEdge(source: EntityNode, target: EntityNode, weight: float, count: int)

Aggregated (implicit) edge between two entities.

ATTRIBUTE DESCRIPTION
source

Entity with the smaller id.

TYPE: EntityNode

target

Entity with the larger id.

TYPE: EntityNode

weight

Aggregated weight ω = Σ decay(δ) over all cooccurrences.

TYPE: float

count

Number of cooccurrence instances.

TYPE: int

implicit_word_network.network._types.CooccurrenceTable dataclass

CooccurrenceTable(source: NDArray[int64], target: NDArray[int64], delta: NDArray[int64], weight: NDArray[float64])

Column-oriented table of all cooccurrence instances of a network.

ATTRIBUTE DESCRIPTION
source

Mention row indices of the first mention of every pair.

TYPE: NDArray[int64]

target

Mention row indices of the second mention.

TYPE: NDArray[int64]

delta

Sentence distances.

TYPE: NDArray[int64]

weight

Decayed weights.

TYPE: NDArray[float64]

Ranking and LOAD weights

LOAD edge weighting and entity-centric ranking queries.

Implements the directed importance weights and the query model of the LOAD graph (Spitz & Gertz, 2016; Spitz, Almasian & Gertz, 2017):

  • load_weight_matrix: ω(x, y) = log(|Y| / |N(x) ∩ Y|) · Σ_i exp(−δ_i(x, y)), where Y is the set of entities of y's type and N(x) the neighbourhood of x.
  • rank_entities: single-entity queries rank by ω(q, x) / ω_max; multi-entity queries use r = c + s with cohesion c(x) = |N(x) ∩ Q| − 1 and the normalised weight sum s(x) = Σ_q ω(q, x) / s_max.
  • rank_sentences: r = c + s with c(x) the number of query entities in the sentence and s(x) = |N(x) ∩ T_Q| / |T_Q| for the union T_Q of the k most important terms of the query entities.
  • rank_documents: c(p) = max c(x), s(p) = Σ s(x) (normalised) over the sentences of a document.

Weighting module-attribute

Weighting = Literal['raw', 'load']

"raw": undirected ω = Σ exp(−δ); "load": directed, type-normalised LOAD weight.

load_weight_matrix

load_weight_matrix(network: ImplicitNetwork) -> csr_matrix

Directed LOAD weights ω(x, y) = log(|Y| / |N(x) ∩ Y|) · Σ exp(−δ).

Row x holds the importance of every neighbour y for x; the matrix is not symmetric. Pairs where x is connected to every entity of y's type get weight 0 (log 1).

rank_entities

rank_entities(network: ImplicitNetwork, query: EntityLike | Sequence[EntityLike], *, label: str | None = None, k: int | None = 10, weighting: Weighting = 'load') -> list[tuple[EntityNode, float]]

Rank entities by their relation to one or more query entities (EVELIN).

PARAMETER DESCRIPTION
network

Source network.

TYPE: ImplicitNetwork

query

One entity or a set of entities.

TYPE: EntityLike | Sequence[EntityLike]

label

Restrict results to this entity type.

TYPE: str | None DEFAULT: None

k

Number of results (None for all).

TYPE: int | None DEFAULT: 10

weighting

Edge weights used for the scores.

TYPE: Weighting DEFAULT: 'load'

RETURNS DESCRIPTION
list[tuple[EntityNode, float]]

(EntityNode, score) pairs sorted by decreasing score. Single-entity

list[tuple[EntityNode, float]]

queries yield scores in [0, 1]; multi-entity queries yield

list[tuple[EntityNode, float]]

c + s ∈ [0, |Q|].

rank_sentences

rank_sentences(network: ImplicitNetwork, query: EntityLike | Sequence[EntityLike], *, k: int | None = 10, n_terms: int = 10) -> list[tuple[SentenceRef, float]]

Rank sentences that describe the query entities (EVELIN sentence queries).

Only sentences containing at least one query entity are returned.

PARAMETER DESCRIPTION
network

Source network.

TYPE: ImplicitNetwork

query

One entity or a set of entities.

TYPE: EntityLike | Sequence[EntityLike]

k

Number of results (None for all).

TYPE: int | None DEFAULT: 10

n_terms

Number of most important terms per query entity used for the term-overlap component.

TYPE: int DEFAULT: 10

rank_documents

rank_documents(network: ImplicitNetwork, query: EntityLike | Sequence[EntityLike], *, k: int | None = 10, n_terms: int = 10) -> list[tuple[DocumentRef, float]]

Rank documents (LOAD pages) for the query entities.

c(p) is the maximum sentence cohesion in the document and s(p) the sum of the sentence term scores, normalised by the maximum over documents. Only documents containing at least one query entity are returned.

Export

Conversion of implicit networks to NetworkX graphs and plain data structures.

entity_node_id

entity_node_id(node: EntityNode) -> str

Stable string identifier of an entity node ("<norm>|<label>").

term_node_id

term_node_id(node: TermNode) -> str

Stable string identifier of a term node ("<text>|<pos>|term").

to_networkx

to_networkx(network: ImplicitNetwork, *, min_weight: float = 0.0, top_k: int | None = None, labels: Sequence[str] | None = None, include_isolated: bool = True, include_terms: bool = False, max_terms_per_entity: int | None = 10, min_term_count: int = 1) -> Graph

Convert the entity layer of a network into a networkx.Graph.

Entity nodes carry kind="entity", text, norm, label and count attributes; edges carry weight and count. Optionally, term nodes (kind="term") are attached to entities with weight equal to the number of shared sentences.

PARAMETER DESCRIPTION
network

Source network.

TYPE: ImplicitNetwork

min_weight

Drop entity edges lighter than this.

TYPE: float DEFAULT: 0.0

top_k

Keep only the heaviest top_k entity edges.

TYPE: int | None DEFAULT: None

labels

Restrict to these entity types.

TYPE: Sequence[str] | None DEFAULT: None

include_isolated

Keep entities without edges.

TYPE: bool DEFAULT: True

include_terms

Add term nodes and entity–term edges.

TYPE: bool DEFAULT: False

max_terms_per_entity

Number of strongest terms per entity to add.

TYPE: int | None DEFAULT: 10

min_term_count

Minimum shared-sentence count of an entity–term edge.

TYPE: int DEFAULT: 1

RETURNS DESCRIPTION
Graph

An undirected graph.

to_dict

to_dict(network: ImplicitNetwork, *, min_weight: float = 0.0, top_k: int | None = None, labels: Sequence[str] | None = None) -> dict[str, Any]

JSON-serialisable representation of the entity layer.

RETURNS DESCRIPTION
dict[str, Any]

A dictionary with config, nodes (id, text, norm, label, count)

dict[str, Any]

and edges (source, target, weight, count) entries.

to_json

to_json(network: ImplicitNetwork, path: str | Path, **kwargs: Any) -> Path

Write to_dict output to a JSON file.

to_edgelist

to_edgelist(network: ImplicitNetwork, *, min_weight: float = 0.0, top_k: int | None = None, labels: Sequence[str] | None = None) -> list[dict[str, Any]]

Flat edge records (one dict per entity–entity edge), e.g. for CSV export.

to_pandas

to_pandas(network: ImplicitNetwork, **kwargs: Any) -> tuple[Any, Any]

Return (nodes, edges) DataFrames (requires the pandas extra).

Storage primitives

Append-only storage primitives used by :class:~implicit_word_network.network.ImplicitNetwork.

Both classes trade a little bookkeeping for amortised O(1) appends, so that networks can grow by thousands of small batches without the quadratic cost of re-concatenating arrays or re-adding sparse matrices on every update.

GrowableArray

GrowableArray(dtype: Any, capacity: int = 64)

1-D NumPy array with geometric capacity growth.

extend copies new values into a pre-allocated buffer that doubles when full; view returns the filled prefix without copying.

from_array classmethod

from_array(values: NDArray[Any]) -> GrowableArray

Wrap an existing array (copied).

extend

extend(values: NDArray[Any]) -> None

Append values (converted to the buffer dtype).

view property

view: NDArray[Any]

The filled part of the buffer (a view, not a copy).

SparseAccumulator

SparseAccumulator(dtype: Any, shape: tuple[int, int] = (0, 0))

Sum of sparse COO contributions with lazy CSR compaction.

Contributions are appended as (rows, cols, data) triplets; the CSR matrix is only rebuilt (summing duplicates) when :attr:matrix is read or when the pending triplets outgrow the compacted matrix, which keeps both the per-update cost and the memory overhead bounded.

from_matrix classmethod

from_matrix(matrix: csr_matrix | coo_matrix) -> SparseAccumulator

Wrap an existing sparse matrix.

resize

resize(shape: tuple[int, int]) -> None

Grow the logical shape (never shrinks).

add

add(rows: NDArray[Any], cols: NDArray[Any], data: NDArray[Any]) -> None

Add data[k] at (rows[k], cols[k]); duplicates are summed.

add_matrix

add_matrix(matrix: csr_matrix | coo_matrix) -> None

Add a whole sparse matrix (its shape must fit the logical shape).

matrix property

matrix: csr_matrix

The compacted CSR matrix.

Decay functions and vectorised primitives

Vectorised building blocks for network construction.

All functions operate on NumPy arrays / SciPy sparse matrices and avoid per-token Python loops. They are the performance core of ImplicitNetwork.

DecayFunction module-attribute

DecayFunction = Callable[[NDArray[np.int64]], NDArray[np.float64]]

Maps an integer array of sentence distances to float weights.

register_decay

register_decay(name: str, function: DecayFunction) -> None

Register a custom decay function under name.

The function receives a 1-D int64 array of sentence distances and must return a float64 array of the same shape.

Example
import numpy as np
from implicit_word_network.network import register_decay

register_decay("gaussian", lambda d: np.exp(-(d.astype(float) ** 2) / 2))

available_decays

available_decays() -> list[str]

Names of all registered decay functions.

resolve_decay

resolve_decay(name: str, *, window: int) -> DecayFunction

Return the decay function registered under name.

"linear" is resolved relative to window as 1 - δ / (window + 1).

RAISES DESCRIPTION
KeyError

If no such decay function exists.

ragged_ranges

ragged_ranges(starts: NDArray[int64], stops: NDArray[int64]) -> tuple[NDArray[int64], NDArray[int64]]

Expand half-open integer ranges into flat (owner, value) arrays.

For every k the values starts[k], ..., stops[k] - 1 are emitted with owner k. Empty ranges contribute nothing.

PARAMETER DESCRIPTION
starts

Range starts.

TYPE: NDArray[int64]

stops

Range stops (exclusive).

TYPE: NDArray[int64]

RETURNS DESCRIPTION
tuple[NDArray[int64], NDArray[int64]]

owner (index k of the range) and value arrays.

window_pairs

window_pairs(group: NDArray[integer], position: NDArray[integer], window: int, *, chunk_size: int = 1000000) -> Iterator[tuple[NDArray[int64], NDArray[int64]]]

Enumerate index pairs (i, j) with i < j inside a positional window.

Two rows pair up when they belong to the same group (document) and 0 <= position[j] - position[i] <= window (sentence distance). The input must be sorted by (group, position). Pairs are produced in chunks of at most chunk_size to bound memory.

PARAMETER DESCRIPTION
group

Group id per row (e.g. document index).

TYPE: NDArray[integer]

position

Position per row (e.g. sentence index).

TYPE: NDArray[integer]

window

Maximum positional distance.

TYPE: int

chunk_size

Maximum number of pairs per yielded chunk.

TYPE: int DEFAULT: 1000000

YIELDS DESCRIPTION
tuple[NDArray[int64], NDArray[int64]]

Arrays (i, j) of row indices.

all_window_pairs

all_window_pairs(group: NDArray[integer], position: NDArray[integer], window: int) -> tuple[NDArray[int64], NDArray[int64]]

Eager variant of window_pairs returning concatenated arrays.

same_group_pairs

same_group_pairs(group_a: NDArray[integer], group_b: NDArray[integer]) -> tuple[NDArray[int64], NDArray[int64]]

Pair every row of a with every row of b sharing the same group value.

Both inputs must be sorted ascending.

incidence_matrix

incidence_matrix(rows: NDArray[integer], cols: NDArray[integer], shape: tuple[int, int], *, dtype: type = int64) -> csr_matrix

Sparse count matrix with one increment per (row, col) pair.

band_structure

band_structure(group: NDArray[integer], window: int) -> tuple[NDArray[int64], NDArray[int64], NDArray[int64]]

Sparsity pattern of the sentence-distance kernel.

For consecutive positions s and s + d (0 <= d <= window) that belong to the same group, both (s, s + d) and (s + d, s) are emitted (the diagonal once).

PARAMETER DESCRIPTION
group

Group id per position, sorted so that groups are contiguous.

TYPE: NDArray[integer]

window

Maximum distance.

TYPE: int

RETURNS DESCRIPTION
tuple[NDArray[int64], NDArray[int64], NDArray[int64]]

rows, cols and delta arrays.

band_kernel

band_kernel(group: NDArray[integer], window: int, values: NDArray[floating] | NDArray[integer], rows: NDArray[int64], cols: NDArray[int64]) -> csr_matrix

Assemble a square kernel matrix from a band_structure pattern.

grow

grow(matrix: csr_matrix, n_rows: int, n_cols: int) -> csr_matrix

Return a copy of a CSR matrix padded with empty rows/columns.