Skip to content

Annotation Module

Typed annotation layer produced by entity extractors and consumed by the network builder.

AnnotatedDocument

implicit_word_network.annotation.AnnotatedDocument dataclass

AnnotatedDocument(id: ID, text: str, sentences: list[Sentence] = list(), tokens: list[Token] = list(), mentions: list[EntityMention] = list(), meta: dict[str, Any] = dict())

A document with sentence, token and entity annotations.

ATTRIBUTE DESCRIPTION
id

Document identifier (copied from the source Document).

TYPE: ID

text

Original text.

TYPE: str

sentences

Sentences in document order.

TYPE: list[Sentence]

tokens

Tokens in document order (each assigned to a sentence).

TYPE: list[Token]

mentions

Entity mentions in document order.

TYPE: list[EntityMention]

meta

Document metadata.

TYPE: dict[str, Any]

Example
from implicit_word_network import GazetteerEntityExtractor

extractor = GazetteerEntityExtractor({"PERSON": ["Feynman"]})
doc = extractor.annotate_text("Feynman taught at Caltech. Feynman loved bongos.")
print(doc.n_sentences)                 # 2
print([m.text for m in doc.mentions])  # ["Feynman", "Feynman"]
print(doc.sentence_text(1))            # "Feynman loved bongos."

n_sentences property

n_sentences: int

Number of sentences.

sentence_text

sentence_text(index: int) -> str

Return the text of sentence index.

tokens_in

tokens_in(index: int) -> list[Token]

Return the tokens of sentence index.

mentions_in

mentions_in(index: int) -> list[EntityMention]

Return the entity mentions of sentence index.

entity_token_mask

entity_token_mask() -> ndarray

Boolean mask over tokens that is True for tokens covered by a mention.

Tokens covered by an entity mention are not used as term nodes when building a network. Mentions with token indices are applied directly; for the others coverage is decided by character overlap, so the method also works for mentions without token indices.

validate

validate() -> None

Check the internal consistency of offsets and indices.

RAISES DESCRIPTION
ValueError

If a sentence, token or mention has invalid offsets or refers to a sentence that does not exist.

Sentence

implicit_word_network.annotation.Sentence dataclass

Sentence(index: int, start: int, end: int)

A sentence as a character span of the document text.

ATTRIBUTE DESCRIPTION
index

Position of the sentence in the document (0-based).

TYPE: int

start

Character offset of the first character.

TYPE: int

end

Character offset one past the last character.

TYPE: int

Token

implicit_word_network.annotation.Token dataclass

Token(text: str, start: int, end: int, sentence: int, is_stop: bool = False, is_punct: bool = False, is_space: bool = False, pos: str = '', lemma: str = '')

A token with linguistic flags used for term selection.

ATTRIBUTE DESCRIPTION
text

Surface form.

TYPE: str

start

Character offset of the first character.

TYPE: int

end

Character offset one past the last character.

TYPE: int

sentence

Index of the sentence containing the token.

TYPE: int

is_stop

Whether the token is a stop word.

TYPE: bool

is_punct

Whether the token is punctuation.

TYPE: bool

is_space

Whether the token consists of whitespace only.

TYPE: bool

pos

Coarse part-of-speech tag (empty when unknown).

TYPE: str

lemma

Lemma (empty when unknown; the lowercased text is used then).

TYPE: str

EntitySpan

implicit_word_network.annotation.EntitySpan dataclass

EntitySpan(start: int, end: int, label: str, score: float = 1.0, text: str = '')

A raw entity prediction (character span) before sentence/token alignment.

Span-based extractors (GLiNER, gazetteers, ...) return these; they are turned into EntityMention objects by align_spans.

ATTRIBUTE DESCRIPTION
start

Character offset of the first character.

TYPE: int

end

Character offset one past the last character.

TYPE: int

label

Entity type.

TYPE: str

score

Confidence in [0, 1].

TYPE: float

text

Surface form (filled from the document text when empty).

TYPE: str

EntityMention

implicit_word_network.annotation.EntityMention dataclass

EntityMention(text: str, label: str, start: int, end: int, sentence: int, score: float = 1.0, token_start: int = -1, token_end: int = -1)

An entity mention aligned to the sentence and token structure.

ATTRIBUTE DESCRIPTION
text

Surface form of the mention.

TYPE: str

label

Entity type (e.g. "PERSON" or "person").

TYPE: str

start

Character offset of the first character.

TYPE: int

end

Character offset one past the last character.

TYPE: int

sentence

Index of the sentence containing the mention.

TYPE: int

score

Extractor confidence in [0, 1].

TYPE: float

token_start

Index of the first covered token (-1 if unknown).

TYPE: int

token_end

Index one past the last covered token (-1 if unknown).

TYPE: int

align_spans

implicit_word_network.annotation.align_spans

align_spans(spans: Sequence[EntitySpan], sentences: Sequence[Sentence], tokens: Sequence[Token], *, text: str) -> list[EntityMention]

Align character spans to sentences and tokens.

Every span is assigned to the sentence containing its first character and to the range of tokens it overlaps. Empty spans are dropped and the result is sorted by document position.

PARAMETER DESCRIPTION
spans

Raw entity spans.

TYPE: Sequence[EntitySpan]

sentences

Sentences of the document (sorted, non-overlapping).

TYPE: Sequence[Sentence]

tokens

Tokens of the document (sorted, non-overlapping).

TYPE: Sequence[Token]

text

Document text (used to fill missing surface forms).

TYPE: str

RETURNS DESCRIPTION
list[EntityMention]

Aligned mentions sorted by (start, end).

RAISES DESCRIPTION
ValueError

If spans are given but the document has no sentences.