Annotation Module¶
Typed annotation layer produced by entity extractors and consumed by the network builder.
AnnotatedDocument¶
implicit_word_network.annotation.AnnotatedDocument
dataclass
¶
AnnotatedDocument(id: ID, text: str, sentences: list[Sentence] = list(), tokens: list[Token] = list(), mentions: list[EntityMention] = list(), meta: dict[str, Any] = dict())
A document with sentence, token and entity annotations.
| ATTRIBUTE | DESCRIPTION |
|---|---|
id |
Document identifier (copied from the source
TYPE:
|
text |
Original text.
TYPE:
|
sentences |
Sentences in document order.
TYPE:
|
tokens |
Tokens in document order (each assigned to a sentence).
TYPE:
|
mentions |
Entity mentions in document order.
TYPE:
|
meta |
Document metadata.
TYPE:
|
Example
from implicit_word_network import GazetteerEntityExtractor
extractor = GazetteerEntityExtractor({"PERSON": ["Feynman"]})
doc = extractor.annotate_text("Feynman taught at Caltech. Feynman loved bongos.")
print(doc.n_sentences) # 2
print([m.text for m in doc.mentions]) # ["Feynman", "Feynman"]
print(doc.sentence_text(1)) # "Feynman loved bongos."
mentions_in
¶
mentions_in(index: int) -> list[EntityMention]
Return the entity mentions of sentence index.
entity_token_mask
¶
Boolean mask over tokens that is True for tokens covered by a mention.
Tokens covered by an entity mention are not used as term nodes when building a network. Mentions with token indices are applied directly; for the others coverage is decided by character overlap, so the method also works for mentions without token indices.
validate
¶
Check the internal consistency of offsets and indices.
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
If a sentence, token or mention has invalid offsets or refers to a sentence that does not exist. |
Sentence¶
implicit_word_network.annotation.Sentence
dataclass
¶
A sentence as a character span of the document text.
| ATTRIBUTE | DESCRIPTION |
|---|---|
index |
Position of the sentence in the document (0-based).
TYPE:
|
start |
Character offset of the first character.
TYPE:
|
end |
Character offset one past the last character.
TYPE:
|
Token¶
implicit_word_network.annotation.Token
dataclass
¶
Token(text: str, start: int, end: int, sentence: int, is_stop: bool = False, is_punct: bool = False, is_space: bool = False, pos: str = '', lemma: str = '')
A token with linguistic flags used for term selection.
| ATTRIBUTE | DESCRIPTION |
|---|---|
text |
Surface form.
TYPE:
|
start |
Character offset of the first character.
TYPE:
|
end |
Character offset one past the last character.
TYPE:
|
sentence |
Index of the sentence containing the token.
TYPE:
|
is_stop |
Whether the token is a stop word.
TYPE:
|
is_punct |
Whether the token is punctuation.
TYPE:
|
is_space |
Whether the token consists of whitespace only.
TYPE:
|
pos |
Coarse part-of-speech tag (empty when unknown).
TYPE:
|
lemma |
Lemma (empty when unknown; the lowercased text is used then).
TYPE:
|
EntitySpan¶
implicit_word_network.annotation.EntitySpan
dataclass
¶
A raw entity prediction (character span) before sentence/token alignment.
Span-based extractors (GLiNER, gazetteers, ...) return these; they are
turned into EntityMention objects by align_spans.
| ATTRIBUTE | DESCRIPTION |
|---|---|
start |
Character offset of the first character.
TYPE:
|
end |
Character offset one past the last character.
TYPE:
|
label |
Entity type.
TYPE:
|
score |
Confidence in
TYPE:
|
text |
Surface form (filled from the document text when empty).
TYPE:
|
EntityMention¶
implicit_word_network.annotation.EntityMention
dataclass
¶
EntityMention(text: str, label: str, start: int, end: int, sentence: int, score: float = 1.0, token_start: int = -1, token_end: int = -1)
An entity mention aligned to the sentence and token structure.
| ATTRIBUTE | DESCRIPTION |
|---|---|
text |
Surface form of the mention.
TYPE:
|
label |
Entity type (e.g.
TYPE:
|
start |
Character offset of the first character.
TYPE:
|
end |
Character offset one past the last character.
TYPE:
|
sentence |
Index of the sentence containing the mention.
TYPE:
|
score |
Extractor confidence in
TYPE:
|
token_start |
Index of the first covered token (
TYPE:
|
token_end |
Index one past the last covered token (
TYPE:
|
align_spans¶
implicit_word_network.annotation.align_spans
¶
align_spans(spans: Sequence[EntitySpan], sentences: Sequence[Sentence], tokens: Sequence[Token], *, text: str) -> list[EntityMention]
Align character spans to sentences and tokens.
Every span is assigned to the sentence containing its first character and to the range of tokens it overlaps. Empty spans are dropped and the result is sorted by document position.
| PARAMETER | DESCRIPTION |
|---|---|
spans
|
Raw entity spans.
TYPE:
|
sentences
|
Sentences of the document (sorted, non-overlapping).
TYPE:
|
tokens
|
Tokens of the document (sorted, non-overlapping).
TYPE:
|
text
|
Document text (used to fill missing surface forms).
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
list[EntityMention]
|
Aligned mentions sorted by |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
If spans are given but the document has no sentences. |