Document Module¶
Corpus ingestion.
Document¶
implicit_word_network.document.Document
dataclass
¶
A single raw text document.
| ATTRIBUTE | DESCRIPTION |
|---|---|
text |
Raw text content of the document.
TYPE:
|
id |
Unique identifier of the document inside its corpus.
TYPE:
|
meta |
Optional dictionary of document-level metadata (e.g. title, publication date).
TYPE:
|
Example
Corpus¶
implicit_word_network.document.Corpus
¶
Corpus(documents: Iterable[Document | str] = (), *, meta: Mapping[str, Any] | None = None)
Ordered collection of Document objects.
Documents can be given as Document instances or as plain strings,
in which case sequential integer identifiers are assigned. Document
identifiers must be unique inside a corpus.
| ATTRIBUTE | DESCRIPTION |
|---|---|
meta |
Optional dictionary of corpus-level metadata.
TYPE:
|
Example
from implicit_word_network import Corpus
corpus = Corpus.from_texts(["First document.", "Second document."])
print(len(corpus)) # 2
print(corpus[0].text) # "First document."
print(corpus.ids()) # [0, 1]
# Load one document per non-empty line
corpus = Corpus.from_txt("documents.txt")
# Load from CSV (needs a text column; other columns become metadata)
corpus = Corpus.from_csv("documents.csv", text_column="text", id_column="doc_id")
from_texts
classmethod
¶
from_texts(texts: Iterable[str], *, ids: Iterable[ID] | None = None, meta: Mapping[str, Any] | None = None) -> Corpus
Build a corpus from an iterable of raw strings.
| PARAMETER | DESCRIPTION |
|---|---|
texts
|
Document texts.
TYPE:
|
ids
|
Optional identifiers, parallel to
TYPE:
|
meta
|
Optional corpus-level metadata.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
Corpus
|
A new corpus with one document per text. |
from_txt
classmethod
¶
from_txt(path: str | Path, *, encoding: str = 'utf-8', delimiter: str = '\n', meta: Mapping[str, Any] | None = None) -> Corpus
Load a plain-text file with one document per delimited block.
Empty blocks are skipped. With the default delimiter every non-empty
line becomes a document; use "\n\n" to split on blank lines.
| PARAMETER | DESCRIPTION |
|---|---|
path
|
Path to the text file.
TYPE:
|
encoding
|
File encoding.
TYPE:
|
delimiter
|
String separating documents.
TYPE:
|
meta
|
Optional corpus-level metadata.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
Corpus
|
A new corpus with sequential integer identifiers. |
from_csv
classmethod
¶
from_csv(path: str | Path, *, text_column: str = 'text', id_column: str | None = None, encoding: str = 'utf-8', delimiter: str = ',', meta: Mapping[str, Any] | None = None) -> Corpus
Load a corpus from a delimited file.
Every row becomes a document. Columns other than text_column and
id_column are stored as document metadata.
| PARAMETER | DESCRIPTION |
|---|---|
path
|
Path to the CSV/TSV file.
TYPE:
|
text_column
|
Name of the column holding the document text.
TYPE:
|
id_column
|
Optional name of the column holding document ids.
TYPE:
|
encoding
|
File encoding.
TYPE:
|
delimiter
|
Field delimiter (
TYPE:
|
meta
|
Optional corpus-level metadata.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
Corpus
|
A new corpus. |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
If a required column is missing. |
from_records
classmethod
¶
from_records(records: Iterable[Mapping[str, Any]], *, text_key: str = 'text', id_key: str | None = None, meta: Mapping[str, Any] | None = None) -> Corpus
Build a corpus from dictionaries (e.g. rows of a DataFrame).
| PARAMETER | DESCRIPTION |
|---|---|
records
|
Mappings with at least a
TYPE:
|
text_key
|
Key holding the document text.
TYPE:
|
id_key
|
Optional key holding the document id.
TYPE:
|
meta
|
Optional corpus-level metadata.
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
Corpus
|
A new corpus. Remaining keys are stored as document metadata. |
add
¶
add(document: Document | str, *, doc_id: ID | None = None, meta: Mapping[str, Any] | None = None) -> Document
Append a document to the corpus.
| PARAMETER | DESCRIPTION |
|---|---|
document
|
A
TYPE:
|
doc_id
|
Identifier to assign when
TYPE:
|
meta
|
Metadata to attach when
TYPE:
|
| RETURNS | DESCRIPTION |
|---|---|
Document
|
The stored |
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
If the identifier already exists in the corpus. |