Skip to content

Document Module

Corpus ingestion.

Document

implicit_word_network.document.Document dataclass

Document(text: str, id: ID = 0, meta: dict[str, Any] = dict())

A single raw text document.

ATTRIBUTE DESCRIPTION
text

Raw text content of the document.

TYPE: str

id

Unique identifier of the document inside its corpus.

TYPE: ID

meta

Optional dictionary of document-level metadata (e.g. title, publication date).

TYPE: dict[str, Any]

Example
from implicit_word_network import Document

doc = Document("Feynman received the Nobel Prize in 1965.", id="feynman")
print(doc.id)         # "feynman"
print(len(doc))       # number of characters

Corpus

implicit_word_network.document.Corpus

Corpus(documents: Iterable[Document | str] = (), *, meta: Mapping[str, Any] | None = None)

Ordered collection of Document objects.

Documents can be given as Document instances or as plain strings, in which case sequential integer identifiers are assigned. Document identifiers must be unique inside a corpus.

ATTRIBUTE DESCRIPTION
meta

Optional dictionary of corpus-level metadata.

TYPE: dict[str, Any]

Example
from implicit_word_network import Corpus

corpus = Corpus.from_texts(["First document.", "Second document."])
print(len(corpus))          # 2
print(corpus[0].text)       # "First document."
print(corpus.ids())         # [0, 1]

# Load one document per non-empty line
corpus = Corpus.from_txt("documents.txt")

# Load from CSV (needs a text column; other columns become metadata)
corpus = Corpus.from_csv("documents.csv", text_column="text", id_column="doc_id")

from_texts classmethod

from_texts(texts: Iterable[str], *, ids: Iterable[ID] | None = None, meta: Mapping[str, Any] | None = None) -> Corpus

Build a corpus from an iterable of raw strings.

PARAMETER DESCRIPTION
texts

Document texts.

TYPE: Iterable[str]

ids

Optional identifiers, parallel to texts. Sequential integers are used when omitted.

TYPE: Iterable[ID] | None DEFAULT: None

meta

Optional corpus-level metadata.

TYPE: Mapping[str, Any] | None DEFAULT: None

RETURNS DESCRIPTION
Corpus

A new corpus with one document per text.

from_txt classmethod

from_txt(path: str | Path, *, encoding: str = 'utf-8', delimiter: str = '\n', meta: Mapping[str, Any] | None = None) -> Corpus

Load a plain-text file with one document per delimited block.

Empty blocks are skipped. With the default delimiter every non-empty line becomes a document; use "\n\n" to split on blank lines.

PARAMETER DESCRIPTION
path

Path to the text file.

TYPE: str | Path

encoding

File encoding.

TYPE: str DEFAULT: 'utf-8'

delimiter

String separating documents.

TYPE: str DEFAULT: '\n'

meta

Optional corpus-level metadata.

TYPE: Mapping[str, Any] | None DEFAULT: None

RETURNS DESCRIPTION
Corpus

A new corpus with sequential integer identifiers.

from_csv classmethod

from_csv(path: str | Path, *, text_column: str = 'text', id_column: str | None = None, encoding: str = 'utf-8', delimiter: str = ',', meta: Mapping[str, Any] | None = None) -> Corpus

Load a corpus from a delimited file.

Every row becomes a document. Columns other than text_column and id_column are stored as document metadata.

PARAMETER DESCRIPTION
path

Path to the CSV/TSV file.

TYPE: str | Path

text_column

Name of the column holding the document text.

TYPE: str DEFAULT: 'text'

id_column

Optional name of the column holding document ids.

TYPE: str | None DEFAULT: None

encoding

File encoding.

TYPE: str DEFAULT: 'utf-8'

delimiter

Field delimiter ("\t" for TSV files).

TYPE: str DEFAULT: ','

meta

Optional corpus-level metadata.

TYPE: Mapping[str, Any] | None DEFAULT: None

RETURNS DESCRIPTION
Corpus

A new corpus.

RAISES DESCRIPTION
ValueError

If a required column is missing.

from_records classmethod

from_records(records: Iterable[Mapping[str, Any]], *, text_key: str = 'text', id_key: str | None = None, meta: Mapping[str, Any] | None = None) -> Corpus

Build a corpus from dictionaries (e.g. rows of a DataFrame).

PARAMETER DESCRIPTION
records

Mappings with at least a text_key entry.

TYPE: Iterable[Mapping[str, Any]]

text_key

Key holding the document text.

TYPE: str DEFAULT: 'text'

id_key

Optional key holding the document id.

TYPE: str | None DEFAULT: None

meta

Optional corpus-level metadata.

TYPE: Mapping[str, Any] | None DEFAULT: None

RETURNS DESCRIPTION
Corpus

A new corpus. Remaining keys are stored as document metadata.

add

add(document: Document | str, *, doc_id: ID | None = None, meta: Mapping[str, Any] | None = None) -> Document

Append a document to the corpus.

PARAMETER DESCRIPTION
document

A Document or a raw text string.

TYPE: Document | str

doc_id

Identifier to assign when document is a string. The next free integer is used when omitted.

TYPE: ID | None DEFAULT: None

meta

Metadata to attach when document is a string.

TYPE: Mapping[str, Any] | None DEFAULT: None

RETURNS DESCRIPTION
Document

The stored Document.

RAISES DESCRIPTION
ValueError

If the identifier already exists in the corpus.

get

get(doc_id: ID) -> Document

Return the document with identifier doc_id.

ids

ids() -> list[ID]

Return document identifiers in corpus order.

texts

texts() -> list[str]

Return document texts in corpus order.

head

head(n: int = 5) -> list[Document]

Return the first n documents in corpus order.

Example data

implicit_word_network.datasets.load_example_corpus

load_example_corpus() -> Corpus

Load the bundled example corpus (six English Wikipedia-style paragraphs about physicists).

RETURNS DESCRIPTION
Corpus

A Corpus with integer ids.

Example
from implicit_word_network import load_example_corpus

corpus = load_example_corpus()
print(len(corpus), corpus[0].text[:60])