Chunking
The chunking module splits structured document content into overlapping,
section-aware pieces suitable for retrieval-augmented generation (RAG)
pipelines. It operates on the json_content representation produced by
doc2mark.load() (with output_format="json"), not on raw
Markdown strings, so that section boundaries, page numbers, and content
types are preserved as metadata on every chunk.
Two sizing modes are supported: character-based (the default) and token-based. Token mode relies on tiktoken and falls back to character counting gracefully when the library is not installed. Install the optional dependency with:
pip install doc2mark[tokenizers]
ChunkingConfig
Configuration dataclass that controls how content is split.
from doc2mark import ChunkingConfig
cfg = ChunkingConfig(
max_chunk_size=1500,
overlap=200,
split_on_heading_level=2,
keep_tables_whole=True,
include_page_markers=False,
size_unit="chars",
encoding_name="cl100k_base",
)
Fields
max_chunk_sizeint, default1500Maximum size of a single chunk, measured in the units selected by
size_unit.overlapint, default200Number of trailing units from the previous chunk prepended to the next chunk to provide context continuity.
split_on_heading_levelint, default2Heading depth at which the document is divided into sections before chunking. Level 1 corresponds to
text:titleitems and level 2 totext:sectionitems in the JSON content.keep_tables_wholebool, defaultTrueWhen
True, a table that would push a chunk overmax_chunk_sizeis kept intact in a single chunk rather than being split across two.include_page_markersbool, defaultFalseReserved for future use.
size_unitLiteral["chars", "tokens"], default"chars"Unit of measurement for
max_chunk_sizeandoverlap."chars"– uses Python’s built-inlen()(fast, no extra dependencies)."tokens"– uses tiktoken with the encoding given byencoding_name. If tiktoken is not installed, a warning is logged and the function silently falls back to character counting.
encoding_namestr, default"cl100k_base"The tiktoken encoding to use when
size_unit="tokens". Common values include"cl100k_base"(GPT-4 / text-embedding-ada-002) and"o200k_base"(GPT-4o).
- class doc2mark.ChunkingConfig(max_chunk_size=1500, overlap=200, split_on_heading_level=2, keep_tables_whole=True, include_page_markers=False, size_unit='chars', encoding_name='cl100k_base')[source]
Bases:
objectConfiguration for the chunking algorithm.
- Parameters:
max_chunk_size (int)
overlap (int)
split_on_heading_level (int)
keep_tables_whole (bool)
include_page_markers (bool)
size_unit (Literal['chars', 'tokens'])
encoding_name (str)
Chunk
Immutable result object returned by chunk_content(). Each
chunk carries the rendered Markdown text together with the section context
and page span that produced it.
Fields
contentstrThe Markdown text of this chunk.
section_titleOptional[str], defaultNoneTitle of the section this chunk belongs to, or
Nonefor content that precedes the first heading.section_hierarchyList[str]Ordered list of ancestor headings, e.g.
["Document Title", "Chapter 2"].page_startOptional[int], defaultNoneFirst page number spanned by items in this chunk (if the source format provides page information).
page_endOptional[int], defaultNoneLast page number spanned by items in this chunk.
content_typesList[str]Set of item type strings present in this chunk (e.g.
["text:normal", "table"]).chunk_indexint, default0Zero-based position of this chunk in the sequence returned by
chunk_content().
- class doc2mark.Chunk(content, section_title=None, section_hierarchy=<factory>, page_start=None, page_end=None, content_types=<factory>, chunk_index=0)[source]
Bases:
objectA section-aware chunk of document content.
- Parameters:
content (str)
section_title (str | None)
section_hierarchy (List[str])
page_start (int | None)
page_end (int | None)
content_types (List[str])
chunk_index (int)
chunk_content
The main entry point for chunking. It accepts the structured
json_content list (a List[Dict] of content blocks) from a
ProcessedDocument and returns an ordered list of
Chunk objects.
Signature
def chunk_content(
json_content: List[Dict[str, Any]],
config: Optional[ChunkingConfig] = None,
) -> List[Chunk]: ...
Parameters
json_contentList[Dict[str, Any]]The structured content blocks obtained from
ProcessedDocument.json_content. Each dictionary has at least"type"and"content"keys. Pass the value returned byload(..., output_format="json").json_content.configOptional[ChunkingConfig], defaultNoneChunking parameters. When
None, aChunkingConfigwith default values is used.
Returns
List[Chunk]An ordered list of
Chunkobjects with sequentialchunk_indexvalues starting from 0.
Notes
The function groups items into sections by scanning for heading items (
text:titleat level 1 andtext:sectionat level 2) up to the depth set bysplit_on_heading_level.Within each section, items are accumulated until the next item would exceed
max_chunk_size; the accumulated items are then flushed as a chunk.When
overlapis greater than zero, the trailing portion of each chunk is prepended to the next chunk (breaking at a word boundary when possible). In token mode the overlap is computed by encoding, slicing, and decoding via tiktoken.Footnotes (
text:footnoteitems) are separated during processing and appended to the chunks that reference them. Unreferenced footnotes are attached to the last chunk.An empty
json_contentlist returns an empty list.You can also call
ProcessedDocument.get_chunks(config)as a convenience shortcut that delegates to this function.
Example – character mode (default)
from doc2mark import load, chunk_content, ChunkingConfig
result = load("report.pdf", output_format="json")
chunks = chunk_content(
result.json_content,
ChunkingConfig(max_chunk_size=1000, overlap=150),
)
for chunk in chunks:
print(f"[{chunk.chunk_index}] {chunk.section_title!r} "
f"(pages {chunk.page_start}-{chunk.page_end})")
print(chunk.content[:120], "...")
print()
Example – token mode
Token mode measures sizes in tiktoken tokens instead of characters.
Install the optional dependency first (pip install doc2mark[tokenizers]).
from doc2mark import load, chunk_content, ChunkingConfig
result = load("report.pdf", output_format="json")
cfg = ChunkingConfig(
max_chunk_size=512,
overlap=64,
size_unit="tokens",
encoding_name="cl100k_base",
)
chunks = chunk_content(result.json_content, cfg)
print(f"{len(chunks)} chunks produced in token mode")
- doc2mark.chunk_content(json_content, config=None)[source]
Split structured document content into section-aware chunks.
- Parameters:
json_content (List[Dict[str, Any]]) – The
json_contentlist from aProcessedDocument.config (ChunkingConfig | None) – Chunking parameters. Uses defaults when
None.
- Returns:
Ordered list of
Chunkobjects.- Return type:
List[Chunk]