- Python 100%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| .letta | ||
| src/dantess_pretokenization | ||
| .codex | ||
| .gitignore | ||
| .python-version | ||
| pyproject.toml | ||
| README.md | ||
| uv.lock | ||
dantess-pretokenization
Simple, extensible pretokenization for Hugging Face datasets with multiple chat formats.
This project turns raw chat or completion datasets into a single Parquet dataset containing:
input_idsattention_masklabels
It is designed for SFT-style preprocessing where some tokens should contribute to loss and others should be masked with -100.
What It Supports
- Chat datasets with configurable conversation and role field names
- Plain completion datasets
- Multiple prompt/chat wrappers:
ChatMLChatML with BOS FixGemma4GLM4Llama3Llama4Mistral V7 TekkenChatMLReasoningChatMLNonReasoningGLM4Reasoning
- Optional reasoning traces on chat data
- Optional system-prompt override
- Optional "train only on the last assistant response" mode
- Interactive tokenization preview before the main run
- Optional push of the final Parquet dataset to the Hugging Face Hub
Installation
Using uv:
uv sync
Then run:
uv run pretokenization --help
Or install with pip:
pip install -e .
Quick Start
Create a config file:
base_model: meta-llama/Llama-3.1-8B-Instruct
sequence_len: 8192
datasets:
- path: HuggingFaceH4/ultrachat_200k
type: dan-chat-advanced
conversation_field: messages
role_field: role
system_prompt: You are a concise assistant.
- path: openwebtext
type: completion
field: text
Run the tokenizer:
uv run pretokenization \
--config config.yaml \
--output output.parquet \
--format Llama3
If you omit --format and your config includes chat datasets, the CLI will open a TUI selector.
Config File
Top-level keys:
| Key | Required | Description |
|---|---|---|
base_model |
Yes | Tokenizer source passed to transformers.AutoTokenizer.from_pretrained(...). |
sequence_len |
No | Maximum sequence length. Defaults to 8192. |
datasets |
Yes | List of dataset entries to process. |
model_type |
No | Present in config parsing but not used by the current CLI. |
tokenizer_type |
No | Present in config parsing but not used by the current CLI. |
Dataset entries are split by type:
dan-chat-advanced: chat/conversation datacompletion: plain text completion data
Chat Dataset Fields
Supported chat dataset keys:
| Key | Required | Default | Description |
|---|---|---|---|
path |
Yes | - | Dataset identifier passed to datasets.load_dataset(...). |
type |
Yes | - | Must be dan-chat-advanced. |
conversation_field |
No | conversations |
Field containing the list of turns. |
role_field |
No | from |
Field inside each turn that contains the role. |
reasoning_field |
No | - | Field name to read reasoning content from. |
reasoning_prefix |
No | format-specific | Override reasoning prefix text. |
reasoning_suffix |
No | format-specific | Override reasoning suffix text. |
reasoning_loss |
No | format-specific | Whether reasoning tokens contribute to loss. |
system_prompt |
No | - | Replaces the existing system message or prepends one. |
train_last_response_only |
No | false |
Masks all assistant turns except the last one. |
Recognized role values are normalized to internal roles:
- system:
system,sys - user:
user,human - assistant:
assistant,model,gpt - tool:
tool,function,environment
Turn content is read from value when present, otherwise from content.
Completion Dataset Fields
| Key | Required | Default | Description |
|---|---|---|---|
path |
Yes | - | Dataset identifier passed to datasets.load_dataset(...). |
type |
Yes | - | Must be completion. |
field |
No | text |
Field containing the raw text to tokenize. |
Expected Input Shape
Chat example:
{
"messages": [
{"role": "system", "content": "You are helpful."},
{"role": "user", "content": "Explain gradient descent."},
{"role": "assistant", "content": "Gradient descent is..."}
]
}
Chat items may also include:
reasoningorreasoning_contenton assistant turnslosson individual turnstoolsorfunctionsfor tool-enabled ChatML reasoning data- assistant
tool_callsorfunction_calls
Completion example:
{
"text": "A plain training sample."
}
Output
The CLI writes a single Parquet file. Each row contains:
| Column | Meaning |
|---|---|
input_ids |
Token IDs for the formatted sample |
attention_mask |
All ones, matching sequence length |
labels |
Training labels, with masked positions set to -100 |
Behavior by dataset type:
- Chat datasets mask non-trainable tokens such as role wrappers, system text, and user text.
- Completion datasets train on every token.
- The final combined dataset is shuffled before being written.
CLI
pretokenization --config CONFIG [--output OUTPUT] [--format FORMAT] [--system-prompt TEXT] [--hf-push REPO] [--seed N] [--debug] [-y]
Flags:
| Flag | Description |
|---|---|
--config |
Path to the YAML config file. Required. |
--output |
Output Parquet path. Defaults to ./output.parquet. |
--format |
Chat format name. Skips the interactive selector. |
--system-prompt |
Global override applied to all chat datasets. |
--hf-push |
Pushes the final dataset to the Hugging Face Hub. |
--seed |
Seed for the final shuffle. |
--debug |
Enables debug logging. |
--yes, -y |
Skips the confirmation prompt after previews. |
Examples:
uv run pretokenization --config config.yaml --format ChatMLReasoning -y
uv run pretokenization \
--config config.yaml \
--output tokenized.parquet \
--format Llama4 \
--system-prompt "You are a careful coding assistant." \
--hf-push your-org/your-dataset \
-y
If you use --hf-push, set HF_TOKEN in the environment first.
Format Notes
ChatMLReasoninghas special handling for reasoning traces, tool definitions, tool responses, and assistant tool calls.ChatMLNonReasoninginjects an empty<think>...</think>block into assistant turns.ChatML with BOS Fixuses the tokenizer's BOS token for the starting sequence.Gemma4includes an in-code warning that the first system message may need manual reasoning markup for Gemma 4 reasoning workflows.- For non-
ChatMLReasoningformats, reasoning content is injected only on the last assistant turn.
Operational Notes
- Source datasets are loaded in streaming mode from the
trainsplit. - The tool previews one example from each dataset before processing unless
-yis used. - Intermediate tokenized shards are written to temporary Parquet files.
- The final combined dataset is materialized in memory for shuffling before the output file is written, so very large runs still need enough RAM.
Development
The project is packaged with Hatchling and includes Ruff configuration in pyproject.toml.