Pretokenization scripts
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-09-20 21:46:34 -04:00
.letta ??? 2026-09-20 21:46:34 -04:00
src/dantess_pretokenization ??? 2026-09-20 21:46:34 -04:00
.codex ??? 2026-09-20 21:46:34 -04:00
.gitignore Diverge and squash forked commits 2026-04-12 15:19:54 -04:00
.python-version Initial commit 2025-10-17 23:20:59 -07:00
pyproject.toml Diverge and squash forked commits 2026-04-12 15:19:54 -04:00
README.md add readme 2026-04-12 15:25:26 -04:00
uv.lock Diverge and squash forked commits 2026-04-12 15:19:54 -04:00

dantess-pretokenization

Simple, extensible pretokenization for Hugging Face datasets with multiple chat formats.

This project turns raw chat or completion datasets into a single Parquet dataset containing:

  • input_ids
  • attention_mask
  • labels

It is designed for SFT-style preprocessing where some tokens should contribute to loss and others should be masked with -100.

What It Supports

  • Chat datasets with configurable conversation and role field names
  • Plain completion datasets
  • Multiple prompt/chat wrappers:
    • ChatML
    • ChatML with BOS Fix
    • Gemma4
    • GLM4
    • Llama3
    • Llama4
    • Mistral V7 Tekken
    • ChatMLReasoning
    • ChatMLNonReasoning
    • GLM4Reasoning
  • Optional reasoning traces on chat data
  • Optional system-prompt override
  • Optional "train only on the last assistant response" mode
  • Interactive tokenization preview before the main run
  • Optional push of the final Parquet dataset to the Hugging Face Hub

Installation

Using uv:

uv sync

Then run:

uv run pretokenization --help

Or install with pip:

pip install -e .

Quick Start

Create a config file:

base_model: meta-llama/Llama-3.1-8B-Instruct
sequence_len: 8192

datasets:
  - path: HuggingFaceH4/ultrachat_200k
    type: dan-chat-advanced
    conversation_field: messages
    role_field: role
    system_prompt: You are a concise assistant.

  - path: openwebtext
    type: completion
    field: text

Run the tokenizer:

uv run pretokenization \
  --config config.yaml \
  --output output.parquet \
  --format Llama3

If you omit --format and your config includes chat datasets, the CLI will open a TUI selector.

Config File

Top-level keys:

Key Required Description
base_model Yes Tokenizer source passed to transformers.AutoTokenizer.from_pretrained(...).
sequence_len No Maximum sequence length. Defaults to 8192.
datasets Yes List of dataset entries to process.
model_type No Present in config parsing but not used by the current CLI.
tokenizer_type No Present in config parsing but not used by the current CLI.

Dataset entries are split by type:

  • dan-chat-advanced: chat/conversation data
  • completion: plain text completion data

Chat Dataset Fields

Supported chat dataset keys:

Key Required Default Description
path Yes - Dataset identifier passed to datasets.load_dataset(...).
type Yes - Must be dan-chat-advanced.
conversation_field No conversations Field containing the list of turns.
role_field No from Field inside each turn that contains the role.
reasoning_field No - Field name to read reasoning content from.
reasoning_prefix No format-specific Override reasoning prefix text.
reasoning_suffix No format-specific Override reasoning suffix text.
reasoning_loss No format-specific Whether reasoning tokens contribute to loss.
system_prompt No - Replaces the existing system message or prepends one.
train_last_response_only No false Masks all assistant turns except the last one.

Recognized role values are normalized to internal roles:

  • system: system, sys
  • user: user, human
  • assistant: assistant, model, gpt
  • tool: tool, function, environment

Turn content is read from value when present, otherwise from content.

Completion Dataset Fields

Key Required Default Description
path Yes - Dataset identifier passed to datasets.load_dataset(...).
type Yes - Must be completion.
field No text Field containing the raw text to tokenize.

Expected Input Shape

Chat example:

{
  "messages": [
    {"role": "system", "content": "You are helpful."},
    {"role": "user", "content": "Explain gradient descent."},
    {"role": "assistant", "content": "Gradient descent is..."}
  ]
}

Chat items may also include:

  • reasoning or reasoning_content on assistant turns
  • loss on individual turns
  • tools or functions for tool-enabled ChatML reasoning data
  • assistant tool_calls or function_calls

Completion example:

{
  "text": "A plain training sample."
}

Output

The CLI writes a single Parquet file. Each row contains:

Column Meaning
input_ids Token IDs for the formatted sample
attention_mask All ones, matching sequence length
labels Training labels, with masked positions set to -100

Behavior by dataset type:

  • Chat datasets mask non-trainable tokens such as role wrappers, system text, and user text.
  • Completion datasets train on every token.
  • The final combined dataset is shuffled before being written.

CLI

pretokenization --config CONFIG [--output OUTPUT] [--format FORMAT] [--system-prompt TEXT] [--hf-push REPO] [--seed N] [--debug] [-y]

Flags:

Flag Description
--config Path to the YAML config file. Required.
--output Output Parquet path. Defaults to ./output.parquet.
--format Chat format name. Skips the interactive selector.
--system-prompt Global override applied to all chat datasets.
--hf-push Pushes the final dataset to the Hugging Face Hub.
--seed Seed for the final shuffle.
--debug Enables debug logging.
--yes, -y Skips the confirmation prompt after previews.

Examples:

uv run pretokenization --config config.yaml --format ChatMLReasoning -y
uv run pretokenization \
  --config config.yaml \
  --output tokenized.parquet \
  --format Llama4 \
  --system-prompt "You are a careful coding assistant." \
  --hf-push your-org/your-dataset \
  -y

If you use --hf-push, set HF_TOKEN in the environment first.

Format Notes

  • ChatMLReasoning has special handling for reasoning traces, tool definitions, tool responses, and assistant tool calls.
  • ChatMLNonReasoning injects an empty <think>...</think> block into assistant turns.
  • ChatML with BOS Fix uses the tokenizer's BOS token for the starting sequence.
  • Gemma4 includes an in-code warning that the first system message may need manual reasoning markup for Gemma 4 reasoning workflows.
  • For non-ChatMLReasoning formats, reasoning content is injected only on the last assistant turn.

Operational Notes

  • Source datasets are loaded in streaming mode from the train split.
  • The tool previews one example from each dataset before processing unless -y is used.
  • Intermediate tokenized shards are written to temporary Parquet files.
  • The final combined dataset is materialized in memory for shuffling before the output file is written, so very large runs still need enough RAM.

Development

The project is packaged with Hatchling and includes Ruff configuration in pyproject.toml.