A small LoRA implementation designed for training, integrated with Transformers models.
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-09-25 15:51:18 -04:00
benchmarks rename :D 2026-09-25 15:20:21 -04:00
examples generalize 2026-09-25 15:51:18 -04:00
src/loralei generalize 2026-09-25 15:51:18 -04:00
tests generalize 2026-09-25 15:51:18 -04:00
.gitignore init 2026-09-19 22:52:43 -04:00
.python-version init 2026-09-19 22:52:43 -04:00
pyproject.toml generalize 2026-09-25 15:51:18 -04:00
README.md generalize 2026-09-25 15:51:18 -04:00
uv.lock generalize 2026-09-25 15:51:18 -04:00

loralei

The base weights can only watch while the optimizer descends onto the adapter.

A small LoRA implementation designed for training, integrated with Transformers models.

The project currently targets the latest Transformers release resolved by uv.

Optimization is provided by the pinned prejac git dependency. It includes regular optimizers such as AdamW and SkipStepAdamW, as well as PoLoRA for paired LoRA factors.

Install

uv sync --dev
uv sync --extra peft --dev  # also run PEFT interoperability tests
uv sync --extra qlora --dev  # enable TorchAO NF4 or int4 QLoRA
import prejac

optimizer = prejac.AdamW(model.parameters(), lr=1e-3)
# or, for explicitly paired LoRA A/B factors:
optimizer = prejac.PoLoRA([(lora_A, lora_B)], lr=2e-4)

Use

from loralei import LoRAConfig, inject_lora, save_pretrained

config = LoRAConfig(
    r=8,
    lora_alpha=16,
    lora_dropout=0.05,
    target_modules=["q_proj", "v_proj"],
)
inject_lora(model, config)

# Native loralei checkpoint:
save_pretrained(model, "adapter")

# A checkpoint loadable by peft.PeftModel.from_pretrained:
save_pretrained(model, "adapter-peft", format="peft")

target_parameters handles packed expert weights used by current MoE models:

config = LoRAConfig(
    r=4,
    target_modules=["q_proj", "v_proj"],
    target_parameters=["experts.gate_up_proj", "experts.down_proj"],
)

For QLoRA, install the qlora extra and quantize the frozen weights of the selected nn.Linear targets:

config = LoRAConfig(
    r=8,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
    quantize_base=True,
)
inject_lora(model, config)

By default, quantize_base stores those base weights as TorchAO NF4Tensors and uses linear_nf4 in forward and backward. To use loralei's fused TileLang kernels for both the forward and input-gradient matmuls, set nf4_kernel="tilelang":

config = LoRAConfig(
    r=8,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
    quantize_base=True,
    nf4_kernel="tilelang",
)
inject_lora(model, config)

On Apple Metal/MPS, this path supports FP16 activations. On CUDA, it supports FP16 and BF16 activations and uses TileLang's generic GEMM lowering, without pinning the kernel to one NVIDIA architecture. Both backends require linear feature dimensions divisible by 32. They keep the TorchAO NF4 weights packed and fuse unpacking and double-scale reconstruction into the matmul. The default remains TorchAO's linear_nf4. The qlora extra installs TileLang. On CUDA, the training runner warms forward and input-gradient kernels for each distinct shape before compilation, benchmarks a small set of valid tile sizes, and reuses the fastest choice. Results are cached per GPU and shape at ~/.cache/loralei/tilelang_nf4_cuda_autotune.json, so later runs skip the search. The first uncached run spends extra startup time compiling and measuring candidates. Set training.tilelang_autotune: false to use the conservative fallback tile instead. Direct kernel users can call warmup_nf4_tilelang_cuda(model, rows) before CUDA graph capture; outside graph capture, a cache miss tunes lazily on the first call. To use TorchAO's CUDA Int4TilePackedTo4dTensor instead, select it explicitly:

config = LoRAConfig(
    r=8,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
    quantize_base=True,
    quantization_type="int4",
    int4_group_size=128,
)
inject_lora(model, config)

Set lora.quantize_base: true in the YAML runner for NF4; set lora.quantization_type: int4 to select int4. NF4 requires each selected weight's element count to be divisible by nf4_block_size * nf4_scaler_block_size (defaults: 64 and 256); smaller models can set both fields in the config. Int4 uses groupwise quantization with group sizes 32, 64, 128, or 256. Its tinygemm kernel requires CUDA compute capability 8.0 or newer and quantizes the frozen weights to bfloat16 before packing. Backward computes input gradients from the packed weights in bounded chunks, adding kernel work compared with NF4's dedicated backward. For packed gated-MoE experts, target_parameters can include both experts.gate_up_proj and experts.down_proj to NF4-quantize their frozen 3-D weight banks. This works with packed expert modules that expose the usual top-k hidden states, indices, and weights and use paired gate/up and down projections, including Mixtral, DeepSeek V3, and Qwen3.5. The expert loop reads only routed 2-D matrices while training per-expert LoRA factors. The expert NF4 matmuls use TorchAO; nf4_kernel selects the backend for ordinary 2-D linear modules. Use lora_dropout: 0 for packed expert targets. The parameter bank path requires both expert targets and currently supports NF4. Its data-dependent routing loop stays eager when the rest of the model is compiled.

Native adapter checkpoints preserve the quantization settings and quantize a dense base again when loaded. Existing native adapter checkpoints still load; new native checkpoints use Loralei filenames. PEFT adapter exports contain the adapter factors and require the caller to prepare the desired base precision separately. Merging NF4-quantized packed expert adapters is not supported because it would restore the full dense expert banks. Merging ordinary quantized linear weights dequantizes those bases; unmerge_lora restores their packed weights. Load int4 checkpoints with the base on CUDA, or pass quantization_device="cuda" to load_pretrained to move selected linears there during loading. Avoid dtype casts on the model after QLoRA injection, since TorchAO's packed tensors do not support changing weight dtype.

Use load_pretrained, merge_lora, and unmerge_lora for loading and reversible inference-time weight folding. To export a normal Transformers model after merging, use merge_and_unload_lora(model) (or merge_lora(model, unload=True)) before calling the model's save_pretrained. PEFT is optional; loralei itself does not import it at runtime.

YAML fine-tuning

The repository includes a small LoRA supervised-fine-tuning loop. It loads a dataset with datasets.load_dataset, applies the tokenizer's chat template, and writes an adapter-only checkpoint:

uv run loralei-finetune examples/finetune.yaml

W&B logging is opt-in. Install the extra and add a wandb block to the YAML configuration:

uv sync --extra wandb
wandb:
  project: loralei
  name: qwen-sft

The runner logs loss, learning rate, token throughput, peak CUDA tensor allocation and allocator reservation across model setup and training, gradient norm (when clipping is enabled), and skipped optimizer steps at training.logging_steps. wandb: false leaves tracking disabled; training.report_to: wandb is also supported.

Tokenization is performed once with datasets.map and reused from the Hugging Face dataset cache. For CPU-heavy input pipelines, set training.preprocessing_num_workers and training.dataloader_num_workers. CUDA runs can also use training.pin_memory: true for asynchronous host-to- device copies. On supported Transformers models, select fused attention with training.attn_implementation: sdpa (or flash_attention_2 when installed). training.torch_compile: true enables PyTorch compilation; benchmark it with fixed sequence lengths or packing enabled. Set training.pad_to_max_length: true to keep batch shapes fixed at training.max_seq_length, which avoids extra compiler graphs when packed sequences have slightly different lengths. With TorchAO QLoRA, the runner compiles decoder layers separately and selectively checkpoints them, preserving expensive matrix products while recomputing cheaper operations. With TileLang QLoRA, it compiles the whole model without activation checkpointing because the fused NF4 matmuls keep the base weights packed. The TorchAO mode requires a decoder with a layers ModuleList; the runner reports an error if it is missing. The runner clears temporary cached CUDA allocations before QLoRA training begins. Peak memory still depends on the model and batch, so measure it before sizing a larger run. For a memory-limited GPU, enable training.gradient_checkpointing to trade some recomputation for lower activation memory. Compiled QLoRA already checkpoints its decoder layers, so that setting adds no further layer checkpointing in this mode.

Qwen3.5 uses a hybrid attention stack. Run the PoLoRA example with its optional attention and Liger kernels:

uv run --extra linear-attention --extra liger loralei-finetune examples/qwen35-9b-polora.yaml

training.fused_shared_input_lora: true enables the common self-attention Q/K/V and MLP gate/up groups. For other architectures, provide a list of groups with parent_name and projection_names, as this example does for its linear attention projections. Shared-input groups require zero LoRA dropout and do not support torch.compile. training.liger_kernel: true uses the Transformers Liger integration to select the patch for the loaded model.

The example uses the measured high-throughput settings: rank-16 all-linear PoLoRA, BF16, packed 2048-token sequences, a per-device batch size of two, configured shared-input LoRA projection groups, Liger RMSNorm and SwiGLU, and Liger's fused linear cross entropy loss. The output head is frozen, and compilation and gradient checkpointing are disabled. This profile needs substantial device memory; reduce the batch size or enable checkpointing if it does not fit. Cut Cross Entropy remains available as a memory-conscious loss option below.

For large-vocabulary models, the finetuning loop can use Apple's memory-efficient Cut Cross Entropy implementation. Install the optional dependency and select it in the training configuration:

uv sync --extra cce
training:
  loss: cut_cross_entropy  # `cce` is also accepted

The runner passes the backbone's final hidden states and output-head weights directly to linear_cross_entropy, so it does not materialize the full logits tensor. The package chooses its CCE or torch_compile implementation for the current platform; set training.cut_cross_entropy.impl to override it.

The dataset can contain messages/conversations, a plain text field, or instruction/response fields. Set data.messages_field or data.text_field when the defaults do not match. The training optimizer can be adamw, skipstepadamw, or polora; the latter uses the paired LoRA factors directly. Set optimizer.kwargs.scale_invariant: true to account for the adapter's lora_alpha/r multiplier (or lora_alpha/sqrt(r) with RSLoRA) in PoLoRA's update budget. See examples/finetune.yaml for the intentionally small configuration surface.

Set training.packing: true to enable best-fit decreasing (BFD) sample packing. Packing tokenizes without truncation, drops complete samples longer than training.max_seq_length, and emits a warning for the dropped samples. For causal language models with packed right-padding, set training.omit_attention_mask_when_packed: true to omit the mask from model input and its device transfer. The padding is a trailing suffix, so causal positions with training targets cannot attend to it; leave this option off for models whose padding semantics require an explicit mask.

Kaitan/dantess pretokenized Parquet files are also supported. Point dataset at the file; rows with input_ids, attention_mask, and labels are detected automatically, so their existing -100 loss masks are preserved:

dataset:
  path: /data/tokenized.parquet

training:
  max_seq_length: 8192