- Python 100%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| benchmarks | ||
| examples | ||
| src/loralei | ||
| tests | ||
| .gitignore | ||
| .python-version | ||
| pyproject.toml | ||
| README.md | ||
| uv.lock | ||
loralei
The base weights can only watch while the optimizer descends onto the adapter.
A small LoRA implementation designed for training, integrated with Transformers models.
The project currently targets the latest Transformers release resolved by uv.
Optimization is provided by the pinned prejac git dependency. It includes
regular optimizers such as AdamW and SkipStepAdamW, as well as PoLoRA for
paired LoRA factors.
Install
uv sync --dev
uv sync --extra peft --dev # also run PEFT interoperability tests
uv sync --extra qlora --dev # enable TorchAO NF4 or int4 QLoRA
import prejac
optimizer = prejac.AdamW(model.parameters(), lr=1e-3)
# or, for explicitly paired LoRA A/B factors:
optimizer = prejac.PoLoRA([(lora_A, lora_B)], lr=2e-4)
Use
from loralei import LoRAConfig, inject_lora, save_pretrained
config = LoRAConfig(
r=8,
lora_alpha=16,
lora_dropout=0.05,
target_modules=["q_proj", "v_proj"],
)
inject_lora(model, config)
# Native loralei checkpoint:
save_pretrained(model, "adapter")
# A checkpoint loadable by peft.PeftModel.from_pretrained:
save_pretrained(model, "adapter-peft", format="peft")
target_parameters handles packed expert weights used by current MoE models:
config = LoRAConfig(
r=4,
target_modules=["q_proj", "v_proj"],
target_parameters=["experts.gate_up_proj", "experts.down_proj"],
)
For QLoRA, install the qlora extra and quantize the frozen weights of the
selected nn.Linear targets:
config = LoRAConfig(
r=8,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
quantize_base=True,
)
inject_lora(model, config)
By default, quantize_base stores those base weights as TorchAO NF4Tensors
and uses linear_nf4 in forward and backward. To use loralei's fused TileLang
kernels for both the forward and input-gradient matmuls, set
nf4_kernel="tilelang":
config = LoRAConfig(
r=8,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
quantize_base=True,
nf4_kernel="tilelang",
)
inject_lora(model, config)
On Apple Metal/MPS, this path supports FP16 activations. On CUDA, it supports
FP16 and BF16 activations and uses TileLang's generic GEMM lowering, without
pinning the kernel to one NVIDIA architecture. Both backends require linear
feature dimensions divisible by 32. They keep the TorchAO NF4 weights packed
and fuse unpacking and double-scale reconstruction into the matmul. The
default remains TorchAO's linear_nf4. The qlora extra installs TileLang.
On CUDA, the training runner warms forward and input-gradient kernels for each
distinct shape before compilation, benchmarks a small set of valid tile sizes,
and reuses the fastest choice. Results are cached per GPU and shape at
~/.cache/loralei/tilelang_nf4_cuda_autotune.json, so later runs skip the search.
The first uncached run spends extra startup time compiling and measuring
candidates. Set training.tilelang_autotune: false to use the conservative
fallback tile instead. Direct kernel users can call
warmup_nf4_tilelang_cuda(model, rows) before CUDA graph capture; outside graph
capture, a cache miss tunes lazily on the first call.
To use TorchAO's CUDA
Int4TilePackedTo4dTensor instead, select it explicitly:
config = LoRAConfig(
r=8,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
quantize_base=True,
quantization_type="int4",
int4_group_size=128,
)
inject_lora(model, config)
Set lora.quantize_base: true in the YAML runner for NF4; set
lora.quantization_type: int4 to select int4. NF4 requires each selected
weight's element count to be divisible by nf4_block_size * nf4_scaler_block_size (defaults: 64 and 256); smaller models can set both
fields in the config. Int4 uses groupwise quantization with group sizes 32,
64, 128, or 256. Its tinygemm kernel requires CUDA compute capability
8.0 or newer and quantizes the frozen weights to bfloat16 before packing.
Backward computes input gradients from the packed weights in bounded chunks,
adding kernel work compared with NF4's dedicated backward.
For packed gated-MoE experts, target_parameters can include both
experts.gate_up_proj and experts.down_proj to NF4-quantize their frozen
3-D weight banks. This works with packed expert modules that expose the usual
top-k hidden states, indices, and weights and use paired gate/up and down
projections, including Mixtral, DeepSeek V3, and Qwen3.5. The expert loop reads
only routed 2-D matrices while training per-expert LoRA factors. The expert
NF4 matmuls use TorchAO; nf4_kernel selects the backend for ordinary 2-D
linear modules. Use lora_dropout: 0 for packed expert targets. The parameter
bank path requires both expert targets and currently supports NF4. Its
data-dependent routing loop stays eager when the rest of the model is compiled.
Native adapter checkpoints preserve the quantization settings and quantize a
dense base again when loaded. Existing native adapter checkpoints still load;
new native checkpoints use Loralei filenames. PEFT adapter exports contain the
adapter factors and require the caller to prepare the desired base precision
separately.
Merging NF4-quantized packed expert adapters is not supported because it
would restore the full dense expert banks. Merging ordinary quantized linear
weights dequantizes those bases; unmerge_lora restores their packed weights.
Load int4 checkpoints with the base on CUDA, or pass
quantization_device="cuda" to load_pretrained to move selected linears
there during loading. Avoid dtype casts on the model after QLoRA injection,
since TorchAO's packed tensors do not support changing weight dtype.
Use load_pretrained, merge_lora, and unmerge_lora for loading and
reversible inference-time weight folding. To export a normal Transformers
model after merging, use merge_and_unload_lora(model) (or
merge_lora(model, unload=True)) before calling the model's
save_pretrained. PEFT is optional; loralei itself does not import it at
runtime.
YAML fine-tuning
The repository includes a small LoRA supervised-fine-tuning loop. It loads a
dataset with datasets.load_dataset, applies the tokenizer's chat template,
and writes an adapter-only checkpoint:
uv run loralei-finetune examples/finetune.yaml
W&B logging is opt-in. Install the extra and add a wandb block to the YAML
configuration:
uv sync --extra wandb
wandb:
project: loralei
name: qwen-sft
The runner logs loss, learning rate, token throughput, peak CUDA tensor
allocation and allocator reservation across model setup and training, gradient
norm (when clipping is enabled), and skipped optimizer steps at
training.logging_steps.
wandb: false leaves
tracking disabled; training.report_to: wandb is also supported.
Tokenization is performed once with datasets.map and reused from the Hugging
Face dataset cache. For CPU-heavy input pipelines, set
training.preprocessing_num_workers and training.dataloader_num_workers.
CUDA runs can also use training.pin_memory: true for asynchronous host-to-
device copies. On supported Transformers models, select fused attention with
training.attn_implementation: sdpa (or flash_attention_2 when installed).
training.torch_compile: true enables PyTorch compilation; benchmark it with
fixed sequence lengths or packing enabled. Set
training.pad_to_max_length: true to keep batch shapes fixed at
training.max_seq_length, which avoids extra compiler graphs when packed
sequences have slightly different lengths.
With TorchAO QLoRA, the runner compiles decoder layers separately and
selectively checkpoints them, preserving expensive matrix products while
recomputing cheaper operations. With TileLang QLoRA, it compiles the whole model
without activation checkpointing because the fused NF4 matmuls keep the base
weights packed. The TorchAO mode requires a decoder with a layers
ModuleList; the runner reports an error if it is missing.
The runner clears temporary cached CUDA allocations before QLoRA training begins.
Peak memory still depends on the model and batch, so measure it before sizing
a larger run.
For a memory-limited GPU, enable training.gradient_checkpointing to trade
some recomputation for lower activation memory. Compiled QLoRA already
checkpoints its decoder layers, so that setting adds no further layer
checkpointing in this mode.
Qwen3.5 uses a hybrid attention stack. Run the PoLoRA example with its optional attention and Liger kernels:
uv run --extra linear-attention --extra liger loralei-finetune examples/qwen35-9b-polora.yaml
training.fused_shared_input_lora: true enables the common self-attention Q/K/V
and MLP gate/up groups. For other architectures, provide a list of groups with
parent_name and projection_names, as this example does for its linear
attention projections. Shared-input groups require zero LoRA dropout and do
not support torch.compile. training.liger_kernel: true uses the Transformers
Liger integration to select the patch for the loaded model.
The example uses the measured high-throughput settings: rank-16 all-linear PoLoRA, BF16, packed 2048-token sequences, a per-device batch size of two, configured shared-input LoRA projection groups, Liger RMSNorm and SwiGLU, and Liger's fused linear cross entropy loss. The output head is frozen, and compilation and gradient checkpointing are disabled. This profile needs substantial device memory; reduce the batch size or enable checkpointing if it does not fit. Cut Cross Entropy remains available as a memory-conscious loss option below.
For large-vocabulary models, the finetuning loop can use Apple's memory-efficient Cut Cross Entropy implementation. Install the optional dependency and select it in the training configuration:
uv sync --extra cce
training:
loss: cut_cross_entropy # `cce` is also accepted
The runner passes the backbone's final hidden states and output-head weights
directly to linear_cross_entropy, so it does not materialize the full logits
tensor. The package chooses its CCE or torch_compile implementation for the
current platform; set training.cut_cross_entropy.impl to override it.
The dataset can contain messages/conversations, a plain text field, or
instruction/response fields. Set data.messages_field or data.text_field
when the defaults do not match. The training optimizer can be adamw,
skipstepadamw, or polora; the latter uses the paired LoRA factors directly.
Set optimizer.kwargs.scale_invariant: true to account for the adapter's
lora_alpha/r multiplier (or lora_alpha/sqrt(r) with RSLoRA) in PoLoRA's
update budget. See examples/finetune.yaml for the intentionally small
configuration surface.
Set training.packing: true to enable best-fit decreasing (BFD) sample
packing. Packing tokenizes without truncation, drops complete samples longer
than training.max_seq_length, and emits a warning for the dropped samples.
For causal language models with packed right-padding, set
training.omit_attention_mask_when_packed: true to omit the mask from model
input and its device transfer. The padding is a trailing suffix, so causal
positions with training targets cannot attend to it; leave this option off for
models whose padding semantics require an explicit mask.
Kaitan/dantess pretokenized Parquet files are also supported. Point dataset
at the file; rows with input_ids, attention_mask, and labels are detected
automatically, so their existing -100 loss masks are preserved:
dataset:
path: /data/tokenized.parquet
training:
max_seq_length: 8192