• Joined on 2026-04-25
Updated 2026-10-09 23:39:21 -04:00
Training to converge early, with as little shown as possible. - Optimizers for deep learning.
Updated 2026-10-09 23:23:17 -04:00
a collection of reusable GPU kernels built with TileLang
Updated 2026-10-04 15:38:43 -04:00
Updated 2026-10-04 12:00:13 -04:00
A small LoRA implementation designed for training, integrated with Transformers models.
Updated 2026-09-25 15:51:27 -04:00
Pretokenization scripts
Updated 2026-09-20 21:52:21 -04:00
Updated 2026-08-05 13:19:15 -04:00
Global workspace interpretability
Updated 2026-07-27 21:27:40 -04:00
Updated 2026-04-27 22:38:52 -04:00
Updated 2026-04-27 22:33:18 -04:00