No description
  • Python 99.4%
  • Shell 0.4%
  • Makefile 0.2%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
fizz fb5b984df8
Some checks failed
Main / Build (push) Has been cancelled
Main / Lint (push) Has been cancelled
Main / Style (push) Has been cancelled
Main / Test (push) Has been cancelled
Main / Type check (push) Has been cancelled
Main / Lint (min Python) (push) Has been cancelled
Main / Release (push) Has been cancelled
paloma
2026-10-04 12:00:03 -04:00
.github Add 18 new BPB tasks for basic_skills, gen, and science/medical categ… (#20) 2026-02-03 15:31:29 -08:00
src paloma 2026-10-04 12:00:03 -04:00
.gitignore Pinning dependencies to known good versions. (#15) 2025-08-12 15:44:35 -07:00
CHANGELOG.md (chore) prepare for release v0.9.0 2026-02-03 15:51:10 -08:00
LICENSE add license 2024-10-28 14:10:16 -07:00
Makefile initial commit 2024-10-28 13:23:30 -07:00
pyproject.toml Add 18 new BPB tasks for basic_skills, gen, and science/medical categ… (#20) 2026-02-03 15:31:29 -08:00
README.md paloma 2026-10-04 12:00:03 -04:00

OLMo-in-loop-evals

Code for in-loop evaluation tasks used by the OLMo training team.

Installation

pip install ai2-olmo-eval

C4 English CE and BPB

c4_en evaluates a deterministic prefix of the English validation split of allenai/c4. It streams the source and prepares the first 1,000 documents by default, without downloading the full corpus. Use C4English(tokenizer, num_samples=100, model_ctx_len=2048) to change the document limit.

import torch
from torch.utils.data import DataLoader
from olmo_eval import C4Metric, build_task

task = build_task("c4_en", tokenizer, model_ctx_len=2048)
loader = DataLoader(task, batch_size=8, collate_fn=task.collate_fn)
metric = C4Metric().to(device)
model.eval()
with torch.no_grad():
    for batch in loader:
        batch = {key: value.to(device) for key, value in batch.items()}
        logits = model(
            input_ids=batch["input_ids"], attention_mask=batch["attention_mask"]
        ).logits  # Adapt this call to your model's forward interface.
        metric.update(batch, logits)
results = metric.compute()  # {"ce_loss": nats/token, "bpb": bits/UTF-8 byte}

Use C4Metric for this task. Labels are already shifted to align with logits; padding is masked with -100. Every text token is scored once, starting from BOS (or EOS when BOS is unavailable), with no EOS target added. Long documents are split into chunks of at most model_ctx_len predictions; adjacent chunks share one context token, and earlier context is discarded. Empty documents are skipped. CE is total negative log likelihood divided by total scored tokens; BPB is total negative log likelihood divided by ln(2) and original text bytes. This uses corpus totals rather than averaging per-document ratios.

Evaluate the entire selected dataset: each document's byte count is recorded on its last chunk, so stopping partway through a document gives an invalid BPB. Distributed metrics sum across ranks; partition samples without duplicates (standard distributed sampler padding would bias the totals). Hugging Face downloads and caches source data on first use.

Paloma CE and BPB

paloma_bpb streams the val split of all 16 sources in allenai/paloma, preparing the first 1,000 documents per source. Each source can also be evaluated separately with paloma_{source}_bpb (for example, paloma_c4_en_bpb or paloma_ptb_bpb). The exported PALOMA_SOURCES tuple lists the source names.

from olmo_eval import Paloma, PalomaMetric, build_task

task = build_task("paloma_bpb", tokenizer, model_ctx_len=2048)
metric = PalomaMetric().to(device)
# Use the same DataLoader and evaluation loop as the C4 example above.

# Select sources, split, and document limit directly:
task = Paloma(tokenizer, dataset_name=["ptb", "wikitext_103"],
              split="test", num_samples=100, model_ctx_len=2048)

Paloma is gated: accept its access terms on Hugging Face and authenticate with a saved Hugging Face token or HF_TOKEN. An explicit token can also be passed to Paloma; tokens are not included in task registrations.

PalomaMetric uses the same corpus CE/BPB, chunking, and complete-pass requirements as C4 above. Results are weighted by original UTF-8 bytes across the selected documents, rather than averaged across sources or domains. Sampling is a deterministic prefix per source and does not guarantee coverage of every domain. This in-loop task does not implement Paloma's full standardized benchmark protocol.

Release process

Steps

  1. Update the version in src/olmo_eval/version.py.

  2. Run the release script:

    ./src/scripts/release.sh
    

    This will commit the changes to the CHANGELOG and version.py files and then create a new tag in git which will trigger a workflow on GitHub Actions that handles the rest.

Fixing a failed release

If for some reason the GitHub Actions release workflow failed with an error that needs to be fixed, you'll have to delete the tag on GitHub. Once you've pushed a fix you can simply repeat the steps above.

Babilong

Babilong is available as zero-shot BPB tasks, with one option per context length: babilong_{length}_bpb_0shot, where length is 0k, 1k, 2k, 4k, 8k, 16k, 32k, 64k, 128k, 256k, 512k, 1M, or 10M (case-sensitive). Each option includes all QAs available at that length: QA1–20 for 0k, QA1–10 for 1k through 1M, and QA1–5 for 10M. Context lengths are never combined.

from olmo_eval import ICLMetric, build_task

# Supply your model's tokenizer.
task = build_task("babilong_4k_bpb_0shot", tokenizer, model_ctx_len=8192)
metric = ICLMetric(metric_type=task.metric_type)
# After metric.update(batch, logits), read metric.compute()["bpb_v2"].

Prompts use input + "\nQuestion: " + question + "\nAnswer:"; only the gold " " + target continuation is scored. bpb_v2 includes the leading space in the UTF-8 byte count and averages BPB across examples. Data downloads on first use and is cached by Hugging Face; it is not bundled with this package. Large lengths require substantial download space and host memory. Set model_ctx_len to accommodate the tokenized passage, question, and answer; as with other tasks, queries exceeding that limit are truncated from the left.

Phonebook

Synthetic retrieval inspired by Jelassi et al., §4.3 is available as phonebook_{length}_bpb_2shot, with the same length labels as Babilong. Each prompt contains random alphabetic names with numbers formatted as 609-323-7777, two answered lookup examples, and a query for a different entry. Only the gold number (with a leading space) is scored with the existing BPB metric. This is a BPB adaptation, rather than the paper's generated-answer accuracy.

task = build_task("phonebook_4k_bpb_2shot", tokenizer, model_ctx_len=8192)

Books are generated locally and deterministically (seed 42, 100 examples by default). Positive lengths budget the book in tokenizer tokens, using 1024 tokens per k and 1024² per M, rounded down to whole entries; 0k uses three entries. The demonstrations, question, and answer add tokens beyond that budget. Set model_ctx_len accordingly to retain the entire book; the standard left truncation applies otherwise. Large configurations require substantial generation time and host memory. For smaller runs or alternate seeds, instantiate olmo_eval.tasks.Phonebook(tokenizer, dataset_name="4k", num_samples=10, seed=42, model_ctx_len=8192) directly. Validation and test use distinct deterministic seeds.