- Python 99.4%
- Shell 0.4%
- Makefile 0.2%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
|
Some checks failed
Main / Build (push) Has been cancelled
Main / Lint (push) Has been cancelled
Main / Style (push) Has been cancelled
Main / Test (push) Has been cancelled
Main / Type check (push) Has been cancelled
Main / Lint (min Python) (push) Has been cancelled
Main / Release (push) Has been cancelled
|
||
| .github | ||
| src | ||
| .gitignore | ||
| CHANGELOG.md | ||
| LICENSE | ||
| Makefile | ||
| pyproject.toml | ||
| README.md | ||
OLMo-in-loop-evals
Code for in-loop evaluation tasks used by the OLMo training team.
Installation
pip install ai2-olmo-eval
C4 English CE and BPB
c4_en evaluates a deterministic prefix of the English validation split of
allenai/c4. It streams the source
and prepares the first 1,000 documents by default, without downloading the full
corpus. Use C4English(tokenizer, num_samples=100, model_ctx_len=2048) to change
the document limit.
import torch
from torch.utils.data import DataLoader
from olmo_eval import C4Metric, build_task
task = build_task("c4_en", tokenizer, model_ctx_len=2048)
loader = DataLoader(task, batch_size=8, collate_fn=task.collate_fn)
metric = C4Metric().to(device)
model.eval()
with torch.no_grad():
for batch in loader:
batch = {key: value.to(device) for key, value in batch.items()}
logits = model(
input_ids=batch["input_ids"], attention_mask=batch["attention_mask"]
).logits # Adapt this call to your model's forward interface.
metric.update(batch, logits)
results = metric.compute() # {"ce_loss": nats/token, "bpb": bits/UTF-8 byte}
Use C4Metric for this task. Labels are already shifted to align with logits;
padding is masked with -100. Every text token is scored once, starting from
BOS (or EOS when BOS is unavailable), with no EOS target added. Long documents
are split into chunks of at most model_ctx_len predictions; adjacent chunks
share one context token, and earlier context is discarded. Empty documents are
skipped. CE is total negative log likelihood divided by total scored tokens;
BPB is total negative log likelihood divided by ln(2) and original text bytes.
This uses corpus totals rather than averaging per-document ratios.
Evaluate the entire selected dataset: each document's byte count is recorded on its last chunk, so stopping partway through a document gives an invalid BPB. Distributed metrics sum across ranks; partition samples without duplicates (standard distributed sampler padding would bias the totals). Hugging Face downloads and caches source data on first use.
Paloma CE and BPB
paloma_bpb streams the val split of all 16 sources in
allenai/paloma, preparing the
first 1,000 documents per source. Each source can also be evaluated separately
with paloma_{source}_bpb (for example, paloma_c4_en_bpb or paloma_ptb_bpb).
The exported PALOMA_SOURCES tuple lists the source names.
from olmo_eval import Paloma, PalomaMetric, build_task
task = build_task("paloma_bpb", tokenizer, model_ctx_len=2048)
metric = PalomaMetric().to(device)
# Use the same DataLoader and evaluation loop as the C4 example above.
# Select sources, split, and document limit directly:
task = Paloma(tokenizer, dataset_name=["ptb", "wikitext_103"],
split="test", num_samples=100, model_ctx_len=2048)
Paloma is gated: accept its access terms on Hugging Face and authenticate with
a saved Hugging Face token or HF_TOKEN. An explicit token can also be passed
to Paloma; tokens are not included in task registrations.
PalomaMetric uses the same corpus CE/BPB, chunking, and complete-pass requirements
as C4 above. Results are weighted by original UTF-8 bytes across the selected
documents, rather than averaged across sources or domains. Sampling is a
deterministic prefix per source and does not guarantee coverage of every domain.
This in-loop task does not implement Paloma's full standardized benchmark protocol.
Release process
Steps
-
Update the version in
src/olmo_eval/version.py. -
Run the release script:
./src/scripts/release.shThis will commit the changes to the CHANGELOG and
version.pyfiles and then create a new tag in git which will trigger a workflow on GitHub Actions that handles the rest.
Fixing a failed release
If for some reason the GitHub Actions release workflow failed with an error that needs to be fixed, you'll have to delete the tag on GitHub. Once you've pushed a fix you can simply repeat the steps above.
Babilong
Babilong is available as zero-shot
BPB tasks, with one option per context length:
babilong_{length}_bpb_0shot, where length is 0k, 1k, 2k, 4k, 8k,
16k, 32k, 64k, 128k, 256k, 512k, 1M, or 10M (case-sensitive).
Each option includes all QAs available at that length: QA1–20 for 0k, QA1–10
for 1k through 1M, and QA1–5 for 10M. Context lengths are never combined.
from olmo_eval import ICLMetric, build_task
# Supply your model's tokenizer.
task = build_task("babilong_4k_bpb_0shot", tokenizer, model_ctx_len=8192)
metric = ICLMetric(metric_type=task.metric_type)
# After metric.update(batch, logits), read metric.compute()["bpb_v2"].
Prompts use input + "\nQuestion: " + question + "\nAnswer:"; only the gold
" " + target continuation is scored. bpb_v2 includes the leading space in
the UTF-8 byte count and averages BPB across examples. Data downloads on first
use and is cached by Hugging Face; it is not bundled with this package. Large
lengths require substantial download space and host memory. Set model_ctx_len
to accommodate the tokenized passage, question, and answer; as with other tasks,
queries exceeding that limit are truncated from the left.
Phonebook
Synthetic retrieval inspired by Jelassi et al., §4.3
is available as phonebook_{length}_bpb_2shot, with the same length labels as
Babilong. Each prompt contains random alphabetic names with numbers formatted as
609-323-7777, two answered lookup examples, and a query for a different entry.
Only the gold number (with a leading space) is scored with the existing BPB metric.
This is a BPB adaptation, rather than the paper's generated-answer accuracy.
task = build_task("phonebook_4k_bpb_2shot", tokenizer, model_ctx_len=8192)
Books are generated locally and deterministically (seed 42, 100 examples by default).
Positive lengths budget the book in tokenizer tokens, using 1024 tokens per k
and 1024² per M, rounded down to whole entries; 0k uses three entries.
The demonstrations, question, and answer add tokens beyond that budget. Set
model_ctx_len accordingly to retain the entire book; the standard left truncation
applies otherwise. Large configurations require substantial generation time and
host memory. For smaller runs or alternate seeds, instantiate
olmo_eval.tasks.Phonebook(tokenizer, dataset_name="4k", num_samples=10, seed=42, model_ctx_len=8192) directly. Validation and test use distinct deterministic seeds.