π BenchmarkingΒΆ
This page explains how to collect and interpret throughput and efficiency metrics for MMIRAGE pipeline runs.
OverviewΒΆ
MMIRAGE includes built-in benchmarking that measures:
wall-clock runtime
throughput (rows per second)
token generation speed per GPU
GPU utilisation
estimated compute cost (GPU-days per billion tokens)
Metrics are collected on each compute node during processing and recorded in the shard state directory (e.g. <state_dir>/shard_<id>/status.json under the stats key).
Enabling benchmarkingΒΆ
Pass --stats to any command that submits or runs shards:
# Local run
mmirage run --config configs/config.yaml --stats
# SLURM array submission
mmirage submit --config configs/config.yaml --stats
# Retry failed shards with stats
mmirage retry --config configs/config.yaml --stats
When --stats is set, MMIRAGE polls GPU utilisation at a fixed interval
during processing and records token counts reported by the SGLang engine.
Viewing collected metricsΒΆ
After a run (or after shards complete), aggregate and display the metrics:
mmirage stats --config configs/config.yaml
This reads per-shard stats files from the state directory and prints a summary to stdout in JSON format:
{
"per_shard": [],
"aggregate": {
"wall_clock_runtime_seconds": 3247.8,
"overall_throughput_rows_per_sec": 12.4,
"tokens_per_sec_per_gpu": 1850.3,
"gpu_days_per_billion_tokens": 0.63,
"mean_gpu_util_pct": 91.2
}
}
Metrics referenceΒΆ
Metric |
Unit |
Description |
|---|---|---|
|
seconds |
Wall-clock time from first shard start to last shard finish |
|
seconds |
Sum of per-shard runtimes (useful even when shards run in parallel) |
|
rows/s |
Dataset samples processed per second |
|
tokens/s/GPU |
Token generation throughput normalised per GPU |
|
GPU-days |
Compute cost estimate: GPU-days needed to generate one billion tokens |
|
% |
Average GPU utilisation during processing |
Interpreting resultsΒΆ
tokens_per_sec_per_gpu is the primary efficiency indicator.
Higher is better. Typical values for a well-configured SGLang engine on modern
hardware range from 1 000 to 5 000+ tokens/s/GPU depending on model size and
batch size.
gpu_days_per_billion_tokens normalises cost across different GPU counts
and runtimes, making it easy to compare runs with different configurations.
mean_gpu_util_pct should ideally stay above 80 %.
Low values (< 60 %) may indicate:
batch_sizeis too small relative to the modelβs throughputheavy I/O overhead between batches
slow JMESPath extraction or prompt rendering
Tuning for throughputΒΆ
If benchmarking reveals low efficiency, try:
Symptom |
Remedy |
|---|---|
Low GPU util, high I/O wait |
Increase |
OOM errors at large batch size |
Reduce |
Low tokens/s/GPU |
Tune |
High variance across shards |
Check dataset skew (very long samples in one shard) |
DataTrove benchmarkΒΆ
MMIRAGE includes a reference config for the DataTrove benchmark dataset:
configs/config_benchmark_datatrove.yaml
Use it to establish a baseline throughput on your hardware and compare against published numbers.
See alsoΒΆ
CLI Reference β
--statsflag andmmirage statscommand detailsSLURM & Cluster Deployment β collecting stats in SLURM mode
Configuration Reference β
extra_engine_argsfor SGLang tuning