Status: ReleasedData ownership: Institute-ownedLicence: CC BY 4.0
AWS L4 (24GB) full-grid benchmark
A one-shot capability snapshot of the engines on a rented AWS L4 24GB instance.
Purpose
A capability snapshot, not a tuned or repeated benchmark.
Source
A single one-shot benchmarking session on a rented L4.
Composition
Downloadable CSV — one-shot rows of (engine, task, metric, value, peak_vram_gb, note): the sovereign text-CUDA Qwen3-30B-A3B q4 full-48 resident decode result (~23 tok/s, 19.73 GB), plus image/video builds that were sovereign but unrunnable on this run.
Known biases
- The text-cuda decode streamed weights from host RAM (not GPU-resident) and is honestly slower than the resident path.
Limitations
- Single run, no variance.
- Image/video builds were sovereign but unrunnable pending portable weight layouts.
Access conditions
The measured snapshot is available for download below, under CC-BY-4.0.
| engine | task | metric | value | peak_vram_gb | note |
|---|---|---|---|---|---|
| text-cuda-moe | Qwen3-30B-A3B q4 full-48 resident decode | tok_per_s | 23 | 19.73 | sovereign (ldd = libcudart only; no cuBLAS/cuDNN/cuTLASS); weights 18.375 GB; parity logits_cosine=1.0 argmax 12/12 |
| image-cuda | SDXL 1024px full-step | status | not-run | sovereign build but unrunnable on this run pending portable weight layouts | |
| video-cuda | Wan2.1 full-frame | status | not-run | sovereign build but unrunnable on this run pending portable weight layouts |
