Skip to content
Octopus Research Institute
Status: ReleasedData ownership: Institute-ownedLicence: CC BY 4.0

AWS L4 (24GB) full-grid benchmark

A one-shot capability snapshot of the engines on a rented AWS L4 24GB instance.

Purpose

A capability snapshot, not a tuned or repeated benchmark.

Source

A single one-shot benchmarking session on a rented L4.

Composition

Downloadable CSV — one-shot rows of (engine, task, metric, value, peak_vram_gb, note): the sovereign text-CUDA Qwen3-30B-A3B q4 full-48 resident decode result (~23 tok/s, 19.73 GB), plus image/video builds that were sovereign but unrunnable on this run.

Known biases

  • The text-cuda decode streamed weights from host RAM (not GPU-resident) and is honestly slower than the resident path.

Limitations

  • Single run, no variance.
  • Image/video builds were sovereign but unrunnable pending portable weight layouts.

Access conditions

The measured snapshot is available for download below, under CC-BY-4.0.

Data preview
enginetaskmetricvaluepeak_vram_gbnote
text-cuda-moeQwen3-30B-A3B q4 full-48 resident decodetok_per_s2319.73sovereign (ldd = libcudart only; no cuBLAS/cuDNN/cuTLASS); weights 18.375 GB; parity logits_cosine=1.0 argmax 12/12
image-cudaSDXL 1024px full-stepstatusnot-runsovereign build but unrunnable on this run pending portable weight layouts
video-cudaWan2.1 full-framestatusnot-runsovereign build but unrunnable on this run pending portable weight layouts
First 3 of 3 rows.
Download · CC BY 4.0