Status: ExperimentalSovereign and Local AIModel Evaluation
Apple Silicon Inference Runtimes
Kernel design, memory movement and bandwidth ceilings for local inference on Apple Silicon and other accelerators — consolidated from earlier Octoryn Research work.
Abstract
This programme studies the hardware-specific limits of local inference: kernel design, memory movement, barrier placement and the bandwidth ceilings that shape real bottlenecks. It is experimental and evidence-first; several core questions remain open (see Open Questions).
Problem & motivation
Runtime performance on accelerators is often attributed to raw compute, but bandwidth ceilings and synchronisation limits frequently dominate — and are poorly characterised for Apple GPUs.
Research questions
- Can Apple GPUs support persistent-kernel global synchronisation safely?
- Where are the occupancy cliffs that change the shape of the bottleneck?
Methods
- Barrier/token replay harnesses and audits
- Bandwidth trace comparison across hardware
- Forward-progress watchdogs under pathological occupancy
Observations
Present-tense observations from internal work. These are not validated results and have not been peer reviewed unless a linked publication says so.
- Observed: bandwidth ceilings reshape bottlenecks more than raw compute deltas suggest.
- Observed: compact "megakernel-lite" designs cut launch overhead but expose scheduling fragility.
Limitations
- No production-safe proof of global synchronisation yet.
- Findings are hardware-specific and internal; not a general benchmark.
Disclosures
- Funding
- Hardware and infrastructure provided by Octopus Core Pty Ltd / Octoryn.
- Conflicts of interest
- Octoryn ships commercial inference tooling; findings are reported independently.
