Skip to content
Octopus Research Institute
Status: ExperimentalSovereign and Local AIModel Evaluation

Apple Silicon Inference Runtimes

Kernel design, memory movement and bandwidth ceilings for local inference on Apple Silicon and other accelerators — consolidated from earlier Octoryn Research work.

Abstract

This programme studies the hardware-specific limits of local inference: kernel design, memory movement, barrier placement and the bandwidth ceilings that shape real bottlenecks. It is experimental and evidence-first; several core questions remain open (see Open Questions).

Problem & motivation

Runtime performance on accelerators is often attributed to raw compute, but bandwidth ceilings and synchronisation limits frequently dominate — and are poorly characterised for Apple GPUs.

Research questions

  • Can Apple GPUs support persistent-kernel global synchronisation safely?
  • Where are the occupancy cliffs that change the shape of the bottleneck?

Methods

  • Barrier/token replay harnesses and audits
  • Bandwidth trace comparison across hardware
  • Forward-progress watchdogs under pathological occupancy

Observations

Present-tense observations from internal work. These are not validated results and have not been peer reviewed unless a linked publication says so.
  • Observed: bandwidth ceilings reshape bottlenecks more than raw compute deltas suggest.
  • Observed: compact "megakernel-lite" designs cut launch overhead but expose scheduling fragility.

Limitations

  • No production-safe proof of global synchronisation yet.
  • Findings are hardware-specific and internal; not a general benchmark.

Disclosures

Funding
Hardware and infrastructure provided by Octopus Core Pty Ltd / Octoryn.
Conflicts of interest
Octoryn ships commercial inference tooling; findings are reported independently.