Skip to content
Octopus Research Institute

Octopus Research Institute

Researching accountable, accessible and sovereign AI.

Octopus Research Institute develops and evaluates the technical foundations required for responsible AI systems in the real world.

Why we exist

Technical capability alone is not enough.

A system can be capable and still be unaccountable, opaque, inaccessible, or dependent on a single provider. We investigate the technical foundations that accountable, accessible, governable and sovereign AI actually require — and we report what we find honestly, including negative results and limitations.

About the Institute
Research areas

What we study

Each area carries an explicit status. A research direction is never presented as a finished product.

  • Governed AI SystemsStatus: Active research

    Agent governance, approval boundaries, human-in-the-loop control, policy enforcement and reversible, fail-closed execution.

  • Evidence, Audit and ReplayStatus: Active research

    Evidence ledgers, tamper-evident records, decision replay, provenance and runtime observability for AI execution.

  • Graph ReasoningStatus: Experimental

    Graph-based reasoning runtimes: planners, transaction logs, evidence-aware inference and constraint enforcement over structured domain knowledge.

  • Sovereign and Local AIStatus: Active research

    Local and on-device inference, model optimisation, edge deployment, privacy-aware inference and multi-provider portability.

  • Privacy-Preserving AIStatus: Active research

    Data minimisation, sensitive-data detection and redaction, controlled model access, data residency and residual-risk evaluation.

  • Healthcare AIStatus: Exploring

    Non-diagnostic, operational clinical-documentation support with human oversight, evidence and governance. No clinical validation is claimed.

  • Accessible interaction, communication support, vision and hearing accessibility, and inclusive AI design.

  • Auslan and Multimodal AIStatus: Exploring

    Camera-based recognition of Auslan (Australian Sign Language) explored as multimodal representation of hands, body, face and timing — not gesture classification. Community-informed and early-stage.

  • Model EvaluationStatus: Active research

    Accuracy, robustness, failure analysis, domain transfer, generalisation, reproducibility and transparent reporting of limitations.

  • Responsible AIStatus: Active research

    Human accountability, governance, safety boundaries, community impact, responsible data use and evidence-backed claims.

  • World ModelsStatus: Experimental

    Deep-perception world models (DAWM): world-model experiments, phase diagnostics, failure analysis and current-belief updates.

Featured research

A selection of active and exploratory programmes. Statuses and limitations are shown on every page.

Research
  • Status: Exploring
    Auslan and Multimodal Accessibility

    Early-stage, community-informed exploration of camera-based Auslan recognition as multimodal representation — not gesture classification, not interpreter replacement.

    • Auslan and Multimodal AI
    • Accessibility Technologies
    • Responsible AI
  • Status: Active research
    Governed Agent Systems

    A single controlled execution boundary for AI agents, with human approval for irreversible actions and policy-based access to tools.

    • Governed AI Systems
    • Evidence, Audit and Replay
    • Responsible AI
  • Status: Active research
    Evidence and Workstate

    Store-untrusting, fail-closed design where verdicts are captured as evidence, contracts are pinned, and work state is replayable and tamper-evident.

    • Evidence, Audit and Replay
    • Responsible AI
  • Status: Experimental
    Graph Reasoning Runtime

    An experimental reasoning runtime separating a planner, a transaction log and an evidence ledger, with constraint enforcement over structured domain knowledge.

    • Graph Reasoning
    • Evidence, Audit and Replay
Latest publications

Recent outputs

All publications
  • AP-2026-0011Architecture paperPeer review: Not peer reviewedEvidence: Experimental
    Apple-Silicon-friendly LLM architecture: substrate laws reverse-engineered from a model bake-off

    Five structurally distinct LLM families were each hand-ported into one custom Apple-GPU decode engine and profiled until each broke under single-stream 4-bit decode on a single M1 Ultra; only one satisfied all five rules. The result is a spec of five rules for fast single-stream decode on unified-memory Apple Silicon: low active params per token, vector-load-aligned quantization, experts big enough to fill the GPU at batch one, uniform attention, and a fusion-friendly non-hybrid layout. Meta-finding: architecture identity does not predict runtime cost; the rules are necessary, not sufficient.

    Published 2026-06-15 · Version 3

  • RN-2026-0012Research notePeer review: Not peer reviewedEvidence: Hypothesis
    Gemma4-26B-A4B on Apple Silicon: a drop-in isomorphism falsified (NO-GO), with a conditional new-port ceiling

    A read-only static deep-dive falsifying that Gemma4-26B-A4B is structurally isomorphic to a mature Qwen3-30B-A3B MoE single-stream inference path on Apple Silicon, able to drop in and inherit its throughput. Ground-truth bf16 shapes break it on four axes: a stacked fused-expert layout, a dense-MLP plus routed-MoE hybrid, GeGLU not SwiGLU gating, and a heterogeneous sliding/full attention geometry with novel scalars. Verdict: NO-GO for a config-swap; only the MoE router carries over. A purpose-built port has a conditional throughput ceiling as a range. Nothing was run; figures are estimates.

    Published 2026-06-15 · Version 3

  • RN-2026-0013Research notePeer review: Not peer reviewedEvidence: Experimental
    A falsify-first root cause for a concurrent 4-bit decode crash: batch-composition KV-pool wipe trips a re-seed precondition

    Falsify-first root cause for a fatal trap that killed a single-resident 4-bit ~30B sparse-MoE decode daemon under concurrent decode on one Apple Silicon machine. An orchestration-free reproducer isolates the trigger: concurrent chat and interleaved chat-plus-embed crash, while embeddings-only and sequential traffic survive, refuting an "embeddings poison decode" guess. The mechanism is a batch-composition-change wipe of resident KV pools tripping a re-seed precondition no steady-state decode meets. Serializing GPU decode prevents it but gives no speedup; the per-stream fix is unimplemented.

    Published 2026-06-15 · Version 3

Open research & reproducibility

Evidence before claims

Where practical, we share code, evaluation setups and honest limitations so results can be interpreted and reproduced. We distinguish what is known, observed, inferred, uncertain and planned. Datasets and models carry explicit licences — and a clear note when a licence is not yet confirmed.

Responsible research
The Octopus ecosystem

Three organisations, three distinct roles

Octopus Core builds the infrastructure. Octopus Research Institute advances the knowledge. Octopus Foundation ensures that progress serves people. The Institute is part of the wider Octopus ecosystem, but it is not a product-marketing department for Octopus Core.

Collaboration

Work with us

We welcome researchers, universities, community organisations, accessibility experts, healthcare researchers, engineers and industry research teams.

Start a collaboration

We do not imply formal university affiliations. Collaborations are described only where configured and verified.