Status: Internal evaluationEvaluation type: Private evaluation
Local vs Hosted Inference on the Local Inference Task Suite
A small internal comparison of local and hosted inference across summarisation, extraction and classification tasks on fixed hardware.
Research question
For which representative tasks is local inference already sufficient on commodity hardware?
Methodology
Each task was run under a fixed prompt and decoding configuration; local and hosted settings used the same inputs. Results are illustrative of our setup, not a ranking.
Metric definitions
- task-success
- Proportion of tasks where output met a pre-registered rubric, judged by the authors. Subjective; see limitations.
Results
| Model | Dataset | Metric | Value | Sample size |
|---|---|---|---|---|
| local (illustrative) | local-inference-task-suite v0.2 | task-success | sufficient on the majority of extraction/classification tasks | few dozen tasks |
| hosted (illustrative) | local-inference-task-suite v0.2 | task-success | preferred on the longer summarisation tasks | few dozen tasks |
Environment
Single workstation; details recorded in the evaluation harness.
Hardware
Commodity desktop GPU (specific model recorded internally).
Limitations
- Sample is too small for statistical significance; no confidence intervals are claimed.
- Success judging is subjective and by the authors.
- Hardware-specific; not a general ranking of local vs hosted inference.
Reproducibility
Tasks, prompts and configuration are available on request via the evaluation harness. Exact hosted-provider behaviour may drift over time.
