Controlled Symbolic Rendering-Invariance via Independent Learned Non-Text Parsers
Ran Tao (Octoryn Research)
Abstract
A controlled symbolic rendering-invariance study, rebuilt after an earlier pass overstated shallow probes. Its central contribution is failure separability: each error class is observable through its own failing control. The work advances to two independent learned non-text parsers (graph and table) mapping distinct renderings of one semantic state to a shared canonical state, which drives downstream scope/candidate execution and state-level contradiction detection with localization. A bounded review closed the cell narrowly, excluding native-multimodal and open-world claims.
Summary
This work was rebuilt from scratch after an earlier, too-fast pass overstated shallow probes. The rebuild reframes the work as a controlled symbolic rendering-invariance cell with explicit failure-attribution scaffolding, then advances it through learned-parser subphases to a narrow, honest closure.
The closing claim is deliberately narrow: controlled learned symbolic rendering invariance under a shared canonical semantic-state scaffold — for the tested graph/table symbolic renderers and bounded controlled-shift families. It is explicitly NOT a native-multimodal, open-world, or learned-ontology claim.
Motivation and reset
The rebuild's first artifact was a review checkpoint, not a closure. The status correction is itself the finding: the prior pass had been accepted too quickly. The checkpoint re-established the claim under a controlled symbolic small-cell definition:
- the same underlying semantic state is presented through distinct controlled renderings (text / graph / table / symbolic diagram);
- a canonical state is recovered across renderings;
- the recovered state supports downstream scope/candidate operations;
- state-level contradiction is observable; and
- surface shortcuts, corrupted states, mismatched state pairs, and missing explicit-alignment coverage all FAIL as controls.
This was a rendering-invariance review checkpoint over engineered controlled renderers/parsers — valid as a diagnostic cell, explicitly not yet a learned substrate. The checkpoint did NOT claim learned rendering invariance, open-world image/audio/video understanding, contrastive-alignment-level cross-modal matching, or learned ontology.
The main value: failure separability
The central contribution of the rebuilt cell is observability and failure separability, not benchmark performance. The cell distinguishes distinct error classes:
- rendering/parse error (corrupted parser, semantic-corruption negative)
- state-convergence error (state equality plus overlap/edit metrics)
- scope-transfer error (rendering-specific namespace and parsed-state ablation negatives)
- state-pairing error (mismatched rendering-state negative)
- surface-difference false conflict (a surface-difference baseline in the contradiction diagnostic)
- explicit-alignment coverage error (alignment-missing-pair negative)
Engineered-parser evidence
State convergence. The shared-parser path reached full triplet convergence and edge overlap versus gold with no residual normalized edit distance. The surface-shortcut baseline collapsed to no convergence / no overlap / maximal edit distance. A deliberately corrupted parser produced no exact-triplet convergence yet high (but imperfect) edge overlap — i.e. partial structure with the intended degradation.
Contradiction. Shared-state contradiction detection reached high precision, recall, and localization. A surface-difference baseline kept full recall but collapsed to near-chance precision and no localization (the intended false-conflict failure mode). An always-no-conflict baseline had no recall.
Cross-rendering scope transfer. Shared-state scope transfer reached high scope-selection accuracy and state accuracy versus gold across renderings with negligible degradation relative to text, while keeping the candidate set effectively narrowed. A rendering-specific-scope negative kept text valid but drove the other renderings' state accuracy to zero (maximal degradation). A parsed-state ablation negative drove all renderings to zero accuracy with the candidate set fully un-narrowed.
Mismatched rendering-state negative. Correctly paired renderings stayed at high state accuracy; mismatched pairs collapsed to zero state accuracy and zero scope-selection with an un-narrowed candidate set.
Explicit alignment control. A shared-state parser and an explicit-alignment lookup control both reached full coverage and high state accuracy; an alignment-missing-pair negative kept text strong but drove the other renderings to zero coverage/accuracy — demonstrating that explicit lookup is separable from the parser path.
Surface perturbation robustness. Nonsemantic perturbations preserved high state accuracy and state-equivalence versus the original (negligible degradation); a semantic-corruption negative drove accuracy and equivalence to zero (maximal degradation).
Verification. Local unit tests passed. A secondary reproduction host reproduced multi-seed summaries for scope transfer, mismatch, alignment, and perturbation; remote compile checks succeeded. Remote limitation: the secondary host's environment lacked the unit-test runner, so the local test run remained authoritative.
Progression to learned parsers
The checkpoint defined the next move: replace one engineered non-text parser with a learned parser while keeping all engineered-parser controls and measuring the engineered-teacher / learned-student / gold-state gap. The closure review inventory shows this was carried through across subphases, each with multi-seed local plus multi-seed remote summaries:
- a learned graph parser cell (parser, scope transfer, contradiction, mismatch, alignment, perturbation controls);
- a learned table parser cell (same control battery);
- learned graph/table convergence to a shared canonical state, plus cross-renderer contradiction, mismatched-pair, and alignment-separation checks;
- a controlled distribution-shift stress battery: shifted generator contract, held-out entity stress, reordered-structure stress, relation-combination stress, and larger-cell stress.
This yields two independent learned non-text renderers (graph and table) mapping to the same canonical semantic state.
Closure review
The closure review was a sequence of bounded review ticks, not an implementation phase.
- The first tick assembled the evidence inventory and an initial closure matrix; every closure criterion was marked "evidence present."
- The second tick ran an automated closure consistency check (with unit tests passing). The check covered the closure criteria across the engineered- and learned-parser subphases and local/remote evidence: all checks passed, none failed, closure not yet claimed, status closure-eligible-for-review.
- The third tick performed the final closure-language and residual-limit review and recorded a controlled, narrow closure decision.
All checked criteria passed: controlled rendering convergence; scope transfer / candidate execution; learned graph parser state recovery; learned table parser state recovery; graph/table/gold triple convergence; semantic contradiction localization; mismatched-provenance rejection; explicit-alignment separation; parser recovery without paired lookup coverage; nonsemantic perturbation robustness; semantic-corruption negative detection; and the distribution-shift gates.
Closure criteria decided "satisfied"
at least two independent learned non-text renderers (graph and table); learned graph/table states converge to a shared canonical state; convergence survives documented controlled shifts; downstream scope execution remains stable; contradiction localized at the semantic-state layer; mismatched provenance rejected; explicit lookup separable from the learned parser path; shortcut and corruption negatives fail correctly; local and secondary-host evidence agree; residual limits explicitly documented.
Residual limits (binding)
The closure boundary is intentionally narrow. It covers: controlled symbolic graph and table rendering, small auditable learned parser cells, canonical engineered semantic state, downstream scope/candidate execution over parsed state, and controlled contradiction/mismatch/alignment/perturbation/corruption/bounded-shift tests.
It does not cover: native multimodal intelligence, general multimodal invariance, open-world image understanding, audio/video understanding, real perception frontiers, learned ontology induction, large-scale unsupervised rendering discovery, arbitrary distribution shifts, downstream cursor/activation-graph execution, or production-ready multimodal architecture.
The evidence still depends on controlled symbolic small-cell worlds, engineered canonical gold state, synthetic graph/table renderers, small learned parser cells trained under teacher/student-style controls, and secondary-host smoke summaries rather than a full remote unit-test run for every subphase.
Final statement
The work is closed only under this controlled-symbolic statement: two independent learned non-text rendering paths (graph and table) map to the same canonical semantic state; that shared state supports downstream scope execution; and contradiction, mismatch, alignment, shortcut, corruption, perturbation, and bounded distribution-shift controls behave correctly across local multi-seed and secondary-host smoke evidence. The next phase (persistent semantic state / cursor / activation-graph execution) may proceed only from this narrow closure; the closed cell must not be retroactively broadened.
Claim boundary
The author's explicit scope — what this work does and does not establish — carried over from the Octoryn Research publishing model.
Proves
- Under a controlled symbolic small-cell world, distinct renderings of the same underlying semantic state can be mapped to a shared canonical state by independent learned non-text parsers.
- The shared recovered state supports downstream scope/candidate execution and supports state-level contradiction detection that localizes the conflict.
- Distinct failure modes (parse error, state-convergence error, scope-transfer error, state-pairing error, surface-difference false conflict, alignment-coverage error) are individually observable, each isolated by its own failing control.
- Convergence and the control battery survive bounded, documented controlled distribution shifts.
- Explicit-lookup alignment is separable from the learned parser path.
Does not prove
- Native multimodal intelligence or general multimodal invariance.
- Open-world image, audio, or video understanding, or real perception.
- Learned ontology induction or large-scale unsupervised rendering discovery.
- Robustness to arbitrary (un-bounded) distribution shifts.
- Readiness of any downstream cursor / activation-graph execution phase as closed evidence.
Applies when
- Renderings are controlled symbolic graph/table forms over a small, auditable cell.
- Canonical gold semantic state is engineered and available as the measured object.
- Learned parser cells are trained under teacher/student-style controls.
- Distribution shift stays within the documented bounded shift families.
Does not apply when
- Inputs are real images, audio, or video, or any open-world perceptual signal.
- The ontology or rendering family must be discovered rather than engineered.
- Shifts fall outside the bounded controlled-shift families tested.
- A claim requires native multimodal or production-ready multimodal architecture.
Authors
- Ran Tao — Investigation, Writing
Cite this
Citation
Tao, R., Octoryn Research. (2026). Controlled Symbolic Rendering-Invariance via Independent Learned Non-Text Parsers (TR-2026-0036). Octopus Research Institute.
BibTeX
@techreport{oritr20260036,
title = {Controlled Symbolic Rendering-Invariance via Independent Learned Non-Text Parsers},
author = {Tao, Ran and {Octoryn Research}},
institution = {Octopus Research Institute},
year = {2026},
note = {Permanent ID TR-2026-0036. Not peer reviewed.}
}Disclosures
- Funding
- Hardware and infrastructure provided by Octoryn / Octopus Core Pty Ltd.
- Conflicts of interest
- Octoryn ships commercial inference and governance tooling; findings are reported independently.
