Status: Active researchPrivacy-Preserving AIResponsible AI
Privacy-Preserving AI Pipelines
Sensitive-data detection, redaction and transformation with data-residency controls and residual-risk evaluation before controlled model access.
Abstract
We study pipelines that detect sensitive data, redact or transform it, keep data within a residency boundary, and evaluate the residual risk remaining after redaction before any model sees the content.
Problem & motivation
Redaction is often treated as binary and complete, but residual re-identification risk usually remains and is rarely measured.
Research questions
- How can residual re-identification risk be estimated after redaction?
- What transformations preserve task utility while reducing disclosure?
Methods
- Build a detection → transformation → residual-risk evaluation pipeline.
- Compare utility/disclosure trade-offs across transformations.
Limitations
- No dataset containing real personal or health data is published or accepted through this site.
- Residual-risk estimates are approximate and model-dependent.
Disclosures
- Funding
- Infrastructure support provided by Octopus Core Pty Ltd.
- Conflicts of interest
- Octopus Core develops commercial privacy infrastructure; findings are reported independently.
- Ethics
- No human-subjects data is collected for this programme through this website.
