Production Infrastructure Audit & Remediation — IAC
Indonesia AIDS Coalition (IAC)
A full audit of production servers running public-health information systems, uncovering critical backup gaps and recurring capacity incidents, then fixing them through a data-driven approach.
- Infrastructure
- Incident Response
- Backup & DR
As the sole IT staff at this organization, I’m responsible for the entire server infrastructure running its public-health information systems - from an e-learning platform and a field-reporting dashboard to a community-monitoring system. Before this audit, infrastructure documentation barely existed.
I ran a full read-only audit across all production servers, mapped what was running where, and standardized everything into consistent documentation. That surfaced several serious gaps that had gone unnoticed - one database backup path had silently stopped running for months, and another production server had no backup mechanism at all.
One of the servers also hit a storage-full incident twice within a short window, even after its capacity had already been doubled in between. Rather than reaching for more capacity as a quick fix, I ran a data-driven root-cause analysis - checking the initial hypothesis against actual file-growth data - before carrying out a large-scale remediation across more than a million production files. Every step was verified with a small-scale trial first, to avoid any risk of corrupting data still in active use by the application.
The outcome: complete infrastructure documentation that didn’t exist before, critical backup gaps identified and being fixed, and automated monitoring now in place across every server so similar issues get caught early - not after capacity is already gone.