2025 — Present
Site Reliability Engineer (self-employed)
- Designed and operated homelab with 13 services in production 24/7 for 365 days, measured uptime 99.7% (3 planned outages, 0 unplanned). Stack: Proxmox VE + ZFS + KVM/LXC + Oracle Cloud free tier.
- Operated public VPS (pentagran.duckdns.org) with Docker, nginx and automatic ACME cert rotation. Serves ~30 static pages + Fastify API with 13 endpoints.
- Built 3 operational AI agents (Lia, Codex, Palomino) with separated roles. Shared memory via Obsidian vault + Git. Estimated ~200 hours/year automated.
- Reduced incident response time by 40% via 23 documented runbooks after each post-mortem. E.g. SSL cert restoration went from 8h (month 3) to 45min (month 8).
- Detected and diagnosed 4 critical failures: expired SSL cert, ZFS ARC consuming all RAM, RAID1 degraded due to dead disk, missing swapfile on Oracle free tier.
13services in prod
99.7%uptime 12m
23runbooks
40%incident time cut
3AI agents
0unplanned downtime