Production doesn't page a ticket queue. It pages you.

Techdome runs live infrastructure for Healthcare, FinTech, AI, and SaaS products where downtime isn't an inconvenience — it's a compliance incident or a lost transaction. We need someone who treats uptime as a personal metric, not a team KPI.

The job, in outcomes:
  • Availability is your scoreboard. You define SLIs/SLOs, own the error budget, and make the call on when to slow down shipping to protect it.
  • Deploys don't cause incidents. You build CI/CD that ships Blue-Green, Canary, and Rolling releases with zero customer-facing downtime — and you built at least one of these pipelines from scratch before, not just configured someone else's.
  • Infrastructure is code, not tribal knowledge. Terraform and Ansible define your environments; nothing gets clicked into existence in a console.
  • You see problems before customers do. Prometheus, Grafana, ELK, Datadog, OpenTelemetry — instrumented well enough that alerts mean something and noise doesn't.
  • Incidents end with a fix, not a Slack thread. You lead response, drive RCA, and turn postmortems into action items that actually close.
  • Cost is engineered, not just monitored. You right-size and forecast capacity instead of reacting to the bill.
  • You use AI to move faster, not to look modern. Alert triage, incident summarization, automated runbooks — if it can be scripted or delegated to a model, it should be.
  • On-call is shared, not survived. You rotate in, and you're the person newer engineers want on the call at 2am.
What gets you in the door:
  • 2+ years running production as an SRE, DevOps, Platform, or Cloud Engineer — real ownership, not observer status
  • Production-grade AWS, Azure, or GCP experience
  • Docker/Kubernetes under actual load, with the scars to prove it
  • Terraform and Ansible (or equivalent IaC) in daily use
  • Built CI/CD pipelines from zero — Jenkins, GitHub Actions, GitLab CI, or similar
  • Solid Linux, networking, and distributed-systems fundamentals
  • Python, Go, or Bash for scripting and automation
  • Shipped Blue-Green, Canary, and Rolling deployments in production, not just in theory
  • Domain background in FinTech, Payments, Healthcare, or another high-availability environment
  • AI tools (Copilot, Claude, Cursor, ChatGPT) are already part of how you work, not a novelty
What sets you apart:
  • You've shipped AI-powered ops tooling — triage bots, auto-summarized incidents, anomaly detection that actually fires correctly
  • You talk SLOs, error budgets, and chaos engineering like it's your first language, because it is
Why this seat, specifically:
You're not the 15th hire on a platform team waiting for tickets. You report close to founders and senior engineering leadership, your infra decisions ship the same week you make them, and the systems you protect handle real healthcare data and real money — not staging environments.