About the Role
Techdome runs live infrastructure for Healthcare, FinTech, AI, and SaaS products — environments where downtime isn't an inconvenience, it's a compliance incident or a lost transaction. We're hiring a Senior SRE who treats uptime as a personal metric, not a team KPI, and who wants direct, high-leverage ownership over production systems handling real money and real patient data — not staging environments. You'll report close to founders and senior engineering leadership, and your infrastructure decisions will ship the same week you make them.

Key Responsibilities
  • Define SLIs/SLOs, own the error budget, and decide when to slow down shipping to protect it
  • Build zero-downtime CI/CD pipelines supporting Blue-Green, Canary, and Rolling releases
  • Manage all environments through Terraform and Ansible — no manual console changes
  • Instrument systems with Prometheus, Grafana, ELK, Datadog, and OpenTelemetry so alerts are actionable, not noisy
  • Lead incident response, drive root cause analysis, and ensure postmortem action items are closed
  • Right-size infrastructure and forecast capacity proactively, rather than reacting to billing
  • Apply AI to alert triage, incident summarization, and automated runbooks
  • Participate in a shared on-call rotation as a dependable, trusted responder

Required Qualifications
  • 2+ years of production ownership experience as an SRE, DevOps, Platform, or Cloud Engineer
  • Hands-on, production-grade experience with AWS, Azure, or GCP
  • Real-world Docker/Kubernetes experience under production load
  • Daily use of Terraform and Ansible (or equivalent IaC tools)
  • Experience building at least one CI/CD pipeline from scratch (Jenkins, GitHub Actions, GitLab CI, or similar)
  • Strong fundamentals in Linux, networking, and distributed systems
  • Proficiency in Python, Go, or Bash for automation and scripting
  • Proven experience shipping Blue-Green, Canary, and Rolling deployments in production
  • Prior work in FinTech, Payments, Healthcare, or another high-availability domain
  • Practical, everyday use of AI tools such as Copilot, Claude, Cursor, or ChatGPT

Preferred Qualifications
  • Experience building AI-powered operations tooling — triage bots, incident auto-summarization, or reliable anomaly detection
  • Fluency in SLOs, error budgets, and chaos engineering as core practice, not theory

Why Join Techdome?
  • Direct reporting line to founders and senior engineering leadership
  • Infrastructure decisions ship the same week they're made — no bureaucratic delay
  • You'll protect systems handling real healthcare data and real financial transactions
  • A team that shares on-call and builds a culture where the SRE on call at 2am is someone people trust