Skills & Languages
Must-have
required
Sql5 Year(s)Scala4 Year(s)Devops5 Year(s)Nosql Databases3 Year(s)Java8 Year(s)Aws8 Year(s)observability5 Year(s)Datadog3 Year(s)
Languages required
English
Description of opportunity
A Staff Software Engineer focused on Reliability is needed to own reliability across the entire platform and drive the practices that ensure system availability, resilience, and observability for mission-critical infrastructure.
You will build reliability from first principles: architecting failover systems, implementing chaos engineering, and improving the observability foundation to maintain 99.9%+ uptime as the company scales into new markets.
As the technical owner of the reliability posture, you will tackle challenges like external service failover, dependency mirroring, and database replication — working alongside highly technical teams across the organization to influence architecture decisions and establish company-wide reliability standards.
This role sits on the Product Foundations team, building the foundational infrastructure that powers a large-scale mobility and commerce platform.
Tech challenge
- Maintain 99.9%+ uptime as the platform scales to new markets
- External service failover, dependency mirroring, and database replication at production scale
Responsibilities
- Own the overall reliability posture for the platform — practices, metrics, and systems that ensure 99.9%+ uptime across all services
- Design and implement automatic failover for critical external dependencies (e.g. SMS/voice and payments providers) with circuit breakers, retry policies, and degraded-mode operations
- Architect and build active-passive or active-active regional deployment strategies with database replication, automated failover, and DNS-based traffic routing — including disaster recovery planning and testing
- Establish comprehensive monitoring using Datadog (or equivalent) for APM, logs, and metrics correlation
- Implement synthetic monitoring, SLO-based alerting, on-call rotation, and escalation policies; build service health dashboards that show customer impact
- Own the incident management process — workflows, tooling, post-mortem culture, runbook automation, and MTTR reduction from detection to resolution
- Drive adoption of resilience patterns across services: health checks, graceful degradation, feature flags, rate limiting, backpressure, and chaos engineering
- Build and maintain local mirrors for critical dependencies — artifact caching, dependency pinning, and vulnerability scanning to prevent build failures from upstream outages
Key requirements
Required
- 10+ years of engineering experience in software engineering, reliability engineering, SRE practices, or production operations at scale
- Expert-level reliability engineering: multi-region architectures, failover automation, circuit breakers, chaos engineering, and disaster recovery
- Production observability at scale — deep experience with monitoring, alerting, tracing, and logging; Datadog or similar APM in high-load environments
- Strong systems thinking — design resilient distributed systems that handle failures, network partitions, and external dependency outages
- Database and data systems knowledge: replication strategies, backup/restore, connection pooling, query optimization; relational and NoSQL experience
- AWS production experience: multi-region deployments, load balancing, DNS-based failover
- Experience with AI-powered development tools (e.g. GitHub Copilot or similar agentic coding tools)
- Expert-level Java and/or Scala — JVM performance, concurrency, and operational characteristics
- Strong technical communication; ability to influence architecture across teams, document complex systems, run post-mortems, and establish org-wide reliability standards
Preferred
- Scala experience
- SRE or Reliability Engineering experience at companies known for operational excellence (e.g. large-scale tech companies or high-growth startups where you built reliability practices from the ground up)
- Incident response leadership: incident management processes, blameless post-mortems, MTTR reduction in production
- Chaos engineering with tools like Chaos Monkey, Gremlin, or similar — including game days and failure injection testing
- Performance optimization: profiling, benchmarking, capacity planning, and system tuning at hyperscale
- Open source contributions or technical writing demonstrating depth in reliability engineering, distributed systems, or production operations
Ideal candidate
- Builds reliability from first principles
- Works alongside highly technical teams to influence architecture and establish company-wide reliability standards
- Excellent technical communication — documents complex systems, conducts post-mortems, and drives reliability standards organization-wide
Benefits
🏥 Medical Insurance🏥 Paid vacation
Additional benefits: Unlimited PTO, 10 holidays, Equipment provided, Stock options negotiable