View all jobs
Closed

Staff Software Engineer focused on Reliability

🇺🇸 New York, United States Fully remote10+ years
Required Candidate location: LATAM or South America

This position is no longer available

Check more opportunities below

Skills & Languages

Must-have required
Sql5 Year(s)Scala4 Year(s)Devops5 Year(s)Nosql Databases3 Year(s)Java8 Year(s)Aws8 Year(s)observability5 Year(s)Datadog3 Year(s)
Languages required
English

Description of opportunity

A Staff Software Engineer focused on Reliability is needed to own reliability across the entire platform and drive the practices that ensure system availability, resilience, and observability for mission-critical infrastructure.

You will build reliability from first principles: architecting failover systems, implementing chaos engineering, and improving the observability foundation to maintain 99.9%+ uptime as the company scales into new markets.

As the technical owner of the reliability posture, you will tackle challenges like external service failover, dependency mirroring, and database replication — working alongside highly technical teams across the organization to influence architecture decisions and establish company-wide reliability standards.

This role sits on the Product Foundations team, building the foundational infrastructure that powers a large-scale mobility and commerce platform.


Tech challenge

  • Maintain 99.9%+ uptime as the platform scales to new markets
  • External service failover, dependency mirroring, and database replication at production scale

Responsibilities

  • Own the overall reliability posture for the platform — practices, metrics, and systems that ensure 99.9%+ uptime across all services
  • Design and implement automatic failover for critical external dependencies (e.g. SMS/voice and payments providers) with circuit breakers, retry policies, and degraded-mode operations
  • Architect and build active-passive or active-active regional deployment strategies with database replication, automated failover, and DNS-based traffic routing — including disaster recovery planning and testing
  • Establish comprehensive monitoring using Datadog (or equivalent) for APM, logs, and metrics correlation
  • Implement synthetic monitoring, SLO-based alerting, on-call rotation, and escalation policies; build service health dashboards that show customer impact
  • Own the incident management process — workflows, tooling, post-mortem culture, runbook automation, and MTTR reduction from detection to resolution
  • Drive adoption of resilience patterns across services: health checks, graceful degradation, feature flags, rate limiting, backpressure, and chaos engineering
  • Build and maintain local mirrors for critical dependencies — artifact caching, dependency pinning, and vulnerability scanning to prevent build failures from upstream outages

Key requirements

Required

  • 10+ years of engineering experience in software engineering, reliability engineering, SRE practices, or production operations at scale
  • Expert-level reliability engineering: multi-region architectures, failover automation, circuit breakers, chaos engineering, and disaster recovery
  • Production observability at scale — deep experience with monitoring, alerting, tracing, and logging; Datadog or similar APM in high-load environments
  • Strong systems thinking — design resilient distributed systems that handle failures, network partitions, and external dependency outages
  • Database and data systems knowledge: replication strategies, backup/restore, connection pooling, query optimization; relational and NoSQL experience
  • AWS production experience: multi-region deployments, load balancing, DNS-based failover
  • Experience with AI-powered development tools (e.g. GitHub Copilot or similar agentic coding tools)
  • Expert-level Java and/or Scala — JVM performance, concurrency, and operational characteristics
  • Strong technical communication; ability to influence architecture across teams, document complex systems, run post-mortems, and establish org-wide reliability standards

Preferred

  • Scala experience
  • SRE or Reliability Engineering experience at companies known for operational excellence (e.g. large-scale tech companies or high-growth startups where you built reliability practices from the ground up)
  • Incident response leadership: incident management processes, blameless post-mortems, MTTR reduction in production
  • Chaos engineering with tools like Chaos Monkey, Gremlin, or similar — including game days and failure injection testing
  • Performance optimization: profiling, benchmarking, capacity planning, and system tuning at hyperscale
  • Open source contributions or technical writing demonstrating depth in reliability engineering, distributed systems, or production operations

Ideal candidate

  • Builds reliability from first principles
  • Works alongside highly technical teams to influence architecture and establish company-wide reliability standards
  • Excellent technical communication — documents complex systems, conducts post-mortems, and drives reliability standards organization-wide

Benefits

🏥 Medical Insurance🏥 Paid vacation

Additional benefits: Unlimited PTO, 10 holidays, Equipment provided, Stock options negotiable

Lost connection

The internet blinked. We’re catching up.

Reload page
An error has occurred. This application may no longer respond until reloaded.Reload 🗙