View all jobs
Closed

Senior Site Reliability Engineer

🇺🇸 San Diego, United States Fully remote8+ years
Required Candidate location: South America or Central America

This position is no longer available

Check more opportunities below

Skills & Languages

Must-have required
Postgresql3 Year(s)Devops4 Year(s)Ci/cd Automation3 Year(s)Aws4 Year(s)Terraform3 Year(s)Datadog3 Year(s)
Nice-to-have
React.js1 Year(s)Typescript1 Year(s)Java2 Year(s)
Languages required
English

About Us


We build software that helps service-based businesses grow and strengthen customer relationships. Our platform combines digital engagement tools, workflow automation, and content-driven solutions that enable small and medium-sized organizations to operate more efficiently and communicate more effectively with their clients.


Our customers rely on our technology to streamline business operations, improve customer engagement, and support long-term growth. We are a remote-first team focused on building reliable, scalable software that delivers meaningful value to our users.


About the Role


We're looking for a Senior Site Reliability Engineer to own the reliability, observability, and operational excellence of our cloud-based platform. This role sits at the intersection of software engineering and infrastructure, helping ensure our systems remain resilient, scalable, and highly available as our business grows.


You'll play a key role in shaping how we monitor, operate, and improve production systems. Working closely with engineering, product, and quality teams, you'll drive initiatives that reduce operational complexity, improve system performance, and strengthen platform reliability.


We seek an engineer who is a:

  • Servant leader – Leads through action, collaboration, and a willingness to tackle challenging problems.
  • Decision maker – Evaluates trade-offs effectively and moves initiatives forward with pragmatism and sound judgment.
  • Communicator – Thrives in collaborative environments and promotes transparency, accountability, and continuous improvement.

What You'll Do

  • Own production reliability by defining service-level objectives, improving operational health, and leading post-incident reviews that drive lasting improvements.
  • Design, build, and maintain observability solutions using Datadog, including dashboards, alerts, and monitoring strategies that provide meaningful operational insights.
  • Develop and manage infrastructure as code using AWS CDK and related cloud technologies.
  • Improve CI/CD pipelines and deployment processes to support rapid, safe, and repeatable software delivery.
  • Partner with engineering teams to optimize performance, scalability, availability, and cloud resource utilization.
  • Automate operational workflows and reduce manual effort through tooling and process improvements.
  • Enhance incident response practices, operational readiness, and on-call effectiveness across the organization.

Essential Skills & Experience

  • 8+ years of experience developing, operating, and supporting cloud-hosted software systems.
  • Strong hands-on expertise with AWS and infrastructure-as-code practices, ideally using AWS CDK, Terraform, or CloudFormation.
  • Deep experience with observability and monitoring platforms, particularly Datadog.
  • Proven experience managing and scaling production-grade databases and distributed systems, with PostgreSQL experience strongly preferred.
  • Strong understanding of modern DevOps practices, CI/CD pipelines, and source control workflows.
  • Experience defining and operating against reliability targets, incident management processes, and operational best practices.
  • Solid understanding of cloud-native and distributed systems architecture.
  • Experience utilizing AI-powered development tools to improve engineering workflows.
  • Excellent communication and collaboration skills.
  • Ability to work independently in a highly autonomous environment.

Nice to Have

  • Experience working with Java-based applications.
  • Familiarity with React and TypeScript.
  • Experience supporting multi-tenant SaaS platforms, communication systems, or high-volume transactional applications.

Our Culture


We value ownership, curiosity, and continuous improvement. Our team embraces collaboration, open communication, and a practical approach to solving complex problems. We encourage individuals to take initiative, challenge assumptions, and contribute ideas that improve both our products and the way we work.


As a distributed team, we prioritize trust, flexibility, and outcomes over process. We believe great work happens when talented people are empowered to do their best work while maintaining a healthy balance between professional and personal commitments.


Work Environment


We operate as a remote-first organization and provide the flexibility needed to support productive, sustainable careers. We foster a culture of accountability, transparency, and mutual respect while maintaining a strong focus on delivering results for our customers and teammates.


Diversity & Inclusion


We are committed to creating an inclusive workplace where everyone feels welcomed, respected, and empowered to contribute. We believe diverse perspectives strengthen our teams and help us build better products for our customers.


We are proud to be an equal opportunity employer and make employment decisions based on qualifications, merit, and business needs, without regard to any protected characteristic under applicable law.

Lost connection

The internet blinked. We’re catching up.

Reload page
An error has occurred. This application may no longer respond until reloaded.Reload 🗙