DevOps Engineer
Full-time
Remote
We seek team members who are creative, proactive, and curious. We welcome diversity of thought and cultivate a culture founded on low power distance and radical candor, where advancement comes from respect and merit, not politics. We're fully remote. You'll set your own hours, with a high value placed on uninterrupted, deep work, and we ruthlessly nix any meeting a message could resolve. If your skills and inclinations align, please apply.
About the role
We're looking for a hands-on DevOps engineer to own the infrastructure, platform tooling, and CI/CD pipelines that keep our systems running. You'll work across Linux and Windows, containers, databases, and a self-hosted service stack running on a major cloud provider, automating everything you can and keeping reliability high.
Our products are built with AI at their core and ship fast, so you'll be building infrastructure that supports a high volume of deployments and keeps a fast-moving engineering team unblocked.
What you’ll do
Before diving into the technical responsibilities, here are the traits we value most:
Candor: You communicate directly and honestly in service of better outcomes.
Conscientiousness: You take ownership, respect teammates, and build things others can rely on.
First-principles thinking: You question assumptions and make decisions grounded in evidence.
In this role, you will:
Own the infrastructure and platform end to end, provisioning it as code and automating manual work away
Build and own CI/CD pipelines, keeping build, test, and deployment fully automated
Own monitoring, alerting, and observability so problems surface before customers feel them
Keep databases healthy and performant: backups, replication, tuning, and reliable analytical workloads
Operate messaging and traffic routing so services communicate reliably and stay available
Harden and tune systems for security, performance, and reliability across Linux and Windows
Design infrastructure that scales as usage and load grow, keeping cost in check
Own on-call for critical systems, diagnosing complex production issues and driving fixes that stop them recurring
Partner with the engineering team to raise deployment velocity and reliability
Who you are
You think of infrastructure as a product, not a chore. You understand that speed, reliability, and cost pull against each other, you enjoy reasoning about the trade-offs, and you automate away manual, repetitive work rather than tolerate it.
You're comfortable being on the hook for systems that can't go down. You can diagnose a failure, ship the fix, and then make sure it never happens the same way twice.
You care about the people around you as much as the systems. You'd rather unblock a teammate than keep the interesting problem for yourself.
Required qualifications
5+ years in a DevOps, platform engineering, or SRE role
Experience with a major cloud provider (AWS, GCP, or Azure)
Advanced Linux administration (RHEL/Debian family) and shell scripting, plus scripting in Python or Bash
Solid hands-on experience with infrastructure as code using Terraform and Ansible
Production experience with container orchestration (Kubernetes, Nomad, or similar); we run Nomad, so hands-on Nomad experience is a plus
Experience with service discovery and secrets management (we use Consul and Vault)
Strong GitLab CI/CD pipeline development and maintenance
Solid PostgreSQL administration (replication, performance tuning), and familiarity with ClickHouse for analytical workloads
Experience with load balancing and reverse proxies (HAProxy, Traefik, or Nginx)
Strong networking fundamentals (DNS, VPNs, service mesh, zero-trust)
Experience administering RabbitMQ: cluster setup, queue management, and monitoring
Experience with Docker and container-based deployments
Experience with observability tooling for metrics, logging, and tracing (Prometheus, Grafana, Loki, or similar)
Comfortable with Windows Server administration
Strong understanding of the full software development lifecycle, from development through production
Nice-to-have skills
Experience with additional caching, messaging, or streaming systems such as Redis or Kafka
Go, for building infrastructure tooling, operators, and CLIs
SRE reliability practices such as SLOs, error budgets, and incident response
Cloud and infrastructure cost optimization (FinOps)
Security and compliance experience such as hardening at scale, SOC 2, or ISO 27001
What we offer
A high-trust, remote-first culture
End-to-end ownership of the infrastructure and platform behind business-critical systems
Close collaboration with the engineering team
A team that values clear thinking, rigor, and direct communication
Room to shape our infrastructure and how we deploy
Competitive compensation based on experience and impact