Backend Engineer (Data and Crawling)
Full-time
Remote
We seek team members who are creative, proactive, and curious. We welcome diversity of thought and cultivate a culture founded on low power distance and radical candor, where advancement comes from respect and merit, not politics. We're fully remote. You'll set your own hours, with a high value placed on uninterrupted, deep work, and we ruthlessly nix any meeting a message could resolve. If your skills and inclinations align, please apply.
About the role
We're looking for a backend engineer focused on web crawling and automation to build API-driven, end-to-end data-acquisition systems that interact with real-world websites using source-specific strategies.
Our crawling systems power other internal services through on-demand REST APIs. They cover a wide range of sources, from openly available websites to authenticated, access-controlled platforms. Depending on the source, they range from high-throughput programmatic access to full browser-based automation for complex, dynamic sites.
This role goes beyond basic browser automation. You'll own the full lifecycle of crawled data, from interaction and extraction through parsing, normalization, indexing, and API delivery, making sure it stays reliable, structured, and usable by the systems that depend on it.
What you’ll do
Before diving into the technical responsibilities, here are the traits we value most:
Candor: You communicate directly and honestly in service of better outcomes.
Conscientiousness: You take ownership, respect teammates, and build things others can rely on.
First-principles thinking: You question assumptions and make decisions grounded in evidence.
In this role, you will:
Design and build end-to-end web crawling and data-acquisition systems
Implement browser-based automation with tools like Playwright or Puppeteer
Build crawlers that adapt to each source, from high-throughput programmatic access to authenticated, session-based browser automation for JavaScript-heavy, dynamic sites
Build robust parsing and extraction pipelines that turn raw web data into structured formats
Design and maintain data normalization, enrichment, and validation workflows
Implement indexing strategies that make crawled data searchable, performant, and reliable
Build backend services and APIs that expose crawled and indexed data to internal consumers
Monitor, debug, and improve crawl correctness, stability, and cost efficiency
Work with other engineers to integrate crawling pipelines into larger product workflows
Contribute to CI/CD, observability, and operational tooling
Who you are
You think of crawling as a system, not a script. You understand that different sources require different approaches, and you enjoy reasoning about the trade-offs between speed, reliability, robustness, and cost.
You're comfortable debugging non-deterministic failures, working with imperfect or inconsistent data, and owning systems end to end, from the first request to the final API response.
You care about data quality and long-term maintainability. You think about schemas, indexing, and downstream consumers as part of the core problem, not an afterthought.
Required qualifications
Strong professional experience with JavaScript/TypeScript and/or Python
Proven experience building production-grade crawling or browser automation systems
Hands-on experience with Playwright, Puppeteer, Selenium, or similar
Experience designing API-driven crawling services
Strong understanding of browser behavior and JavaScript execution, and of sessions, cookies, headers, and authentication flows
Experience building parsing, normalization, and data-processing pipelines
Backend experience building services and REST and/or GraphQL APIs
Experience with relational and/or NoSQL databases
Proficiency with Git and collaborative development workflows
Nice-to-have skills
Experience designing resilient crawling strategies for large, complex, dynamic, or authenticated platforms
Search and indexing systems (Elasticsearch, OpenSearch, or similar)
Distributed or queue-based processing systems
Rust for performance-critical components
Experience integrating AI/ML services into extraction, enrichment, or classification workflows
What we offer
A high-trust, remote-first culture
End-to-end ownership of complex, business-critical systems
Close collaboration with the engineering team
A team that values clear thinking, rigor, and direct communication
Room to influence architecture and technical direction
Competitive compensation based on experience and impact