Backend Engineer (Data and Crawling)

Full-time

Remote

We seek team members who are creative, proactive, and curious. We welcome diversity of thought and cultivate a culture founded on low power distance and radical candor, where advancement comes from respect and merit, not politics. We're fully remote. You'll set your own hours, with a high value placed on uninterrupted, deep work, and we ruthlessly nix any meeting a message could resolve. If your skills and inclinations align, please apply.

About the role

We're looking for a backend engineer focused on web crawling and automation to build API-driven, end-to-end data-acquisition systems that interact with real-world websites using source-specific strategies.

Our crawling systems power other internal services through on-demand REST APIs. They cover a wide range of sources, from openly available websites to authenticated, access-controlled platforms. Depending on the source, they range from high-throughput programmatic access to full browser-based automation for complex, dynamic sites.

This role goes beyond basic browser automation. You'll own the full lifecycle of crawled data, from interaction and extraction through parsing, normalization, indexing, and API delivery, making sure it stays reliable, structured, and usable by the systems that depend on it.

What you’ll do

Before diving into the technical responsibilities, here are the traits we value most:

  • Candor: You communicate directly and honestly in service of better outcomes.

  • Conscientiousness: You take ownership, respect teammates, and build things others can rely on.

  • First-principles thinking: You question assumptions and make decisions grounded in evidence.

In this role, you will:

  • Design and build end-to-end web crawling and data-acquisition systems

  • Implement browser-based automation with tools like Playwright or Puppeteer

  • Build crawlers that adapt to each source, from high-throughput programmatic access to authenticated, session-based browser automation for JavaScript-heavy, dynamic sites

  • Build robust parsing and extraction pipelines that turn raw web data into structured formats

  • Design and maintain data normalization, enrichment, and validation workflows

  • Implement indexing strategies that make crawled data searchable, performant, and reliable

  • Build backend services and APIs that expose crawled and indexed data to internal consumers

  • Monitor, debug, and improve crawl correctness, stability, and cost efficiency

  • Work with other engineers to integrate crawling pipelines into larger product workflows

  • Contribute to CI/CD, observability, and operational tooling

Who you are

You think of crawling as a system, not a script. You understand that different sources require different approaches, and you enjoy reasoning about the trade-offs between speed, reliability, robustness, and cost.

You're comfortable debugging non-deterministic failures, working with imperfect or inconsistent data, and owning systems end to end, from the first request to the final API response.

You care about data quality and long-term maintainability. You think about schemas, indexing, and downstream consumers as part of the core problem, not an afterthought.

Required qualifications

  • Strong professional experience with JavaScript/TypeScript and/or Python

  • Proven experience building production-grade crawling or browser automation systems

  • Hands-on experience with Playwright, Puppeteer, Selenium, or similar

  • Experience designing API-driven crawling services

  • Strong understanding of browser behavior and JavaScript execution, and of sessions, cookies, headers, and authentication flows

  • Experience building parsing, normalization, and data-processing pipelines

  • Backend experience building services and REST and/or GraphQL APIs

  • Experience with relational and/or NoSQL databases

  • Proficiency with Git and collaborative development workflows

Nice-to-have skills

  • Experience designing resilient crawling strategies for large, complex, dynamic, or authenticated platforms

  • Search and indexing systems (Elasticsearch, OpenSearch, or similar)

  • Distributed or queue-based processing systems

  • Rust for performance-critical components

  • Experience integrating AI/ML services into extraction, enrichment, or classification workflows

What we offer

  • A high-trust, remote-first culture

  • End-to-end ownership of complex, business-critical systems

  • Close collaboration with the engineering team

  • A team that values clear thinking, rigor, and direct communication

  • Room to influence architecture and technical direction

  • Competitive compensation based on experience and impact

Submit application for Backend Engineer (Data and Crawling)