Remote (United States) Remote (Global)

Ant-Tech is hiring a Research Crawling Engineer

Responsibilities

  • Develop and sustain web crawlers capable of operating across a wide range of domains at scale
  • Architect high-performance, resilient data collection systems handling millions to billions of URLs daily
  • Navigate anti-bot defenses, rate limiting, and JavaScript-rich content during crawling
  • Create data processing workflows for cleaning, removing duplicates, filtering, and standardizing content
  • Assemble and manage datasets used in research and machine learning model development
  • Track and assess crawl efficiency, data coverage, and quality metrics with rapid iteration
  • Partner with research teams to ensure data collection supports modeling objectives
  • Enhance data infrastructure for optimal cost, speed, and dependability

Requirements

  • Proficient in one or more programming languages such as Go, Rust, Python, Java, or C++
  • Proven track record in developing web crawlers or large-volume data processing systems
  • Strong grasp of HTTP protocols, network fundamentals, and browser operations
  • Experience with distributed computing and parallelized processing techniques
  • Worked with datasets ranging from terabytes to petabytes in size
  • Skilled in diagnosing and resolving issues in unstable or adversarial environments

Nice to Have

  • Background in NLP workflows or preparing datasets for machine learning
  • Knowledge of data used in large language model pretraining or retrieval systems
  • Hands-on experience with headless browsers like those controlled via Chrome DevTools Protocol, Playwright, or Puppeteer
  • Understanding of proxy networks, IP rotation, and managing large volumes of HTTP requests
  • Experience evaluating or benchmarking data quality
  • Operational experience deploying workloads on cloud platforms or bare-metal servers

Benefits

  • You’ll receive a competitive salary, benefits and equity package.

Compensation

Competitive salary, benefits and equity package

Work Arrangement

Remote (Worldwide)

Team

Lean, high-output team operating remotely with a focus on impact and mutual improvement

What This Role Involves

  • Operating at the boundary of scale and reliability
  • Adapting to constantly changing web environments
  • Balancing throughput, coverage, and data quality
  • Owning end-to-end data acquisition pipelines

Evaluation Criteria

  • Ability to design systems that scale without degrading quality
  • Practical problem-solving under real-world constraints
  • Speed of iteration and ownership
  • Measurable improvements in data coverage, quality, or efficiency

Example Projects

  • Build a distributed crawler for a continuously updated, high-quality web project
  • Design a system to classify and filter billions of pages for pretraining
  • Extract structured data from dynamic, JS-heavy sites at scale
  • Improve deduplication and quality scoring across multimodal datasets

Why Work With Us

  • Opportunity. We are at the forefront of developing a web-scale crawler and knowledge graph that improves access to public web data and extends the value of AI to the people.
  • Culture. We're a lean team with a high bar. We come to work not to be comfortable, but to find out what we're capable of and to do work that matters. We're not calling for people who keep things moving. We're calling for people who make everyone around them better.
  • We prioritize low ego and high output. This is a fully remote team.
  • Compensation. You’ll receive a competitive salary, benefits and equity package.
Required Skills
Distributed Systems
About company
Ant-Tech
A headhunter agency in France specializing in recruitment services across various industries.
All jobs at Ant-Tech Visit website
Job Details
Department Engineering
Category data
Posted 22 days ago