Responsibilities
- Design and implement distributed web scraping systems using Python and Scrapy to efficiently gather large-scale datasets.
- Extract data from dynamic, JavaScript-driven websites by leveraging browser automation tools to simulate user interactions.
- Deploy and manage containerized scraping workloads on Kubernetes, ensuring efficient scaling and system resilience.
- Establish data integrity through strict JSON Schema definitions and enforce validation using Pydantic to detect inconsistencies.
- Process raw web data and build optimized pipelines for indexing and storage using Elasticsearch.
- Develop end-to-end workflows that coordinate data discovery, extraction, validation, and storage stages reliably.
- Implement advanced evasion techniques including proxy rotation, session management, and browser fingerprint manipulation to bypass anti-bot defenses.
Work Arrangement
Remote (Worldwide)
Other
Must be eligible to obtain a security clearance.