About the Role
Mindrift is seeking a highly skilled Senior Python Data Scraping Engineer to join the Tendem project, driving specialized data scraping workflows within a hybrid AI + human system. This freelance role is based in Amman, Jordan, and is ideal for an experienced engineer who thrives on building robust, scalable data extraction pipelines that power AI-driven solutions.
Core Responsibilities
- Design, develop, and maintain advanced Python-based web scraping scripts and frameworks for large-scale data extraction.
- Build and optimize data scraping workflows that integrate seamlessly with Mindrift's hybrid AI + human platform.
- Identify and resolve anti-bot mechanisms, CAPTCHAs, IP blocks, and other scraping obstacles using rotating proxies, headless browsers, and session management techniques.
- Ensure data quality, accuracy, and consistency across all scraping outputs through validation and cleaning pipelines.
- Collaborate with AI/ML teams to deliver structured, well-formatted datasets that feed directly into machine learning models and training pipelines.
- Monitor scraping performance metrics, implement error handling, and ensure high uptime and reliability of all data collection processes.
- Develop and maintain automation scripts for scheduling, monitoring, and reporting on scraping jobs.
- Stay current with the latest web scraping libraries (Scrapy, BeautifulSoup, Selenium, Playwright) and adapt tools as project needs evolve.
- Document scraping architectures, data schemas, and standard operating procedures for internal knowledge sharing.
Required Qualifications
- 5+ years of professional experience in Python development with a strong focus on web scraping and data extraction.
- Deep proficiency in Python libraries and frameworks including Scrapy, BeautifulSoup, Selenium, Playwright, Requests, and lxml.
- Strong understanding of HTML, CSS, XPath, and JavaScript for navigating and parsing complex web structures.
- Experience working with REST APIs, GraphQL endpoints, and parsing JSON/XML data formats.
- Solid knowledge of database systems (PostgreSQL, MongoDB, Redis) for storing and managing scraped data.
- Familiarity with proxy management, user-agent rotation, rate limiting, and evasion of anti-scraping technologies.
- Experience with version control (Git) and collaborative development workflows.
- Strong problem-solving skills and the ability to work independently in a freelance/remote environment.
Preferred Skills
- Experience with cloud platforms (AWS, GCP, or Azure) for deploying and scaling scraping infrastructure.
- Familiarity with containerization tools such as Docker and orchestration with Kubernetes.
- Knowledge of data pipeline tools (Airflow, Luigi, Prefect) for workflow automation and scheduling.
- Exposure to AI/ML data annotation workflows and understanding of how scraped data feeds model training.
- Experience with headless browser automation and browser fingerprint randomization.
- Proficiency in data cleaning and transformation using Pandas, NumPy, or similar libraries.
- Understanding of web scraping ethics, robots.txt compliance, and legal considerations around data collection.
What Makes This Opportunity Stand Out
- Work at the intersection of AI and human intelligence on the innovative Tendem project.
- Freelance flexibility — manage your own schedule while contributing to cutting-edge AI data initiatives.
- Collaborate with a forward-thinking team at Mindrift, a leader in AI training and data solutions.
- Gain hands-on experience building production-grade scraping systems that directly impact AI model performance.
- Opportunity for long-term engagement and growth as the Tendem project scales.