ApplySarthi

Senior Data Engineer (Web Scraping)

Jobgether

Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.

Got this interview? Our apps help you get the job.

Skills named in this job

Read from the description itself, not inferred.

This role on the market

4 open scraping roles across 3 companies are on ApplySarthi right now.

What scraping roles keep asking for: Playwright (100%), Selenium (100%), HTML (75%), AWS (50%), JavaScript (50%), Python (50%), REST APIs (50%), Stakeholder management (50%) — counted across their open postings here.

Data Engineer jobs in India · Remote Data Engineer jobs · AWS jobs · Docker jobs · HTML jobs · JavaScript jobs

Jobgether has 4,246 open roles listed here.

Counted across 14 company job boards, updated as roles open and close.

Preparing for this interview

Interviews for scraping roles keep coming back to Playwright, Selenium, HTML, AWS. Practise those questions before you sit with Jobgether.

Questions you are likely to be asked

  1. Why do you want to join Jobgether?
  2. What is your experience with Observability? Tell me one thing you learned the hard way.
  3. When would you not use machine learning for a problem?
  4. Walk me through a model you built, from the data to how it was used.
  5. How did you know your model was actually good, and not just good on your test set?

Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.

Practise the Senior Data Engineer (Web Scraping) at Jobgether interview free →

Accountabilities:: Own the development, deployment, and ongoing operation of web-scraping and web-data ingestion pipelines. Design and establish scalable web-scraping frameworks with reusable patterns for extraction, scheduling, storage, monitoring, validation, and failure handling. Build, maintain, and improve reliable production scrapers for both new and existing data sources. Investigate websites and determine the most appropriate acquisition method, including APIs, direct HTTP requests, HTML parsing, browser automation, or third-party tooling. Evaluate build-versus-buy options for scraping infrastructure and external services, considering capabilities, reliability, cost, operational complexity, and risk. Ensure web-data acquisition activities appropriately account for internal policies, website terms, robots.txt, access restrictions, privacy, and intellectual-property considerations, escalating unclear situations when required. Diagnose and resolve scraping challenges related to website changes, dynamic content, authentication, sessions, rate limits, concurrency, and other operational constraints. Integrate scraping workloads into scalable data-platform and lakehouse architectures. Improve scheduling, monitoring, storage, validation, and operational support for scraping workloads. Use AI-assisted engineering tools where appropriate while maintaining a thorough understanding of, and accountability for, the code being delivered. Support production workloads through monitoring, debugging, maintenance, and continuous improvement. Contribute to a remote engineering environment through code reviews, documentation, ticket-based workflows, and knowledge sharing. Requirements Demonstrated professional experience building and operating production web-scraping systems at scale . Proven ability to independently take substantial scraping projects from initial investigation through implementation, deployment, and ongoing production support. Strong production-level Python engineering skills, with experience developing maintainable applications rather than standalone scripts. Hands-on experience with scraping technologies such as Requests/httpx, BeautifulSoup, Scrapy, Playwright, or Selenium . Strong practical understanding of HTTP, HTML, APIs, JavaScript-rendered websites, and browser/network behavior. Experience addressing common scraping challenges including pagination, authentication, sessions, retries, rate limiting, concurrency, and proxies. Strong understanding of data pipelines, data quality, and how collected data should be validated, stored, and consumed by downstream systems. Experience deploying, monitoring, and supporting production workloads in a cloud environment. Strong debugging, analytical, and problem-solving abilities, with the judgment to make effective engineering decisions independently. Comfortable working within a remote engineering team and participating in code reviews, documentation, and ticket-based development workflows. Experience with AWS is desirable. Familiarity with lakehouse or data-lake architectures, particularly Apache Iceberg , is a plus. Experience with PySpark or other distributed data-processing technologies is beneficial. Familiarity with Docker and containerized workloads is advantageous. Experience with Terraform or other infrastructure-as-code tools is a plus. Familiarity with Grafana or comparable observability platforms is desirable. Experience operating high-volume or distributed crawling systems is beneficial. Experience evaluating or operating commercial scraping, proxy, or browser-infrastructure services is a plus. Experience implementing automated scraper testing, canary runs, or source-drift detection is desirable. Exposure to legal, compliance, privacy, or data-governance processes related to web-data acquisition is advantageous. Strong ownership, autonomy, documentation, communication, and collaboration skills. Benefits Fully remote position within a remote-first technology team. Opportunity to take ownership of a critical web-data acquisition capability and influence its architecture and operating standards. Senior, hands-on role with substantial autonomy across investigation, engineering, deployment, and production support. Work on scalable data pipelines and modern lakehouse architectures supporting research and data products. Exposure to cloud infrastructure, distributed processing, observability, browser automation, APIs, and production scraping technologies. Opportunity to establish reusable engineering patterns and improve the reliability and scalability of data ingestion. Collaboration with a distributed engineering team through code reviews, documentation, and structured workflows. Environment that supports independent problem-solving, technical ownership, and continuous improvement. Fully remote setup available across the relevant distributed team environment.

Match this job to your CV

ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.

Check my match →

Similar open roles

Need answers during your interview? Try Live Sarthi.

Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.

Try Live Sarthi free →

A Windows app, from the same team as ApplySarthi.

Listed on lever · posted 2026-09-30. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.