AI Data Engineer
pst.ag
Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.
Got this interview? Our apps help you get the job.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
7,366 open data roles across 591 companies are on ApplySarthi right now, most of them in Bengaluru (435), Hyderabad (326), Mumbai (157).
- Senior Data Engineer – Data Analytics & BIJobgether
- Sr. Data Center Operations EngineerZscaler
- Sr Staff Outbound Product Manager - Data & AnalyticsServiceNow
- Senior Data Scientist, Product AnalyticsHandshake
- Account Executive / Account Manager, Data Sales at KalshiKalshi
What data roles keep asking for: AWS (25%), Python (23%), SQL (23%), Machine learning (12%) — counted across their open postings here.
CI/CD jobs · CSS jobs · Docker jobs · LLMs jobs
pst.ag has 4 open roles listed here.
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for data roles keep coming back to AWS, Python, SQL, Machine learning. Practise those questions before you sit with pst.ag.
Questions you are likely to be asked
- Why do you want to join pst.ag?
- What is your experience with LLMs? Tell me one thing you learned the hard way.
- When would you not use machine learning for a problem?
- Walk me through a model you built, from the data to how it was used.
- How did you know your model was actually good, and not just good on your test set?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the AI Data Engineer at pst.ag interview free →Specification-Driven Extraction Engineering: * Design and maintain declarative extraction specifications—using Pydantic models, JSON schemas, or domain-specific languages—that describe exactly which fields to capture, their types, and validation rules. * Implement pipelines that translate these specifications into executable extraction plans, leveraging both classical (Scrapy, Playwright) and AI-augmented (LLM-based semantic parsing) backends. * Build reusable specification libraries for recurring data types (product prices, tariff codes, regulatory texts) to accelerate onboarding of new sources. Autonomous & Self-Healing Systems: * Deploy self-healing spiders that automatically detect website layout changes and repair themselves using Model Context Protocol (MCP) servers (e.g., Scrapy MCP Server, Playwright MCP). * Integrate semantic extraction (Scrapy-LLM, custom LLM pipelines) to eliminate selector brittleness—spiders rely on field descriptions, not fragile XPaths. * Hands-on experience building AI agents and orchestration systems. * Orchestrate complex, multi-step browsing workflows with agentic frameworks (BMAD/TEA, AutoGPT-like agents) that reason about page state, adapt to anti-bot measures, and correct their own behaviour in real time. Platform Thinking & Reusability: * Move beyond one-off scrapers: build a component-based extraction platform where selectors, login handlers, and pagination logic are shared, versioned, and tested. * Implement monitoring, alerting, and automatic rollback for failed extraction runs. * Champion ethical crawling by design—rate limiting, robots.txt respect, and compliance with GDPR/CCPA are built into the specification layer, not retrofitted. Collaboration & Continuous Innovation: * Partner with data scientists and domain experts to refine extraction specifications for complex, unstructured domains (e.g., legal texts, tariff classifications). * Evaluate and pilot emerging tools to push automation coverage beyond 90%. * Document and evangelise specification-driven best practices across the engineering organisation. Qualification: * Bachelor’s degree in Computer Science * 3+ years of experience in web scraping or data extraction Required Skills: * Proficiency with Python * Experience with specification-Driven Extraction * Experience with LangChain, LangGraph, LlamaIndex, AutoGen * Hands‑on use of Scrapy‑LLM, Scrapy MCP Server, or similar systems that decouple field definitions from page structure * Familiarity with frameworks that give LLMs browser control (Playwright + MCP, BMAD/TEA) to handle complex, non‑deterministic crawling tasks. * Design and implement autonomous data extraction agents that can make decisions about source selection, retry logic, and parsing strategies * Classical Scraping Fundamentals * Data Validation & Storage – Ability to define validation rules within specifications and land clean data into SQL/NoSQL databases or data lake * Basic API integration and authentication flows. * HTTP, DOM, XPath, CSS. Nice to Haves: * Contributions to open-source scraping or AI-automation projects. * Contributions to open-source scraping or AI-automation projects. * Familiarity with data privacy engineering (GDPR, CCPA) baked into specification design. * DevOps light – Docker, CI/CD for testing extraction specifications. Mindset & Approach (Non-Negotiable): * Strong belief that the future of scraping is declarative, not imperative. * Candidate rather write a schema that says “extract the price” than debug an XPath when a website redesigns. * Looking to shift from “code that scrapes” to “systems that understand extraction”
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on wellfound · posted 2026-09-21. ApplySarthi collects openings and links to application pages; the role is advertised by pst.ag, not by us.