Senior Data Engineer – Data & Context Platform
StackGen
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
2,640 open platform roles across 402 companies are on ApplySarthi right now, most of them in Bengaluru (125), Hyderabad (42), Pune (31).
- Senior SDE - PlatformJobgether
- Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - FederalServiceNow
What platform roles keep asking for: AWS (31%), Kubernetes (30%), Python (24%), Observability (24%), GCP (21%), System design (20%), Azure (19%), CI/CD (17%) — counted across their open postings here.
Data Engineer jobs in the United States · Data Engineer jobs in San Francisco · Remote Data Engineer jobs · AWS jobs · Azure jobs · CI/CD jobs · Data modelling jobs
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for platform roles keep coming back to AWS, Kubernetes, Python, Observability. Practise those questions before you sit with StackGen.
Questions you are likely to be asked
- Why do you want to join StackGen?
- What is your experience with Kubernetes? Tell me one thing you learned the hard way.
- How did you know your model was actually good, and not just good on your test set?
- Tell me about a time the data was messy or wrong. What did you do?
- How would you explain your model's result to someone who is not technical?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Senior Data Engineer – Data & Context Platform at StackGen interview free →### About StackGen StackGen delivers an Agentic Infrastructure Platform powered by Aiden, its AI agent that enables platform engineering, DevOps, and SRE teams to move from manual processes to intent-driven automation. Our platform enables autonomous infrastructure across four pillars: building, governing, healing, while maintaining compliance and security standards across multiple cloud environments. StackGen serves enterprise and fast-growing customers and is based in the San Francisco Bay Area, with a globally distributed team. ### The Role Agents are only as good as the context they can reason over. Today that context is scattered across cloud provider APIs, IaC state, Kubernetes clusters, telemetry systems, CI/CD pipelines, and ticketing tools. Each speaking a different dialect, changing on its own schedule, and describing the same underlying resource in a different way. We are hiring a Senior Data Engineer to build the layer that solves this problem: the ingestion pipelines, the canonical data model, and the graph and retrieval system that Aiden's agents query. This is a foundational build, not maintenance of something that already works, and the architectural decisions made here will shape the platform for years. This is a hands-on individual contributor role. You will design the system and write the code that runs it. ### What you will do **Data ingestion and pipelines** Build ingestion pipelines across heterogeneous infrastructure sources: cloud provider APIs, Terraform/OpenTofu state, Kubernetes, observability backends, source control, and ticketing systems. Handle the operational reality of production pipelines: incremental sync, backfill, rate limits, retries, and graceful degradation when an upstream source is unavailable. Make the batch-versus-streaming call per source rather than applying one pattern everywhere. Design the canonical schema that unifies how infrastructure, services, ownership, and events are represented across sources. Design for schema evolution so that new sources and entity types can be absorbed without a migration crisis each time. **Context graph and retrieval** Model the relationships between infrastructure, services, ownership, deployments, and incidents so agents can traverse them to answer real operational questions. Build the retrieval layer on top: graph queries, vector search over unstructured artifacts, and whatever hybrid approach proves out in practice. Evaluate and select the storage technologies, with the tradeoffs argued explicitly rather than assumed. **Data quality, freshness, and observability** Define and enforce staleness contracts. Agents acting on stale or incorrect context is worse than agents with no context at all. Instrument the layer so data quality, coverage, and freshness are measurable rather than anecdotal, and build the validation and alerting that catches drift early. **Multi-tenancy, scale, and security** Build tenant isolation into the storage and query paths from the start, not as a later retrofit. Apply access control and data-handling practices appropriate to customer infrastructure metadata, working with security to keep controls enforceable. **Cross-functional collaboration** Partner with the agent, platform, and product teams to understand what context agents actually need, and make the layer usable enough that other engineers build on it without needing you in the loop. Write design docs and tradeoff memos that let the team engage with the architecture, and raise the bar through code review and shared practice. ### What we're looking for * 4+ years building production data systems, with substantial experience owning data pipelines end to end: ingestion, transformation, orchestration, and the 3am operational reality. * Hands-on experience with data normalization, entity resolution, or master data management in messy multi-source environments where the same entity has different names and different levels of completeness in each system. * Strong data modeling judgment. You can explain why you would choose a graph, relational, document, or hybrid store for a given problem, and you have been wrong about one before and learned from it. * Proficiency in Python (Go is a plus) and SQL well beyond the basics. * Experience designing systems from a blank page, including making decisions with incomplete information and revisiting them when reality disagrees. * Working familiarity with cloud and DevOps ecosystems — AWS/GCP/Azure resource models, Terraform, Kubernetes APIs. You should find the infrastructure domain genuinely interesting, not just the pipelines. * Clear written communication. Design documents, tradeoff memos, and code that explains why rather than what. * Comfort in a startup environment with ambiguity and high ownership. * Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience. ### Nice to have * Graph databases in production (Neo4j, Neptune, Dgraph, Memgraph, or similar), including a clear view of where they stop being the right answer. * Vector databases and retrieval systems, particularly RAG patterns serving agents rather than chat interfaces. * Streaming or CDC experience (Kafka, Debezium, Flink) and the judgment to know when a batch is the better call. * Multi-tenant SaaS data architecture. * Data quality, lineage, or catalog tooling (Great Expectations, OpenLineage, DataHub, or similar). * Exposure to agentic AI frameworks and an understanding of how context quality affects agent behavior. * Experience with infrastructure or CMDB-style data models (CAI, OCSF, or internal equivalents). ### Why StackGen * High ownership. You will set the standard for how data and context work at StackGen, as the engineer who owns this layer. * Build the foundation of an AI-driven infrastructure platform from the ground up, on a greenfield problem with real enterprise customers already depending on the outcome. * Work directly with engineering leadership and influence architecture across the platform. * A collaborative engineering culture that values continuous learning, deep technical work, and short paths from problem to fix.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on wellfound · posted 2026-09-16. ApplySarthi collects openings and links to application pages; the role is advertised by StackGen, not by us.