ApplySarthi Match jobs to your CV

Member of Technical Staff | Observability & Reliability

Jobgether

Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.

Skills named in this job

Read from the description itself, not inferred.

This role on the market

138 open observability roles across 58 companies are on ApplySarthi right now, most of them in Bengaluru (6), Hyderabad (3), Delhi NCR (2).

What observability roles keep asking for: Observability (79%), Kubernetes (37%), System design (36%), Elasticsearch (25%), AWS (25%), Python (23%), Product management (20%), Go (19%) — counted across their open postings here.

Member of Technical Staff jobs in the United States · Remote Member of Technical Staff jobs · AWS jobs · GCP jobs · Kubernetes jobs · Observability jobs

Jobgether has 3,838 open roles listed here.

Counted across 14 company job boards, updated as roles open and close.

Preparing for this interview

Interviews for observability roles keep coming back to Observability, Kubernetes, System design, Elasticsearch. Practise those questions before you sit with Jobgether.

Questions you are likely to be asked

  1. Why do you want to join Jobgether?
  2. What is your experience with Kubernetes? Tell me one thing you learned the hard way.
  3. Tell me about a problem you solved at work that you are proud of.
  4. Tell me about a time you disagreed with your manager. What happened?
  5. Where do you want to be in three years?

Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.

Practise the Member of Technical Staff | Observability & Reliability at Jobgether interview free →

Accountabilities: Evolve and maintain the observability platform covering logs, metrics, traces, alerting, and system health across cloud and customer-hosted dataplanes. Ensure every environment reports critical operational information, including active releases, health status, heartbeats, logs, metrics, and usage to the central control plane. Implement telemetry collection within customer Kubernetes environments using outbound-only connectivity models. Detect and investigate differences between desired infrastructure or deployment state and what is actually running in each environment. Monitor the health and availability of deployment and runtime agents, including ephemeral workloads such as Ray clusters supporting batch inference. Define and maintain Service Level Objectives (SLOs), establish actionable alerting, and contribute to error-budget practices. Lead or participate in incident response and postmortems, identifying improvements that reduce recurring failures and mean time to recovery (MTTR). Coordinate incident resolution across internal teams and customers when fixes involve customer-managed environments. Optimize telemetry pipelines to reduce redundant data, control infrastructure costs, and improve the signal-to-noise ratio of operational information. Write production-quality code, review technical changes, and take operational ownership of the systems you build. Requirements Deep professional experience with OpenTelemetry and modern observability platforms or backends. Hands-on experience defining SLOs, working with error budgets, designing actionable alerts, and managing production incidents. Strong experience with Kubernetes and infrastructure-as-code tools such as Terraform and Helm. Experience operating software across distributed or customer-hosted environments where infrastructure and connectivity may not be fully under your control. Strong software engineering fundamentals, with experience producing maintainable, production-ready code and conducting effective code reviews. Willingness to participate in operational ownership, troubleshooting, incident response, and continuous reliability improvements. Strong analytical and problem-solving abilities, with an ability to investigate complex distributed-system behavior. Experience communicating clearly across engineering teams and, when required, working directly with external customers or stakeholders. Experience with GCP/GKE or AWS/EKS is an advantage. Familiarity with multi-node or multi-cluster ML workloads in production is a plus. Experience deploying software to customer-hosted Kubernetes environments, including Helm-based deployments and outbound-only connectivity, is valuable. Experience in financial services or other regulated environments is an additional advantage. Benefits Fully remote role based in Brazil. Full-time position within an engineering-focused environment. Opportunity to own critical observability and reliability systems with direct impact on production availability. Work across cloud and customer-hosted Kubernetes environments, providing broad exposure to distributed infrastructure. Opportunity to work with modern observability, Kubernetes, infrastructure-as-code, and ML infrastructure technologies. High degree of technical ownership, with responsibility for both building and operating the systems you develop. Exposure to complex reliability challenges involving real-time services, batch workloads, customer environments, and distributed systems. Opportunity to contribute to incident management, platform architecture, and long-term reliability practices.

Match this job to your CV

ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.

Check my match →

Similar open roles

Need answers during your interview? Try Live Sarthi.

Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.

Try Live Sarthi free →

A Windows app, from the same team as ApplySarthi.

Listed on lever · posted 2026-09-23. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.