C
IT(Software/Hardware)

Engineer III - Data + ML Platform

CrowdStrike

Role Summary

CrowdStrike is looking for an Engineer III - Data + ML Platform to join our infrastructure engineering team in Bengaluru under a hybrid model. In this role, you will own the reliability, performance profiling, and operational stability of the mission-critical ML platforms that power our Falcon cybersecurity ecosystem—processing over 3 trillion events each day. You will debug large-scale distributed computing frameworks (Ray, Spark, Kubernetes), troubleshoot GPU scheduling bottlenecks, build deep observability tooling, and lead incident post-mortems across our training and inference clusters.

Key Responsibilities

  • Platform Reliability & Debugging: Diagnose root causes for production failures, kernel crashes, memory leaks, and distributed scheduling bottlenecks across Ray, Apache Spark, and Kubernetes.
  • Distributed Compute Optimization: Profile and optimize high-throughput distributed ML workloads across Kubernetes and public cloud environments (AWS, GCP, Azure, or OCI), maximizing GPU and HPC cluster efficiency.
  • Workflow & Tooling Maintenance: Troubleshoot and scale core ML pipeline components, including JupyterHub, SLURM, MLflow, Kubeflow, and Apache Airflow.
  • Observability & SRE Operations: Architect distributed tracing, real-time health checks, and alerting runbooks with Prometheus and Grafana to maintain stringent latency, MTTR, and SLA targets.

Key Qualifications

  • Education: Bachelor's or Master's degree in Computer Science, Software Engineering, or a related technical discipline.
  • Experience: 5+ years of hands-on experience in distributed systems engineering and debugging ML platforms in production.
  • Core Stack: Advanced Python debugging in Linux/Unix environments, deep production troubleshooting with Ray or Apache Spark, and container orchestration using Kubernetes and Docker.
  • Nice to Have: Hands-on experience with SLURM cluster management, CUDA/GPU profiling, chaos testing, open-source ML platform contributions, and large-scale model inference pipelines.
  • Professional Competencies: Disciplined operational triage, blameless post-mortem leadership, proactive mentoring, and an automation-first mindset.

About CrowdStrike

CrowdStrike is a global cybersecurity pioneer that redefined endpoint and cloud security with the world’s most advanced cloud-native, AI-driven platform built to stop breaches.

Important Candidate Notice

This job posting and its descriptions are owned by CrowdStrike . We provide this overview to help developers discover opportunities. For applications, responses, or follow-ups, please connect with CrowdStrike directly. Job Central Hub does not process applications or review candidate queries.

Profile Match Score

Sign in to see how well your skills and experience align with this job.

Sign In
Overview
Expected Salary ₹28 - ₹50 Lakhs/Annum
Experience Required 5 - 9 Years
Last Date to Apply Continuous Recruitment
Posted On 05-09-2026
Mandatory Skills
Python Ray Apache Spark Kubernetes Docker Distributed Systems Linux Prometheus Grafana
Preferred Skills
SLURM Kubeflow MLflow Apache Airflow JupyterHub AWS GCP CUDA Chaos Engineering