AIOps Leader
Airbus · Bangalore Area
Job description
Job Description: Job Description - JR Job Description Date: August 2026 Role : AIOps Leader Number of positions : 1 Description: We are seeking a seasoned, hands-on AIOps Leader (7+ Years of Experience) to serve as the principal technical architect and lead builder for our centralized AI Product Support & Operations function. Operating within a high-growth AI PSL (Product/Service Line), you will design, architect, and execute the end-to-end automation strategy that transforms raw operational chaos into scalable, self-healing, and data-driven infrastructure. In this role, you will bridge software engineering, site reliability, and AI/ML architectures. You will lead the creation of intelligent diagnostic pipelines, custom RAG-driven knowledge tools, self-healing systems, and automated triage engines. You will work closely with cross-functional leadership, L1/L2 support teams, and platform engineering to systematically eliminate operational toil, optimize MTTR, and build proactive anomaly detection mechanisms across our AI ecosystem. Qualification & Experience: Education: Bachelor’s or Master’s degree in Computer Science, Software Engineering, Information Technology, or a related quantitative field. Overall Experience: 7+ years of hands-on experience across Software Engineering, Site Reliability Engineering (SRE), DevOps, or Systems Operations—with a focused concentration on cloud infrastructure automation and AI/ML operational tooling. Platform Specialization: At least 3+ years architecting and running operations directly within major cloud ecosystems (AWS or GCP), including native AI/ML compute platforms. Key Responsibilities 1. Architecture & Advanced AI/ML Automation Self-Healing Infrastructure & Workflows: Architect, build, and deploy auto-remediation routines, script-based diagnostic runners, and event-driven automation triggers that autonomously resolve platform issues. LLM & RAG Systems Engineering: Design, implement, and maintain advanced Retrieval-Augmented Generation (RAG) knowledge tools, vector databases, and LLM utilities that index telemetry, historic logs, and RCAs for instant incident context. Bot & Agent Development: Lead the development and production rollout of conversational AI agents, custom webhooks, and self-service bots integrated into ticketing engines to automate Tier-1 and Tier-2 resolutions. 2. Observability, Telemetry & Predictive Analytics Observability Architecture: Build enterprise-grade telemetry ingestion workflows, automated log scraping, and context-enrichment pipelines that dynamically append system metrics directly to incident tickets upon creation. Predictive Anomaly Detection: Configure and tune real-time predictive alerting, log-pattern analysis, and AI-driven monitoring models across AWS, Azure, or GCP microservices. Operational BI & Analytics: Architect and own centralized executive and operational dashboards (e.g., ServiceNow, Datadog) tracking MTTR velocity, system uptime, defect density, ticket deflection rates, and SLA/CSAT compliance. 3. Operational Engineering & L1/L2 Empowerment Toil Elimination: Continuously audit support operational bottlenecks across product teams, transforming high-volume manual intervention points into production-grade, single-click, or fully autonomous workflows. Log & Metadata Standardization: Standardize system log outputs, stack-trace formatting, and tagging taxonomy across all AI products to ensure platform telemetry remains machine-readable for AI engines. Technical Mentorship: Guide L1/L2 support engineers on best practices for automation, code-based triage, and log parsing. Technical EssentialsAWS & GCP Native AI Platforms: Deep hands-on experience orchestrating production AI/ML workflows on AWS (Bedrock, SageMaker AI, OpenSearch, AWS Lambda) or GCP (Vertex AI, Vertex AI Agent Builder, Cloud Run, BigQuery). Cloud Infrastructure & Infrastructure-as-Code (IaC): Advanced experience writing and managing cloud provisioning scripts using Terraform, AWS CloudFormation, or Google Cloud Deployment Manager to deploy auto-scaling, resilient operations environments. Containerization & Orchestration: Production experience managing microservices via Kubernetes (EKS/GKE) and Docker to support agentic AI workers, vector indexing engines, and automated micro-tasks. Observability & Cloud Telemetry: Proven capability to configure full-stack observability across cloud environments using AWS CloudWatch or GCP Cloud Logging/Monitoring to trigger automated alerts and log enrichment. Advanced Automation & Scripting: Strong engineering capability in Python, Go, or Shell to build autonomous cloud functions (AWS Lambda/GCP Cloud Run), self-healing infrastructure scripts, and custom ITSM connectors (Jira/ServiceNow APIs). Enterprise Generative AI Stack: Production execution experience deploying RAG (Retrieval-Augmented Generation) architectures using cloud vector engines (Amazon Bedrock Knowledge Bases, GCP Vertex AI Search, Pinecone, or Qdrant) for automated incident context retrieval. Soft Skills & Behavioral AttributesTechnical Leadership & Influence: Proven ability to architect complex automation systems independently, establish technical standards, and convince cross-functional teams to adopt new operational models. Root-Cause Obsession: Relentless drive to solve underlying architectural flaws rather than applying temporary operational workarounds. Cross-Functional Bridge: Ability to speak fluidly with product developers, AI researchers, enterprise clients, and technical support teams. Thrives in Ambiguity: Highly adaptable problem solver capable of establishing order, structure, and code standards in unstructured environments. Nice-to-Have / Added AdvantagesLinguistic Skill: Professional working proficiency or full fluency in French (spoken and written). Certifications: ITIL v5 Foundation or Managing Professional, Certified Agile Service Manager (CASM), PMP, or introductory Cloud/AI certifications (AWS Certified Cloud Practitioner, Azure AI Fundamentals). AI Bot Implementation Experience: Direct hands-on experience implementing conversational AI, RAG-based internal search, or auto-remediation bots within support portals. Success MetricsMTTR (Mean Time to Resolution) - Reduction in overall incident MTTR via automated routing, diagnostic bots, and clear escalation paths . Ticket Deflection & Reduction - Percentage of Tier-1 issues resolved via self-service bots, automated workflows, and root-cause resolution. SLA & First Contact Resolution - High compliance across overall uptime, response, and resolution SLAs for all AI products. Operational Centralization - Onboarding of existing and newly building AI products into the unified service delivery framework. Stakeholder & CSAT Score - Satisfaction scores from internal product units, business leadership, and end-users. This job requires an awareness of any potential compliance risks and a commitment to act with integrity, as the foundation for the Company’s success, reputation and sustainable growth. Company: Airbus India Private Limited Employment Type: Permanent------- Experience Level: Professional Job Family: Digital By submitting your CV or application you are consenting to Airbus using and storing information about you for monitoring purposes relating to your application or future employment. This information will only be used by Airbus. Airbus is committed to achieving workforce diversity and creating an inclusive working environment. We welcome all applications irrespective of social and cultural background, age, gender, disability, sexual orientation or religious belief. Airbus is, and always has been, committed to equal opportunities for all. As such, we will never ask for any type of monetary exchange in the frame of a recruitment process. Any impersonation of Airbus to do so should be reported to emsom@airbus.com. At Airbus, we support you to work, connect and collaborate more easily and
Verified and listed by ActiveJobs. Applications are made directly on Airbus's own career page — we never sit in the middle.