
Senior Data Engineer
Amgen · India - Hyderabad
Job description
Career CategoryEngineeringJob Description Join Amgen’s Mission of Serving PatientsAt Amgen, if you feel like you’re part of something bigger, it’s because you are. Our shared mission-to serve patients living with serious illnesses-drives all that we do. Since 1980, we’ve helped pioneer the world of biotech in our fight against the world’s toughest diseases. With our focus on four therapeutic areas -Oncology, Inflammation, General Medicine, and Rare Disease- we reach millions of patients each year. As a member of the Amgen team, you’ll help make a lasting impact on the lives of patients as we research, manufacture, and deliver innovative medicines to help people live longer, fuller happier lives. Our award-winning culture is collaborative, innovative, and science based. If you have a passion for challenges and the opportunities that lay within them, you’ll thrive as part of the Amgen team. Join us and transform the lives of patients while transforming your career. Senior Databricks Platform Engineer About Amgen Amgen harnesses the best of biology and technology to fight the world’s toughest diseases, and make people’s lives easier, fuller and longer. We discover, develop, manufacture and deliver innovative medicines to help millions of patients. Amgen helped establish the biotechnology industry more than 40 years ago and remains on the cutting-edge of innovation, using technology and human genetic data to push beyond what’s known today. About the RoleAmgen’s Data Platform team is seeking a senior, hands-on Databricks Subject Matter Expert to lead the design, evaluation, engineering, governance, and enterprise enablement of capabilities on the Databricks Data Intelligence Platform. The Data Platform team owns the enterprise platform architecture, governance model, policies, engineering standards, guardrails, lifecycle, observability, cost management, security controls, feature evaluations, and reusable platform services that enable data engineering, analytics, BI, data science, machine learning, and generative AI teams. This is a senior individual-contributor and technical-leadership role. The successful candidate will combine deep Databricks platform expertise with strong AWS, security, governance, FinOps, observability, automation, and AI/ML knowledge. The role will translate business and technology needs into secure, scalable, reusable, observable, and cost-efficient platform capabilities. This role enables delivery teams through paved-road patterns, frameworks, APIs, automation, reference implementations, architecture reviews, and expert guidance. It does not own the architecture or implementation of business-specific data pipelines. What you will doRoles & Responsibilities:Serve as the Databricks platform lead, shaping the platform roadmap, reference architecture, standards, operating model, guardrails, lifecycle strategy, and capability backlog in partnership with Enterprise Data Architecture, security, governance, operations, and delivery teams. Evaluate Databricks features and third-party technologies through structured assessments and PoCs. Assess architectural fit, security, privacy, compliance, interoperability, performance, reliability, cost, supportability, and user experience before recommending adoption, restriction, deferral, or rejection. Lead architecture reviews covering account, workspace, metastore, catalog, networking, storage, compute, SQL, ML/AI, dev/test/prod isolation, regional deployment, serverless and classic compute, capacity, upgrades, high availability, disaster recovery, and RTO/RPO requirements. Build reusable frameworks, accelerators, platform APIs, custom templates, provisioning workflows, policy-as-code controls, and reference implementations using Terraform, Databricks SDKs and REST APIs, the Databricks CLI, and Declarative Automation Bundles-formerly Databricks Asset Bundles. Design and operationalize integrations between Databricks and third party services like Collibra, enabling enterprise metadata, classification, ownership, lineage, business glossary, data-quality, AI-asset, and governed-tag capabilities. Define and enforce security guardrails covering identity federation, SSO/SCIM, OAuth, service principals, least privilege, segregation of duties, Unity Catalog access controls and ABAC, secrets, encryption, private connectivity, egress controls, auditing, and data-exfiltration prevention. Lead Databricks FinOps, including tagging, allocation, showback or chargeback, budgets, alerts, compute policies, serverless usage policies, cost-anomaly detection, and model or token cost controls. Optimize spend through rightsizing, Photon, autoscaling, auto-termination, workload isolation, and appropriate compute selection. Establish platform observability through SLOs, KPIs, telemetry, dashboards, and alerts for availability, reliability, jobs, queries, compute, SQL warehouses, capacity, security, adoption, and cost, using system tables, audit logs, billing data, lineage, data-quality monitoring, inference tables, and MLflow tracing. Enable governed ML, GenAI, RAG, and agent capabilities, including feature and model lifecycle, Model Serving, AI Search, foundation-model access, evaluation, monitoring, and LLMOps. Define responsible-AI guardrails for models, agents, prompts, retrieval components, MCP tools, and external providers. Improve platform reliability through root-cause analysis, operational reviews, automated remediation, runbooks, resilience testing, and elimination of recurring incidents. Drive user enablement through onboarding, documentation, reference solutions, workshops, office hours, and best-practice communities. Translate strategic-program requirements into reusable platform capabilities and measure improvements in adoption, developer experience, governance, security, reliability, and cost. What we expect of youBasic Qualifications and Experience: Master’s or Bachelor’s degree in computer science or engineering field and 9 to12 years of relevant experience Must-Have Skills: Demonstrated experience owning or technically leading an enterprise Databricks platform, extending beyond notebook or pipeline development into architecture, administration, governance, security, lifecycle management, reliability, and enablement. Broad hands-on Databricks knowledge, with expert depth across several areas such as account and workspace administration, serverless and classic compute, Apache Spark, Delta Lake, Unity Catalog, Lakeflow Jobs and Pipelines, Databricks SQL, Photon, AI/BI, MLflow, and Model Serving. Proven ability to design enterprise account, workspace, metastore, catalog, dev/test/prod, workload-isolation, promotion, high-availability, disaster-recovery, and capacity strategies, and to establish enforceable platform standards and lifecycle controls. Strong Unity Catalog and security expertise, including privileges, ownership, managed and external storage, governed tags, lineage, workspace bindings, row filters, column masks, ABAC, SSO/SCIM, OAuth, service principals, secrets, encryption, private connectivity, auditing, and data-exfiltration controls. Strong AWS experience with IAM, VPC, PrivateLink, S3, EC2, KMS, CloudWatch, CloudTrail, Secrets Manager, and STS. Familiarity with EKS, Lambda, Glue, EMR, and RDS is beneficial. Proven cost-management and observability experience, including compute rightsizing, serverless adoption, SQL warehouse optimization, Photon, budgets, tagging, showback or chargeback, system billing data, system tables, operational telemetry, dashboards, alerts, and cost-anomaly investigation. Strong understanding of Databricks AI/ML foundations, including MLflow, Models in Unity Catalog, feature engineering or Feature Store, batch and real-time inference, Model Serving, inference tables, model monitoring, MLOps, RAG, agent evaluation, AI security, and LLM cost and performance considerations. Strong Python, PySpark, SQL, automation, and platform API skills. Hands-on experience wi
Verified and listed by ActiveJobs. Applications are made directly on Amgen's own career page — we never sit in the middle.