Sr. Product Manager - Observability
Workday · USA, CA, Pleasanton
Job description
Your work days are brighter here. We’re obsessed with making hard work pay off, for our people, our customers, and the world around us. As a Fortune 500 company and a leading AI platform for managing people, money, and agents, we’re shaping the future of work so teams can reach their potential and focus on what matters most. The minute you join, you’ll feel it. Not just in the products we build, but in how we show up for each other. Our culture is rooted in integrity, empathy, and shared enthusiasm. We’re in this together, tackling big challenges with bold ideas and genuine care. We look for curious minds and courageous collaborators who bring sun-drenched optimism and drive. Whether you're building smarter solutions, supporting customers, or creating a space where everyone belongs, you’ll do meaningful work with Workmates who’ve got your back. In return, we’ll give you the trust to take risks, the tools to grow, the skills to develop and the support of a company invested in you for the long haul. So, if you want to inspire a brighter work day for everyone, including yourself, you’ve found a match in Workday, and we hope to be a match for you too. About the Team The Data Platform and Observability Engineering (DPOE) team is building Workday's next-generation, multi-petabyte scale Observability Platform. We own the libraries, distributed services, and infrastructure that power ingestion, storage, and query across the observability stack — Iceberg, ClickHouse, Tempo, Mimir, Grafana, S3, Kafka, and Elasticsearch — serving traces, metrics, and logs for every workload at Workday. How we help our customers: Every engineering, SRE, and product team at Workday depends on us to see inside their systems — from a single service's latency spike to a cross-service cascading failure impacting thousands of customers. By owning the full observability data lifecycle at petabyte scale, we give internal teams the speed and confidence to find and fix issues before they affect Workday's customers. Our roadmap directly shapes how the company detects, diagnoses, and eventually predicts operational issues at scale — turning observability from a reactive debugging tool into a proactive, AI-assisted safety net for every workload running on Workday's platform. About the Role We are looking for a hands-on, technical Senior Product Manager (P4) to own and drive our Observability strategy, with a strong emphasis on Distributed Tracing. You will define the product vision for how engineers understand, debug, and optimize complex distributed systems, with a particular focus on building AI-powered detection and triage capabilities that reduce time-to-detect and time-to-resolve production issues. This role requires someone comfortable diving deep into technical architecture discussions, reading code/traces, and partnering closely with engineering and applied ML teams to ship technically sound, high-impact products. What You'll DoOwn the product vision, strategy, and roadmap for Observability, with a primary focus on Distributed Tracing capabilities (trace context propagation, sampling strategies, span analysis, service maps, latency/error analysis, etc.) Define and drive the roadmap for AI-enabled anomaly detection, including specifying requirements for statistical and ML-based detection methods (e.g., time-series forecasting, seasonality-aware baselining, change-point detection, multivariate anomaly detection across correlated metrics/traces/logs) Partner with ML engineering to define model requirements, evaluation metrics (precision/recall, false-positive rate, alert-to-incident ratio), and feedback loops for continuous model improvement Define requirements for LLM-based root cause summarization and triage assistance — e.g., generating human-readable incident summaries from raw trace/log/metric data, ranking probable root causes, suggesting remediation runbooks based on historical incident patterns Specify how confidence scores, explainability, and human-in-the-loop review are surfaced in the triage workflow so on-call engineers can trust and act on AI-generated recommendations Partner closely with engineering teams to define technical requirements, evaluate architectural trade-offs, and make hands-on contributions to product decisions (e.g., reviewing design docs, understanding OpenTelemetry/OTel standards, tracing protocols, and instrumentation approaches) Define and track success metrics (detection precision/recall, MTTD, MTTR, alert noise reduction, triage automation rate, trace coverage) to measure product impact Conduct customer and internal stakeholder research to identify pain points in debugging, alert fatigue, and incident triage Write clear, detailed product requirements, user stories, and specs; work closely with design, engineering, and data science to bring them to life Stay current on industry trends in observability (OpenTelemetry, eBPF-based tracing, service mesh telemetry) and applied AI/ML for anomaly detection, incident correlation, and LLM-based operational tooling About You Basic Qualifications 8+ years of product management experience, with meaningful time spent in Observability, Monitoring, APM, or related infrastructure/developer tooling domains Deep, demonstrable expertise in Distributed Tracing concepts and technologies (e.g., OpenTelemetry, Jaeger, Zipkin, trace sampling, span/context propagation) Strong technical background — comfortable reading code, understanding system architecture, APIs, and engaging directly in technical design discussions Concrete, hands-on understanding of AI/ML techniques applied to detection and triage, including:Anomaly detection methods (statistical thresholds, time-series forecasting models, change-point/seasonality detection) Alert correlation and clustering techniques (embedding/vector similarity, graph-based dependency analysis, unsupervised clustering) LLM-based summarization and reasoning for incident root-cause analysis and runbook suggestion Model evaluation frameworks (precision/recall, false-positive/false-negative trade-offs, drift monitoring) Experience partnering directly with ML engineering teams to translate detection and triage requirements into shippable models and features Proven track record of shipping complex, technical products from concept to launch Excellent written and verbal communication skills; ability to translate complex technical and ML concepts for varied audiences Strong analytical and data-driven decision-making skills Other QualificationsPrior experience as a software engineer, SRE, ML engineer, or data scientist before transitioning to product management Familiarity with observability platforms (Datadog, New Relic, Honeycomb, Grafana, Splunk, Dynatrace, etc.) and their AI/ML-driven detection features Direct experience productionizing ML models (e.g., anomaly detectors, classifiers, LLM-based tools) for operational/incident-management use cases Familiarity with vector databases/embedding search, graph-based correlation techniques, or LLM prompt/agent design for triage automation Background in large-scale, cloud-native, microservices-based systems Workday Pay Transparency Statement The annualized base salary ranges for the primary location and any additional locations are listed below. Workday pay ranges vary based on work location. As a part of the total compensation package, this role may be eligible for the Workday Bonus Plan or a role-specific commission/bonus, as well as annual refresh stock grants. Recruiters can share more detail during the hiring process. Each candidate’s compensation offer will be based on multiple factors including, but not limited to, geography, experience, skills, job duties, and business need, among other things. For more information regarding Workday’s comprehensive benefits, please click here. Primary Location: USA.CA.Pleasanton Primary Location Base Pay Range: $168,000 USD - $252,000 USD Additional US Location(s) Base Pay Range: $140,600 USD -
Verified and listed by ActiveJobs. Applications are made directly on Workday's own career page — we never sit in the middle.