
Senior Data Engineer
Procter & Gamble · CINCINNATI GENERAL OFFICES
Job description
Job Location CINCINNATI GENERAL OFFICES Job Description As a Senior Data Engineer, you will be a technical leader and expert, responsible for architecting, designing, and implementing highly scalable and robust cloud-based data and analytics platforms (DAP) and complex data pipelines. You will drive the strategy for acquiring, cleansing, transforming, and publishing critical data assets from diverse enterprise and external sources. Beyond building cutting-edge data solutions, you will act as a principal liaison, partnering deeply with senior business stakeholders, solution architects, and analytics leaders to define technical roadmaps, significantly influence data architecture, and establish enterprise-wide engineering standards and best practices. We seek individuals who are not only masters of current technologies but are also visionary, continuously exploring and integrating emerging data engineering paradigms and tools to push the boundaries of what's possible. Key Responsibilities Strategic Technical Leadership:Lead the architectural design and implementation of complex, large-scale data solutions, ensuring scalability, performance, security, and cost-efficiency. Partner with senior business stakeholders and product owners to deeply understand strategic business objectives and translate them into architectural blueprints and technical roadmaps for data platforms. Influence and drive the overall data strategy, architecture, and technology choices across multiple teams or domains. Advanced Data Platform & Pipeline Development:Architect, build, and optimize highly resilient, performant, and secure ETL/ELT pipelines on modern cloud data platforms, handling petabyte-scale data volumes and real-time processing requirements. Design and implement advanced data integration patterns, connecting complex enterprise systems, third-party services, streaming sources, and APIs, ensuring high data availability and reliability. Drive the adoption of advanced data processing techniques (e.g., stream processing, graph databases, data mesh principles). Mentorship & Community Building:Serve as a primary technical mentor and subject matter expert for a team of data engineers, providing guidance on complex technical challenges, architectural decisions, and career development. Lead code reviews, design discussions, and technical workshops, fostering a culture of excellence and continuous improvement. Champion and evolve our enterprise-wide engineering standards, best practices, and governance for data (e.g., data quality frameworks, testing automation, CI/CD pipelines, security protocols, documentation standards, data observability). End-to-End Ownership & Operational Excellence:Take ultimate end-to-end ownership for critical data solutions, from strategic inception and architectural design through implementation, deployment, advanced monitoring, performance tuning, and incident response for production systems. Implement robust data quality frameworks, observability solutions, and anomaly detection to ensure the highest integrity and reliability of data assets. Innovation & AI Integration:Proactively evaluate, prototype, and integrate cutting-edge technologies, including advanced Generative AI models and sophisticated agentic systems, to dramatically enhance developer productivity, automate complex tasks, and create novel data solutions. Act as a thought leader in the responsible and ethical application of AI in data engineering, ensuring best practices for security, privacy, and bias mitigation. Lead initiatives for continuous learning and knowledge sharing across the broader engineering organization. Modern Development Practices:Master modern development tools and practices, including advanced IDE features, sophisticated Git strategies (e.g., monorepos, gitflow), infrastructure as code (IaC), and advanced CI/CD pipelines tailored for data platforms. Job Qualifications Required: Education: Bachelor's or Master's degree in Computer Science, Data Engineering, or a closely related quantitative field. Experience: 5+ years of progressive experience in data engineering, with a significant track record of designing and delivering large-scale, complex data platforms and pipelines. Technical Leadership: Proven experience leading technical projects, mentoring senior and junior engineers, and influencing architectural decisions across multiple teams. Advanced Python & SQL: Expert-level proficiency in Python and SQL for complex data manipulation, optimization, performance tuning, and advanced analytics. Deep Cloud Expertise: Expert-level understanding and hands-on experience with at least one major modern cloud platform (Azure preferred, and/or GCP), including deep knowledge of their data services (e.g., Azure Synapse, Databricks, Data Factory, Event Hubs, Data Lake Storage; or GCP BigQuery, Dataflow, Pub/Sub, Cloud Storage). Distributed Processing Mastery: Extensive hands-on experience and deep understanding of distributed data processing technologies (e.g., Spark, PySpark, Dask), including performance optimization, cluster management, and resource allocation for petabyte-scale data. Data Modeling & Architecture: Expert-level knowledge of advanced data modeling techniques (dimensional, Kimball, Inmon, data vault, data mesh concepts), data warehousing principles, and data lake architectures. Ability to design highly optimized and flexible data schemas. API & Integration Expertise: Proven ability to architect and implement complex data integrations with a wide array of systems, including advanced API integrations, message queues (e.g., Kafka, Azure Event Hubs), and enterprise-grade data transfer protocols. DevOps & MLOps for Data: Extensive experience with modern development tools, CI/CD pipelines, infrastructure as code (Terraform, ARM templates), and best practices for deploying, monitoring, and managing data and machine learning pipelines in production. AI Integration & Responsible AI: Demonstrated practical experience and leadership in leveraging Generative AI tools and agentic systems to accelerate development and solve complex data problems. Deep understanding of responsible AI principles, including data privacy, security, and ethical considerations. Communication & Influence: Exceptional communication, presentation, and interpersonal skills, with the ability to articulate complex technical concepts to both technical and non-technical senior stakeholders and influence strategic decisions. Strategic Ownership: Demonstrated ability to drive initiatives from conception to completion, taking full architectural and operational responsibility for critical data assets. Continuous Innovation: A profound curiosity and passion for continuous learning, staying abreast of industry trends, and proactively evaluating and adopting emerging data technologies. Preferred: Azure Specialization: Deep expertise and certifications in Azure data services (e.g., Azure Databricks, Azure Synapse Analytics, Azure Data Factory, Azure Stream Analytics). Advanced Data Governance: Experience implementing robust data governance, master data management (MDM), and data lineage solutions. Real-time Processing: Hands-on experience with real-time data streaming and processing frameworks (e.g., Kafka, Spark Streaming, Flink). Software Engineering Background: Strong software engineering fundamentals (design patterns, clean code principles, microservices architecture) applied to data platforms. NoSQL/Graph Databases: Experience with NoSQL databases (e.g., Cosmos DB, MongoDB) or graph databases (e.g., Neo4j) for specialized data use cases. Advanced Certifications: Professional or Expert-level certifications (e.g., Azure Data Engineer Expert, Databricks Certified Data Engineer Professional, Google Cloud Professional Data Engineer). Machine Learning/MLOps: Experience collaborating with or supporting MLOps initiatives and integrating data pipelines with ML models. Compensation for roles at P&G varies dependi
Verified and listed by ActiveJobs. Applications are made directly on Procter & Gamble's own career page — we never sit in the middle.