Data Platform Architecture & Modernization
· Lead the design and implementation of scalable, cloud-native data architectures on Databricks (Delta Lake, Unity Catalog, Lakehouse patterns).
· Own and execute the migration strategy from legacy Azure SQL Server to Databricks, including schema translation, ETL/ELT re-platforming, data validation, and cutover planning.
· Define data modeling standards (medallion architecture, star/snowflake schemas) and ensure consistency across all pipelines and domains.
· Evaluate and recommend tools, frameworks, and platforms to support long-term data strategy and organizational goals.
· Collaborate with Solution and Enterprise Architects to review and approve new data architecture designs.
Streaming & Real-Time Data Engineering
· Architect and implement Kafka-based streaming pipelines for real-time data ingestion, transformation, and delivery.
· Design event-driven architectures and streaming topologies using Kafka Streams, ksqlDB, or Spark Structured Streaming on Databricks.
· Establish patterns for schema management (Confluent Schema Registry), consumer group strategy, offset management, and dead-letter queuing.
· Ensure streaming pipelines meet SLA requirements for latency, throughput, and fault tolerance.
Optimization, Quality & Standards
· Debug and optimize Spark jobs, Delta Lake tables, and SQL workloads for performance, cost efficiency, and maintainability.
· Lead code reviews focused on senior engineers to enforce standards, best practices, and technical quality.
· Manage pipeline quality, data models, and CI/CD delivery workflows; guide teams on continuous improvement.
· Promote strong data management practices — data quality, lineage, observability, and governance.
· Identify opportunities to improve service delivery methods, processes, and resource utilization.
Leadership, Mentorship & Strategy
· Mentor data engineers at all levels, with particular emphasis on developing senior talent.
· Establish and evolve data engineering standards, best practices, and management of technical debt.
· Stay current on data platform trends (Databricks releases, Kafka ecosystem, open table formats) and contribute to long-term architectural vision.
· Develop plans for data security, disaster recovery, backup, business continuity, and archiving across the Lakehouse.
AI, Agentic Capabilities & Intelligent Data Products
· Design and enable AI and agentic capabilities on the Databricks platform, including Databricks Genie for natural language data exploration and self-service analytics.
· Architect data foundations — clean, governed, well-documented Delta tables — that power Genie spaces, AI/BI dashboards, and LLM-driven data agents.
· Collaborate with ML and AI teams to build and maintain feature stores, vector stores, and retrieval-augmented generation (RAG) pipelines on Databricks.
· Evaluate and integrate emerging agentic frameworks (LangChain, Mosaic AI Agent Framework) to automate data workflows and enable intelligent data products.
· Define governance and observability standards for AI-driven data pipelines, ensuring reliability, auditability, and responsible AI practices.
Cross-Functional Collaboration
· Partner with project managers and business leaders on initiatives involving enterprise data.
· Collaborate across teams to influence and strengthen data engineering practices organization-wide.
· Work with Analytics, ML, and product engineers to design and deliver end-to-end data solutions that meet business needs.