Data Pipeline Development & Operations
• Design, build, and maintain ETL/ELT pipelines, building upon and further optimizing our existing medallion architecture (Bronze → Silver → Gold) to move data between source systems (Salesforce CRM, HubSpot, NetSuite, Stripe, DealHub, LMS) and our Databricks data warehouse.
• Build pipelines using PySpark and SQL in Databricks notebooks, following established development standards for naming, project structure, and layer-appropriate transformations.
• Own the data sync layer between Databricks and HubSpot — enrichment flows inbound to HubSpot (license status, renewal dates, subscription state, firmographic data) and marketing engagement data flowing back to Databricks (email events, workflow enrollment, lifecycle changes).
• Build and maintain Exchange layer pipelines that curate data for external system consumption, formatting and validating data to meet target system requirements.
• Build and maintain scheduled batch jobs and event-driven integrations using APIs (REST, webhooks, OAuth).
• Monitor pipeline health, set up alerting for failures and data quality degradation, and own incident response when syncs break.
• Maintain documentation of data flows, integration architecture, and troubleshooting runbooks.
• Build and maintain dimensional models in Databricks (fact tables, dimension tables, bridge tables) following our data warehouse object type definitions and naming standards.
• Work in collaboration with stakeholders and data analysts to build curated, business-ready tables and datamarts that apply business logic, KPI calculations, and aggregations optimized for analytics and campaign activation.
• Implement identity resolution and deduplication logic to produce unified customer profiles from multiple source systems.
• Establish data validation rules, quality checks, and monitoring to ensure accuracy and freshness of data flowing into marketing systems.
• Normalize disparate data sources into clean centralized schemas with proper type enforcement, deduplication, and null handling.
Marketing Data & Segmentation Support
• Ensure the data infrastructure supports audience segmentation, including firmographic, behavioral, and engagement signals.
• Build the data layer that powers lifecycle marketing - triggered campaigns, dynamic journey branching, and personalization based on enriched customer profiles.
• Support marketing and demand gen teams with reliable, accessible data for building audience targets in HubSpot.
• Maintain data flows for email deliverability, subscription management, and suppression list synchronization.
• Build and maintain API integrations between marketing, sales, and operational systems using Python and SQL.
• Implement field-level transformation logic, sync orchestration, and error handling for system-to-system data flows.
• Support website form and lead capture data flows - ensuring clean handoff from web properties into HubSpot and Databricks.
• Work with third-party enrichment providers (firmographic, intent, technographic) to integrate enrichment data into automated workflows.
• Build and maintain the data infrastructure that supports campaign attribution, channel performance analysis, and funnel reporting.
• Ensure accurate data for conversion analytics, lead source tracking, and marketing ROI measurement.
• Support centralized reporting by routing marketing engagement data back into Databricks for cross-functional analysis.