Ingestion pipeline product: define the stages, validation gates, and quality checks that data passes through from partner arrival to catalog-ready; own the platform requirements that make this repeatable across modalities and verticals
Metadata generation: own the product decisions around what metadata gets extracted or generated at ingestion, including transcripts, tags, confidence scores, schema inference, at what threshold, and how it gets stored and surfaced
QA standards and tooling: define what “catalog-ready” means, build the tooling that enforces it, and get into the data directly to validate that standards are being met; you’ll run queries and review pipeline outputs, not just read dashboards
Cross-vertical consistency: work with vertical stakeholders to translate their “what does ready mean for our vertical” requirements into consistent platform-level standards that don’t require custom engineering per deal
30 days: Ramp: Build a clear understanding of Protege’s current data ingestion workflow, including how raw partner data moves from arrival to catalog-ready. Get hands-on with pipeline outputs, schemas, metadata, validation checks, and QA processes so you understand where quality risk shows up in practice. Build context with engineering, vertical stakeholders, GTM / delivery, and DataLab on where ingestion quality is most manual, inconsistent, or risky today.
60 days: Take Ownership: Own the first clear version of what “catalog-ready” means across the ingestion pipeline, including validation gates, metadata requirements, QA standards, and readiness criteria. Translate the highest-priority ingestion quality gaps into product requirements engineering can build against.
90 days: Operate Independently: Own the roadmap for improving ingestion quality, metadata generation, QA tooling, de-identification workflows, and catalog readiness. Create a repeatable operating rhythm for reviewing pipeline outputs, quality signals, and ingestion risks with the right cross-functional partners