Own the research, technical evaluation, ingestion, cleansing, standardization, computation, storage, and servicing of financial market data, covering securities master data, real-time and historical market quotes, fundamentals, corporate actions, indices, and product/risk data as required by the business. For content-type data such as announcements, news, and research reports, own source ingestion, raw retention, and stable delivery to the knowledge engineering pipeline.
Design scalable unified data models and ingestion frameworks that handle varying market conventions for trading calendars, time zones, currencies, security identifiers, listing relationships, lifecycle events, and data corrections, enabling rapid onboarding of new markets and sources.
Build batch-stream unified data pipelines centered on Flink, continuously optimizing latency, throughput, query performance, stability, and cost, while supporting consumer-facing trading products, research and analysis, and AI use cases.
Establish data quality and service-level frameworks, taking ownership of completeness, accuracy, timeliness, consistency, and traceability. Build capabilities for automated reconciliation, anomaly detection, monitoring and alerting, raw data replay, backfill, and disaster recovery.
Evaluate the coverage, quality, stability, revision mechanisms, and technical compatibility of various data sources—including vendors, exchanges, APIs, file feeds, and compliant collection. Collaborate with product, procurement, legal, and compliance teams to define boundaries for usage, display, derivation, storage, and redistribution, and drive reasonable primary/backup source and fallback strategies.
Partner with trading product, data platform, AI engineering, and algorithm teams to jointly define data semantics, metric definitions, and service contracts, ensuring that the same stock facts can be used consistently and reliably across different products.
Drive data engineering efficiency and technical quality improvements, including metadata management, data lineage, automated testing, CI/CD, task orchestration, capacity governance, and AI-assisted development.