[Responsibilities]
· Data operations: own day-to-day operations of data platforms/pipelines capacity, stability, upgrades, deployments, and recovery drills to sustain high availability and low latency.
· Data collection: design/manage multi-source ingestion (exchanges, internal and external systems), protocol parsing, and robust retry mechanisms.
· Develop rule-based and statistical data quality checks (completeness, uniqueness, time alignment, anomaly detection, error handling).
· Implement automated remediation, reconciliation workflows, and historical backfilling.
· Establish monitoring and alerting frameworks to ensure trusted, production-grade datasets.
· End-to-End pipelines: plan and maintain scalable ETL/ELT including scheduling, caching, partitioning, modelling, schema evolution, and lineage to support both batch and real-time streaming.
· Enforce data access controls, encryption, auditing, and classification to comply with internal policies and external regulatory requirements (including PII management).
· Apply Infrastructure-as-Code, data versioning, data tests, and CI/CD to improve predictability and reduce manual risk.
· Contribute to embedded GenAI and LLM-powered data applications for enterprise analytics, reconciliation, and internal productivity use cases.
· Partner with analytics and product teams to operationalize AI-driven data solutions.