Design, build, and maintain agentic systems that turn a pointer to a biomedical data source (web page, S3 path, GitHub repository, a table in a paper) into a reviewed, versioned dataset. The agent’s output is committed pipeline code and its test suite, so re-running it later reproduces the same dataset. Develop LLM-based pipelines for data cleaning, normalization, and formatting across diverse data modalities (e.g., molecular, genomic, clinical, literature)