•
Perform model validations on GenAI and agentic AI model use cases across The Hartford’s functional areas and lines of business to ensure models are performing effectively and efficiently
− Ensure model calculations, algorithms, and methods are accurate and appropriate for intended use
− Design and build challenger solutions and/or testing methods for tasks such as summarization, question answering, search, data synthesis, LLM-as-a-judge etc.
− Review and assess quantitative and qualitative testing techniques to ensure model accuracy, robustness, and reliability
− Assess key data inputs, assumptions, prompt engineering, and context engineering for accuracy and appropriateness
− Design and execute testing approaches appropriate for non-deterministic AI systems, including variability analysis, confidence intervals, and statistical evaluation methodologies.
− Deliver effective challenge to key modeling elements such as inputs, calculations, outputs, conceptual soundness, monitoring & controls, documentation, etc.
− Identify findings and recommendations, including impact analysis, to mitigate model risk and compile clear and concise model validation reports
•
Perform governance accountabilities related to findings tracking, remediation testing, and validation
•
Assist in enhancing the existing GenAI model validation framework to include standardized evaluation metrics for performance and reliability, deployment of model validation tools for increased efficiency, and ensure continued alignment with regulatory standards
•
Strengthen partnerships with Data Science teams to keep model risk practices aligned with the proliferation and sophistication of modeling, promote proactive risk management, and share best practices.
•
Pro-actively stay informed with advancements in AI/ML, GenAI, agentic systems, and regulatory expectations for emerging technologies and of department initiatives, deliverables, and reporting
•
Assist with the evaluation and testing of cutting-edge tools, such as VertexAI/Google Agent Development Kit (ADK), LangChain/LangGraph, RAG frameworks, HuggingFace, OpenAI APIs, etc.
•
Assess solution alignment with internal AI Governance standards, Model Risk Management policies, and emerging regulatory expectations.