AI Systems Ownership & Feature Delivery
• Take complete ownership and deliver major AI engineering features within agreed timelines
• Own AI output quality, structure, and predictability across all user-facing AI interactions
• Design, implement, and maintain output-type–based AI systems, including segmentation, routing, and enforcement
• Ensure consistent output structure and formatting across different LLMs for the same request type
• Integrate and orchestrate multiple LLM providers via OpenRouter, managing model selection, fallback strategies, and cost optimisations
• Design and orchestrate tool-using / agentic AI workflows — defining clean tool contracts (including MCP-based tools), function-calling interfaces, and reliable AI-to-system integrations
• Build and maintain complex, multi-step LLM workflows — including with orchestration frameworks such as LangChain or LlamaIndex — for advanced reasoning, context reuse, and retrieval
Prompt Engineering & Experimentation
• Design and manage production prompt systems with dynamic prompting, context injection, and conditional logic
• Own the deployment and release of LLM experiments, prompt management, and Langfuse-based evaluation pipelines
• Run A/B tests across models, analyse results, and present data-driven impact assessments of AI features and experiments
• Monitor AI system metrics, quality signals, latency, and release health using Langfuse and other observability tools
• Deep-debug complex LLM chains using Langfuse traces — identifying bottlenecks and optimising for cost, latency, and context-window usage — and build output-scoring to root-cause hallucinations and logic errors
Code Quality & Production Reliability
• Write clean, scalable, and maintainable TypeScript code across the Next.js / Node.js stack
• Build reliable backend logic for AI systems, with strong error handling, request validation, fallback flows, and predictable behavior in production — including reliable tool execution and AI-to-service integrations
• Ensure high code quality through testing, code reviews, and clear engineering standards
• Monitor, troubleshoot, and improve production performance, reliability, and system health
• Drive maintainability and technical quality through solid architecture, refactoring, and disciplined release practices