Shaun McIlroy
Ai Product Quality Lead @Help Scout
Signup · Get unlimited contacts
WORK HISTORY
Ai Product Quality Lead @Help Scout
Lead evaluation and reliability efforts for Help Scout’s AI-driven product initiatives, establishing how large language model behaviour is evaluated in real customer environments. I partner closely with AI Services, Customer, Product, and Engineering teams to investigate system behaviour, evaluate AI performance, and translate ambiguous signals into clear product decisions.My work spans qualitative and quantitative evaluation, investigation of real-world AI behaviour, and the design of frameworks that help teams assess reliability, identify risks, and improve the consistency and safety of AI-powered features.Key Contributions• Lead evaluation and reliability work for AI-driven features, designing frameworks that measure large language model behaviour and help teams understand where AI responses succeed, fail, or require safeguards before release.• Investigate real-world AI system behaviour using customer sessions and product data, identifying failure patterns and translating those insights into prompt improvements, product changes, and clearer quality signals.• Built evaluation workflows for large language model responses, enabling teams to systematically test prompts and assess output quality as new features are developed.• Define reliability and safety boundaries for emerging AI features, identifying scenarios where AI responses should escalate to humans or require additional safeguards.• Partner with Product and Engineering teams to translate complex system behaviour into actionable insights that guide product decisions and improve the reliability of AI capabilities.• Establish feedback loops that allow teams to monitor AI behaviour over time and evaluate the impact of prompt, model, and product changes.
EDUCATION
University of Portsmouth
BSc. Computer Games Technology, Computer Games and Programming Skills
SKILLS
ABOUT SHAUN MCILROY
I investigate how complex product and AI systems behave in real-world environments, establishing how large language model behaviour is evaluated in production systems.My work sits at the intersection of AI systems, product reliability, and real-world customer behaviour. I focus on understanding how systems perform in practice and helping teams translate those insights into improvements that make products more dependable.My background spans advanced product support, technical investigation, and AI product quality. I specialize in analysing large language model behaviour and turning customer-reported issues and real usage patterns into actionable insights for product and engineering teams.At Help Scout I lead evaluation and reliability efforts for AI-driven features, establishing the evaluation practices used to measure model behaviour and improve the reliability of AI systems in production. This has included building evaluation frameworks for AI-generated responses, analysing real customer sessions to uncover systemic issues, and iterating on prompts and system behaviour to improve output accuracy, consistency, and safety. My work focuses on grounding AI behaviour in real customer environments so teams can ship new capabilities with greater confidence.Over the past 18 months this approach has expanded into broader AI work across the organization. I helped establish the quality signals teams use to understand where AI systems succeed, fail, or require safeguards before scaling new features. As AI product development continues to expand across engineering teams, I’m leading efforts to create practical guidance that helps teams design, test, and iterate on AI features with greater independence and confidence.Across my roles I’ve become known for connecting customer insight with deep system investigation. I work closely with Product, Engineering, and customer-facing teams to translate complex system behaviour into clear explanations and practical next steps, ensuring decisions are informed by both data and real-world usage.Earlier in my career I built a strong investigative foundation in advanced technical support roles, focusing on complex product behaviour, root-cause analysis, and translating customer insight into engineering action. I also designed the onboarding and coaching structure for our Advanced Triage team, helping teammates develop the investigative skills required to diagnose complex technical issues.
This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.