AI Evaluation • Rubric Review • Visual & Document Understanding
Matthew Arthur Elliott
AI data evaluator and human rater with an Actuarial Science background — reviewing model outputs against versioned rubrics, tagging failure modes with severity, and writing the structured rationale a modeling team can actually train on.
Positioning
Quantitative rigor applied to human evaluation.
I evaluate AI-generated content for a living. Across TELUS Digital AI, Appen and DataAnnotation I apply multi-page guidelines to long batches of model output — comparing paired responses, flagging hallucinations and instruction-following failures, and writing rationale that points to the specific rubric clause rather than to taste.
The habit comes from an Honours BSc in Actuarial Science and SOA Exams P and FM: check the assumption, verify the work, and be able to defend the conclusion step by step. Document review is familiar ground — my first analyst role at Mercer was cross-referencing benefit policy documents line by line to surface discrepancies between editions.
I also build with these models. Four live AI applications, shipped independently, mean I have seen how outputs fail in production — which makes the failure modes easier to name, categorize and escalate as regression cases.
Core Competencies
Evaluation strengths at a glance.
How I Review
A repeatable method, not a gut reaction.
Consistency across a 50-row batch matters more than brilliance on any single item.
Built With AI
Three live applications, shipped solo.
Building with these models is what makes their failure modes recognizable.
Experience
Evaluation, building and document review.
Credentials