AI Evaluation • Rubric Review • Visual & Document Understanding

Matthew Arthur Elliott

AI data evaluator and human rater with an Actuarial Science background — reviewing model outputs against versioned rubrics, tagging failure modes with severity, and writing the structured rationale a modeling team can actually train on.

matthewaelliott@hotmail.com
Remote — US authorized•Available 20+ hrs / week, async•Fayetteville, NC

Positioning

Quantitative rigor applied to human evaluation.

I evaluate AI-generated content for a living. Across TELUS Digital AI, Appen and DataAnnotation I apply multi-page guidelines to long batches of model output — comparing paired responses, flagging hallucinations and instruction-following failures, and writing rationale that points to the specific rubric clause rather than to taste.

The habit comes from an Honours BSc in Actuarial Science and SOA Exams P and FM: check the assumption, verify the work, and be able to defend the conclusion step by step. Document review is familiar ground — my first analyst role at Mercer was cross-referencing benefit policy documents line by line to surface discrepancies between editions.

I also build with these models. Four live AI applications, shipped independently, mean I have seen how outputs fail in production — which makes the failure modes easier to name, categorize and escalate as regression cases.

Core Competencies

Evaluation strengths at a glance.

Model Output Evaluation
Rubric-Based Annotation
Severity & Policy Tagging
Paired Response Comparison
Inter-Rater Calibration
Hallucination Detection
Instruction-Following Analysis
Visual Document Understanding
Image Reasoning Evaluation
Structured Written Rationale
Batch Auditing & Drift Reporting
Failure-Mode Documentation

How I Review

A repeatable method, not a gut reaction.

Consistency across a 50-row batch matters more than brilliance on any single item.

Paired Comparison

Judge two model responses against the same prompt, select the stronger output, and write a rationale that names the specific rubric clause driving the decision — not a vague preference.

Severity & Safety Tagging

Label hallucinations, instruction-following failures, unsafe content and policy violations with structured tags and consistent severity levels across long batches.

Document & Visual Review

Apply multi-page rubrics to document annotation and image reasoning tasks — verifying that text, layout, diagrams and visual elements actually support the claimed answer.

Calibration & Escalation

Calibrate against gold-standard examples, audit batches for rubric drift, and route ambiguous prompts back to the program team when the prompt itself is the defect.

Built With AI

Three live applications, shipped solo.

Building with these models is what makes their failure modes recognizable.

SkyJourn

AI Reflection & Decision Support App

Three-mode AI application built and iterated through structured output testing — prompts refined by evaluating generated responses for usefulness, consistency and instruction-following.

Try SkyJourn →

SellScript

AI Listing Optimizer with Scoring

Rewrites product listings and returns before-and-after scoring with annotated explanations — a working example of rubric-style scoring applied to generated content.

Try SellScript →

ClosedWon

AI Sales Follow-Up Generator

Generates three personalized follow-ups with timing logic and the reasoning behind each — built and tuned by comparing candidate outputs side by side.

Try ClosedWon →

Experience

Evaluation, building and document review.

2025–Present

AI Data Evaluator & Research Analyst

TELUS Digital AI / Appen / DataAnnotation

Evaluate AI-generated responses against versioned guidelines for logical consistency, factual accuracy, relevance and instruction-following. Current work includes a human-editing trajectory collection program — editing and judging reference documents against reference, description and inference instructions, and rating instruction clarity itself. Identify hallucinations, unsupported conclusions and reasoning gaps, then document them in structured, evidence-based feedback. Conduct multi-source research to verify claims and separate accurate information from plausible-sounding error.

2016–Present

Founder, Operator & AI Tool Developer

SkyStarter Creative

Independently designed, tested and deployed four live AI-powered web applications using Claude Code, Cursor and ChatGPT. Refined prompts and workflows through iterative output evaluation — comparing generations, scoring quality and documenting recurring failure modes. Day-to-day exposure to how models fail sharpens the review work.

2012

Health & Benefits Analyst

Mercer Canada / Marsh McLennan

Reviewed and cross-referenced employee benefit policy documents and handbook editions against prior-year versions, identifying discrepancies and material changes for consultant review — line-by-line document comparison under a professional compliance and confidentiality framework.

Credentials

Education & setup.

Honours BSc, Actuarial Science

University of Western Ontario, 2011. Probability, regression, time series, loss models and survival analysis.

SOA Exams P & FM

Society of Actuaries Probability and Financial Mathematics — both passed.

Remote work setup

Native English speaker. Dedicated home workspace, high-speed internet, desktop and mobile test devices. Reliable async availability well beyond a 10-hour weekly minimum, US work authorized.

Contact

Available for AI evaluation and human-rater contract work.

Available as an independent specialist contractor for remote programs across document understanding, image reasoning, safety review and rubric-based annotation.

matthewaelliott@hotmail.com(click to copy)