Score the deck against the reference, not against a mood.
I grade AI-generated pitch decks for Ethos against a fixed reference — 1–7 per task, with written rationale — the way a creative director reviews work before it ships: narrative first, craft second, taste last and only when it changes the outcome.
Grade AI-generated decks the way a Creative Director reviews work before it ships.
For Ethos, this role sits where design critique, narrative judgment and business communication overlap. The task isn't to say what looks good — it's to place a generated deck on a fixed 1–7 scale against a reference, and write the rationale that would let another grader land on the same number.
Read the reference first
Establish what the top of the scale looks like for this task before looking at the generated deck — the reference is the anchor, not a suggestion.
Diagnose slide function
Assess whether each slide has a clear job: context, evidence, proof, decision, objection handling or call to action — same lens as a client deck.
Place it on the scale
Score 1–7 against the reference, not against an abstract ideal — a deck that matches the reference's ceiling scores at the ceiling, nothing more.
Write the rationale
Justify the number with the exact slide or claim that drove it — the score has to be reproducible from the sentence, not just defensible in conversation.
Eight lenses I use to place a deck on the 1–7 scale.
Repeatable enough to hold across graders, flexible enough for investor decks, sales decks, founder stories, category narratives and internal strategy presentations.
Reference match
How close does this deck land to the reference's ceiling on the same brief?
Executive narrative
Does the deck make the business case in a sequence a busy decision-maker can follow?
Slide function
Does each slide have one clear role, or is it carrying multiple conflicting messages?
Takeaway title
Does the title state the point of the slide, not merely the topic?
Visual hierarchy
Can the audience scan the page and understand what matters first, second and third?
Evidence quality
Are claims supported with numbers, market signals, customer insight, proof points or operating logic?
Audience fit
Is the language calibrated to investors, sales buyers, internal teams or strategic partners?
Rationale reproducibility
Could a second grader land on the same score from my written rationale alone?
How I'd score a generic AI-generated deck against a strong reference.
The deck states the market, the product and the ask. Slide titles name topics ("Our Brand", "The Market") rather than claims. The reference deck for this task titles each slide as a decision-ready statement backed by one proof point.
- Covers the required sections — nothing is missing structurally.
- No slide states a takeaway a reader could act on.
- Evidence exists but isn't attached to the claim it should prove.
- Falls well short of the reference's title logic and evidence placement.
Written rationale
The deck is structurally complete but not decision-ready: titles name topics instead of claims (Slide 2 "The Market" vs. the reference's "Category grows 34% while premium grows 61%"), and the one proof point that would carry Slide 4 sits in the appendix instead of on the slide. This is a narrative-and-evidence-placement gap, not a missing-content gap, so it lands mid-scale rather than at the floor.
- Reference for this task scores 7: every title is a claim, every claim has its proof point on-slide.
- A 5 would require rewriting titles as claims across all six slides.
- A 6–7 would additionally require moving the strongest proof point on-slide.
Before and after: from decorative deck copy to a claim the reference would reward.
What Ethos requires before the first task, and where it stands.
Stated plainly rather than glossed over: the offer is signed, the gate is two paid steps inside Feather, and the account isn't provisioned yet on Ethos's side. Nothing below is inflated to look further along than it is.
Offer signed
Ethos Expert Opportunity — Creative Director / Art Director. USD 80/h, cap USD 1,600/week, monthly pay cycle (10th–9th, paid the 14th). Signed 01-Sep-2026.
Feather account access
The contractor ID isn't provisioned in Ethos's Microsoft tenant — a login failure on Ethos's side, not a credentials or browser issue. Reported and confirmed with Ethos support; still unresolved.
Training course (Feather)
Paid onboarding step required before the first graded task. Cannot start until Feather access is provisioned.
AI Slop Screener
Paid qualifying screener required before task access — the same slop-detection discipline this case is built on, formalised as a gate.
Confidentiality separation
Dedicated Chrome profile for Ethos work, kept out of the macOS user that runs Mercor's Insightful tracker — the two contracts never share a desktop capture surface.
Scoring discipline
1–7 criteria-based scoring is the same muscle as blind IWC and Decanter judging — fixed scale, evidence-first, no drift between one task and the next.
What unblocks this
Ethos has to provision the contractor ID on their tenant. Once Feather opens, the two paid onboarding steps above are the only thing standing between this profile and the first graded task.
Why this profile is believable from a real career.
The case connects public-facing experience in premium communication, brand systems and commercial strategy with a private AI-evaluation layer. It doesn't pretend to be a pure data scientist — it positions the high-value human judgment labs are short of.
Decks & brand systems
Grand Crew Studio is the natural bridge: brand manuals, sales narratives, proposals, portfolio pages and client-facing strategic documents, owned end to end.
Business communication
Export management, agency leadership and published editorial work back the ability to explain complex commercial ideas clearly for different audiences.
Fixed-scale judgment
IWC and Decanter judging convert directly into Ethos's 1–7 grading: blind, criteria-based, reproducible from the written rationale.
The presentation work I'm strongest at reviewing.
Ethos-facing keywords, copied exactly from the archetype bank — Presentation Design Reviewer / Pitch Deck Reviewer archetype, plus the Creative Director / Art Director title as offered.
Work_AiLabClosing_retrato.jpg
The judgment behind the score.
I'm Flor Gómez — brand strategist, founder of Grand Crew Studio and Master of Wine Stage 2 candidate. I've built pitch decks, brand manuals and sales narratives for real clients, judged blind at Decanter and IWC, and I grade AI decks the same way: against a fixed reference, with the number and the reason attached.
Flor Gómez · Barcelona · English & Spanish · flor-gomez.com
You've seen the sample grading. If that's the standard you need, contact me.
Deck grading against a reference, pitch and investor narrative review, executive communication evaluation. Tell me the project and I'll reply with fit, scope and rate within one business day.
Prefer to write? Use the contact form · Related case: AI Writing & Editorial QA (Mercor) · Portfolio: flor-gomez.com
Contact me.
Goes straight to my inbox — no middlemen. Tell me what you need, paste a link if you have one, and I'll reply within one business day.
Got it — thank you.
Your message is on its way to me. I'll reply within one business day. If it's time-sensitive, you can also book a 15-min call.