You cannot interview your way to certainty about a GTM engineer. The role is too new, the tooling changes monthly, and the vocabulary is cheap — anyone who has watched four hours of YouTube can say “waterfall enrichment” and “Claygent” in a sentence that sounds correct. The fix is not a harder conversation. It is a 90-minute timed build, screen-shared, scored against a fixed rubric, with a real list and a real sending tool at the end of it. Enrich 200 rows, waterfall three providers, write one Claygent prompt, push the qualified segment to Smartlead. Ninety minutes tells you more than three interviews and a portfolio deck.
Why the resume cannot carry this hire
Start with the base rate of misrepresentation. 80% of US hiring managers say candidates’ resumes don’t match their real-world skills at least sometimes, and 34% say it happens all the time or often. That is not a GTM-specific problem, but GTM engineering is where it bites hardest — the job is a stack of named tools, and named tools are the easiest thing in the world to list.
Credentials do not rescue you either. Only about 1,000 people hold a Clay certification, and Clay has paused new submissions while it reworks the program. Meanwhile, seven independent bootcamps teaching GTM engineering have graduated over 2,500 students. Do the arithmetic on your own pipeline: for every certified operator you see, you will see several people with structured tutorial exposure and no production reps. Tutorial exposure is now the market default. It is not a differentiator, and it is not evidence.
Then there is the honesty problem in the interview itself. A 2Q25 Gartner survey of 3,000 candidates found 6% admitted to interview fraud — either posing as someone else or having someone else pose as them. Six percent is small until you remember you are hiring one person. And the same survey found the most AI-assisted part of the application is the assessment: 29% of candidates generate text for answers to assessment questions. That single number kills the take-home. A build you cannot watch is a build you cannot attribute.
A take-home tells you what someone’s tools can do. A live build tells you what your next hire can do.
Why a job-sample build, specifically
The selection-validity literature was substantially revised in recent years, and the structured, job-sample end of the spectrum came out on top: structured interviews lead with a mean validity of .42, with an 80% credibility interval running from .18 to .66. Read the interval, not just the mean. The same method can be excellent or nearly worthless depending on how tightly it is designed and scored. That variability is the whole argument for a written rubric and identical tasks for every candidate — an unscored “walk me through your Clay setup” conversation lands at the .18 end.
And the work is genuinely technical, which is why watching it matters. In an analysis of 1,000 GTM engineering job postings, SQL and Python each appeared in 38% of them. Roughly two in five of these roles expect someone who can reason about data structure, joins and failure states — not just click through a template. Ninety minutes of build surfaces that reasoning. A resume bullet hides it. If you are calibrating the rest of the loop, our GTM engineer hiring guide covers scorecards, comp and leveling around this exercise.
The 90-minute build, spec’d
Give the candidate a Clay seat on your workspace, a 200-row list of accounts in your actual ICP, and access to a Smartlead sandbox. Share the brief 10 minutes before the call — not 24 hours before. Screen recorded, camera on, no pausing.
| Minutes | Task | What you are actually testing |
|---|---|---|
| 0–10 | Import the 200-row list, dedupe, normalize domains | Data hygiene instincts; do they check before they enrich |
| 10–35 | Find company + contact data, waterfall three providers in priority order | Cost discipline, provider sequencing, conditional run logic |
| 35–55 | Write one Claygent prompt against a defined qualification question | Prompt specificity, output schema, hallucination control |
| 55–75 | Build the qualification filter and the personalization variable | Judgment about what actually belongs in a first line |
| 75–90 | Push the qualified segment to Smartlead, mapped to fields | Downstream thinking, field mapping, error handling |
One brief, every candidate, same list, same clock. That consistency is what pulls the exercise toward the top of the validity range instead of the bottom.
The Claygent prompt is the tell
If you only have time to watch one segment, watch minutes 35 to 55. A weak operator writes a prompt like “find out if this company sells to enterprises.” A real operator specifies the source to check, constrains the output to a fixed set of values, tells the agent what to return when it cannot find the answer, and then spot-checks five rows against the live site. The difference is visible in about 90 seconds.
The rubric
Score each dimension 0–3. Do not average in your head afterward — fill the grid live, in the call.
| Dimension | 0 — Absent | 1 — Tutorial | 2 — Operator | 3 — Owner |
|---|---|---|---|---|
| Data hygiene | Enriches raw input | Dedupes when reminded | Dedupes and normalizes unprompted | Flags list quality issues and quantifies them |
| Waterfall design | One provider | Three providers, no conditions | Conditional runs, cheapest-first sequencing | Explains cost per verified record and where to stop |
| Claygent prompt | Vague, open-ended | Works on happy-path rows | Constrained output, defined null case | Adds a validation column and audits samples |
| Qualification logic | Filters on one field | Filters on stated criteria | Builds a scored segment | Challenges your criteria with evidence from the data |
| Smartlead push | Cannot complete | Pushes with mapping errors | Clean mapped push | Handles bad rows, plans the fallback sequence |
| Narration | Silent | Describes clicks | Explains tradeoffs | Explains tradeoffs and what they would cut for time |
Eighteen points available. Anything at or above 13, with no zeros, is a genuine operator. Below 9 is someone repeating a tutorial. The zeros matter more than the total — a candidate scoring 15 with a zero on the Smartlead push has never shipped a campaign end to end.
What each output tier tells you
| Tier | What they shipped in 90 minutes | What it means | Decision |
|---|---|---|---|
| Tier 1 | Full pipeline live, waterfall with conditions, audited Claygent column, clean Smartlead push, and a note on cost per record | Has run this in production under a deadline. Will pay back inside a quarter | Move to offer stage the same week |
| Tier 2 | Enrichment and waterfall solid, Claygent prompt works but is unconstrained, push completes with minor mapping fixes | Real hands-on reps, thin on agent design. Coachable in weeks | Hire if you have a senior operator to review their work |
| Tier 3 | Enrichment works, single provider, Claygent output unchecked, no push | Template-level fluency. Knows where the buttons are, not why | Pass for a senior req; possible junior hire on a team with structure |
| Tier 4 | Stalls on import or waterfall, narrates in vocabulary rather than actions, asks to “finish this offline” | Tutorial exposure only, or the work was never theirs | Pass |
The Tier 4 signature to watch for is the request to finish offline. Given the 29% assessment-assistance rate, an offline finish is not a scheduling accommodation. It is a different test.
Rules that keep the test honest
Camera on, screen shared, one continuous session. This is your identity control as much as your ability control. With 6% of candidates admitting to interview fraud, the person who builds must be visibly the person who applied.
Let them use AI — out loud. Banning ChatGPT is unrealistic and unrepresentative of the job. Ask them to narrate what they prompt and why. Strong candidates use AI as a syntax accelerator and still know what correct output looks like. Weak candidates paste and pray.
Use your real ICP list. A generic list lets them run a rehearsed motion. Your messy list — inconsistent domains, holding companies, a few dead accounts — is where judgment shows.
Pay for the time. Ninety minutes of scored production work deserves a stipend. It also raises completion rates among employed senior candidates, who are exactly the pool you want.
Score before you debrief. Fill the grid independently, then compare. Talking first is how a fixed rubric quietly becomes an unstructured impression.
Slotting the build into a loop that closes
A 90-minute exercise only helps if the rest of the process moves in days, not weeks. The sequence that works: 20-minute screen for scope and comp, the scored build, then a 30-minute panel on collaboration and roadmap ownership. Three touchpoints, one week, one decision. Every additional stage after a Tier 1 build is risk you are adding, not removing.
The build also tells you what to pay. A Tier 1 operator who can defend cost per verified record is not competing with your junior ops req — check the range against our salary benchmarks before you scope the role, because underpricing a Tier 1 is the fastest way to lose one. If the same candidate is also touching your warehouse or your CRM layer, our data warehouse and Salesforce recruiting practices work the same problem from the adjacent direction, and marketing operations recruiting covers the sending-and-lifecycle side.
This is why spray-and-pray sourcing fails on GTM engineering specifically. The keyword surface is huge and nearly free to fake, so volume sourcing produces a pipeline you cannot read. A pre-vetted shortlist where every candidate has already shipped a scored build is a different exercise entirely — see how we run it in GTM engineer recruiting, or point strong operators you already know at our GTM engineering talent network.