What Did the AI Actually Earn?

The Attribution Discipline Inside the Agentic GTM Stack

The question shows up in a planning conversation. A co-founder glancing at the tools line. An investor asking what the stack produces. Your own discipline the week the API bill goes up. It sounds like “what did the AI actually earn,” and at most companies the honest answer is a shrug, because every AI tool claims impact and almost none can show it in a form a skeptic accepts.

I launched a kit today built for that exact moment, and what follows is the long version of how its numbers get made. The short version lives on the offer page. This version is for the person who wants to check the math before trusting it, which is the right instinct and exactly the buyer I had in mind.

The kit in one paragraph

The Agentic GTM Stack is a self-deploy kit: two working sales agents on one shared attribution layer, sold as a single product because the shared layer is the point. InboxCopilot reads your inbound email, qualifies leads against your criteria, and drafts replies. EnrichmentMessenger takes a CSV of prospects, enriches the rows it can verify, and writes first-touch messaging for net-new pipeline. Every action either agent takes logs as a touchpoint in a Postgres schema you own, and SQL views assign revenue credit under rules built to survive scrutiny. You buy it, you deploy it in an afternoon, and nobody schedules a call, because there is no call to schedule.

The one rule everything else serves

The model assigns credit. It does not prove causation.

Every number the views produce is a credit assignment under stated rules, applied the same way every time. That distinction sounds like a concession until you sit across from a skeptic, where it becomes the whole advantage. A receipt survives questioning. An inflated claim does not.

Three rules do the work.

AI credit caps at 50% of any deal. Your humans keep at least half, always, including deals where no human touch was logged. An empty human-touchpoints table is a logging gap, not evidence the AI did everything. Deals involve calls nobody logged, relationships that predate the database, a founder answering email at midnight. The cap is the model admitting what it cannot see.

The split inside the AI share is proportional and boring. Each agent's share equals its fraction of eligible touchpoints on that deal. No first-touch bonuses, no U-shaped curves, no time decay, because those weighting schemes claim to know which touch mattered most, and at your deal volume the data cannot support that claim. Boring is what auditable looks like.

Nothing counts twice. A touchpoint belongs to exactly one agent. Shares on one deal sum to at most 100%. Sourced pipeline and attributed revenue are separate ledgers that never sum, anywhere, by anything.

And when a touchpoint is ambiguous, meaning the system could not tie it to a lead with confidence, it earns nothing. No partial count, no estimate standing in. The views round down on purpose, because a defensible small number beats an impressive contested one.

The strictest number in the kit

EnrichmentMessenger's headline number is sourced pipeline: opportunity value on leads that did not exist in any form before one of its runs created them. That is provenance in the database, not a judgment call. The strictness gets enforced at the front door: a CSV row with no email address never becomes a lead, because a row that cannot be checked against your existing leads could mint a phantom net-new record, and one phantom poisons the cleanest number the product has. The adapter refuses the row before the agent ever sees it.

(That rule exists because I nearly shipped the phantom. During pre-launch verification one agent failed exactly this way, creating a lead with nothing behind it. The defect, the fix, and the re-verification are written up separately for anyone who wants the engineering version.)

The comparison views state their own limits

Assisted deals next to unassisted deals, same period, same pipeline. Your baseline before deploy next to your pilot after. Both views ship with honesty rules baked into their output. Percentages render only when both sides have at least 20 closed opportunities; below that you get counts and a flag that says the sample is too small for rates. Early on you will see that flag a lot, and the flag showing up is the product working, because a rate quoted off nine deals is a rate that lies with confidence. The cohort view also states its own confound: these deals were not randomized, so it answers how AI-touched deals performed. It does not answer what those deals would have done untouched, and it says so in its own output.

The section most vendors would delete

The kit ships with METHODOLOGY.md, the document you hand to whoever asks the question. It is numbers built to survive a CFO, whether or not you have one yet. All sales are final once your delivery email sends, which means you cannot inspect the repo before checkout. So I am publishing the document's most important section here, in full, unsoftened. If it reads like a dealbreaker, the kit is not for you, and you just saved $249.

From METHODOLOGY.md, “What these numbers cannot claim”:

Read this section before you present anything, because the person you present to may know these limits already, and naming them first is how you keep the room.

These numbers cannot prove the AI caused any deal to close. Credit shares are an accounting convention, applied consistently and capped deliberately.

They cannot measure lift. Lift requires a counterfactual, the same deals untouched, and that counterfactual does not exist. The cohort and baseline views are context for judgment. They are not a controlled experiment, and dressing them up as one would burn the credibility this whole system exists to protect.

They cannot guarantee an outcome. The system measures what the agents did and assigns conservative credit. If the agents produce nothing worth crediting, the views will say that too, plainly, and that answer is the product doing its job.

They cannot see touches that were never logged. Untracked human work makes the AI look relatively larger. The 50% cap exists to blunt exactly that distortion, and logging your human touches shrinks it further.

And they will not claim more because the sample is small. Small samples get counts and flags, never percentages.

Every claim in that section is implemented in the SQL views you deploy. Nothing in the document describes an aspiration. You can read the SQL yourself to check, and I hope some of you do.

No case studies, and a dashboard instead

This product launches with no case studies, and I am not going to invent any. A measurement product that borrowed someone else's numbers to sell itself would be confessing something. Its first proof will be its own attribution data from my own deployment, published as it accrues at the live dashboard. At launch, that dashboard shows the minimum-sample flag doing its job, because no deals have closed yet. I am publishing the empty state on purpose. It is the honesty pitch, demonstrated.

Where to go from here

The kit is $249 founding for the first 25 buyers, $349 list after, one time, self-deploy: offer page. Plan on an afternoon if you are the RevOps hire or the operator wearing that hat. If nobody at your company has ever created a Supabase project, this is not your kit, and the offer page will tell you the same thing.

The question is coming. The answer is a number with its rules attached.