Writing

Hands-On Beats Rollout: What Prompt-a-Thons Taught Me About Scale

We delivered the training. Attendance was genuinely excellent. The completion metrics were the kind you put on a slide. And for a meaningful chunk of the population, almost nothing changed about how they actually worked.

That gap — between excellent training metrics and unchanged behavior — is the most useful thing I've learned in this work, because it forced us to stop trusting a whole category of number.

Thousands of training touchpoints. Thousands through foundational sessions, thousands more registered for advanced prompting. Every one of those is a real person who showed up. But attendance measures exposure, and exposure is not capability. Someone can sit through an excellent hour on prompting and return to a desk where nothing about their Tuesday has changed.

What actually moved the needle

The format that broke the pattern was the Prompt-a-thon. The structure is simple to the point of feeling insufficient: put teams in a room, in the real tool, working on their own actual work, for a focused session, with coaches circulating.

Not a demo. Not a case study about a fictional company. The thing that was already on their plate that week.

What happened in Atlanta made the case better than any deck I could have built. Teams moved from basic prompts to real business use cases inside a single session. Real workflows improved in hours — not weeks, not next quarter. And the demand that came out of it wasn't for more general training. It was for business-specific use-case coaching, which is a completely different and much better request.

The signal

When you put the right tools in people's hands, capability scales fast. Hands-on, business-led, outcome-focused. That's how adoption actually happens.

Why hands-on wins

Three reasons, and I think they generalize well beyond this particular tool.

The transfer problem disappears. Ordinary training asks people to learn in one context and apply in another. That translation step is where most of the value leaks out, quietly and invisibly. If the learning happens on the actual work, there's nothing to transfer. They're already there.

The first failure happens with a coach nearby. This one matters more than I expected. Almost everyone's second or third attempt with a new AI tool is disappointing — that's just the shape of the learning curve. Alone at a desk, that disappointment becomes a durable conclusion: this doesn't work for my kind of work. Very hard to reverse later. In a room with a coach, it becomes a thirty-second correction and a small win. Same moment, opposite outcome.

People see peers succeed. A vendor demo proves the tool works for the vendor. Watching the person two desks over — same role, same constraints, same institutional skepticism — get a real result in ten minutes proves something else entirely. Most of the adoption curve is social, and we consistently underinvest in that.

What this means for measurement

If exposure metrics don't predict capability, you need metrics that do. This is why I've pushed hard for measuring across four layers rather than one:

  • Adoption — are people using it, and are they still using it in week twelve rather than week two?
  • Behavior — has anything changed about how the work gets done, or is the tool bolted onto an unchanged process?
  • Outcomes — is there a business result underneath the activity?
  • Sentiment — do people believe this is for them, and would they say so honestly?

Any single layer will mislead you. Adoption without behavior means people are using AI to do the old process slightly faster. Outcomes without sentiment means you have results and a workforce quietly waiting for the initiative to pass. Sentiment without adoption means everyone likes the idea and nobody has changed anything.

You need the set. Read together, they tell you something reasonably close to the truth.

The uncomfortable implication

Hands-on doesn't scale the way training scales, and there's no honest way around that. You can put five thousand people through a session. You cannot put five thousand people through a Prompt-a-thon at the same quality.

So the model has to be different. Broad training establishes the floor — language, policy, basic fluency, permission. Hands-on sessions create the people who go back and change things, who become the local proof for their own team, who make the tool credible in a way no enterprise communication can.

Which means the scarce resource isn't licenses or training capacity. It's coaching capacity — people who can sit with a team, on their real work, and get them to a first genuine win. Almost nobody budgets for that, and I think it's the single most common reason these programs stall at the awareness stage and never get past it.


If you're running enablement somewhere and seeing the same gap between training metrics and real behavior change, I'd be glad to compare notes — find me on LinkedIn or send me a note.

← All writing Newer: The Dream-Build Loop: Capability Starts With a 45-Minute Walk

Keep going

More on the human side of this

Shorter thinking on LinkedIn, longer arguments here.