Nathan Livni

Education & AI Measurement

Corporate training program evaluation is a well-established, if not glamorous field. However, it has a lot to teach the much newer field of agent evaluation.

Smile sheets

Most teams measure their AI agents the way bad training programs measure learning, by counting who liked it. This is called a smile sheet. You run a workshop, hand out a survey at the end, 80% of people check “useful,” and you file it as a win. But nothing changes in the real world, and nobody can say why the program didn’t improve performance.

An approval meter resting on a textbook beside a checked survey, with its lead unplugged

In 1959, Donald Kirkpatrick named this problem and his four levels of training evaluation are still the clearest way to ask whether your education actually works. And when you map this to agent measurement, you find that you stall out at the same points, for similar reasons.

The four levels applied to agents

Amplitude built a five-step maturity model for agent measurement, L0 through L4 which runs from raw tracing up through evals, semantic intelligence, behavioral analytics, and revenue impact. It came to similar conclusions Kirkpatrick reached sixty years earlier, albeit from a different starting point.

L1: ReactionL2: LearningL3: BehaviorL4: Results
A “thumbs up” tells you whether a user liked the agent’s answer. Almost nobody leaves one. When Amplitude studied 27,000 of its own agent sessions, only about 2.2% of users gave explicit feedback. You can’t read the mood of one user in fifty and call it product health.

Also, two users can have completely different experiences and look identical on a reaction metric. Maybe one got their problem solved while another gave up. Both sent the same number of messages.
In training this is a post-test. For an agent it’s an eval. Did the response answer the question, stay on task, respect the user’s constraints.

This is where LLM-as-judge scoring and quality rubrics live, and it’s a real step up from “did they like it.” It still doesn’t answer what the agent is worth.
A good post-test score means nothing if the learner never changes how they work, and the same holds for an agent. Did the user take the next action, come back the following week, adopt the feature the agent pointed them to?

This data only exists if the agent’s events sit in the same stream as everything else the user does in your product.
Retention, conversion, expansion, cost. In Kirkpatrick’s world this is the level executives always asked for and most training teams can’t produce. 

For agents it’s “what is this thing worth,” and answering it means tying every agent interaction back to revenue.

The data gets harder to interpret and connect as you climb, so teams stop at Level 1, declare success, and never find out whether the thing worked. Instead, they  ship the “thumbs up” button because it’s easy, then go hunting for meaning in feedback that 98% of users never leave.

Design backward from Level 4

Kirkpatrick’s most useful lesson is that you measure from the bottom up, but you design from the top down:

  1. First decide the L4 business result first…

  2. then work back to the L3 behavior that drives it..

  3. then the L2 learning that enables the behavior…

  4. and then finally the L1 reaction.

The world is a messy place and correlation is not causation. But if you don’t start with your end goal in mind, how will you create a system to get there. Amplitude’s research indicated that a great first agent session might impact roughly 3x long-term retention. The finding sits at Level 4, driven by a Level 3 behavior.

Of course, you must be able to tie the agent’s quality scores to the user’s downstream actions and business results. Starting from “we want 3x retention,” the question stops being “did they smile” and becomes “what does a first session that earns a second one look like.”

You’re saying that satisfaction is unimportant?

I’m saying that if you were a medical doctor and you needed a patient to adhere change their behavior to adhere to a medication it would go like this, think of which is most important:

  • The medication is having a positive effect because it’s the right medication (L4)
  • The patient takes the medication every day (L3)
  • The patient understood your directions (L2)
  • The patient enjoyed coming in to the doctor (L1)

What kind of a doctor would measure primarily L1? If anything L1 is bolted on as an afterthought because at the end of the day it’s about positive results and good health. I assure you, it wasn’t always that way. Doctors dealt with this by professionalizing and creating standards for medical care.

Takeaways

If your agent dashboard is mostly measuring response counts and thumbs up, you’re running a smile sheet and stuck at L1. While it’s the easiest data to collect, it’s also the not going to tell if your agent is worth what you’re paying for it.

Some training teams have learned not to trust end-of-workshop surveys and instead focus on behavior change and business result. Customer enjoyment is only a means to an end.

If you could keep only one measurement of your agent, would it be how users reacted (L1) or if it drove business results (L4)? Most teams have already chosen one.