Is your AI sales agent actually working, or just busy? Open the agent's weekly summary and you will usually find the same three numbers: messages sent, leads worked, records updated. All three can go up every single week without a single new person ever getting a real reply. That is the gap founders and sales leaders keep tripping over: a dashboard full of motion, and no way to tell from the dashboard alone whether any of it produced something real.
This matters more once the agent is not just drafting copy for a human to send. Today's AI SDR agents send, log, and move on to the next lead on their own, so a bad week can look exactly like a good one unless something outside the agent checks its work. Before you can measure that, it helps to know what enrichment data actually is, since "validated" only means something once you know what is being checked. This guide walks through the one-line test for telling busy from working, the four numbers worth tracking instead, and a 30-day plan to start checking your agent's claims from chat.
The Busy Agent Trap: Why a Full Activity Log Does Not Prove Anything
An activity log only tells you the agent was active, not that anything it touched was real. A contact can be added, messaged, and marked "worked" without anyone ever confirming that contact still exists at that company, under that title, today.
Why a Full Week on the Dashboard Can Still Be an Empty One
- Every send, touch, and field update logs itself automatically, so volume climbs whether or not any of it landed.
- A stale or duplicate record counts the same as a fresh, correct one on an activity chart.
- The agent's own dashboard has no reason to flag its own miss, since nothing outside it is checking.
"An AI agent sends 4,000 messages. Works 900 leads. Researches 300 accounts. Updates 700 CRM records. Success? Maybe." -- Rania Kuraa, LinkedIn
The One-Line Test: Busy or Working?
- Ask: could this number go up if the agent just did more, with nothing outside it confirming any of it? If yes, it is an activity number.
- Ask: does confirming this number require checking a source the agent does not control? If yes, it is a working number.
- A number that fails both is worth dropping from the scorecard entirely.

Activity Numbers vs. Checked Outcomes: What Actually Belongs on a Scorecard
An activity number comes straight from the agent's own log. A checked outcome only counts once something outside that log confirms it. The difference decides whether your scorecard is measuring effort or measuring results.
| What the Agent Reports | What You Should Actually Trust |
|---|---|
| Accounts researched | Meetings booked with a matched, current company |
| Sequences launched | Pipeline that survives a data check |
| CRM records updated | Contacts confirmed current, not just added |
| Leads worked | Pipeline that still holds up at close |
Reading This Table: The One Pattern to Flag First
- The left side is easy to inflate. It rewards doing more, not doing it right.
- The right side only moves once an independent source agrees.
- A left column climbing while the right column stays flat is the first thing to raise in a pipeline review, not the last.
Why a Generic QA Score Does Not Settle the Question
Most teams grade an AI agent with either a human-SDR coaching rubric or the vendor's own internal quality score, and neither one checks against a source outside the agent's control. That leaves the agent grading its own homework.
The Borrowed-Standard Problem
- A QA rubric built for coaching a human SDR was never designed to catch a stale record an agent worked anyway.
- A vendor's opaque internal score has no incentive to surface a miss that makes the agent look bad.
- Neither standard tells you whether the company or the title behind a "qualified" contact is still accurate today.
"Most quality programs run on borrowed standards, a generic QA scorecard, or an AI vendor's own opaque score." -- Jon Odalen, LinkedIn
What "Independent" Actually Means Here
- The check has to come from a data source the agent itself cannot edit or influence.
- It has to run at the same volume the agent worked, not a hand-picked sample of ten records.
- It has to answer a yes-or-no question: is this contact, at this company, still real right now.
A Real Qualified Meeting vs. a Meeting That Just Got Booked
A real qualified meeting has three things confirmed before it ever hits the calendar: the right company, the right current title, and a clear next step. A meeting that is just booked is missing at least one. Volume on the calendar cannot tell the two apart by itself.
The Three-Part Check Before a Meeting Counts
- Company confirmed against a real, matched profile, not just an email domain that happens to look right.
- Title confirmed current. A stale title is the single most common reason a "qualified" meeting turns into a wasted 30 minutes.
- A specific next step logged, not a vague "exploring options" note.
What Skipping This Check Actually Costs
- Meeting-booked counts keep climbing while close rates stay flat.
- Sales reps stop trusting agent-sourced meetings and start manually re-qualifying everything, which erases the time the agent was supposed to save.
- Leadership sees a full calendar and a flat forecast, and the agent takes the blame for a measurement gap it did not create.
Give the Agent One Job Before You Try to Grade It
Name a single outcome the agent owns, in one segment, before you write a single success metric. Not every go-to-market motion is the same, and grading a bundle of tasks as if they were one job is how measurement turns into a vague monthly report nobody trusts.
"Not all go-to-markets are the same... it's treated as if every GTM action is identical." -- nerddiva, Agent Insight, LinkedIn
Writing the Agent's One-Page Job Description
- Pick one outcome: for example, booked and verified first meetings in one segment, not "outbound in general."
- Name the source that will confirm that outcome before the agent starts running, not after a leader asks why numbers look off.
- Set a ceiling on volume the agent cannot exceed without a matching rise in the checked-outcome number.
- Cross-check the job against a simple broken-process-vs-broken-agent checklist, since a vague job often hides a process gap that no metric will fix.
Turn the Job Description Into a Config, Not a Slide Deck
Write the job as a small config object that lives next to the agent's other settings. Tying a volume ceiling to a verified-outcome floor means a spike in raw activity with no matching rise in checked outcomes shows up immediately, instead of surfacing three weeks later in a pipeline review.

