Outbound

Replies Up, Calendar Empty. How to Judge Outbound Reply Quality Before You Rewrite the Email.

Outbound reply quality decides whether replies turn into meetings. Sort replies into five piles, check three numbers, and learn if it is the list or the copy.

Vibe Prospecting team9 min readAugust 13, 2026
Replies Up, Calendar Empty. How to Judge Outbound Reply Quality Before You Rewrite the Email.

TL;DR

  • A reply saying wrong person and a reply asking for pricing sit in the same column of most outbound reports. That is why the number climbs while the calendar stays empty.
  • Sort every reply into five piles: interested, not interested, wrong person, out of office or handoff, do not contact. Then divide the interested pile by messages sent, never by opens.
  • Strong campaigns land 1.5% to 3% of sends in the interested pile, and 4% to 5% on a tight list. For every reply of any kind, Belkins put the average at 0.45% across 7,530,489 cold emails sent through 2025.
  • If more than 15% of your replies say wrong person, or hard bounces clear 2%, the list is what broke and no rewrite reaches it.
  • A funding round is worth roughly 5x to 7x normal reply performance in the first 24 to 72 hours and about 1.2x by Day 7, so the age of your reason for writing belongs in the audit.
  • Run the audit in a chat window in thirty minutes, then rebuild the list side in the next message: ask, preview 5 records with a cost estimate, then build. Powered by Explorium Enterprise Business Data.

Six weeks into a campaign, the number everyone quotes is the reply rate, and it is climbing. The calendar is not. Outbound reply quality is the gap between those two facts, and it opens because a note saying "wrong person, try finance" and a note saying "can you send pricing" land in the same column of the same weekly report. Every time the list gets looser, the column gets taller.

It is worth naming where that road ends. Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027 on rising cost and unclear value. An outbound agent whose only score is total replies has no number worth defending when somebody asks what it bought.

So here is how to read a week of replies honestly: five piles instead of one column, the three numbers that say whether to fix the list or the words, what a good result actually looks like, and a thirty-minute version of the whole audit you can run tonight in a chat window.

Two Replies, One Column in the Weekly Report

A brush-off and a pricing request are recorded as the same event, so the quickest way to lift total reply rate is to write to more people who fit you worse. That is not a reporting quirk. It is an incentive, and agents follow incentives faster than people do.

The Arithmetic of a Looser List

Same 6,000 messages, two ways of choosing who gets them:

What changedTotal repliesShare asking a real questionInterested replies per 6,000 sends
More titles, wider company sizes480 (8.0%)15%72 (1.20%)
One role, one recent event300 (5.0%)40%120 (2.00%)

The looser list wins the report by three percentage points and loses the month by 48 real conversations. Nobody in the meeting can see that, because only one of those four columns is on the slide.

Three Habits That Inflate the Number

  • Counting auto-replies. An out-of-office is a mail server following a rule, not a person forming an opinion.
  • Treating a handoff as a yes. Being passed to a colleague is routing work you now owe someone, not interest.
  • Adding titles whenever replies dip. Each extra title lowers the share of recipients who could ever buy, which is exactly why the total goes up.

Sort Every Reply Into Five Piles

Five is the smallest set of piles where every reply names one owner and one fix: wrong person, not interested, out of office or handoff, do not contact, interested. Then one division does the rest, interested replies over messages sent, never over opens.

What Each Pile Tells You

PileSounds likeWhat it tells youWhose job
Wrong person"Try procurement"The list chose the recipient badlyWhoever builds lists
Not interested"No need right now"Right person, wrong moment or offerWhoever picks the moment
Out of office or handoff"Back Monday, ask Dana"Routing. Never a yesWhoever owns the sequence
Do not contact"Remove me"Suppression missed a nameFix it the same day
Interested"Send pricing" or "Tuesday works"List and words both landedNobody. Do more of this

Where the Piles Have to Live

  • Write the pile onto the person and the company record, not only inside the sending tool where it dies at the end of the campaign.
  • Sort the same day. HBR's lead-response study put the odds of qualifying someone roughly 7x better when first contact happens inside the opening hour.
  • Push wrong person and do not contact into suppression before the next batch goes out, not after a complaint arrives.

Is It the List or the Words? Three Numbers Settle It

Wrong-person share and hard bounce rate point at the list. The split between interested and not interested among right-person replies points at the words. The age of your reason for writing points at the moment.

  • The list: wrong person above 15% of replies, or hard bounces above 2%. No rewrite touches either one.
  • The words: wrong person under 10%, bounces under 2%, and fewer than a fifth of replies landing in the interested pile.
  • The moment: a healthy spread across the piles, but a median reason-for-writing older than a week.
  • Two of the three live in the data underneath the agent, which is why editing prompts keeps producing so little. That layer is the one to look at, and enrichment for outbound agents covers what it owes you.
Diagram pairing wrong-person replies above 15% with a list problem and interested replies under 20% with a copy problem

The Delete-the-Reason Test

Cut the sentence that explains why you wrote today, then read what is left. If it still looks like a reasonable note to that specific person, send it. If it falls apart without the funding round or the new VP, the event was your excuse to write rather than their reason to answer, and it will land in the not interested pile every time.

What Good Looks Like, and Which Denominator You Used

Strong campaigns put 1.5% to 3% of messages sent into the interested pile, 4% to 5% on a genuinely tight list, with a quarter to half of all replies classed interested. Anything much above that is usually a denominator problem rather than a triumph.

Benchmarks Worth Writing Down

  • Belkins recorded a 0.45% average reply rate across 7,530,489 cold emails sent through 2025, moving between 0.35% in December and 0.54% in February.
  • Who you wrote to moved that further than any rewrite did: 0.72% at companies of 0 to 10 people against 0.22% at 10,000 and above.
  • From the same team's phone data, 4.6% of conversations booked a meeting, roughly 370 dials each. Useful context before deciding email is uniquely broken.

Two Numbers, Two Denominators

  • The 0.45% is every reply divided by 7.5 million sends. The 1.5% to 3% band is interested replies divided by sends on a focused campaign. Putting them on the same axis makes both meaningless.
  • Well-timed messages sit in a third band. Notes sent on a leadership change have been measured at a 14% reply rate, per Explorium's work on timing outreach to recent company activity.
  • Anything measured against opens is not a reply rate. Report it that way once and every later number is suspect.

Addresses That Never Had a Chance

Hold every address the checker did not return as valid, and keep Gmail spam complaints under 0.10% with a hard stop well before 0.30%. Some share of your not-interested pile is not an opinion at all. It is mail that never reached a human.

What the Six Statuses Mean

  • Hunter's email checker returns one of valid, invalid, accept_all, webmail, disposable or unknown, plus a 0 to 100 confidence score, documented field by field.
  • accept_all and unknown are holds, not sends. Those servers say yes to every address, so a yes proves nothing about the mailbox behind it.
  • Drop invalid and disposable for good rather than for this campaign. Route webmail to a human when the business you sell to always has its own domain.
"A product whose entire pitch is fixing outbound could not work out that cold emailing a direct competitor's CEO, from a domain Gmail had already flagged..." r/salesdevelopment, July 2026

Numbers That Should Trigger a Stop

NumberFineLook into itStop sending
Hard bouncesUnder 2%2% to 5%Above 5%
Wrong-person share of repliesUnder 10%10% to 15%Above 15%
Gmail spam complaintsUnder 0.10%0.10% to 0.29%0.30% and above
Median age of your reason for writingUnder 72 hours3 to 7 daysOver 14 days
Daily volume to Gmail per domainUnder 5,000At the bulk-sender lineBulk volume with DMARC failing

Past 5,000 messages a day to Gmail you are a bulk sender in Google's terms, and SPF, DKIM and DMARC all have to pass. None of that is optional once an agent is doing the sending for you.

Old News Reads Like No News

Reasons expire. A funding round is worth roughly 5x to 7x normal reply performance in the first 24 to 72 hours, about 1.2x by Day 7, and nothing at all by Day 14.

Reply lift curve highest inside the first 24 to 72 hours, fading through day 7 and gone by day 14, with the send window shaded in mint

How Long Each Reason Lasts

  • Funding: 24 to 72 hours, with performance near 400% of baseline inside the first 48.
  • A new executive in the seat you sell to: 30 to 60 days, 4x to 5x, with 14% reply rates recorded.
  • A hiring push: 14 to 30 days. A new tool going live: 7 to 21 days. Both around 3x to 4x.
  • Good fit and nothing happening: a 1% to 2% floor. That is the number your timed sends have to beat to earn the extra work.

Why Stale Reasons Land in the Wrong Pile

  • A Day-10 funding note reads like a template because by then everyone has sent one. It lands in not interested instead of interested.
  • Batches of 50 records cannot refresh a 5,000-company list inside three days, so throughput quietly becomes a reply-quality feature.
  • Reasons to skip decay too. A company that cut staff last week should not enter the run at all, whatever else it looks like on paper.
  • A handful of activity types drive most of the pipeline agents source, which is why choosing which reasons you act on beats collecting more of them. Explorium's read on intent data for AI sales agents sets out which ones earn their place.

The Thirty-Minute Reply Audit, Run in Chat

Export last week's replies, paste them into a chat, and ask for the five piles with counts. You get the honest number before anyone reopens a template. Two asks, in order.

Text
Here are last week's outbound replies as a CSV with columns: company, title, reply text.

Sort every reply into exactly one of five piles: interested, not interested, wrong person,
out of office or handoff, do not contact.

Return a count and a share of all replies for each pile, plus the interested count divided
by 6,000 messages sent. Then list every company in the wrong person pile with the title we
wrote to next to the title the reply pointed us toward.

The second list that comes back is the actual fix. It names the titles your list rules keep choosing and the titles your buyers keep pointing you toward, which is a rule change rather than a rewrite. With Vibe Prospecting connected to the same chat, the rebuild happens in the next message:

Text
Using Vibe Prospecting: for these 40 companies, find the person who owns revenue operations
today and return an email status for each one.

Skip any company that announced layoffs in the last 30 days, and flag any company whose last
funding or leadership change is older than 30 days so I can decide whether it is still worth
writing about.

Preview 5 records with the cost estimate before you build the full list.

Preview first, then build. A weak segment costs you five sample records and a cost estimate instead of a sending domain and a quarter. That preview step is the closest thing outbound has to a dry run, and the platform behind it is what keeps the company facts, the people and the recent activity on one timestamp.

Getting Vibe Prospecting Into That Chat

Add Vibe Prospecting from the Connectors Directory inside Claude (claude.ai, then Settings, then Connectors) or from the same directory in ChatGPT. That takes one click and no config file. On Claude Code, the Vibe Prospecting plugin does the same job. A free Explorium account is enough to run both asks above, and no call is required to get one.

Where Vibe Prospecting Fits

Reply quality is decided before the send, by who is on the list and why you are writing today. Vibe Prospecting is the layer that sets both, in the same chat window where you just ran the audit.

One Connection Behind Every Pile

  • 150M+ company profiles and 800M+ people profiles drawn from 50+ premium sources, so "matched" means one thing across the whole audit instead of three things across three tools.
  • 18 categories of recent company activity and 80+ types of it arrive next to the plain company facts: size, industry, location, the tools they run.
  • Reasons to skip, such as layoffs or an executive leaving, travel on the same feed as the reasons to write. That is what stops an agent messaging a company it should leave alone this month.

Enough Throughput to Audit the Whole List

  • Up to 1,000 entities per call, handled on the server at 100 QPS, so a full send list is refreshed inside a 72-hour window rather than sampled.
  • 97.8%+ company match accuracy is the upstream number that decides whether wrong-person replies are a matching problem or a writing problem.
  • Chat tools that pull every record into the model's context stall somewhere around 20 to 100 prospects. That is a sample, and a sample cannot tell you what your reply piles look like.

Spend You Can Point at Interested Replies

  • Free account, first list in minutes, no seat tax and no procurement cycle to sit through.
  • One credit pool across every call, which cuts agent-workload spend 30% to 60% against per-seat pricing, so renewal spend maps to interested replies.
  • Five sample records and a cost estimate before any credits move. Powered by Explorium Enterprise Business Data, and worth comparing side by side against whatever you run today.
"Rich and hard-to-find customer data, which helps in targeting the right audience and reaching decision-makers directly." An Explorium reviewer on G2, via a public review roundup

If You Change One Thing This Week

Retire total reply rate from the weekly report and put two lines in its place: interested replies over messages sent, and wrong-person share of replies. Everything else on this page is detail that hangs off those two.

  • Show the five piles as percentages with wrong person on its own line. Nobody argues with a pile they can read.
  • Report cost per interested reply. Cost per send flatters exactly the behaviour you are trying to stop.
  • Track the median age of your reason for writing and raise a flag past 72 hours.
  • Show how many contacts you are holding at accept_all or unknown, so the list gets repaired instead of sent.
  • Read the whole thing as one system rather than a copy problem: the architecture of an outbound engine shows which layer each number belongs to.

Total reply rate never told you whether outbound worked. Two lines and five piles will, and both of them are built before a single message leaves.

Fix the list side before the next batch sends. Connect Vibe Prospecting in chat
FAQs

Frequently Asked Questions

Get Started Banner

Get Started for free

Sign Up
Outbound Reply Quality: Replies Up, Calendar Empty