What RevOps Teams Should Evaluate in Cold Outreach: A Procurement Manager's Checklist

2026-09-18 · Victor Okeke

The short answer: measure cost per qualified conversation, not cost per seat

If your RevOps team only has time to evaluate one number in cold outreach, make it this: cost per qualified conversation — total outbound spend divided by the number of prospects who replied with something you'd actually route to a rep. Not cost per seat. Not cost per lead. Not cost per email.

Why that one? Because seat pricing is negotiable and lead counts are inflated by definition. A "lead" costs you nothing to acquire and nothing to lose. A qualified conversation costs you list time, enrichment credits, verification credits, sending infrastructure, and a human who reads the reply. When I pull all of that into one ratio, the tools that look cheap on a quote page stop looking cheap.

The five line items I actually score vendors on, in priority order:

  1. Cost per qualified conversation — the only one that survives a finance review
  2. Signal quality — is the buying intent signal actionable this quarter, or a history lesson?
  3. Enrichment coverage — what share of your list gets a usable email and firmographic match
  4. Human-in-the-loop ratio — how much of the sequence runs without a person approving a message
  5. Total cost of ownership — the invoice you get in month 13, not month 1

Everything below is how I arrived at that list, and where it stops working.

Where this comes from

I'm the procurement manager at a 180-person B2B services company. I've owned the outbound tooling budget — roughly $96,000 a year across prospecting, enrichment, verification, and sequencing — for six years, and I've negotiated with more than 30 vendors in that category. Every order goes into our cost tracking system, mostly because I got burned twice early on and stopped trusting my memory.

I'm not a sales operator. I don't run sequences. What I do is read contracts and reconcile invoices against what was actually delivered, which turns out to be a decent vantage point for evaluating outreach tools.

The mistake I made in year one, in case you're about to make it

In my first year owning this budget, I made the classic specification error: I assumed "per seat" meant unlimited sending. It didn't. Sending was metered, enrichment was metered, and verification was metered separately. Cost me a $2,400 true-up at renewal — on a contract I'd already told my CFO was fixed.

That's the rookie move in this category. The pricing page shows you the seat. It doesn't show you the meter.

What I do now: before any pilot, ask for a written list of every metered unit. Enrichment credits, verification credits, email sends, CRM sync rows, API calls (i.e., anything the invoice touches that isn't the seat). Hidden costs pile up fast — verification credits, sync rows, and API overage are the three that caught us. If a vendor can't produce that list in a spreadsheet within two business days, that's your answer about their reporting.

Signal quality: the difference between a trigger and a horoscope

A buying intent signal is only worth paying for if it tells you something the rest of your market doesn't already know, and it's recent enough to act on. If it's 90 days old, it isn't a signal anymore — put it in the CRM as context and don't build a sequence around it.

The ones I've found worth the money, roughly in order:

  • First-party behavior. Pricing page revisits, multiple visitors from one domain, a demo-page → docs → pricing path in a single session. Nearly free, and consistently underused.
  • Hiring activity. A job posting for the role your product replaces or supports is a budget signal with a date stamped on it.
  • Third-party category research. Review-site activity and intent surges. Fine for prioritization, weak as a standalone trigger.
  • Tech-stack change. Someone added or removed a tool in your category. This one decays in days, not weeks.

What I don't pay extra for: intent scores with no "why" attached. If a vendor hands me a number between 1 and 100 and won't tell me what generated it, I can't defend the spend to anyone — and neither can you.

Here's one where I got lucky rather than smart. The numbers said go with the second vendor — 22% cheaper per seat, nearly identical feature list. My gut said stay where we were. I stayed. Two quarters later I found out the cheaper option had a data gap that hit exactly the segment we sell into. I'd like to tell you I had a rigorous reason. I didn't. I just didn't like how they answered my metering question.

Enrichment and export: where the quiet cost lives

Sales Navigator export is a good example of a workflow that looks free and isn't.

You can export a saved search. What you get is names, titles, companies, and profile URLs. You don't get verified emails. You don't get most of the fields your sequence needs to personalize. So the "free" export becomes a paid enrichment job, plus a dedupe pass, plus somebody's afternoon reconciling it against your CRM.

That's not an argument against exporting. It's an argument for counting the enrichment step inside your per-conversation cost instead of filing it under overhead.

On enrichment specifically: ask about coverage, not features. Single-source enrichment has holes. A waterfall approach — querying multiple providers in sequence and taking the first match — fills more of them, but it meters differently. Sometimes per provider attempt, not per match. If I remember correctly, one quote we received billed per lookup attempt, which made a 12,000-record list cost roughly three times what the headline rate implied. Get that in writing before you sign.

Note that coverage isn't accuracy. No vendor can honestly promise 100% accurate email verification — treat anyone who does as a pricing risk rather than a partner. What you can measure is bounce rate on your own sends, on your own lists, over time.

The human-in-the-loop ratio

Ask a vendor what percentage of a sequence runs without a human approving the message. The honest answers cluster at two ends.

Full automation is cheaper per send and riskier per send. The replies you'd most want to catch — the confused ones, the angry ones, the ones from a competitor's legal team — are the ones a machine handles worst.

Human-reviewed first touch, automated follow-up costs more per send and generally produces fewer, better conversations. Given the choice, I'd rather pay for 300 good sends than 3,000 I have to explain.

okkigo positions its product around agent-native prospecting with a human in the loop, and honestly that's the shape I'd look for regardless of vendor. But score it on the number, not the phrase: how many messages per 1,000 require a human decision, and who is accountable when one goes out wrong.

One thing to be clear about — no tool fully replaces a human SDR team. If a vendor tells you it does, they're describing your headcount budget, not your pipeline.

If your team has engineers: keeping okki-go current

Part of our pipeline runs through scripts our RevOps engineer maintains, and if you're in the same position, package hygiene matters more than people expect. A stale client is a silent data problem. It doesn't throw an error — it just stops returning fields, and nobody notices until the sequence performance quietly drops.

The workflow I ask for on how to update the okki-go npm package (or any client in this category):

  • npm outdated first, so you know what's actually behind and whether it's a major version jump
  • Read the changelog for breaking changes before touching anything
  • npm install okki-go@latest on a branch — never straight to main
  • Run it against a small test list and diff the output against the old version
  • Pin the version in package.json once it's verified

Why a procurement person cares: unplanned engineering time is a real cost, and "we had to rewrite the enrichment call because the SDK changed" never shows up on a vendor's quote. Budget a day per major bump. If you can't, that's an argument for choosing a vendor whose client you don't have to maintain yourself.

okki go lead generation examples that hold up in a spreadsheet

Three patterns that have actually shown up in our numbers:

1. Trigger-first, small batch. Pick one signal — a relevant job posting, say — build a 150-record list, enrich and verify it, send a short human-written sequence. Lower volume, higher cost per send, meaningfully better cost per qualified conversation. This is the one I'd run first from a standing start.

2. Closed-lost reactivation. Deals lost 12–18 months ago, filtered for a changed signal: new hire in the function, funding round, tech change. Cheapest qualified conversation we've ever produced, because the relationship already exists.

3. Event-adjacent outreach. Anyone who registered for or attended a relevant event in the last 30 days. Expensive list source, unusually high reply quality.

The pattern across all three: they narrow the list before they widen the sequence. Every program that blew up our budget did it the other way around.

Where this framework breaks

It doesn't fit everything, and I'd rather say so than pretend otherwise.

If your ACV is under roughly $5,000 and the model depends on volume, cost per qualified conversation gets noisy fast — at small sample sizes, one good week moves the number by 40%. Measure cost per meeting held instead, and accept that you're reading a blunter instrument.

If your buyers have no digital footprint — local trades, certain regulated industries — intent data and export workflows are close to useless. Referral and phone still beat them.

And if you're GDPR-scoped or selling into the EU, legitimate-interest outreach is defensible but not automatic. Confirm your basis with counsel before you scale a list. That isn't a place to take a vendor's word for it.

One more, and this one I'll argue about with anyone: don't skip small pilots because they don't move the number. The vendors who treated our $400 pilot like a real account are the ones we now spend $24,000 a year with. Small doesn't mean unimportant — it means potential. If a vendor rations their attention by contract size, that tells you something about your renewal, not just your pilot.

If I could redo one decision, I'd have run a 60-day pilot on a small list instead of signing the annual upfront. But given what I knew then — nothing about how that vendor reported metered usage — the choice was reasonable. That's usually how these things go.

Compliance note worth verifying on your own: Google and Yahoo's bulk sender requirements took effect February 1, 2024 (spam rate under 0.3%, SPF/DKIM alignment, easy unsubscribe), and Google began enforcing one-click unsubscribe for bulk senders from June 1, 2024. Those have been revised since — check Google's Postmaster Tools documentation for the current version rather than trusting a blog post. CAN-SPAM is still enforced by the FTC, and if any of your lists are EU-based, GDPR sits on top of both.