# The five steps of outbound, and which one an agent closes

> Outbound splits into five steps. Across 3,706 sourced people and 1,651 invitations from our own profiles, a model closed one of them end to end. Here is the split, with numbers.

Page: https://re-vault.io/blog/five-steps-of-outbound/
Published: 2026-09-02 · Dmitry Ostrovtsev, Re:Vault


Between 29 July and 2 September 2026 our own sending profiles sourced 3,706 people, sent 1,651 connection requests, collected 344 accepts and 72 replies. A model with tools ran the whole thing: it picked the people, wrote every line, queued every send and answered most of the replies. That is a useful sample for a question founders keep asking us in different words. If an agent can do outbound, which part of it does the agent actually finish?

Outbound splits into five steps: choose the signal, source the people, verify them, write the message, handle the reply. In our logs the model closed one of those five end to end. It does most of the work in three others, with a person holding a specific piece in each. The fifth waits on arithmetic that no single run can produce. Every figure below carries its denominator, and the splits inside them run on small samples, so read them as direction and rerun the same cuts on your own data.

## Step one: the signal is chosen by the calendar

A signal is the reason you are writing this week: a company posted a sales role, someone changed jobs, a lease came up, a round closed. Picking between them looks like a judgment call, which is exactly the kind of thing a model is good at. The obstacle is arithmetic.

Since 29 July we ran 225 batches. The median batch carried 6 sends. Six batches out of 225 crossed 30 sends. Three collected three or more replies.

Rolled up to the signal level, the picture is the same shape. Our default list with no event attached: 859 sends, 19 replies, 2.2%. A hiring event: 428 sends, 10 replies, 2.3%. A role change: 230 sends, 4 replies, 1.7%. Those three rates sit inside each other's noise. Ranking them after any single run produces a confident answer built on four replies.

So the model proposes and the calendar decides. Our rule is a floor of 30 sends and 3 replies in a slice before that slice is allowed to change anything, and one change a week, applied on Monday against the full week of data. A model running at nine in the morning has a run. The decision needs a quarter. The practical version of this for a two person company: write the signal down when you start using it, tag every message with it, and refuse to have an opinion about it until the counter passes thirty.

## Step two: sourcing works when the rejections are written down

Sourcing is finding the actual humans behind the signal. This one a model closes, on a condition that took us a while to see.

Of everything the sourcing runs looked at since 29 July, 526 candidates were rejected with a reason code recorded at the moment of the rejection. Grouped:

- 240 wrong person or wrong company, 45.6%
- 101 with no checkable fact available, 19.2%
- 67 whose record failed verification, 12.7%
- 64 already in our world, 12.2%
- 25 on the exclusion list, 4.8%
- 29 other, 5.5%

That ledger is the whole trick. A rejection reason exists for about one second, while the model is looking at the profile and deciding. Written down then, it becomes a filter that gets sharper every week: 45.6% wrong person told us our title list was too loose, and the fix was a role check on the profile itself, because a job title on a list is a claim about last year. Left unrecorded, that second is gone, and by Friday you have a pile of names with no memory of why the other four hundred were skipped.

Two things make the ledger work. Every rejection carries a code from a fixed vocabulary, so the reasons add up across weeks. And the dedup runs against everyone ever touched from that sending profile, where this week's list is the smallest part, because the expensive duplicate is the person you wrote to in March.

## Step three: verification decides whether steps four and five exist

Verification asks two questions about a sourced name. Is this person still in that job at that company, and is there one fact about them that can be checked by a stranger?

Of the 3,706 people sourced in this window, 2,336 carried a checkable fact, 63%. Two thousand and thirty four of them reached a first touch, 55%. The rest sit parked, excluded or waiting, and that gap is the honest cost of the step.

The fact requirement is the part worth copying. A fact has a source: a post the person wrote, a page on their product, a job ad on their own careers page, the signal itself. A description with no source is a guess, and a guess spends the single first impression you get with that person. Our runs record where each fact came from, and a lead that arrives without one goes back to the pool rather than into a template.

Verification is reading a lot of pages and comparing them, which is the part of this job a model does fastest. The piece a person holds is the definition: what counts as a fact, what counts as a source, and which roles are in scope. We learned the last one by writing about a job ad to the person who had posted it, and that person's job was forwarding applications to the hiring manager.

## Step four: writing is the step a model closes end to end

This is the one. Given a verified person, a fact with a source and a house style, the model writes the message. We moved first touches off individual review in July, and what replaced that review is a set of checks that run on every draft before it can be queued.

Since 29 July those checks ran 14,717 times against messages to people and caught 650 problems, one for every twenty three checks. The largest categories:

- assembled from a template rather than for this person: 273 caught of 437 checked
- ending without a direct question: 265 of 2,605
- the same text about to go to two different people: 122 of 1,412
- a stock phrase from the banned list: 101 of 3,492

The 273 out of 437 is the number to sit with. Left alone, a model writing at volume settles on a skeleton it liked once and refills the slots. Frequency is the tell: banning a phrase moves the sameness into the next phrase, so the check that holds is a limit on repetition inside a batch. Ours: no two messages share a closing line, no five word sequence appears in more than 30% of a batch.

The second half of closing this step is a stop rule, and we learned it by watching a loop misbehave. Since 4 August our logs hold 1,702 rejected drafts. They concentrate on 11 people, and 689 of them belong to a single person. The writing loop had a check and a blank where the counter should be, so it rewrote the same doomed draft several hundred times, each attempt costing a model call and earning the same rejection. A writing agent needs a ceiling: three attempts, then park the person and say so out loud. A quality gate with a counter beside it is the difference between a stop and a bill.

## Step five: the reply is a race, and the classifier is the boundary

Since 29 July, 167 inbound messages arrived tied to a person we had written to. We answered 142 of them. The median answer landed 18 minutes after the message. One hundred and thirteen were answered inside two hours, and six took longer than 48 hours.

Eighteen minutes is the argument for putting an agent on replies at all. It lands while the person is still inside the tab they typed from, and that is where a two line exchange turns into a conversation.

The boundary sits in the classifier. Answered by the machine: agreeing a day and a time, sending the booking link, factual questions about how the work runs, and polite closes. Handed to a person as a draft: price, discounts, contract terms, incoming pitches, anything negative, and anything the classifier reads as ambiguous. Two rules keep that safe. Ambiguity always resolves toward the human. And the queue keeps moving while it waits, because the safe default goes out on its own and a person's answer changes course from there.

Booking is the exception inside the exception. The mechanics are automatable and we automate them. The judgment about who gets a calendar slot stayed with a person, because a wrongly booked call costs an hour out of a founder's week and a wrongly declined one costs a customer.

## What to run on Monday

Five checks, all on data you already have.

1. Count sends per signal for the last month. Any slice under 30 sends and 3 replies is a story you are telling yourself, so mark it unknown and keep sending.
2. Open your sourcing process and add a rejection log with a fixed list of reason codes. One week of it will tell you which end of your list is broken.
3. Take twenty messages you sent last month and check whether each one names a fact that a stranger could verify, with the source written down beside it. Our pool ran at 63%.
4. Read the last lines of your last batch. Two identical closings in one batch is a defect of the batch, and it is the earliest visible sign of drift.
5. Measure the median time between a reply arriving and your answer going out. If it runs in hours, that single number is the cheapest thing on this list to fix.

All five run on the same idea: outbound is five separate jobs with five separate failure modes, and the automation question deserves a separate answer at each one. If you want to run the machinery yourself with your own model, our stack exposes the sourcing tools directly and that door is described on the [MCP page](/mcp/). If you would rather have it operated for you, with the ledgers, the checks and the counters already built, that is our work, and the levels are laid out on our [levels and pricing page](/levels/).


---

Re:Vault runs LinkedIn outreach end to end for B2B companies: finds the buyers, writes in the client's voice, handles replies, books meetings. $2,000/month. MCP access for your own Claude: from $199/month.

- Book a 30-minute call: https://cal.com/dmitry.o/re-vault-gtm
- Email: dmitry@re-vault.io
- MCP access: https://re-vault.io/mcp/
