An AI outbound engine, built node by node
How I built and run DriveU.auto's outbound engine — seven n8n workflows, 120+ nodes, one operator.
What this is
I'm DriveU.auto's first GTM hire — I founded and ran the outbound and RevOps function as a team of one. DriveU sells an autonomous retrofit kit for terminal tractors to container terminals worldwide: a narrow ICP, six-figure enterprise deals, long multi-stakeholder cycles.
This page is about the system I built to run that function: an AI outbound engine covering enrichment, ICP scoring, contact discovery, personalized outreach generation, human approval, and metrics. It is an internal pilot under active iteration, deployed and live since April 2026. No spec, no engineering support — I built it myself, one node at a time.
The problem
One operator, a global market. The account map came first: before any automation, I researched the total addressable market by hand and built a proprietary database of 600+ organizations — it drove the company's entire pipeline and account strategy, and produced the deals. The engine came later, built on top of that foundation.
What I needed from automation was not more accounts. It was research depth and outreach preparation at a pace one person can't sustain manually — without giving up control over what actually reaches a buyer.
Architecture
n8n is the orchestration layer: chained workflows connected by webhooks, Slack events, and scheduled runs. Slack is the user interface — review cards with buttons, corrections made by replying in a thread, and a free-text router: type an account name in the channel and enrichment starts. The spreadsheet the team already used is the state machine: every workflow reads and writes it, so state survives restarts and needs no queue infrastructure. That was a deliberate adoption decision — ship into the tools the team already works in; the workflow layer stays replaceable.
The workflow chain. Approving a review card is a fan-out event: it writes state back, fires the next account's enrichment, and starts contact discovery for the approved one — one click keeps the belt moving. The metrics feed and error notifier run alongside everything.
A fixed pipeline, not an autonomous agent
The obvious 2026 shape for this system is an agent: give a model tools and a goal, let it decide the steps. I built that variant of the enrichment workflow too — and chose the pipeline. Fixed steps mean every step is auditable: I can point at any output and say which node produced it, from which inputs, at what cost. When something goes wrong, that property is worth more than flexibility.
The LLM is a bounded component
Bounded in both directions. Where deterministic code wins, the model was removed: the search-query builder started as an LLM call and is now plain code — cheaper, faster, and it deleted an entire JSON-parse failure mode. Data gaps are computed from empty fields in the record, not from a model imagining what's missing.
Where synthesis is the job, the model runs — wrapped in control:
- The search API's own AI-generated answer field is dropped entirely. I caught it fabricating a partnership claim about a terminal we were researching, extrapolated from an unrelated press release. Raw sources only.
- Search results are filtered by token-matching the account's name, with stopword and abbreviation handling — terminal names often carry abbreviations that match several unrelated ports.
- The ICP safety net: code re-parses the model's own scoring table, validates each factor against the legal rubric values, and recomputes the total deterministically. If the model's arithmetic is wrong, the report is corrected in every place the score appears — and an audit row records what changed.
- Confidence is deterministic, not vibe-based: "high" requires five or more distinct cited sources and no unresolved conflicts.
- Conflicting sources are surfaced, not resolved silently. A surfaced conflict is more useful than a silent guess.
One enrichment run. The single LLM call sits between deterministic preparation and deterministic validation — code decides what the model sees, and code checks what the model returns.
Human in the loop, by design
Every enriched account lands as a review card in Slack with two buttons: approve, or request corrections. Approval is idempotent — the handler re-reads the sheet before doing any work, so a double-click or a click on a stale card produces a polite "already approved" reply and zero side effects.
Corrections are conversational. Click the button, then reply in the thread in plain language. A 45-second debounce batches multiple replies into one surgical rewrite — and the debounce clock is a timestamp column in the sheet, so it survives restarts with no queue infrastructure. The rewrite contract is strict: apply only the stated corrections, user facts override web sources, and contradicted citations get marked as superseded — never deleted. The updated report replaces the original Slack message in place, so there is exactly one card per account and the whole review conversation stays attached.
Just as deliberate is what I have not automated yet:
- Sending is manual. Outreach messages are generated per contact and queued behind an accept button; a human sends them. The first touch is the most expensive moment in outbound — it stays human until the system has earned more.
- Reply handling is not live. I built a reply handler, pulled it from rotation, and queued a rewrite. Capabilities enter the loop in stages, behind trust.
- Net-new lead discovery is not started. Sourcing today means pulling accounts from the researched database, not finding new ones.
Cost and observability
Cost is an architecture property, not an afterthought. One account per execution — easier to reason about, easier to debug, and cheap enough that there's no pressure to batch: about $0.09 per enriched account, about $0.06 per corrections round. Models are selected per task on cost/performance (OpenAI GPT / Claude models), benchmarked and re-benchmarked as pricing moves.
Failures from any workflow route to a central error notifier in Slack. That notifier taught its own lesson: a provider outage once flooded it with identical alerts, so the scheduled reads got retry-and-continue hardening — deliberately not applied to the interactive paths, where a failure must surface immediately.
On top of the engine's own data sits a read-only metrics feed powering a live GTM dashboard: funnel, conversion, velocity, and per-unit cost. The system reports on itself.
What's next
Sequenced multi-touch outreach, the reply-handler rewrite, and CRM integration — the workflow layer is replaceable by design, so wiring the engine into a CRM is an integration task, not a rebuild. Each of these follows the same staging rule as everything above: nothing goes autonomous until the human-reviewed version of it has earned that.
Contact
If you're building or hiring around AI-native GTM systems, I'd like to talk — dor@tzabari.dev · LinkedIn · more about me.