n8n is the right tool for a linear workflow: pull a lead from a form, enrich it, write it to a CRM, notify a human. It stops scaling when the work turns conversational, since branching, state and model evaluation are hard to reason about on canvas and per-execution pricing makes them expensive. That crossover is where custom code costs less
We have built both. The real-estate inventory system we shipped runs routing that would be a small n8n canvas if it were simpler, and it is not simpler. Below is where we draw the line, the failure modes that force it, and the decision table we use when a client walks in with "can you just do this in n8n?"
What is n8n actually good at?
n8n shines at straight-through processing: a trigger, a transformation, a destination, no human choices in the middle. Typical cases we agree are n8n territory: a Google Sheet row becomes a HubSpot contact; a Property Finder enquiry lands in a WhatsApp thread and a CRM; a daily digest aggregates several sources and emails it. Each step is a fixed node, the data moves in one direction, and the failure is visible in the canvas
This is the class of work the tool was built for, and it is often the fastest correct answer. For a one-directional job with five to fifteen nodes, an n8n flow is cheaper to build and easier for a non-engineer to inspect than the equivalent code. Choosing custom code there is ceremony
When does n8n stop scaling?
The tool stops scaling when the workflow stops being a flow and becomes a conversation. Three changes mark the boundary
State is the first. A linear flow carries each document through once. A conversational agent keeps a thread open: it remembers what the customer already answered, which offer they saw, whether they are qualified yet. That state has to live somewhere the flow can query mid-run, and every node that reads and writes it adds coupling you debug from the canvas
Branching is the second. Real lead qualification forks: an Arabic-speaking buyer on WhatsApp follows one path, a caller with a budget cap follows another, a tenant doing a renewal follows a third. Each branch doubles the routing nodes, and the evaluator that decides which branch to take is itself a node that can misfire. The canvas stops being a map and becomes a maze
Evaluation is the third. An agent that must decide whether to escalate to a human, or whether an extraction is confident enough to file, is not running logic you can see in a node. It is running a model call with a threshold, and the threshold needs iteration. Iterating on model output is a code loop, not a canvas operation
What are the failure modes of an n8n agent?
The most reliable signal is visual. When your canvas has more routing nodes than work nodes (more switches, conditionals, and error handlers than actual data-transforming steps), the flow has stopped being a pipeline and started pretending to be a program. n8n can express branching, but it expresses it as a sprawl that is hard to test, hard to version, and hard to explain to the next person who opens the editor
Testing is the quiet killer. A linear flow you test with one happy-path sample and a couple of error samples and you are done. A branching conversational agent needs a test set: a hundred conversations with known outcomes, so a change to a prompt or a threshold does not silently break qualification. Building that test harness inside a canvas is painful. In code it is a CI job that runs on every pull request
Per-execution pricing is the cost killer. n8n's hosted product bills a subscription that covers a fixed number of workflow executions (n8n pricing), so a conversation that takes many model round-trips consumes several executions for one customer. We estimated a conversational WhatsApp qualifier running inside n8n cloud at several times the on-demand token cost of the same conversation in custom code, because every hop between nodes counts once per run. For a channel that answers hundreds of leads a day, that markup is material.
The decision table: n8n or custom?
This is the artifact we actually use in scoping calls. Ask the four questions; the answers place the workflow on one side of the line
| Question | n8n (workflow tool) | Custom agent |
|---|---|---|
| Does the data move one way, once? | Yes | Overkill |
| Does the conversation branch and hold state across turns? | Sprawl, hard to test | Natural fit |
| Do you need a test suite run in CI before changes ship? | Painful in canvas | Standard practice |
| Is per-execution volume large enough that pricing matters? | Plan execution cap | Pay per model call |
| Does it touch customer data or money where a bug is costly? | Fine for low blast radius | Own the failure surface |
| Will a second engineer need to reason about it in six months? | Canvas can be read | Code can be reviewed |
The heuristic that shortcuts all six questions: count the nodes. When routing and error-handling nodes outnumber the nodes that actually transform data, the workflow has crossed into agent territory. That is the point to stop extending the canvas and start writing code
How much does each approach cost in Dubai?
The honest answer is that n8n is cheaper to start and a custom agent is cheaper to run, and the crossover depends on conversation volume and how many edge cases you have. For a linear intake flow doing a few hundred records a month, n8n is the right spend. For a WhatsApp qualification bot answering several hundred leads a day, with Arabic and English branches and a human-escalation loop, the recurring execution costs and the test burden push the total toward a custom build
We put the build-side AED numbers for both paths in how much AI automation costs in Dubai, and the ongoing model, hosting and retainer costs in what AI systems cost to run after launch. The running-cost line is where the two curves cross
In the inventory CRM we shipped, the routing that decides which agent sees a lead and whether a change needs human approval sits in about a hundred and fifty lines of code, versioned and tested. The same logic as an n8n canvas would run to two screens of nodes and a debug session every time someone touched it
Can n8n and custom code share one system?
Yes, and this is the configuration we increasingly recommend instead of an either-or. Use n8n for the parts that are genuinely linear (the integrations, the schedules, the CRM writes, the notifications) and give a custom service ownership of the parts that branch, hold state, or evaluate model output
A concrete split: n8n watches the webhook and the queue, calls the custom conversation service when a message arrives, and writes the resolved outcome back to the CRM. The custom service owns the conversation logic, the escalation decision, and the test set. The boundary question has one rule: state and branching live in code, plumbing lives in n8n
This matches how we think about agent systems overall: capture and act are happy as workflow nodes, reason and route are happier as tested code
When does the crossover happen for a Dubai business?
Look at the failure surface. A Dubai brokerage with portal leads from Property Finder and Bayut and bilingual WhatsApp qualification is running a decision-heavy, stateful, per-conversation workload. That is custom-agent territory, and the creep toward it is exactly where the agency-vs-custom decision also lands: customer-facing and multi-system work earns the bigger build
Wherever your conversations touch personal data (buyer names, phone numbers, Emirates ID on a document), the storage and deletion paths fall under the UAE PDPL (Federal Decree-Law No. 45 of 2021), whichever side of the line you pick. A canvas and a codebase both need a data-flow map; that is not the tool's problem, it is the system's.
The takeaway
Use n8n for the linear plumbing and reach for a custom agent the moment conversations branch, hold state, or need evaluation you can test. The reliable boundary is node count: when routing nodes outnumber work nodes, the canvas has stopped earning its keep. The cheapest path is usually both: n8n for the integrations, tested code for the conversation
Book the free audit: we map your workflow, flag the nodes that are doing program work without you noticing, and show you where the crossover sits before you build the wrong thing
FAQ
Is n8n good enough for AI automation?
For linear workflows — intake, enrich, write to CRM, notify — n8n is often the fastest correct answer. It stops scaling when conversations branch, state gets complex, or you need model evaluation loops
When do you outgrow n8n?
When the workflow needs multi-turn reasoning, per-tenant configuration, sub-second latency at volume, or tests you can run in CI. The signal: your canvas has more routing nodes than work nodes
Is a custom AI agent more expensive than n8n?
Up front, yes; over time it depends. A custom build costs more to start but carries no per-execution pricing and handles edge cases n8n flows paper over
Can n8n and custom code coexist?
Yes, and often should: n8n handles integrations and schedules while a custom service owns the conversation logic. The boundary is where state and branching live
Sources
- n8n pricingn8n.io
- Federal Decree-Law No. 45 of 2021uaelegislation.gov.ae
