Building AI Agents That Don't Fall Apart
You've wired up an LLM call, handed it a couple of functions, and watched it do something genuinely impressive — then watched it do something baffling on the next three inputs. This post is about the gap between those two runs: what an agent actually is, where the leverage is, and the guardrails that turn a demo into something you'd let near production data.
What an agent actually is
An agent is a model in a loop with tools and a termination condition. That's it. Three parts:
- The loop — you call the model repeatedly instead of once, feeding each result back in
- The tools — functions the model can request by name, with arguments it produces
- The termination condition — the rule that decides when to stop calling
I like this definition because it's the one that tells you what to engineer. Every fancier thing people build — planners, reflection passes, memory layers, swarms of subagents — is a variation on those three knobs. When an agent misbehaves, the fix is almost always in one of them: the loop ran too long, the tool was badly designed, or nothing told it to stop.
One turn of that loop is small: the user's goal plus everything observed so far goes to the model; the model either answers or asks for tools; a tool runner executes the calls and feeds the results back in as observations for the next turn.
The core loop, in code
Here is the whole thing, with nothing clever in it. Provider APIs differ in naming, but every tool-calling API is this shape. The line that makes it a loop is the last one in the body: observations go back in as a user turn.
def run_agent(goal: str, tools: dict, max_steps: int = 12) -> str:
messages = [{"role": "user", "content": goal}]
for _ in range(max_steps):
response = call_model(messages=messages,
tools=[t.schema for t in tools.values()])
messages.append({"role": "assistant", "content": response.content})
calls = [b for b in response.content if b["type"] == "tool_use"]
if not calls:
return response.text # the model answered instead of asking
messages.append({"role": "user", "content": [
{"type": "tool_result", "tool_use_id": c["id"],
"content": execute(tools, c["name"], c["input"])} for c in calls]})
raise Halt(f"no answer after {max_steps} steps")Four things terminate a run, and you want all four wired from day one:
- The model stops requesting tools and answers
- The model calls an explicit
submit_answertool — better than (1) when you need a structured result - The step budget or wall-clock deadline expires
- Something outside cancels it — a user hitting stop, a supervisor killing the job
Most "agent" problems are workflows in disguise
This is the most useful opinion in the post, so it goes early: if you can draw the steps on a whiteboard and they're the same for every input, you don't need an agent — you need a function. An agent buys you exactly one thing, the ability to decide the next step at runtime based on what it found in the last step. That is a real superpower for open-ended tasks, and it is also what you pay for in nondeterminism, cost variance, latency variance, and debuggability. Paying that price to execute a fixed five-step sequence is a bad trade, and a lot of "agentic" software is doing exactly that.
| Shape | Fits when | What you give up |
|---|---|---|
| Deterministic pipeline | Steps are known ahead of time and identical for every input | Adaptivity — a novel input falls off the path with no recovery |
| Model-routed branches | A small, enumerable set of paths; the model just picks one | Anything outside the branches you wrote |
| Single agent with tools | The number and order of steps genuinely depend on mid-run findings | Determinism, bounded cost, easy step-through debugging |
| Multi-agent / subagents | Subtasks are separable and each needs its own large context | Shared context; you add handoff loss and coordination bugs |
Work down that table, not up. The honest test: write the deterministic version first. If you can't — because step three depends on what step two found, in a way you can't enumerate — you have an agent-shaped problem. Otherwise you have a workflow with an expensive if statement in it.
Tools are where the leverage is
Prompt tweaks hit diminishing returns fast. Tool design doesn't. If your agent is flailing, the tools are the first place to look — the model can only be as competent as the interface you gave it.
- Name tools like an API someone else has to use.
search_ordersbeatsquery.read_filebeatsfs. The model has no documentation beyond the name, description, and schema, and the name is what it pattern-matches on first. - The description is a prompt, so write it like one. Say what the tool does, when to reach for it, when not to, what it returns, and how it relates to its neighbours. A description that disambiguates against the other tools is worth more than one that describes this tool perfectly in isolation.
- Make illegal calls unrepresentable. Every constraint you push into the schema is a class of failure you never handle: enums instead of free-text status strings, explicit required fields, formats and bounds on everything. A schema is cheaper than a validation round-trip.
- One well-designed tool beats five overlapping ones. If two tools both plausibly answer "how do I find a customer's recent orders", the model will pick wrong some of the time — and the failure is silent, because the wrong tool returns a perfectly valid empty result. Overlap also compounds: on a large flat tool list, selection gets less reliable and every schema burns context on every single turn. Prefer one tool with a well-chosen parameter over three near-synonyms.
- Return what the model needs to pick the next step. Not a raw API dump. Trim fields it can't use, cap the size, and include the identifiers it will need for a follow-up call.
The description on search_orders is what that looks like written out — it disambiguates against its neighbours before it describes itself:
Find orders for a single customer, newest first. Use this when you have a customer_id and need their order history or a specific order_id. Do NOT use this to look up a customer by name or email — call find_customer first to resolve an id. Returns at most
limitorders with id, status, total_cents and placed_at; call get_order for details.
The schema carries the rest: customer_id required, since an ISO-8601 date, limit an integer bounded to 1–50, and status an enum of pending, shipped, delivered, cancelled rather than a free-text string.
Validate at the tool boundary, and return errors the model can use
The schema you ship is a strong hint, not a guarantee. Models emit arguments that are the right shape but wrong: a date in the wrong format, an id it invented, a required field quietly dropped. Treat every tool input as untrusted and parse it before it touches anything real.
Parse the raw arguments against a real schema before anything runs — refund_order's RefundInput takes an order_id matching /^ord_[a-z0-9]{8}$/, a positive integer amount_cents, a reason from a four-value enum, and an idempotency_key UUID — and hand that same schema to the model as its JSON Schema. What matters is the failure path:
async run(raw: unknown): Promise<ToolResult> {
const parsed = RefundInput.safeParse(raw);
if (!parsed.success) {
// Not a throw. This is a message addressed to the model.
return { is_error: true, content:
`invalid_arguments: ${formatIssues(parsed.error)}. An order_id looks ` +
`like "ord_8fa21b3c" — get one from search_orders, do not guess.` };
}
return { is_error: false, content: await refunds.create(parsed.data) };
}The pattern that matters: tool failures are observations, not exceptions. An exception ends the run. A tool result that says what went wrong gives the model a chance to fix it, which is the entire point of running a loop. Only crash out for things the model genuinely cannot resolve — auth failures, a dead dependency, a budget hit.
Error text is prompt text, so write it that way. The good version below names the field, states the rule, echoes the bad value, and points at the next action; the bad one just invites another guess.
Bad: ValueError: invalid literal for int() with base 10: 'last week'
Good: invalid_arguments: `since` must be an ISO-8601 date such as
"2026-04-01"; you passed "last week". Call get_current_date if
you need today's date to compute a relative range.Handle malformed calls at the same boundary. A tool name that doesn't exist gets a tool result listing the valid names — not a stack trace. Unparseable JSON arguments get a result saying exactly that. In both cases the loop keeps turning and the model usually recovers on the next step.
Planning: three shapes, and when each earns its keep
Interleaved reasoning and acting (ReAct-style) is the default, and it's the loop above: think, call a tool, look at the result, think again. It adapts well because every decision is conditioned on real observations. It also has no global view — on long tasks it drifts, forgets the original constraint, and rediscovers things it already knew twenty turns ago.
Plan-then-execute has the model produce an explicit plan first, then work the steps. The plan is a reviewable artifact, which is worth a lot: you can show it to a human before anything executes, run independent steps in parallel, and often let a cheaper model execute a well-specified step. The catch is that a plan written before any evidence exists goes stale the moment reality disagrees, so you need an explicit replan trigger — a step failing, or a check that the remaining steps still make sense.
Reflection adds a critique pass over the output before it ships. It reliably catches format problems, missed requirements, and skipped steps. It is much weaker as a correctness check: a model grading its own work with no external signal tends to agree with itself. Attach reflection to something checkable — tests that run, a schema that validates, a linter, a query that either returns rows or doesn't — and it becomes genuinely useful.
My default: start with the plain loop. Add plan-then-execute when a human needs to see the plan first, or when steps parallelize. Add reflection only when you have an objective signal to reflect against.
Context and memory are two different things
Working context is the transcript in the current request — everything the model can see right now. It's finite, and more importantly it degrades before it's full: as a transcript grows, instructions in the middle carry less weight than the ones at the ends, and stale tool output crowds out fresh. Long agent runs die of this.
Durable memory is state that survives the run: files, database rows, a scratchpad the agent reads and writes. It's external, addressable, and re-readable.
The rule that follows: treat working context as a cache, not as the store. Anything you'd be upset to lose belongs outside the transcript.
Four techniques, in the order I reach for them:
- Cap tool output at the boundary. A tool that can return a hundred kilobytes should return the first page plus a cursor. The model asks for more if it needs more. Cheapest win available, and most people skip it.
- Give the agent a scratchpad.
write_noteandread_notebacked by a file. The agent records its plan and findings as it goes, so state is addressable rather than buried in turn 14 of the transcript — and it survives compaction. - Compact old turns. Summarize the early transcript, keep recent turns verbatim.
- Pass state forward explicitly. When one step hands off to another, put the state in the message. Never rely on it being remembered because it was mentioned earlier.
One gotcha before you write that compaction step: a tool_result has to keep the tool_use that produced it, so the split point between the summarized head and the verbatim tail has to walk backwards until it no longer orphans a tool call.
Then ask the model to summarize the work so far for a successor who cannot see those turns, and name what has to survive:
- the original goal
- decisions made, and why
- identifiers discovered, and records modified
- what failed, and what was already ruled out
- and none of the raw tool output
Compaction is lossy and the loss is silent, so the summary prompt is doing real work. Keeping identifiers and dead ends is what stops the agent from repeating an hour of exploration.
Multi-agent: what you actually buy
A subagent is a fresh context window with its own tools and its own termination condition, invoked by a parent as if it were a tool. The parent hands it a scoped task, a subset of the tool list filtered against an allowlist of what subagents may touch, and its own step budget. Only its final report crosses back.
Two things make this worth the trouble: context isolation — a subagent can read forty files and hand back three paragraphs, and the parent never carries the forty files — and parallelism on genuinely independent subtasks.
Everything else is cost. The parent only sees what the subagent chose to report, so handoffs are lossy in the way a rushed code review is lossy; a subagent that misunderstood its brief returns a confident, wrong summary the parent has no way to check. Debugging becomes distributed, and you're paying for several loops instead of one. So: delegate when a subtask consumes a lot of context and produces a small answer — search, exploration, verification, "read these and tell me which ones matter". Don't delegate when the subtask needs the parent's full context to be done correctly, and don't let two subagents write to the same state concurrently. That's a distributed-write problem, and you did not set out to build a distributed system.
Failure modes and the guardrails that stop them
| Failure mode | What it looks like | Guardrail |
|---|---|---|
| Runaway loop | Same tool, same args, forever | Step budget and wall-clock deadline, always both |
| Tool thrashing | Ping-ponging between two tools, no progress | Hash name and args; on repeat, return "you already called this" |
| Hallucinated arguments | Plausible ids that don't exist | Validate against the real store; return an error naming how to get a real one |
| Hallucinated tool name | Calls a tool you never defined | Tool result listing the valid names, loop continues |
| Context rot | Coherent early, incoherent deep into a long run | Capped tool output, compaction, scratchpad |
| Error cascade | One bad observation poisons every later step | Make failures explicit in the transcript; abort after N failures of one tool |
| Confident destruction | Deletes or refunds the wrong thing, cheerfully | Approval gate, dry-run mode, idempotency keys |
Most of these collapse into a single wrapper around tool execution: one object, held for the length of the run, that checks the step budget and a monotonic deadline before each step, hashes the tool name plus its sorted arguments so that a third identical call returns repeated_call instead of running again, routes tools marked as requiring approval through a prompt and returns denied_by_user when it is refused, times the call out with a message telling the model to narrow the request, and counts errors per tool so that three failures of one tool halt the run instead of poisoning the rest of it.
Two details worth calling out. Idempotency keys belong in the schema of every mutating tool — the model supplies one, your backend deduplicates on it, and a retried refund charges once. Approval gates should key on tool plus arguments, not on the tool alone: run_shell with git status is not run_shell with rm -rf, and gating the whole tool just trains people to click approve without reading.
One more thing to know before you ship: agents fail at the trajectory level, not only at the final answer. A run can land on a correct result after two destructive detours, and a pass/fail on the output alone will happily call that a success. Whatever evaluation you build has to look at the steps.
Sandboxing agents that touch real things
The moment your agent can run a shell, write files, or make outbound requests, its blast radius is your security model.
- Give the agent its own credentials, scoped down. Read-only by default, write only where it must, never your own token. If a tool can reach exactly one bucket and one schema, a compromised loop can reach exactly one bucket and one schema.
- Confine the filesystem by resolved path. String checks on the raw argument lose to
..and symlinks. - Allowlist shell commands; never denylist them. Denylists lose to shell features you didn't think of. And build argv arrays — never interpolate model output into a shell string.
- Allowlist egress. An agent with unrestricted network access is an exfiltration path and a proxy for whatever it happens to read.
- Treat every tool result as untrusted input. A fetched page, a document, an issue comment can all contain text aimed at your agent. Permissions are decided by your code, never by something the transcript asked for — no tool result should ever be able to widen what the agent may do.
- Log every tool call with its arguments. When a run goes wrong you need the exact sequence, not a summary of it.
ROOT = Path("/srv/agent/workspace").resolve()
def safe_path(candidate: str) -> Path:
# resolve() collapses '..' and follows symlinks — check after, never before.
p = (ROOT / candidate).resolve()
if not p.is_relative_to(ROOT):
raise ToolError(f"path_denied: {candidate} is outside the workspace.")
return pHuman-in-the-loop is the guardrail of last resort, and the one that actually holds. Gate anything irreversible, anything expensive, and anything a third party will see. Keep the gates rare enough that people still read them.
TL;DR checklist
- Write the deterministic version first; build an agent only when you genuinely can't
- Step budget and wall-clock deadline from the first commit, not after the first runaway
- Spend your time on tool names, descriptions and schemas — that's where the accuracy is
- One well-scoped tool beats several overlapping ones; keep the list small and non-overlapping
- Validate every tool input; return errors that name the field, the rule, and the next action
- Tool failures are observations, not exceptions — let the loop recover
- Cap tool output, compact old turns, keep durable state in a scratchpad outside the transcript
- Delegate to subagents only when the subtask eats context and returns little
- Idempotency keys on every mutating tool; approval gates on tool plus arguments
- Confine paths after resolution, allowlist shell and egress, never let a tool result grant permissions
None of this is exotic. It's the same discipline as any other distributed system that calls unreliable dependencies: bound the work, validate at the edges, make failures recoverable, and don't let a component do more than its job requires. The model is just the least predictable dependency you've ever integrated — build around it accordingly.