# How to Tell If You Need an AI Agent or a Workflow (7 Smells) [2026]

> Most “AI agents” in production are deterministic workflows wearing an LLM costume. Here’s the teardown rubric I use to simplify designs, cut cost, and make on-call sane.

- Canonical: https://www.kunalganglani.com/blog/ai-agent-workflow-smells
- Author: Kunal Ganglani
- Published: 2026-10-04 · Updated: 2026-10-04
- Category: AI and Machine Learning · Tags: ai-agents, architecture, cost-optimization, llmops, agent-orchestration

## TL;DR

How to tell if you need an AI agent or a workflow comes down to control. A workflow is a fixed set of steps you can predict and test. An agent is when the model decides what to do next, often in a loop, and that flexibility can get expensive fast. Many “agents” are really just simple scripts that call a model too often. This post gives you a quick checklist to spot fake agents, a scoring rubric to decide when a loop is justified, and a step-by-step refactor plan to collapse an agent into a smaller, cheaper workflow that’s easier to run on-call.

How to tell if you need an AI agent or a workflow is mostly about admitting an uncomfortable truth: **a lot of “agents” are just if-statements with a GPU bill.** A workflow is a bounded, predefined set of steps. An AI agent is a system where a model decides what to do next, usually in a loop, by choosing tools and updating its plan. This matters right now because 2026 tooling makes it absurdly easy to ship something that demos well, then quietly destroys your latency, reliability, and spend in production.

## How to tell if you need an AI agent or a workflow: the 30-second checklist

If you only read one section, read this. Here’s my practical rubric for **how to tell if you need an AI agent or a workflow**.

![black calculator beside black pen on white printer paper](https://cdn.sanity.io/images/vzekdneq/production/3f027bdb1b73ec6b2e38888a091749e141225422-1200x675.webp)

Build a **workflow** when:

- Your state space is bounded (you can list the states on a whiteboard in 15 minutes).
- Inputs and outputs have stable schemas (you can write a JSON Schema and keep it for months).
- Branching factor is low (typically \<= 5 distinct routes).
- Tool calls are deterministic (same input should hit the same tool 99% of the time).
- You care about P95/P99 latency and predictable failure modes.
Use an **agent loop** when:

- The task is open-world (the next step depends on information you don’t have yet).
- Tool choice is ambiguous and depends on interpretation, not rules.
- You expect surprises and want the system to adapt without you shipping code every week.
- Recovery requires reasoning over partial failures (not just retry/backoff).
Then, if you _do_ need an agent: keep the loop tiny. Budget it. Instrument it. Treat it like a distributed system.

## What is the difference between an AI agent and a workflow?

The cleanest line I’ve seen comes from Anthropic: **workflows orchestrate LLMs/tools through predefined code paths; agents dynamically direct their own process and tool usage**. That’s from [Anthropic Engineering](https://www.anthropic.com/engineering/building-effective-agents).

![person using calculator at desk with coffee mug](https://cdn.sanity.io/images/vzekdneq/production/a6ceb1854083d6a284caea25f4d731f606eb8585-1200x675.webp)

Here’s how I explain it to senior engineers without the hype tax:

- A **workflow** is a state machine with guardrails. You can read the code and know the maximum number of steps, the tools involved, and the allowed transitions.
- An **AI agent** is a controller that chooses actions. There’s still structure, but the model is deciding which branch to take, which tool to call, and when to stop.
> If you can replace the “planner” with a `switch` statement and nothing breaks, you never had an agent. You had a workflow with an expensive router.

### A quick comparison table (workflow vs agent)

| Dimension | Workflow (deterministic orchestration) | Agent (model-driven loop) |
| --- | --- | --- |
| Control flow | Predefined states/transitions | Chosen dynamically by model |
| Step count | Bounded (you can cap it tightly) | Often unbounded unless you budget it |
| Debuggability | Logs map to code paths | You need trace + “why next action” |
| Failure modes | Known, testable, repeatable | More variance, more weird edge cases |
| Cost shape | Mostly linear and predictable | Can explode via loops + retries |
| Best for | Extraction, routing, CRUD-ish ops, fixed pipelines | Open-ended tool use, exploration, ambiguous plans |

Most production systems want boundedness. Predictable step count. Predictable blast radius. The boring stuff.

## When (and when not) to use agents

I’m going to be blunt: **start with a workflow, earn your way into an agent.** That’s not a philosophy. It’s an on-call survival strategy.

![a person sitting at a desk with a calculator and a notebook](https://cdn.sanity.io/images/vzekdneq/production/dede9ea01258feebc50ff9ba8398ea717c85c9de-1200x675.webp)

Anthropic says basically the same thing: “find the simplest solution possible, and only increase complexity when needed,” because agentic systems trade latency and cost for task performance. Same source: [Anthropic Engineering](https://www.anthropic.com/engineering/building-effective-agents).

Teams ignore this because “agent” demos better than “workflow.” Then they pay twice: 1) once in API/GPU spend, 2) then again when someone has to reproduce a one-off “agent decision” from last Tuesday with zero useful traces.

### Use an agent when the branching factor is real

My favorite proxy is the **branching factor**: how many plausible next actions exist at each step.

- If your system has **2–5** routes (say “billing”, “account access”, “bug report”), you can almost always do deterministic routing with a small classifier prompt.
- If it’s routinely **10+** routes, and the right choice depends on messy semantics and partial context, an agent loop starts earning its keep.
Concrete example: “triage an incoming support ticket” is usually a workflow with some LLM extraction. “Investigate an incident across logs, metrics, and deploy history” is where agents can shine, because you don’t know the search path upfront.

### Avoid agents when you can define the world

I avoid agents when:

- the steps are known,
- the tools are stable,
- the output format must be correct,
- and the blast radius is high.
That’s anything involving money, permissions, or production writes.

If you’re building [AI agents](/pillars/ai-agents) that can mutate state, you need to be painfully honest about whether you’re buying flexibility or just injecting variance into your system.

## Building blocks, workflows, and agents (the teardown model)

Whenever someone shows me an “agent architecture,” I immediately try to reduce it to three primitives:

1. **Routing**: decide which path to take.
1. **Extraction/transform**: turn messy input into structured data.
1. **Tool execution**: call APIs, run queries, perform side effects.
In practice, most “agents” are those three things, with an LLM loop slapped in the middle because it feels modern.

My 2026-era rule: **use LLMs for the parts that are actually language problems, and use deterministic code for everything else.** Stop paying tokens to simulate `if/else`.

### The part competitors skip: deterministic gates beat “smarter” models

On this site, I run a 7-agent publishing pipeline with a deterministic SEO quality gate. After shipping 261+ posts through it, the annoying truth is: **deterministic gates before LLM review catch more issues than just upgrading the review model.** Cheap validators beat “more thoughtful” agents when what you need is correctness.

That same principle applies to agent designs. If you can validate, constrain, cap, or short-circuit, do it. Every time.

### A minimal architecture ladder

Here’s the ladder I use in planning docs because it gives teams a sane, incremental path:

- **Level 0: Pure workflow** (no LLM): rules, state machine, templates.
- **Level 1: Workflow + single LLM call**: classification or extraction.
- **Level 2: Workflow + tool calling**: model chooses arguments, not tools.
- **Level 3: Budgeted agent loop**: model chooses tools, retries, and substeps.
- **Level 4: Multi-agent**: only if you can’t solve it with Level 3.
Most orgs jump from Level 1 straight to Level 4 because “agents” sell internally. Then everyone acts surprised when the thing is expensive and flaky.

If you’re serious about production AI, climb the ladder deliberately.

## The 3 fake-agent failure modes (and what to build instead)

Dimitris Kyrkos has a teardown framing I like because it calls out the three most common fake-agents: “a model call where a regex would do,” “vector search where SQL would do,” and “an autonomous agent where a decision tree would do.” (I’m not linking DEV. Not because it’s useless. Because it’s not an authoritative domain, and I try to keep citations clean.)

Let’s turn each one into something you can use in a design review.

### Failure mode 1: a model call where a regex would do

If your “agent” is basically:

- detect intent,
- extract one ID,
- normalize a value,
- map to a known action,
you probably need:

- a parser,
- a validator,
- and a workflow.
Concrete example I see constantly: extracting invoice numbers, policy IDs, order IDs.

- Regex + checksum validation: ~0ms compute, deterministic.
- LLM extraction: **hundreds to thousands of tokens**, plus retries when parsing fails.
If you want a compromise, do **workflow + structured extraction**:

- one model call,
- output constrained to a schema,
- deterministic validation.
This is exactly what OpenAI is pushing with [OpenAI’s Structured Outputs documentation](https://platform.openai.com/docs/guides/structured-outputs): constrain responses to a JSON Schema so the “LLM part” becomes a predictable component.

If your agent loops because “sometimes the JSON is wrong,” that’s not an argument for more autonomy. It’s a sign you didn’t design the boundary.

### Failure mode 2: vector search where SQL would do

This one is expensive because it sneaks in an entire subsystem: embeddings generation, a vector database, chunking, and then a retrieval loop.

Use vector search when:

- the query is semantic (“find similar issues to this”),
- the corpus is unstructured text,
- and you can tolerate approximate results.
Use SQL (or filtered key-value lookups) when:

- you know the fields,
- you can filter deterministically,
- you need exactness (compliance, billing, auth).
A dead-simple heuristic:

- If you can write the query as `WHERE customer_id = ? AND status IN (...) AND created_at > ...` then embeddings are probably overkill.
- If your query is “what does the user mean here?” then semantics matter.
Also: vector search does not magically fix bad knowledge hygiene. I’ve written about this in the context of [RAG](/glossary/rag) and retrieval-augmented generation systems. Retrieval quality is a product problem as much as it is an infra problem. See my [RAG evaluation metrics](/blog/rag-evaluation-metrics-retrieval-quality) writeup if you want the gritty version.

### Failure mode 3: an autonomous agent where a decision tree would do

If the “agent” makes the same decisions for the same inputs most of the time, you can probably replace it with:

- a deterministic router,
- a decision tree,
- or a state machine.
This shows up constantly in “agentic customer support” and “agentic IT helpdesk” prototypes.

A strong smell is when the agent prompt includes rules like:

- “If it’s about billing, do X. If it’s about shipping, do Y.”
That’s literally a decision tree. You do not need a planner for that.

If you genuinely need flexibility, separate concerns:

- deterministic routing for known cases,
- a small agent loop for unknown cases,
- human-in-the-loop when risk is high.
## A practical decision checklist (with measurable signals)

Competitors love “principles.” Engineers need **thresholds**.

I use a scoring rubric. It’s not science. It’s just consistent. And it gives you something concrete to argue about in a design review.

Score each axis 0–2:

- **Boundedness** (0 = bounded states, 2 = open-world)
- **Schema stability** (0 = stable JSON, 2 = messy/unstructured)
- **Branching factor** (0 = \<=5 routes, 2 = \>=15 routes)
- **Tool uncertainty** (0 = one obvious tool, 2 = many plausible tools)
- **Failure recovery complexity** (0 = retry works, 2 = reasoning needed)
Total score (0–10):

- **0–3**: workflow.
- **4–6**: workflow + LLM (router/extractor/tool args) with strict validation.
- **7–10**: budgeted agent loop.
Concrete example:

- “Turn an email into a Jira ticket with correct fields”
  - boundedness 0, schema stability 0, branching 0–1, tool uncertainty 0, recovery 1. Score ~2. Workflow.
- “Investigate why a build is flaky across CI logs + recent merges”
  - boundedness 2, schema stability 2, branching 2, tool uncertainty 1–2, recovery 2. Score ~9. Agent loop.
If you want to make this painfully real, treat “agent loop” as a budget:

- max steps: **3–8**
- max tool calls: **3–10**
- max wall time: **15–60 seconds** depending on UX
If you can’t put numbers on it, you’re not designing. You’re hoping.

Here’s the official demo I like for explaining these concepts to non-specialists:

[Watch: AI Agents vs Mixture of Experts: AI Workflows Explained](https://www.youtube.com/watch?v=4-FH09AMsp0)

## Refactor an AI agent to a workflow: a migration playbook

Most teams don’t need to “kill the agent.” They need to **collapse it**.

When I refactor, I aim for this end state:

1) deterministic control flow 2) schema-constrained extraction 3) tool calls with contract tests 4) a tiny, budgeted loop only where ambiguity remains

### Step 1: Replace planner text with a state machine

Take the actions your agent can do and name them. If you can’t name them, you can’t test them.

- `CLASSIFY_INTENT`
- `EXTRACT_FIELDS`
- `FETCH_ACCOUNT`
- `CREATE_TICKET`
- `REQUEST_HUMAN_APPROVAL`
- `DONE`
This is “workflow first.” In agent terms, you’re making the action space explicit.

### Step 2: Constrain outputs using JSON Schema

If your model output feeds a tool, you want it parseable. Every time.

Use schema-constrained outputs so “formatting” stops being the reason you added more steps.

OpenAI’s guide is the clean starting point: [OpenAI (Structured Outputs)](https://platform.openai.com/docs/guides/structured-outputs).

The engineering point: treat the model as an untrusted component that must satisfy a contract. Same as any other upstream service.

### Step 3: Move “tool selection” to deterministic routing when possible

A lot of “tool selection” is fake complexity.

- If tool choice is a pure function of intent, hardcode it.
- If tool choice depends on one extracted field, route on that field.
Only let the model choose tools when:

- the mapping changes frequently,
- or the space is too messy to encode,
- or the cost of a wrong tool call is low.
If you want to go deeper here, I’ve written about protocol boundaries in [MCP vs function calling in agents](/blog/mcp-vs-function-calling-agents).

### Step 4: Keep a minimum viable agent loop (only where it earns its keep)

If you still need an agent loop after Steps 1–3, make it boring.

- **Step cap**: hard max steps (start with 5).
- **Tool budget**: max tool calls (start with 8).
- **Token budget**: max input/output tokens per step.
- **Stop conditions**: “done”, “needs human”, “budget exceeded”.
- **Reason for next action**: log a short string for why it’s choosing the next step.
This is where ReAct actually matters. The whole point is interleaving reasoning and actions so the model updates its plan based on tool results. The paper reports **absolute success rate improvements of 34% (ALFWorld) and 10% (WebShop)** with only 1–2 in-context examples. That’s from [Shunyu Yao](https://arxiv.org/abs/2210.03629).

The catch: ReAct-like loops are great when the environment is uncertain. They’re wasteful when your environment is basically CRUD.

### Step 5: Validate, test, and add fallbacks

This is the part that separates “cool demo” from “can we page this?”

I treat an agent refactor like any other reliability project:

- **Schema validation** for model outputs.
- **Tool contract tests**: if the tool returns weird stuff, the system should degrade, not spiral.
- **Golden-file evals** for prompts: store inputs, expected structured outputs, and diffs.
- **Fallback paths**: if budget exceeded, route to human or return a safe partial result.
If you want a deeper testing playbook, I’ve already built a bunch of pieces in public:

- [AI engineering evals: regression gates](/blog/ai-engineering-evals-gates)
- [How to do non-deterministic AI system testing](/blog/non-deterministic-ai-testing)
- [How to do agent tool call failure testing](/blog/agent-tool-call-failure-testing)
## Risks and limitations: the stuff that blows up at 2 a.m.

Agents fail in ways workflows usually don’t. Not because they’re evil. Because you gave a stochastic component control flow.

IBM calls out two agent risks I care about operationally: **infinite feedback loops** and **computational complexity**. They also push best practices like activity logs and human supervision. See IBM (What are AI agents?).

Here are the production-grade limitations I’d add.

### Loop amplification is the real cost killer

Agent costs don’t grow linearly. They blow up when the loop starts retrying.

- planner step says “call tool A”
- tool A fails (timeout)
- agent “reflects” and retries
- now you have another planner step, more context, more tokens
A 5-step agent with 2 retries can become **12–15 model calls** fast.

If you’re trying to reason about LLM cost, model selection matters. But loop shape matters more.

### Security blast radius expands with tool use

The moment your agent can call tools, you inherit:

- permissions,
- audit logs,
- egress controls,
- prompt injection risk.
If you’re shipping tool-using agents, you should already have a threat model for prompt injection and a plan for [AI security](/pillars/ai-security-safety). This isn’t paranoia. It’s basic engineering.

My practical recommendation: if the tool can mutate production state, require human approval. I’ve covered approval patterns in [HITL tool approval patterns](/blog/tool-approval-patterns-ai-agents) and spend controls in [AI agent kill switch spend limits](/blog/ai-agent-kill-switch-spend-limits).

## Best practices: budgets, tracing, and boring reliability

A real agent needs to be observable. Full stop.

The minimum I want:

- a trace per run,
- spans per tool call,
- token/cost per step,
- a reason string for each decision.
On this site I run weekly automated GSC feedback loops across 261+ posts. That’s convinced me that **feedback loops matter more than one-time “prompt tuning.”** You need a system that tells you when things drift, not a heroic prompt that works until it doesn’t.

For agent systems, the equivalents look like this:

- **Budget alarms**: alert when median steps goes from 3 to 6.
- **Outcome tracking**: success rate is necessary but not sufficient. Track “human escalation rate” and “tool error rate.”
- **Caching**: cache stable intermediate results. Don’t pay for the same extraction twice.
Based on the benchmark data I maintain at [kunalganglani.com/llm-benchmarks](https://www.kunalganglani.com/llm-benchmarks), throughput differences across runtimes and hardware tiers are routinely **multi-x**. And that gap shows up directly as agent latency because agents serialize work. If you’re running locally, throughput becomes your budget. If you’re paying per token, loop shape becomes your budget.

If you want a concrete framework for designing control flow, see my writeup on [AI agent control flow patterns](/blog/ai-agent-control-flow-patterns) and deeper observability in [OpenTelemetry instrumentation for AI agents](/blog/opentelemetry-ai-agents-instrumentation).

## The bet I’m making (and the challenge)

By the end of 2026, I think “agent” will stop meaning “LLM loop that calls tools” and start meaning “a workflow engine with a model-powered router and a strict budget.” The teams that win will treat agentic AI like distributed systems engineering, not prompt craft.

If you’re building an agent today, do one uncomfortable exercise: **delete the planner.** Replace it with a state machine and a schema-constrained extractor. If it still works, ship the simpler version. Keep your GPU for something that actually deserves it.

Photo by Christian Brok on Unsplash.

## FAQ

### What is the difference between an AI agent and a workflow?

A workflow follows predefined code paths. You can usually list the states, transitions, and maximum number of steps. An AI agent is when the model dynamically decides the next step and which tools to use, often inside a loop, which makes behavior less predictable.

### When should I use an AI agent?

Use an agent when the task is genuinely open-ended: tool choice is ambiguous, the next step depends on information you discover mid-run, and hardcoding the branching logic would be brittle. It’s most valuable when the environment is uncertain and the agent can learn from tool results.

### When should I avoid using an AI agent and use a deterministic workflow instead?

Avoid agents when the state space is bounded, outputs need to be reliably structured, and latency/cost must be predictable. If you can replace the “planner” with a simple decision tree or state machine without losing capability, a workflow will be cheaper and more reliable.

### How do I know if my AI agent is actually just a script?

If it makes the same choices for the same inputs most of the time, it’s probably a workflow in disguise. Another sign is prompts full of “if X then do Y” rules, which are better implemented as deterministic routing plus schema validation. Check your logs: repeated loops and retries are a tell that you’re paying for uncertainty you didn’t need.

### How do I make LLM outputs reliable and parseable?

Use schema-constrained structured outputs so the model must produce valid JSON that matches your contract. Then validate it deterministically and fail fast when it’s wrong, instead of adding more agent steps. This turns the model into a component with a clear interface, not a free-form text generator.

### What are common failure modes of AI agents in production?

The big ones are loop amplification (retries and reflection steps that silently multiply cost), tool-call brittleness (one flaky API causes spirals), and poor observability (you can’t explain why the agent acted). Agents also expand your security surface area because tool access and permissions become part of the system design.
