A team at a 30-person fintech startup wires up a LangChain agent to answer support tickets: check account balance, look up a transaction, escalate if fraud-flagged. Day two in production, one agent run calls the “check balance” tool 41 times in a row before the request times out at 60 seconds. Nobody touched the code. The model just got stuck restating the same thought and re-invoking the same tool, and the framework had no built-in ceiling to stop it. That’s not a bug report you’ll find in the docs. It’s the default behavior of an orchestration loop with no guardrails, and it’s why understanding what’s actually happening under
AgentExecutor.invoke()matters more than knowing the API surface.
The Core Mechanism
Strip away the abstractions and both frameworks are doing the same basic thing: turning a sequence of LLM calls, tool calls, and data transformations into a graph, then walking that graph. LangChain’s LCEL (LangChain Expression Language) builds this as a static DAG — you pipe a prompt into a model into an output parser using the | operator, and the shape of that graph is fixed at definition time. A chain built this way runs the same three or four nodes every time, in the same order, regardless of input. That’s a chain, and it’s the easy 80% case.
Agents are different, and this is where most of the confusion (and most of the production incidents) live. An agent doesn’t have a fixed graph. Instead, it runs a loop: the LLM receives the user’s goal plus a scratchpad of everything done so far, decides what to do next, and the framework parses that decision, executes it, appends the result to the scratchpad, and calls the LLM again. This is the ReAct pattern (Reason + Act) that both LangChain’s AgentExecutor and LlamaIndex’s ReActAgent implement almost identically under different class names. The graph is decided at runtime, one node at a time, by the model itself.
LlamaIndex frames this slightly differently at the retrieval layer. Its QueryEngine abstraction splits the pipeline into a retriever (fetch candidate nodes from a vector index), a node postprocessor chain (rerank, filter by similarity threshold, deduplicate), and a response synthesizer (combine retrieved chunks into a final answer, using strategies like compact — stuff everything into one prompt — or tree_summarize — recursively summarize in a binary-tree pattern when context won’t fit). Each stage is swappable, which is the actual value proposition: you’re not buying “RAG,” you’re buying a slot machine of interchangeable retrieval and synthesis strategies you can A/B test without rewriting glue code.
The part vendors don’t advertise: the output parser sitting between “LLM emits text” and “framework executes an action” is doing enormous, fragile work. The LLM outputs a string. Something has to turn Action: search_transactions\nAction Input: {"account_id": "4471"} into an actual Python function call with actual arguments. That parser is a regex or a JSON-mode contract, and it breaks constantly — trailing commas, an extra sentence before the JSON blob, a tool name the model slightly misremembers. When it breaks, the default behavior in most setups isn’t to fail loudly. It’s to catch the exception, shove “Invalid format, try again” back into the scratchpad, and loop. That’s your 41-call incident.


