Ask a language model for an essay in one prompt, and it writes from the first word to the last without a backspace key. Andrew Ng uses that image to explain why single prompts top out. The fix everyone reaches for next is a loop: let the model call a tool, look at the result, and go again. That works for a demo. Then someone asks you to pause it for approval, resume it after a crash, or explain why it searched the same thing nine times, and a plain loop gives you nowhere to hook any of that in.

Graph engineering grew out of those requests. By the end of this article, you’ll know what the term means, why agent builders moved from prompts to loops to graphs, how a graph-based agent works, and where the approach costs more than it returns.

What is graph engineering?

Graph engineering is the practice of designing an AI agent system’s control flow as an explicit graph. Nodes do the work, edges decide what runs next, and a shared state object carries everything the nodes need.

  • Nodes. A node is one unit of work: a model call, a tool call, or a plain function that formats data.
  • Edges. An edge says which node runs after another. A conditional edge picks the next node at runtime, based on the current state.
  • State. The state is a single record that every node reads from and writes to, such as the message history, a draft, or a retry count.

The term is new and informal. It took off on social media in 2026, helped by a two-hour compilation video that stitches together an Andrew Ng talk and a DeepLearning.AI course on LangGraph. The ideas underneath are older and have other names: Anthropic calls fixed paths workflows, and the AlphaCodium paper calls the design work flow engineering.

The word “graph” also shows up in knowledge graphs, which store facts as entities and relationships for an agent to look up. A knowledge graph models what the agent knows. Graph engineering models what the agent does next.

An analogy: the restaurant kitchen

A busy kitchen runs on the same structure. Each station (grill, fry, plating) is a node that does one job. The ticket that travels with the order is the state; every station reads it and adds to it. The expediter at the pass is the conditional edge, deciding whether a plate goes out, back to the grill, or to the chef for a look.

The ticket rail does what graph frameworks call checkpointing: it keeps a copy of where every order stands. If the grill cook walks off mid-shift, the next cook reads the ticket and picks up where things stopped. The analogy breaks in one place: kitchen stations never decide on their own to send a dish somewhere new. In an agent graph, the model inside a node can steer which edge fires next, so two runs of the same graph can take different paths.

Why agent builders moved from prompts to loops to graphs

Each step in this progression fixed the previous step’s biggest limit and exposed a new one.

One prompt: no way to revise

A single prompt gets one pass. The model can’t search for a fact it lacks, check its own output, or fix a mistake it notices halfway through. Ng’s argument for agentic workflows is that iteration (outline, research, draft, critique, revise) beats one pass on most real tasks. In his Sequoia talk, he groups the patterns that make iteration work into four: reflection, tool use, planning, and multi-agent collaboration.

A loop: iteration without visibility

The agent loop gives the model iteration. The model picks an action, your code runs it, the result goes back into the prompt, and the cycle repeats until the model says it’s done. The ReAct paper formalized this reason-then-act cycle, and the coding agents you use daily run a version of it.

The loop’s control flow lives inside a while statement, though, and that causes trouble once the agent matters:

  • You can’t stop it at a specific step for a human to approve a risky tool call.
  • You lose everything if the process dies on step 14 of 20.
  • You can’t test the “critique” step on its own, because it only exists as a branch inside the loop.
  • You can’t see the path it took without reading logs line by line.

A graph: the loop with its arrows drawn

A graph keeps the iteration and makes the control flow a data structure. Every arrow you used to bury in if statements becomes a named edge. Because the framework now knows where one step ends and the next begins, it can save the state at each boundary, stop there, and pick up again later. Harrison Chase’s framing in the LangGraph course is that these agents are defined by cyclical graphs, and LangGraph exists to describe and orchestrate that control flow.

How a graph-based agent works

The smallest useful agent graph has two nodes and one decision. The model node asks the language model what to do. If the model requests a tool, a conditional edge routes to the action node, which runs the tool and loops back. If the model answers, the edge routes to the end.

flowchart TB START((Start)) --> LLM[Model node: ask the LLM] LLM --> DECIDE{Tool call requested?} DECIDE -- Yes --> ACTION[Action node: run the tool] ACTION --> LLM DECIDE -- No --> END((End))

Here is that graph in LangGraph, the Python library most graph-engineering material uses. The code is short because the library handles the loop; your job is to declare the pieces. If you don’t read Python, skip to the four mechanisms below; each one stands on its own.

import operator
from typing import Annotated, TypedDict

from langgraph.checkpoint.memory import MemorySaver
from langgraph.graph import END, START, StateGraph


class AgentState(TypedDict):
    # operator.add appends new messages instead of overwriting the list
    messages: Annotated[list, operator.add]


def call_model(state: AgentState) -> dict:
    reply = model.invoke(state["messages"])
    return {"messages": [reply]}


def run_tools(state: AgentState) -> dict:
    calls = state["messages"][-1].tool_calls
    return {"messages": [execute(call) for call in calls]}


def wants_tool(state: AgentState) -> bool:
    return bool(state["messages"][-1].tool_calls)


graph = StateGraph(AgentState)
graph.add_node("llm", call_model)
graph.add_node("action", run_tools)
graph.add_edge(START, "llm")
graph.add_conditional_edges("llm", wants_tool, {True: "action", False: END})
graph.add_edge("action", "llm")

agent = graph.compile(checkpointer=MemorySaver(), interrupt_before=["action"])

model and execute stand in for your model client and tool runner. Four mechanisms in that snippet do the work a hand-written loop can’t.

State has merge rules

Each node returns only the fields it changed, and the framework merges them into the shared state. The operator.add annotation tells it to append messages rather than replace them. Merge rules matter more as graphs grow: a “draft” field should be overwritten by each revision, while a “sources” field should accumulate.

Checkpoints save the state after every node

The checkpointer writes a snapshot of the state each time a node finishes, like the kitchen’s ticket rail. Those snapshots let you resume a thread after a crash and run many conversations at once, each under its own thread ID. They also let you time travel: load an earlier snapshot, edit it, and replay the graph from that point down a different branch.

Interrupts put a human between two nodes

interrupt_before=["action"] stops the graph each time the model asks for a tool, before the tool runs. A person can inspect the pending call, approve it, or rewrite its arguments in the saved state, then resume. In a plain loop, that pause point would be a hand-rolled input prompt that dies with the process.

Cycles are allowed and bounded

The edge from action back to llm is a cycle, which is what makes this an agent instead of a pipeline. Frameworks enforce a step limit (LangGraph calls it a recursion limit) so a confused model can’t loop forever on your credit card.

Common graph shapes

Most production agent graphs combine a handful of recurring shapes. Anthropic’s “Building Effective Agents” catalogs the fixed-path ones, and the LangGraph course covers the more autonomous ones.

Prompt chaining A straight line of nodes, each refining the last output. Add a check node between steps to stop bad drafts early.
Routing A classifier node sends each input down one of several specialized branches, such as billing, refunds, or technical support.
Parallelization Several nodes run at once on the same input, then a join node merges or votes on their results.
Evaluator and optimizer A generator drafts, a critic grades, and a conditional edge loops back until the critic passes it or the step budget runs out.
Plan and execute A planner writes a task list, worker nodes execute it, and a check node decides whether to replan or finish.
Supervisor One model routes work to sub-agents, each of which can be its own graph with private state.

The course also points past these shapes to tree search, where an agent tries several actions, reflects on each, and follows the most promising branch. The Language Agent Tree Search paper combines planning, acting, and reflection that way, at a higher cost in model calls.

A worked example: the essay writer

The LangGraph course closes with an essay writer that layers several of these shapes. A planner outlines the essay. A research node searches the web for the outline’s claims. A generator writes a draft. A conditional edge then checks a revision counter: under the limit, the draft goes to a reflection node for critique, then to a second research node that finds evidence for the critique, then back to the generator.

flowchart TB PLAN[Planner: write an outline] --> RESEARCH[Research the plan] RESEARCH --> GEN[Generate a draft] GEN --> CHECK{Revisions left?} CHECK -- Yes --> REFLECT[Reflect: critique the draft] REFLECT --> CRITIQUE[Research the critique] CRITIQUE --> GEN CHECK -- No --> DONE((Final essay))

This is what the viral video means by agents that “rewrite themselves.” The agent rewrites its draft, many times, under a critic’s direction, the draft, critique, and revise loop that the Self-Refine paper showed improves output. Its code and its graph stay the same.

How graph engineering connects to other agent practices

Graph engineering decides the order of calls. Neighboring practices decide what happens inside each one. Context engineering decides what goes into each model call, and the graph’s state is where that context lives between calls. Tool design decides what the action node can do; protocols such as the Model Context Protocol (MCP) and packaged agent skills plug in there.

Evaluation gets easier with a graph, because you can feed the critique node fixed inputs and grade it alone, the way AI evals test any model feature. Knowledge graphs and graph databases can serve as long-term memory that a research node queries. The fundamentals of graph databases apply there, while the control-flow graph stays in your application code. At the other end of the spectrum, coding agents are mostly agent loops with a large toolset, the structure that graphs extend.

Trade-offs and limits

Graphs give you control at the cost of extra structure. Beware of these potential failures before committing to one.

  • Up-front design cost. You must decide the nodes and edges before you know which ones the task needs. Anthropic’s advice is to start with the simplest solution, often one well-crafted prompt, and add steps only when measurements show they help.
  • Rigidity. A fixed graph handles the paths you drew. Inputs that need a path you didn’t draw get forced down the nearest wrong one. Fully autonomous loops handle novel tasks better, at the cost of predictability.
  • Cost and latency multiply. Every reflection cycle adds model calls. A four-revision essay graph makes about 15 calls where a single prompt makes one, and its latency grows the same way.
  • Framework coupling. Graph libraries change fast, and their state, checkpoint, and interrupt APIs leak into your code. Keep node functions plain so you could move them to another runner.
  • Debugging still needs traces. A graph shows which path ran. It can’t tell you why the model in a node chose badly; that still takes traces and evals.

Common misconceptions

  • “You need LangGraph to do graph engineering.” The graph is a design. A dictionary of functions, a state object, and a routing function in a loop is a graph. Libraries add checkpoints, interrupts, and visualization, which you’d otherwise build yourself.
  • “More agents means better results.” Headlines about going from one prompt to 100 agents sell scale. Each added node adds cost and latency. Add one when an eval shows it improves results.
  • “Graph engineering means building knowledge graphs.” The two share a word and nothing else. An agent can use both, or neither.
  • “Self-rewriting agents modify their own code.” In this material, rewriting means a reflection loop revising output. Systems that rewrite their own prompts or graphs exist in research, but they’re a separate and much riskier idea.
  • “A graph makes the agent reliable.” A graph makes the agent’s possible paths explicit and bounded. Each node’s output is still a model’s output and still needs evaluation.

Conclusion

Graph engineering takes the loop at the heart of every agent and draws it as a map. Once the control flow is data instead of branches buried in a while statement, the framework can save it, pause it, replay it, and show it to you.

The mental model to keep is a progression. A prompt gives one pass, a loop gives iteration, and a graph gives iteration you can inspect and steer. Reach for a graph when you need checkpoints, human approval, or testable steps, and stay with the simpler options when you don’t.

Next steps

References