Prompt, Context, Loop, Harness: Four Layers, Four Different Failures
Prompt engineering. Context engineering. Loop engineering. Harness engineering. They're not interchangeable. They're four layers of the same system, and each one answers a different question about how you get useful, reliable behavior out of a model.
The 30-second version
Next time something goes wrong with an AI feature, ask which of these it was:
- It was worded badly or misunderstood the ask. That's a prompt problem.
- It didn't know something it should have known. That's a context problem.
- It needed several steps and had to decide on its own when to stop. That's a loop problem.
- The steps were designed fine but execution broke. A tool failed silently, a retry never fired, nothing caught the error. That's a harness problem.
If you remember nothing else: wording is the prompt, knowledge is the context, steps and stopping are the loop, plumbing is the harness.

| Layer | The question it answers | Typical failure |
|---|---|---|
| Prompt | How do I phrase this one request well? | Vague or misread answer to a single message |
| Context | What does the model actually have in front of it right now? | Confident, fluent, and wrong because the data was missing, stale, or buried |
| Loop | What should happen across steps, and when is it done? | Spins in circles, never finishes, or declares victory too early |
| Harness | What machinery makes each step run reliably? | A tool fails silently, a retry never fires, nobody notices |
Prompt engineering: the wording of one exchange
This is the one everyone already half-understands, so I'll keep it short. Prompt engineering is optimizing a single message to a model: how you phrase the instruction, what examples you give, whether you ask for step-by-step reasoning, what format you demand back. It lives entirely inside one turn.
Think of a support chatbot. Someone asks a question, the model answers, and a person reads the response and decides what happens next. The team behind it spends its time refining the system prompt and few-shot examples until the answers are consistently good. That's prompt engineering doing its job. It's about the quality of a single response, not what happens after.
Context engineering: what the model can actually see
This is where a lot of the confusion starts, because context engineering looks like prompt engineering from the outside. You're still assembling something and sending it to the model. But the question is different. Prompt engineering asks "how do I phrase this instruction well." Context engineering asks "what information does the model have access to right now, and is it the right information."
In practice that means curating everything that fills the context window: retrieved documents, conversation history, tool definitions, memory files, results from other tools. A RAG pipeline that pulls in only the three relevant paragraphs instead of dumping a whole wiki into the prompt is context engineering. The failure it guards against isn't badly worded instructions. It's a model that's articulate, well-instructed, and confidently wrong because it's reasoning over the wrong information, or too much of it. Anthropic puts it well: context is a finite resource, and the job is deciding what earns a spot in it on every turn.
Loop engineering: designing the system, not the sentence
This is the one that sparked the "wait, what does that even mean" reaction in the first place. As I use the term, loop engineering is designing a workflow where the model is called repeatedly, checks its own progress, and decides what to do next without a person in between each step. (The term is still new and people draw its edges differently, so treat this as my working definition, not an official one.)
A loop needs a few things a single prompt doesn't: a trigger (what kicks it off), a goal (what "done" means), a set of available actions, a way to verify an action actually worked, and somewhere to keep memory across steps. A concrete example is a competitor-monitoring agent that runs every Monday morning. It scrapes a handful of sites, hands the content to a research step, checks the findings against relevance criteria, retries with a broader search if results are thin, trims the summary to a word limit, and emails the briefing. Nobody touches it.
Here's how I think about the difference from prompting: a prompt is one brick, a loop is the floor plan. You can have beautifully worded prompts inside a loop that's structurally unsound, with no retry logic, no way to tell success from failure, and no exit condition, and the whole thing still falls over. The skill is less about phrasing and more about systems design. What happens when a step fails? How does the loop know it's finished? Where does it stop itself before burning your token budget going in circles?
Harness engineering: the scaffolding that keeps the loop honest
This is the newest layer, and the one people conflate with loop engineering most, since both are about the system around the model rather than any single message. My distinction: loop engineering designs the workflow, meaning which steps happen, in what order, toward what goal. Harness engineering builds and hardens the machinery that executes it: tool calling, retries, timeouts, checkpoints, error recovery, sandboxing, logging.
LangChain gives a compact formula for the whole idea: agent = model + harness. The harness is every piece of code, configuration, and execution logic that isn't the model itself. Birgitta Böckeler, writing on Martin Fowler's site, breaks this down for coding agents: a harness works through "guides" (feedforward controls that steer the agent before it acts) and "sensors" (feedback controls that observe after it acts and help it self-correct, like a test suite or a linter). None of that is about wording a prompt better. It's infrastructure.
One honest caveat: the boundary between loop and harness is blurry. LangChain's own write-up treats loop patterns as part of the harness, and Böckeler describes building a harness for a coding agent as a specific form of context engineering. So read my four layers as a way to ask better questions, not as four boxes with hard walls.
Why the layers get tangled
These layers fail differently, and treating a harness problem like a prompt problem wastes real time. If a careful prompt still returns a stale number, it may not be a prompt failure at all. It could be context (the model never had the current data), a loop (nothing re-checked the number before sending it), or a harness (the retrieval tool silently failed and nobody caught it). "Just improve the prompt" is the reflexive fix for all of it, and it's often the wrong one.
Two adjacent terms get tangled into the same conversations. "Agentic" usually describes behavior, a system that plans, acts, and adapts with some autonomy, not a specific technique. So "agentic" and "loop engineering" aren't synonyms: agentic is the what, loop engineering is one way to build the how. "Orchestration" often gets used almost interchangeably with harness engineering in vendor content, but it's really the narrower piece that coordinates multiple steps or multiple agents, not the full stack of retries, sandboxing, and monitoring.
None of these terms are precise the way big-O notation is precise. The industry is still arguing about the boundaries, and you'll see the lines drawn differently depending on whose blog you're reading. But the four-layer shape is a useful mental model. Next time someone tells you they "don't prompt anymore, they just run loops," you'll know which layer they've moved up to, and which question to ask about the layers they didn't mention.
-amrita
Read later
Prompt vs. context
- Effective context engineering for AI agents (Anthropic): context as a finite resource, and how to decide what goes in on each turn.
- Prompts vs. Context (Drew Breunig): the moment "context engineering" took off, including Andrej Karpathy backing it over "prompt engineering."
Loops
- Building effective agents (Anthropic): workflows (predefined code paths) versus agents (the model directs its own process), and when each fits.
- everything is a ralph loop (Geoffrey Huntley): the "program the loop" mindset, from the person behind the Ralph Wiggum loop pattern.
Harnesses
- The Anatomy of an Agent Harness (LangChain): "agent = model + harness" and the parts a harness needs.
- Harness engineering for coding agent users (Birgitta Böckeler, martinfowler.com): guides, sensors, and building an outer harness for a coding agent.
- Harness engineering: leveraging Codex in an agent-first world (OpenAI): a team shipping a product with no hand-written code, and the environments and feedback loops that made it work.
- Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents - A Source-Code Study of Eleven Systems (Barbaste et al.): a source-code study of eleven coding-agent harnesses, covering their architecture, recurring design patterns, and how they evolved.
Member discussion