Part 2 of the series “Beyond Vibe Coding: The Engineering Blueprint for Reliable AI Agents”
🗺️ The Series Roadmap: Beyond Vibe Coding
In 2026, software development is drowning in fragmented buzzwords: Harness Engineering, Loop Engineering, Context Engineering, Spec-Driven Development, Evaluators, and Sandboxes. Are they competing frameworks? Redundant prompt techniques? Or something fundamentally new?
This three-part series unifies these concepts into a single, cohesive architectural blueprint for building software with AI without writing code manually:
- Part 1: From Vibing Chaos to Reliable Agents: The What, Why, and How of Harness Engineering
We diagnose the hangover of vibe coding, unpack the paradigm shift of shifting developer interventions left, establish our bedrock formula (Agent = Model + Harness), explore the Lewis Hamilton motorsport analogy, and deconstruct the deterministic components that make up a production developer environment. - Part 2: Loop Engineering in the Harness: Autonomous, Self-Correcting Systems (This Article)
We fire up the dynamic runtime engine of the harness: deconstructing the core loop components, establishing hybrid evaluation rubrics (deterministic tests + LLM goal fidelity + tool trajectory verification), and applying the Ratchet Principle. - Part 3: Spec-Driven Development For-the-Win: Disposable Code, Living Blueprints, and Building Like a Boss (Coming Soon)
We explore Spec-Driven Development (where code is disposable and specifications are the single source of truth), master Google Antigravity’s unified harness and hybrid slash-command workflows, and deploy enterprise-ready agents to Google Cloud.
The “Generate and Fix” Trap: Stop Being the Human Correction Loop
Howdy friends, and welcome back!
In Part 1, we established the bedrock formula of modern agentic engineering:
Agent = Model + Harness
We deconstructed the core deterministic components of the harness: hierarchical context rules (AGENTS.md / GEMINI.md), living version-controlled state, curated Model Context Protocol (MCP) tools, on-demand skills, sandboxing, and observability. This is how we engineered the race car chassis.
But wait… How does the work actually get done?
In simple, naive AI interactions, developers issue a prompt, get some output, and then typically course-correct. Maybe the developer copies the output from the terminal into the chat and asks for a fix.
Notice what is happening here: we have an iterative cycle of action, verification, and course-correction. You cannot escape it. But in an ad-hoc vibe coding workflow, you are the loop. You act as an exhausted organic clipboard, forcing yourself to micromanage, inspect, and babysit every single turn. If you’re doing that, you haven’t automated much; you’re just paying a subscription to watch a terminal like a nervous ferret.

Loop Engineering automates that execution cycle.
Instead of forcing a human to watch and correct every turn, the loop enables our agent to cycle through action, verification, and course-correction autonomously within the harness. The agent executes the code, runs the automated test suites, evaluates its work against strict rubrics, refines its hypotheses, and makes corrections… On its own!
Crucially, this does not mean ejecting the human entirely. It means moving from continuous babysitting to strategic governance: Human-in-the-Loop does NOT mean Human-in-Every-Turn. The agent executes the rapid, deterministic feedback cycle autonomously, and only pauses to escalate to the human conductor at predefined escalation points — such as hitting persistent evaluation plateaus, encountering ambiguous requirements, or proposing irreversible, high-risk actions.

A Crucial Distinction: The Development Workbench, Not Runtime Infrastructure
When we talk about loop engineering, the goal is not usually to ship autonomous, self-modifying while-loops into a live, customer-facing production application. The value of harness engineering and loop engineering is that it provides a disciplined feedback mechanism during the development phase to build software that is production-grade. The loop is the automated forge where code is hammered, tested, and tempered; what you actually deploy to the cloud is resilient, production-ready software.
In tokenomics terms, this aligns to so-called Spec Spend (as explored in my upcoming enterprise whitepaper, The Agentic Revolution). Burning tokens in the developer harness to explore hypotheses, run tests, and self-correct is a high-value, high-leverage investment. Spending a couple of dollars in developer Spec Spend to forge a hardened application prevents catastrophic regressions and uncontrollable costs in production Run Spend.
Core Components of an Engineered Developer Loop
To move beyond fragile trial-and-error, an engineered agentic loop requires distinct, interconnected components:

1 — Goal (The Anchor): A clear, unambiguous statement of what the agent must accomplish. In long-running autonomous runs, context compression can cause the model to drift. A resilient harness restates the core goal on every single iteration, keeping the agent firmly anchored to its original objective.
2 — Context (The Environment & Grounding Data): Everything the agent knows about the world before taking an action. This context extends far beyond prompt templates; it is the total information environment available to the loop:
- Standing Architectural Rules & Guardrails: Hierarchical instructions, coding standards, and operational boundaries drawn from GEMINI.md, AGENTS.md, and activated Agent Skills.
- Workspace Data & Local Fixtures: Source code, schema files, seed datasets, sample payloads, and test mocks residing directly in the repository.
- Dynamic Knowledge & Database State: Grounding information retrieved on demand via MCP tools or APIs, such as querying living documentation from (say) Confluence or Google Drive, inspecting issues on GitHub, or introspecting schemas and retrieving data from databases.
💡 Decoupling Engineering Guardrails from Application Specs
A crucial architectural discipline in production harnesses is separating the Application Goal Specification from the Engineering Discipline. Naive prompt engineering often conflates the two, burying feature requirements beneath paragraphs of repetitive coding hygiene instructions. In reality, engineering guardrails (e.g. Python standards, mandating TDD approaches, Pydantic input parameterisation) should be independent of any single application spec. They represent reusable, global, organisation-, domain- or team-level engineering policies.
Centralising them in persistent workspace configurations (GEMINI.md, AGENTS.md) and reusable Agent Skills means you define them once and govern them centrally across all projects. The feature specification stays lean and domain-focused.
3 — Tools (The Levers): The restricted set of actions the agent is permitted to execute, including sandboxed shell commands and MCP server tools. The harness strictly limits tool access to reduce blast radius and limit tool choice paralysis.
4 — Evals (The Objective Judge): The scoring mechanisms that evaluate the progress of each iteration. Without automated evals, the agent is flying blind. A robust harness evaluates across three dimensions:
- Deterministic Output Checks: Mechanical unit tests, integration suites, and static linters (pytest, ruff, codespell) that verify the compilation, syntax, and functional correctness of the generated artifacts.
- Tool Trajectory & Process Verification: Inspecting how the agent reached the result. Did it invoke only approved tools in the expected sequence? Did it respect tool schemas without hallucinating arguments? Did it avoid thrashing loops or calling forbidden APIs?
- Semantic LLM-as-a-Judge Analyses: Using models to evaluate the qualitative requirements — crucially including Goal Fidelity (verifying whether the solution actually delivers what was requested) and architectural cleanliness. While semantic autoraters can use rubric scales or qualitative feedback, framing them as binary yes/no assertions is often a very effective approach to eliminate subjective score drift.
5 — Memory (Loop History & Scratchpad): The structured log of what has been attempted across preceding iterations. The agent records what worked, what failed, and the progression of evaluation scores over time. In a well-engineered harness, this persistent scratchpad prevents the model from repeating known dead-ends or thrashing in circular loops.
6 — Iteration Discipline (The One-Change Rule): This is about applying the scientific method to autonomous iteration. This discipline applies specifically to changes evaluated for semantic value. When iterating against semantic autoraters, the loop instructs the agent to modify one thing at a time before re-scoring. If multiple semantic changes are introduced at once, causal observability collapses: if a rubric score drops, the agent cannot tell which alteration caused the regression. Conversely, purely mechanical deterministic fixes (resolving compiler syntax, fixing lint warnings from ruff or codespell, and formatting) can and should be batch-remediated in a single turn.
7 — Escalation Protocol (The Safety Valve): Explicit operational triggers that halt the loop and request human intervention. For example, if an agent encounters conflicting specifications, hits three consecutive evaluation failures, or proposes an irreversible action (e.g. deleting cloud infrastructure), it must not hallucinate a guess. It must package its findings, present a reasoned recommendation, and yield control to the human conductor.
8 — Stopping Condition (The Finish Line): Clear, programmatic exit criteria that tell the loop when execution is complete. Such as:
- 100% of deterministic checks and semantic rubric assertions passing with zero lint warnings.
- Target state verification: the critique model explicitly certifies that the initial goal has been fully met.
- A convergence plateau where successive iterations produce negligible score improvements.
- Hard circuit breakers, such as reaching a maximum iteration count (e.g. 5 loops) or a predefined token spend limit.
Is It a Loop Bug or a Harness Bug? When an autonomous agent goes off the rails during development, we often waste days trying to fix the wrong things.
Use this simple litmus test:
- If you can resolve the issue by refining the prompt, sharpening the goal, updating the rubric, or adjusting the stopping condition → It is a Loop bug.
- If resolving the issue requires changing tool definitions, adjusting sandbox permissions, modifying execution timeouts, or intercepting runtime errors → It is a Harness bug.
Don’t attempt to solve a harness problem by writing a more pleading prompt. Fix the layer that actually broke!

The Ratchet Principle: Engineering Continuous Improvement
What happens once you diagnose an edge case or failure that wasn’t caught by your initial harness?
This brings us to the operational crown jewel of harness engineering: The Ratchet Principle.

When an AI agent stumbles, makes an architectural mistake, or introduces a subtle bug, the amateur response is to re-prompt the model in the chat: “No, don’t do that, use the new API method instead.”
That is an utter waste of a lesson. The moment that conversation ends, that context evaporates into thin air. In tomorrow’s session, on a new branch, or when a teammate runs a similar task, the agent will walk directly into the exact same trap.
Under the Ratchet Principle:
Whenever an agent makes a mistake, you update the harness, not just the prompt.
- Did the agent use an outdated API parameter? Update the workspace GEMINI.md or add an authoritative type schema.
- Did the agent forget to handle a specific error code? Add a permanent check to the binary evaluation rubric.
- Did the agent create an unformatted commit? Add a pre-commit hook to your Git harness.
- Did the agent struggle with a Google Cloud product? Load the corresponding skill from the google-cloud-developer plugin.
Like a mechanical ratchet, every single failure permanently tightens the system. The harness clicks forward one notch. Once a rule or rubric check is committed to your repository, that specific failure mode is eliminated forever. Over time, your development harness becomes an impenetrable repository of institutional engineering wisdom.
Why Loops Work: The Power of Feedback
Why go through the effort of building a multi-component loop instead of relying on a single, well-crafted prompt?
I created an application that demonstrates building an application agentically, and compares with and without harness and loop, side-by-side.

We’ll deep-dive into this shortly. But for now, I’ll give you the key findings from this application, which align with industry findings. Essentially, if you have a harness with a loop, these are the outcomes:
- You use fewer tokens.
- The overall time to get to a solution is dramatically reduced.
- The human input required is massively reduced.
- The overall cost to produce the solution is significantly reduced.
- The quality of your solution, and its alignment to your requirements, is much better.
- Most importantly — the solution is more likely to actually work.
TL;DR — single-pass prompting is gambling that an LLM will get dozens of interconnected variables right on the first try.
In contrast, wrapping the model in an automated loop mirrors the exact rhythm experienced software engineers use every day: propose a targeted change, run the tests and evals, feed the diagnostics straight back into the next turn, and iterate until the solution passes. The loop keeps the model honest, converges on working code dramatically faster, and stops intermediate errors from compounding into an unrecoverable mess.
Three Traps to Avoid
However, naive loops without rigorous harness constraints may fall prey to some common loop traps:
1 — The Premature Exit:
The agent runs superficial tests, assumes everything works because there were no syntax errors, and prematurely declares victory.
Neutralised by: Mandatory TDD gates and non-negotiable passing thresholds on automated test suites.
2 — The Sycophantic Yes-Agent
We all know that LLMs can be hopelessly sycophantic.
Q: “Do I look good in this shell suit?”
A: “Why yes, Dazbo! You look so hunky wearing that.”

When using an LLM to rate the output of a loop, the model easily falls into the same trap — uncritically praising the worker agent’s output.
Neutralised by: Structuring evaluations as strict, binary assertions and verifiable code artifacts.
3 — The Bottomless Cash Pit
An unachievable rubric threshold or a brittle stopping condition traps the agent in a recursive loop, repeatedly modifying the same file and burning through hundred-dollar token budgets without making forward progress. While developer Spec Spend is meant to be exploratory, an unconstrained loop turns healthy investment into an uncontrolled financial bonfire.

Neutralised by: Stall detection, plateau convergence checks, and hard iteration circuit breakers.
The Hybrid Evaluation Rubric (Deterministic + Semantic)
One of the biggest mistakes teams make when building agent loops is relying on fluffy, subjective evaluation prompts: “Rate the quality of this code on a scale of 1 to 5.”
Models are notoriously inconsistent when asked for qualitative ratings; what scores a 4 on Tuesday might score a 2 on Wednesday. In contrast, binary rubrics convert these fluffy judgments into repeatable deterministic results. So let’s just turn the evaluations into verifiable yes/no questions!
Furthermore, as Google AI leaders like Shubham Saboo have pointed out in Finally, Loop Engineering Explained when contrasting loops with traditional prompt engineering, reliable loops must automate what is checkable — and that begins with testing whether the generated output actually achieves the goal that was set.
The goal is not just passing unit tests and being clean of linting errors. We also need to be sure that the solution actually delivers the feature requested.
For this reason, our evaluation rubric should combine deterministic checks with semantic checks. The semantic checks are executed by a critique agent or LLM-as-a-judge autorater.
And we must evaluate tool trajectory alongside code execution. You cannot evaluate only the final artifact; the harness must verify how the agent reached the result. Tool trajectory evaluation operates across two complementary layers:
- Deterministic Trajectory Evaluation (Mechanical Trace Inspection): Code-level validation inspecting the execution trace without using an LLM. Did the agent invoke only authorised tools from its allowlist? Did arguments strictly satisfy JSON parameter schemas? Did it follow mandatory prerequisite ordering (e.g. running the test suite before attempting to commit code)? Did it stay within maximum tool call budgets rather than entering an unconstrained loop?
- Semantic Trajectory Evaluation (Reasoning & Efficiency): Evaluated by a critique agent or LLM judge. Did the agent select the most appropriate and efficient tool for the problem, or did it take an absurdly convoluted detour (e.g. running fifteen individual file-read calls instead of a single targeted grep)? When a tool returned an unexpected error or empty result, did the agent reason soundly about the failure and adapt its strategy, or did it flail repetitively?
An agent that arrives at a working file after making ten redundant, unapproved tool calls or taking dangerous exploratory detours is an operational failure.

The Anatomy of a Production-Grade Evaluation Rubric
Here is an simple example; a 10-check rubric illustrating how deterministic and semantic checks — across both code and tool trajectories — work together:
# Evaluation Rubric
## Part A: Deterministic & Static Checks (Automated via Test Runners, Linters & Trace Inspectors)
| ID | Check (1 = True, 0 = False) | Evaluation Method | Result |
|:---|:---|:---|:---:|
| 1.1 | All unit tests in `tests` pass cleanly? | `pytest` runner | [ ] |
| 1.2 | `uvx ruff check .` and `uvx codespell` report zero errors or warnings? | Static analysis CLI | [ ] |
| 1.3 | Function signatures use PEP 585 standard type hints (e.g. `list[str]`, `dict[str, Any]`)? | Linter | [ ] |
| 1.4 | All database queries use parameterised inputs (no raw f-strings or string concatenation)? | Static analysis / AST | [ ] |
| 1.5 | **Deterministic Trajectory Conformance:** Did the agent invoke only authorised tools, satisfy parameter schemas, respect tool call limits, and follow mandatory prerequisite ordering (e.g. test before commit)? | Harness Trace Inspector (Code) | [ ] |
## Part B: Semantic Yes/No Analyses (Evaluated via LLM-as-a-Judge / Critique Agent)
| ID | Binary Semantic Assertion (1 = True, 0 = False) | Evaluation Method | Result |
|:---|:---|:---|:---:|
| 2.1 | **Target State Completion (Goal Fidelity):** Does the generated solution demonstrably deliver all functional capabilities and explicit deliverables declared in the goal statement, without omitting mandatory acceptance criteria? | Semantic Judge | [ ] |
| 2.2 | **Scope Discipline (Anti-Hallucination):** Does the solution achieve the goal without generating unrequested speculative abstractions, hallucinating unused helper modules, or modifying files outside the agreed boundary? | Semantic Judge | [ ] |
| 2.3 | **Trajectory Efficiency & Tool Reasoning:** Did the agent choose sensible, direct tool pathways to solve the problem rather than taking convoluted detours or repeatedly retrying failed tool strategies without adaptation? | Semantic Judge | [ ] |
| 2.4 | **Architectural Integrity & Interface Boundaries:** Does the code maintain clean boundaries, preventing internal database implementation details or raw connection strings from leaking into public interfaces or caller errors? | Semantic Judge | [ ] |
| 2.5 | **Design Rationale & Intent:** Does the module-level docstring clearly explain the architectural *why* (why this pattern was chosen over alternatives) rather than merely restating what the code does? | Semantic Judge | [ ] |
**Passing Criteria:** 10 / 10 (100% pass required before code can be committed).
**Iteration Policy:** Failures on code, test suites, or sub-optimal tool trajectories feed diagnostic feedback into the next turn for autonomous remediation.
**Hard Stop (Human Escalation):** The harness only aborts and alerts the developer if the agent breaches security boundaries (e.g. unapproved destructive commands), enters an unrecoverable tool thrashing loop, or reaches the maximum iteration limit (5 loops).
Notice how this structure enforces discipline across every dimension:
- Checks 1.1–1.5 verify mechanical hygiene (syntax, compilation, test passage, and deterministic tool trajectory allowlists/orderings).
- Checks 2.1–2.2 verify Goal Fidelity — comparing the candidate output directly against the initial user intent. If the agent forgot batch processing or wandered off into speculative refactoring, Check 2.1 or 2.2 drops to 0, and the loop refuses to exit.
- Check 2.3 verifies semantic tool trajectory efficiency — ensuring the agent didn’t burn tokens on irrational intermediate actions.
- Checks 2.4–2.5 verify semantic architecture, error safety, and living documentation quality.
By presenting the agent with this exact markdown rubric, the model can self-score its work after every turn. It doesn’t wonder whether its output is “good enough”; it knows with precision whether it satisfied the criteria and achieved the goal.
No Arbitrary Point Ceiling
While our example shows 10 checks for clear illustration, let’s avoid any misconceptions: an evaluation rubric does not have an arbitrary 10-point ceiling. In practice, your rubric can have whatever point total you want. For example, you might also include:
- Frontend contract checks: Verifying TypeScript interface alignment, component contracts, and state store mutations.
- API wiring verification: Confirming route registration, middleware attachment, payload serialisation, and HTTP status code mappings.
- Latency bounds: Asserting that mock or integration calls stay strictly within defined latency budgets (e.g. 95% in under 250ms).
- Security invariants: Enforcing parameterised queries, sanitised inputs, absence of leaked credentials or stack traces, and zero-trust egress policies.
Side-by-Side Demo: The Harness Engineering Workbench
Theory is cheap. Let’s prove it!
To demonstrate the quantifiable difference between raw vibe coding and disciplined loop engineering, I built an open-source, interactive, split-screen evaluation workbench: the Harness Engineering Demo.

The workbench pits two autonomous tracks against each other in real time, each running concurrently with the same model. (The demo uses Gemini 3.8-Flash out-of-the-box, but this is configurable.)
- The Unharnessed Track — Vibe Coding: A raw single prompt containing the user specification is dispatched directly to the model. We supply no additional workspace rules or guardrails, and no specialised skills. There is no autonomous loop and no iterative self-healing. We evaluate the output once, at the very end.
- The Harnessed Track — Loop Engineering: An autonomous multi-turn loop executes in an isolated workspace. We provide the same specification prompt as before. But this time we also provide additional context and guardrails, we equip with a few relevant skills, and we evaluate each iteration with a 14-point rubric.
What is the demo actually building? Well, the specification is for the non-trivial task of building a “Cosmic Trivia & Strategy Conquest Game”, a full-stack sci-fi game featuring a FastAPI backend, a directed graph of galactic sectors, turn-based combat, dynamic AI trivia generation, and an interactive browser-based star map with real-time HUD metrics and combat dialogues.
The Harness: Context, Skills, and Living Memory
Let’s start by looking at the supporting context we supply. This would typically be part of your global reusable context (e.g. AGENTS.md or GEMINI.md) that you use as an engineer, as a team, or even as an organisation. For example:
# Engineering Context & Harness Guardrails
## Persona & Role
You are a Principal Software Engineer and Systems Architect. You produce modular,
production-grade, and self-documenting code following strict engineering standards.
You practice test-driven development, enforce strict interface contracts,
and design systems with clear separation of concerns.
## 1. Environment & Coding Standards
- **Python Version**: Python 3.13 standard.
- **Type Hinting**: Strict PEP 585 (use built-in `list[str]`, `dict[str, Any]`, `tuple[int, ...]`,
never use deprecated `typing.List` or `typing.Dict`).
- **Scope Discipline**: Only create required files specified in the task goal or architecture.
Do not generate speculative helper files, hallucinated third-party dependencies, or dead code.
## 2. Test-Driven Development (TDD)
- **Mandatory Test Suite**: Author a comprehensive unit test suite in `tests/` before or
alongside implementation covering core domain logic, edge cases, and state transitions.
- **Verification Gate**: Verify that all unit tests pass cleanly via `pytest` before
declaring any task complete.
## 3. Backend Architecture & Security Guardrails
- **Parameterisation**: All API request bodies and input parameters must use
structured Pydantic models with input validation. Never accept raw unvalidated
dictionaries or execute dynamic evaluation (`eval`, `exec`).
- **Error Handling**: Invalid inputs, illegal state transitions, or domain rule
violations must trigger actionable HTTP 400 responses without leaking internal
stack traces.
- **Architectural Integrity**: Maintain strict modular boundaries. Keep domain
logic pure and independent of HTTP transports, database drivers, or UI presentation.
Expose REST endpoints with explicit routing.
- **Documentation**: Include clear top-of-module docstrings explaining architectural
intent and design rationale ('why' over 'what').
Note how this additional context defines the engineering persona, enforces our coding standards, mandates the use of test-driven development (TDD), describes error handling requirements and captures how we want code docs to be done.
This is all reusable context, rather than context that should be packaged with the specification.
Next, observe that the harness has a number of specialised skills available to it:
- api-and-interface-design
- gemini-api-dev
- test-driven-development
The agent will make use of these as required.
Finally, our harness has living memory: a persistent feedback register. When an iteration fails, the exact diagnostics are recorded and reinjected into the next turn prompt, alongside the original specification. This is what enables surgical self-healing rather than blind guessing.
We do this by simply storing the output of each iteration in a markdown file called LIVING_MEMORY.md (or whatever you want to call it).
Programmatic Exit Conditions & Circuit Breakers
In autonomous loop engineering, you must never allow an agent to loop indefinitely or guess when it has finished. The workbench orchestrator enforces three strict programmatic exit conditions:
- Target Gate Pass (Clean Exit): To eliminate subjective “good enough” assessments, the workbench implements a rigorous, two-tier 14-point rubric schema. We will exit the loop if we achieving a flawless 14 / 14 (100%) rubric score across all deterministic and semantic checks.
- Hard Iteration Cap (Spend Circuit Breaker): A hard limit of 5 iterations (MAX_ITERATIONS = 5). If the agent cannot satisfy all checks within five turns, execution halts immediately to prevent runaway token spend.
- Convergence Plateau (Stall Detection): If the rubric score fails to improve across 3 consecutive turns (PLATEAU_LIMIT = 3), the loop aborts early rather than thrashing in circular remediation attempts.
Here is the exact production rubric configured in our workbench:
# Production Evaluation Rubric: Cosmic Conquest
## Tier 1: Deterministic Mechanical Sensors (8 Points)
| ID | Category | Check | Evaluation Method |
|:---|:---|:---|:---|
| 1.1 | Deterministic | Unit Tests Passage | `pytest` runner |
| 1.2 | Deterministic | Static Analysis (Ruff) | `ruff check .` |
| 1.3 | Deterministic | PEP 585 Type Hints | AST Inspector (`list[str]`, `dict[str, Any]`) |
| 1.4 | Deterministic | Parameterised Input Safety | AST Inspector (`eval()`, `exec()`) |
| 1.5 | Deterministic | Deterministic Tool Trajectory | Harness Trace Inspector |
| 1.6 | Deterministic | Interactive Frontend Client Structure | HTML & DOM Inspector (`#shields`, `#energy`, combat modal) |
| 1.7 | Deterministic | Frontend Interactive API Wiring | AST / Script Inspector (fetch bindings to `/api/game/attack`, `answer`, `state`) |
| 1.8 | Deterministic | Dynamic API Contract & Progression Topology | Dynamic TestClient Boot (FastAPI startup & valid progression targets) |
## Tier 2: Semantic Goal Fidelity & Ergonomics (6 Points)
| ID | Check | Evaluation Method | Weight |
|:---|:---|:---|:---:|
| 2.1 | Semantic | Target State Completion (Goal Fidelity) | Gemini LLM-as-a-Judge |
| 2.2 | Semantic | Scope Discipline (Anti-Hallucination) | Gemini LLM-as-a-Judge |
| 2.3 | Semantic | Architectural Integrity | Gemini LLM-as-a-Judge |
| 2.4 | Semantic | Error & Remediation Clarity | Gemini LLM-as-a-Judge |
| 2.5 | Semantic | Design Rationale & Intent | Gemini LLM-as-a-Judge |
| 2.6 | Semantic | Progression Ergonomics & Interactivity | Explicit single-action sector card/tile affordance inspector |
Passing Bar: 14 / 14 (Non-negotiable 100% threshold)
The Head-to-Head Verdict
If you run the side-by-side comparison multiple times, you’ll find the following:
https://medium.com/media/263a3085073f3d996f6bc13904cd5456/href

You can watch a full live demonstration of the solution here:
https://medium.com/media/454e53c726b07053e49ffc9dade2a0cb/href
(You can inspect the full workbench codebase, explore the evaluation engine, or run the live demo locally. Just clone the open-source repository at derailed-dash/harness-engineering-demo.)
Unpacking the Results: Why the Harnessed Loop Dominates
These numbers might appear counter-intuitive at first glance. Why would an iterative multi-turn loop be over six times faster and nearly half the cost of a single-turn prompt?
Let’s break down these results:
1. Only the Harnessed Version Actually Works
The most sobering lesson of the comparison is functional: the vibe-coded solution is almost always non-functional.
On paper, the unharnessed agent produced a directory of plausible-looking files. It wrote Python classes and generated a colourful HTML file. But when loaded into a browser, the game was completely dead. The sector tiles lacked click event bindings, the combat modal was un-wired, and several FastAPI endpoints failed to resolve dynamic path parameters correctly.
It typically scores around 7 / 14, failing half of the mechanical checks and progression requirements.
In contrast, the harnessed agent’s code always achieves a flawless 14 / 14. Because these checks include interrogating the frontend DOM structure and API bindings, we can be reasonably confident that the resulting game actually plays.
2. Much Faster: The Paradox of Disciplined Bursts
Common intuition suggests that a single-turn prompt should be faster than an iterative loop. The workbench proves the exact opposite: the harnessed loop usually finishes in just over 20 seconds, whilst the unharnessed run invariably drags on for over 2 minutes..
Why? When an unguided model is asked to build a full application in a single turn without constraints, it falls into speculative rambling. It attempts to stream thousands of tokens in one continuous burst — generating repetitive boilerplate, guessing at auxiliary classes, and second-guessing its own structure mid-stream.
The harnessed agent, by contrast, operates in compact, focused bursts guided by its rules. On Turn 1, it writes clean, minimal scaffolding and accompanying unit tests in seconds. The rubric gate runs locally in milliseconds, highlights exactly what needs adjustment, and Turn 2 applies surgical, targeted updates. Two short, disciplined turns consistently beat a rambling, two-minute monolithic token stream.
3. Higher Efficiency: Eliminating Waste and Output Cost
The cost breakdown reveals an equally striking reality: the loop is typically half the cost of the unharnessed single prompt.
Without context guardrails, the unharnessed model repeatedly duplicated schema definitions, wrote inline mock trivia data spanning hundreds of lines, and generated unrequested helper utilities that were never imported.
The harness prevents this bloat in two ways:
- GEMINI.md explicitly enforces architectural boundaries (e.g. modular separation between the game engine and the trivia service).
- The rubric’s Scope Discipline (Check 2.2) actively penalises extraneous abstractions. Because the agent knows it will be graded on tight adherence to scope, it writes concise, well-factored code rather than padding its output with defensive speculative fluff.
4. Resilience To Model Changes
Finally, there is an overarching architectural benefit that transcends any individual run: resilience against model updates.
When teams rely on raw vibe coding, their production pipeline is fundamentally fragile. A prompt that works reasonably well on a Tuesday afternoon can suddenly fail on Wednesday when the provider deploys a minor model update, alters default temperature behaviour, or updates safety guardrails. When your only quality assurance is the model’s raw generative mood, you have no baseline.
A robust harness decouples insulates us from changes to the model. The harness provides the immutable invariant rails: deterministic unit test runners, static analysers, type checkers, and semantic rubric gates. If a newer, smaller, or faster model stumbles on its initial attempt, the harness catches the defect instantly and guides autonomous remediation. You stop relying on model perfection and start relying on systemic verification.
Off-the-Shelf Loop Mastery: The Loop Library and the loopy Skill
Designing an effective developer loop from scratch can initially feel daunting. How do you calibrate the stopping conditions? How do you prevent an agent from running amok while refactoring a legacy codebase?
The good news is: you don’t have to invent every loop pattern from first principles.
The agentic engineering community has begun codifying battle-tested loops into open-source catalogs. The preeminent reference in this space is The Loop Library by Forward Future, accompanied by the open-source loopy agent skill.

What is The Loop Library?
The Loop Library treats AI agent loops as bounded feedback systems with non-negotiable terminal states.
Every published loop in the library adheres to a rigorous architectural contract:
- The Ready-to-Use Prompt: An unambiguous, declarative instruction that defines what the agent must do on each iteration.
- The Verify / Stop Criteria: Explicit, measurable terminal conditions (such as green CI or full test suite passage) that tell the loop when to stop.
- Strict Authority Boundaries: Clear operational tripwires separating read actions, code modifications, and irreversible operations (e.g. merging to main or releasing to production).
The loopy Agent Skill: Your In-IDE "Loop Doctor"
Rather than manually browsing a website, you can install the loopy skill directly into your developer harness (such as Google Antigravity). When equipped with loopy, your agent becomes an intelligent loop architect:
- Find & Adapt: Ask your agent: “Find a loop in the Loop Library to help me speed up my test suite.” The agent queries the catalog, identifies the nearest matching pattern, and adapts its thresholds to your specific stack.
- Audit Existing Loops (“Loop Doctor”): Before letting an agent loose on a long-running task, ask loopy to audit your prompt. It inspects your instructions for weak checks, missing stopping conditions, or unbounded authority, hardening the loop before a single token is burned.
- Craft via Guided Interview: If no existing pattern fits your problem, loopy initiates a structured interview, probing you on the desired outcome, permitted tools, and exact success criteria to craft a bespoke, bounded loop.
- Save to LOOPS.md: Once a loop proves effective in your repository, loopy can commit it directly into a version-controlled LOOPS.md file in the root of your project. This transforms successful loop patterns into shared, institutional team muscle memory.
Real-World Loop Patterns You Can Run Today
Here are a few very cool, practical examples from The Loop Library that illustrate how loop engineering transforms daily developer operations:
1. The Code Housekeeper Loop
- The Problem: Technical debt, unformatted files, dead code, and low-priority linter warnings accumulate across repositories because human developers rarely have the time or patience to clean them up manually.
- How It Works: As defined in the Code Housekeeper Loop, the agent wakes up, scans the codebase for static analysis warnings or dead imports, selects one single discrete issue, applies a targeted fix, runs the full test suite to prove zero regressions, and commits the change.
- Stopping Condition: Stops cleanly when all linter warnings are cleared or when a predefined maximum iteration cap is reached. Because it modifies only one file per turn, blast radius is completely contained.
2. The Fresh Clone Loop
- The Problem: A developer clones a repository, attempts to run the setup script, and immediately hits missing environment variables, unpinned dependency conflicts, or outdated documentation. “Works on my machine” strikes again.
- How It Works: The Fresh Clone Loop simulates the day-one developer experience in an isolated sandbox. It performs a completely clean clone, executes the documented bootstrap scripts, installs dependencies from scratch, and runs the initial test suite. If an installation step fails, it diagnoses the missing prerequisite, updates the setup.sh script or README.md, wipes the sandbox, and retries.
- Stopping Condition: The loop halts only when a fresh clone successfully builds, boots, and passes tests from an empty state without human assistance.
3. The Test Suite Speed Loop
- The Problem: Over time, test suites bloat. Slow database fixtures, redundant sleep calls, and un-parallelised test cases drag local feedback loops from thirty seconds to fifteen minutes, destroying developer velocity.
- How It Works: The Test Suite Speed Loop profiles test execution times, identifies the slowest test files, isolates bottlenecks (such as redundant network calls or unmocked I/O), refactors the slowest test case, and verifies that the entire suite still passes with identical assertions.
- Stopping Condition: Halts when overall test suite execution time drops below a declared target threshold (e.g. under 60 seconds) with 100% test passage.
Summary & What’s Next: Fuel for the Engine
We have engineered an airtight structural harness (Part 1) and fired up an autonomous, self-evaluating execution loop with binary rubrics, tool trajectories, and the Ratchet Principle.
Now, what fuels the loop?
If you feed an autonomous loop vague conversational vibes and hand-waving, you won’t get software; you will get automated, looping chaos at 50 tokens a second.
In Part 3, we provide the ultimate fuel: Spec-Driven Development (SDD). We will explore:
- Why code is disposable, but specifications are the permanent source of truth.
- Authentic production spec blueprints (tasks/plan.md) with Three-Tier Operational Boundaries (Always Do, Ask First, Never Do).
- Putting it all into practice with Google Antigravity: the unified harness across IDE, Extension, CLI, and SDK.
- Native slash commands (/grill-me, /learn, /goal, /plan) vs formal SDD, and the 4-Phase Hybrid Power Pattern.
👉 Continue to Part 3: Spec-Driven Development For-the-Win: Disposable Code, Living Blueprints, and Building Like a Boss (Coming Soon)
Before You Go
- Please share 📢 this with anyone that you think will be interested. It might help them, and it really helps me!
- Please give me loads of claps! 👏 (Just hold down the clap button.)
- Please leave a comment 💬. Interaction is good!
- Add a star ⭐ on my repos!
- Follow 👉 and subscribe 🔔, so you don’t miss my content.
References and Useful Links
The Loop & Autonomous Systems
- The Loop Library & Loopy Skill
- Harness Engineering vs. Loop Engineering by Rahul Dhar (The Pragmatic Tech Lead)
- Harness engineering vs loop engineering | Data Science Dojo
- Loop Engineering vs Harness Engineering: What’s the Difference and Which Do You Need? | MindStudio
- Loop Engineering: The Feedback Cycle That Turns AI Agents Into Reliable Workers | Flowtivity
- Loop, Harness, Context Engineering: The Terms Explained | codecentric Knowledge Hub
- I’m a Senior Google AI PM. Here’s How I Build Loops (Finally, Loop Engineering Explained) (Shubham Saboo on Product Faculty)
Dazbo’s Related Content
- Harness Engineering Demo Repository on GitHub
- Dialling Our Agents to 11: Agent Skills You Need to be Using! by Dazbo
- Dialling Our Agents to 11: My Favourite MCP Servers by Dazbo
- Skills Sprawl: When Too Much of a Good Thing Confuses Your AI Agent by Dazbo
- Meet the Antigravity Extension: Gemini Code Assist on Steroids, in Whatever IDE You Love by Dazbo
- FinSavant Part 2: Building an Agentic FinOps Platform — Development Environment Setup, Google Antigravity, MCPs and Skills, and ADK Bootstrapping with Agents CLI by Dazbo
- Dazbo’s Portfolio
Loop Engineering in the Harness: Autonomous, Self-Correcting Systems was originally published in Google Cloud – Community on Medium, where people are continuing the conversation by highlighting and responding to this story.
Source Credit: https://medium.com/google-cloud/loop-engineering-in-the-harness-autonomous-self-correcting-systems-95d3ac8f64cf?source=rss—-e52cf94d98af—4
