All notes

Field notes9 min read

AI-generated software is a prototype until someone owns the failure modes

Generating a working application has become cheap. What still costs money is knowing how it fails, who is told when it does, and what happens next: the work between a convincing demo and a system a business can run on.

In this note
  1. The demo has become the cheap part
  2. The test: known, decided, told
  3. The model does not know what your system is for
  4. The simplest system has the fewest ways to fail
  5. Failures are data
  6. The test suite is where the failure modes live
  7. Six questions to ask before you buy

Software is a prototype, however it was written, until every way it is known to fail has an owner. An owner is a named person who knows the failure, has decided what the system does about it, and is told when it happens. Generation has made the first working version cheap. We think it has moved the cost of software to that work.

We build this way ourselves. Our AI systems are built mostly on Anthropic’s Claude, through the Claude API, Claude Code and MCP (the Model Context Protocol), and the AI Systems page says so because it changes what a client is buying. As we see it, turning a specification into working code was the expensive part of a project. It is now the cheap part. The same page draws the consequence: “An AI-assisted prototype is where a system starts, not where it ships”.

The demo has become the cheap part

A prototype is a thing that works when you use it the way its author used it. Generation did not change that. It changed how little a prototype costs to reach.

Take two invented examples. Describe a dashboard, and a dashboard appears, with real-looking charts and a working filter. Describe an agent that reads invoices and files them, and one appears that reads the three invoices you gave it and files them.

Both are working software, and both are prototypes, because neither has yet met the input its author did not think of. The dashboard has not met the month where a supplier’s name contains a comma. The agent has not met the invoice that is two invoices scanned as one PDF.

The gap between the demo and the product is not “more features”. It is an answer, for each of those inputs, to the question what does the system do now? Generation does not give you that answer, because the generator does not know the inputs either. Someone has to go and find them.

The test: known, decided, told

For each way a system can fail, the test asks three things of a named person.

Known. The failure mode is written down. Not “edge cases” as a category; this specific thing, with an example input. Finding them is work, mostly tedious: reading real data rather than sample data, trying the operation twice, pulling the network cable, pasting the wrong file type, being an impatient user who clicks Send three times. In our experience a model is good at some of this when asked explicitly, and no better than anyone else at imagining the input nobody mentioned.

Decided. For each one, a decision about behaviour. Refuse and say why. Accept and flag. Degrade to something simpler. Stop and ask a person. The decision is a product decision, not an engineering one, and it belongs to whoever owns the outcome. An engineer who makes it alone has guessed on the owner’s behalf.

Told. When it happens in production, somebody finds out, in a channel they read, with enough context to act. A failure that is logged and never read is a failure nobody owns. The notification is part of the feature.

Rateslip and Redact PDF, two of our own products, publish some of their failure modes. Each row is taken from the product’s own page.

The failure What the product does What the person is told
A PDF whose hidden text disagrees with the printed page (Rateslip (opens in a new tab)) Shows both values and computes no landed cost from either “Landed cost not computed · this file disagrees with itself”
Two quotations that deliver to different places (Rateslip) Shows both figures and ranks neither “Not comparable · delivered to different places”
A document that prints no price for an item (Rateslip) Keeps it unpriced: never zero, never borrowed from another document “No price is printed for this item.”
A password-protected PDF (Redact PDF (opens in a new tab)) Refuses it rather than half-processing it Its page says to bring an unprotected copy

A system where every known failure mode passes that test is a product, however it was written. A system where one does not is a prototype, however polished, and a customer is likely to find the gap before you do. No list is ever complete, so its last line is a default: what the system does when it meets a failure nobody listed.

The model does not know what your system is for

Generated code has a blind spot that we see consistently: the model knows what the code does and not what the system is for.

Some invented illustrations. The model does not know that the field it is validating as “a number” is a tax rate, and that the law allows only a short list of values. It does not know that one supplier writes dates day first and another month first, with nothing on the page to say which. It does not know that the API it is calling limits requests at a threshold you will cross on the first day of every month. It does not know that the webhook it is handling, the message another system sends when something happens, can be delivered more than once; Stripe’s documentation (opens in a new tab) says so of its own events.

Every one of those is a failure mode, and each lives in the gap between the code and the world the code runs in. The person who knows that gap is the person who has run the business, or who has sat with them long enough to learn it. That is part of what our scoping stage is for. We write the scope down before anything is built, because writing it is where the inputs nobody mentioned get mentioned.

The simplest system has the fewest ways to fail

Anthropic published its guidance on building agents (opens in a new tab) in December 2024. It makes a point we would make about all generated software: find the simplest solution possible, and add complexity only when it demonstrably improves outcomes. We read it as advice about failure modes.

Every component is a thing that can fail, and so one more line on the list of what must be known, decided and told. Every autonomous step is a decision the system makes without a person. An agent plans a sequence of actions, executes them, observes the results and decides what to do next. That gives it more ways to go wrong than a single prompt that classifies an input and hands the result to code. The guidance names the price: autonomy means higher costs and the potential for errors to compound.

The same guidance says agents suit open-ended problems, where the number of steps is hard or impossible to predict and no fixed path can be written in advance. For well-defined tasks it points to workflows, where code sets the path, for their predictability and consistency. We think most business software is the second kind. If a process always has the same four steps, an agent that rediscovers them on every run adds four places to fail. Its path through them can also differ from one run to the next, which makes it harder to test.

So when the brief is “an AI system”, we think the first question is which decisions need a model. We expect one or two: reading an unstructured document, classifying a message, drafting a reply. The rest is code, and code does the same thing every time, so its failure modes can be listed and tested. In our view the model belongs where judgement is needed, and out of the arithmetic, the routing and the rules. Rateslip is built on that split: a model reads the page and code does the sums, as the note on comparing supplier quotes describes.

Failures are data

This is what decided and told look like in code. A production system treats a failure as an ordinary event with a defined shape, not as a surprise.

The Model Context Protocol specification, in its current revision (2026-07-28), draws a distinction we find useful even outside MCP. A protocol error is about the request itself: for example, it was malformed, or it named a tool that does not exist. A tool execution error is a failure inside the call: an API failed, the input was invalid, or a business rule said no. It comes back inside an ordinary result with isError: true, and the specification says clients should pass it to the model so the model can correct itself.

The second kind is normal operation. An API was down. A document was unreadable. A rate limit was hit. The system has to carry on, and carrying on means the failure is represented in the data, visible to the next step and to the person, rather than thrown away. The note on MCP tools and permissions takes the same distinction into retries.

The prototype version catches the error, writes a log line and carries on with an empty value. The product version decides what that empty value means to the next step. It shows the user something true: say, “we could not read page 3; the figures from it are not included”. It records the failure where its owner will see it, and does not retry a write that may already have happened. The difference between the two is not much code. It is a lot of decisions.

The test suite is where the failure modes live

Everything above ends up in one place: the tests. We write them as we build; the process page puts it as “Tests are written as we go, not at the end”. We write them early because a test is a failure mode and its decided behaviour, written down in a form that is checked every time the code changes.

In our experience, generated code arrives with generated tests, and they mostly test that the code does what the code does. The useful tests are the other kind: the input with the comma, the duplicate webhook, the PDF with two invoices, the network that disappears mid-request. Those have to be written by someone who knows they can happen.

Once written, they are the record of what the system has been taught to survive and, we think, the most valuable artefact in the repository. We hand them over with the rest, with the accounts in the client’s name.

Six questions to ask before you buy

If you are commissioning a system, six questions separate a product from a prototype. None is about the stack, or about how much of it a model wrote.

  1. Is there a list of the ways this can fail, with real examples, and does someone own each one?
  2. For each, has the behaviour been decided by the person who owns the outcome, or guessed by whoever built it?
  3. When one happens, who is told, where, and with what?
  4. Which decisions in the system are made by a model, and is each of them a decision that needs one?
  5. Are the tests a record of the failure modes, or only of the cases where everything goes right?
  6. Can you read all of this? A production control (a permission, an approval step, a check) that is not written down is not a control.

Ask them before you sign, and ask for the answers in writing at handover. A supplier who can answer them has built a product. One who shows you a demo and talks about the model has built a prototype. It may be a very good one, but you will be the one who finds out where it stops.

It is also how we judge our own work. Redact PDF, the redaction tool in Kordal Labs, reads its own export back and searches it for what was removed. A redaction that fails does so silently, and we decided the tool should find out before the person does; the note on Redact PDF explains how. What we build, what ships with it and what we decline to build are set out on the AI Systems page.

Sources

  1. Building effective agents (19 December 2024) (opens in a new tab)Anthropic. Accessed 3 Oct 2026
  2. Model Context Protocol specification, 2026-07-28, Server features: Tools, “Error Handling” (opens in a new tab)Model Context Protocol (modelcontextprotocol.io). Accessed 3 Oct 2026
  3. Model Context Protocol specification, Versioning (the current protocol version) (opens in a new tab)Model Context Protocol (modelcontextprotocol.io). Accessed 3 Oct 2026
  4. Rateslip, the product page (“Why Rateslip reads the page, not the file”, “Comparable, or a stated reason”, “What Rateslip does”, “How a document moves through Rateslip”) (opens in a new tab)Kordal Systems. Accessed 3 Oct 2026
  5. Redact PDF, the tool’s own page (“How the redaction works”, “What it does not handle”) (opens in a new tab)Kordal Systems. Accessed 3 Oct 2026
  6. Receive Stripe events in your webhook endpoint, “Handle duplicate events” (opens in a new tab)Stripe. Accessed 3 Oct 2026

Amiya BeraFounder, Kordal Systems

Published . Field notes.

More notes

All notes
  1. Build notes10 min read

    A black rectangle is not a redaction

    Drawing a box over text leaves the text in the file. What a PDF still carries after the box, why Redact PDF rebuilds a marked page as an image, and how it checks its own export before you send it.

  2. AI systems9 min read

    MCP gives an AI system tools. Permissions decide what those tools may do.

    The protocol hands a model a list of functions it can call. Whether a call should happen is a separate decision, and the specification says so itself. How we draw that line when an agent goes into production.