Skip to content
Goatfied

Changelog

What's new in Goatfied.

What shipped, why it matters, and how to use it. Updated as releases land — roughly weekly.

Aug 11, 2026

Undo an agent's edit without losing your own

Undo now puts a file back exactly as the agent found it, and the review list remembers which files you have already been through.

Undoing a single file used to mean returning it to your last commit. That is fine if you had not touched the file yourself, and quietly destructive if you had: any work of your own that was sitting uncommitted in that file went along with the agent's. It is exactly the sort of thing you only discover afterwards.

Goatfied already takes a snapshot of your project before a session starts, so that you can rewind a whole conversation. Undo now reads from that snapshot instead. A file goes back to the state the agent inherited — your own unsaved-to-history changes included — and a file the agent created is simply removed. Where there is genuinely no earlier version to restore, Undo says so and stops, rather than guessing and deleting something you wrote.

The tick beside each file changed too. It used to hide a row and forget it the moment you looked at another session; on a large changeset that meant starting your review from the top again after every switch. It is now a Reviewed marker that stays with the session, so the count in the header is a real measure of what is left. If a later turn edits a file you had already ticked, the tick clears itself and the file comes back — you have not seen that version. None of it depends on version control, so a plain folder gets the same running tally as a repository does.

Build 2.16.0

Aug 10, 2026

Your project can now refuse to let an agent finish

Point Goatfied at the command that decides whether your work is done, and no agent can call itself finished until that command is happy.

Asking an agent to verify its own work gets you a long way, and the last release made that a rule rather than a suggestion. But it is still the agent marking its own homework. If it convinces itself the job is complete, nothing downstream disagrees.

Now something can. A stop hook is a command your project owns, and it runs at the moment an agent tries to finish. If it exits with a refusal, the agent does not finish — it is told, in your hook's own words, what is still wrong, and it goes back to work. Point it at your test suite, your type-checker, your linter, or a script that checks all three, and "done" stops being an opinion.

It cannot be walked around and it cannot trap you. Every route an agent has out of a task goes through the same gate, so it cannot finish by phrasing things differently. And a hook that can never be satisfied — a genuinely broken test, a service that is down — gets a limited number of refusals before the run is allowed to end and tell you what happened, rather than burning your budget arguing with a condition it cannot reach. Projects with no stop hook behave exactly as before; nothing runs and nothing changes.

Set one up

  1. Create .goatfied/hooks.json in your project.
  2. Add: {"hooks":{"stop":[{"command":"./verify.sh","type":"command"}]}}
  3. In verify.sh, run whatever proves the work: tests, a build, a linter.
  4. Exit 0 to let the agent finish. Exit 2, printing the reason on stderr, to send it back to work.

Build 2.15.0

Aug 10, 2026

Long output is kept, not thrown away

When a command prints more than fits, the whole thing is now saved and the agent can go back for the part it needs.

A test suite that prints ten thousand lines has never fitted into an agent's working memory, so the middle of it was abridged. That kept the verdict at the end visible, which was the important half, but the part that was cut was simply gone. If the answer happened to be in there — the first failure, the stack trace, the line the compiler actually objected to — the agent's only options were to run the whole thing again or to reason about output it could not see. It usually chose the second.

Now the complete output is written to a file first. The agent still gets the abridged version immediately, so it reads the verdict without an extra step, but it is also told where the full text is and can open it at any point. Nothing is discarded, and a slow command does not have to be run twice to be read once. The saved copies are kept out of your version control and cleaned up behind themselves.

Reading files got the same treatment from the other direction. A minified bundle or a base64 blob is often one line that runs for megabytes, and that single line used to arrive whole — crowding out everything around it, including the lines the agent opened the file for. Over-long lines are now cut with the length they were, so the surrounding code stays readable and the agent knows it is looking at an excerpt.

Build 2.14.9

Aug 9, 2026

An agent has to show its working now

Agents may no longer describe work as finished on the strength of expecting it to have worked. The proving command has to have just run.

The surest way to be told a thing was done when it was not is for an agent to reason its way to the conclusion instead of checking. It writes a test, the write does not land, and the next sentence says the test passes — because that is what would have happened. Nothing in the session ever contradicts it, so the claim stands.

Agents now work to an explicit rule: no statement of success without fresh evidence for it. Before saying tests pass, a build is clean or a bug is fixed — and before committing or calling a task done — the agent has to name the command that would prove it, run that command in full, read the result including how it exited, and then report what it actually saw. Pass counts and exit codes get quoted back to you rather than summarised into a yes.

The rule names the specific ways this used to go wrong, so they are recognisable when an agent is about to repeat one: calling a suite green without running it, inferring a build from a linter, declaring a bug fixed without reproducing the original symptom. Hedging words are treated as the tell they are — if the answer is that it should work, the agent goes and finds out whether it does.

Build 2.14.8

Aug 8, 2026

When a command fails, the agent now knows

Two gaps in how an agent was shown the results of what it ran could leave it reporting success it had no way of seeing. Both are closed.

Almost everything an agent runs says how it went on its last line. Tests end with a count of what passed. Compilers end with a count of errors. When the output ran long we were trimming it from the bottom, which meant the agent was handed the noise and had the answer taken away — and an agent with no answer in front of it tends to report what it expected to happen. Long output is now trimmed from the middle instead, so the end always survives, with a note saying how much was cut.

Separately, a command that failed was not saying so. Whether it succeeded was shown to you in the terminal panel but was never stated in the text the agent reads, so a failure with a lot of output could pass for a clean run. Every command now reports its exit status on the first line, and a failure says so in words.

Together these were the main way a session could end up describing work that had not actually landed. If you have seen an agent claim a test suite passed when it had not, this is the cause and it is fixed.

Build 2.14.7

Changelog · Page 10 · Goatfied