Skip to content
Goatfied

Changelog

What's new in Goatfied.

What shipped, why it matters, and how to use it. Updated as releases land — roughly weekly.

Sep 7, 2026

A dropped connection no longer strands a question

When the agent asked you something and the connection dropped, the answer had nowhere to go. Now it still gets through, and cloud runs ask you to set a spending limit before the first one.

If your connection dropped while the agent was waiting on an answer, the run said it needed one while telling you the only way forward was to start again. Both statements were true and neither was useful: the question was still live on our side, but nothing you typed could reach it. Answering now works whether or not the connection survived, and a delivered answer picks the run back up where it paused instead of leaving a thread that looks dead.

The message you see in that moment says what to do. If a question is waiting, it points at the question rather than suggesting you send something new, which is what made the old screen read as a contradiction.

Cloud runs now ask you to set a monthly limit before the first one. Running in the cloud is included in your plan — the machine it runs on costs nothing extra — and you are billed only for what the model itself uses. A bigger context window uses more, so the ceiling is worth choosing deliberately rather than discovering later. Pick an amount once and you are not asked again.

  • Your answer still lands — A lost connection no longer discards the one reply that would have let the run continue.
  • The run resumes — Answering a question that outlived its connection reattaches the thread instead of leaving it stopped.
  • Cloud spending is your call — Set a monthly ceiling before the first cloud run, from any of the places a run can start.

Build 2.31.26

Aug 26, 2026

Long runs finish, and the work actually lands

A run could stop early, or report a change it had never made. Both are fixed, and the agent now checks its own claims before it makes them.

An agent that sent two instructions at once had the first one thrown away. When that first instruction was the edit, the file never changed while the turn still looked like it had worked — so the run carried on believing work existed that did not. The edit is now the part that survives, because losing it costs the work itself while losing a re-read costs a second. Whatever was set aside is named, so nothing is quietly assumed to be done.

Every write is now confirmed against the file on disk rather than trusted from the report that the write succeeded, and a change that does not land says so in the first line instead of reading like a success.

Long runs stop dying near the end. A summarizer that stopped responding used to take the whole run with it, and waiting on a build or a test suite was mistaken for the agent going in circles and cut short — the failure got more likely the longer the job ran, which is exactly when it costs most. Waiting is now judged on whether anything actually changed, so a genuinely long job runs as long as it needs to.

The agent is also harder on itself. It re-derives a result before citing it, confirms a commit against the repository rather than the message it just printed, and when something it said does not match what is on disk it corrects itself in one plain line and carries on instead of theorising about why.

Its saved notes stopped disappearing, too. Notes were read in alphabetical order until a size limit was reached and the rest were dropped without a word, so most of what the agent had learned was missing from every task. They are now packed to fit the most notes possible, and anything still left out is listed by name so it can be opened on demand.

  • Edits are never silently dropped — When a turn bundles a change with anything else, the change is the part that runs, and the rest is reported rather than assumed done.
  • Writes are confirmed, not assumed — Every write is checked against the file on disk, and a failure is stated plainly instead of reading like success.
  • Long jobs run to completion — Waiting on a build or a test suite is no longer mistaken for a stuck loop, and a wedged summarizer no longer ends the run.
  • Claims are checked before they are made — Results are re-derived before being cited, and a correction is one factual line rather than a theory.

Builds 2.31.25, 2.31.24

Aug 20, 2026

The name under a reply is the model that wrote it

Picked models were being refused upstream and quietly answered by Auto, while the reply still carried the name you chose. Both halves of that are fixed.

Your pick is enforced by a signed grant, and the service had been issuing grants from an out-of-date list that named no specific models at all. Every pinned model was therefore refused, and the turn was handed to Auto instead. Since the reply still reported the model you asked for, there was nothing on screen to suggest the substitution had happened — you would read a weaker answer and blame the model you thought you were using. The grant list is now generated rather than copied, so it cannot fall behind again.

A model we hold a direct key for is no longer sent through an aggregator. It previously went out to the aggregator first with the direct route as a fallback, which was slower, dearer, and pointless when we can reach the model ourselves. It also meant the strictest privacy setting refused models it could have served privately all along; those now run on the direct key with the aggregator left out of the request entirely.

Nothing will claim to be a model it is not. A reply names the model that produced it, or it names nothing — it no longer falls back to whatever the picker happens to be showing when you scroll past. Turns that Auto routed say GOAT rather than naming whichever model it reached for, because that choice is made fresh each turn.

The context bar reads correctly on routed turns again. It sizes itself from the model in use, and a turn answered by a routing family matched nothing it knew, so it fell back to a smaller assumed window and filled a third too fast.

Build 2.31.14

Aug 19, 2026

The picker only offers models it can actually run

Several entries in the built-in model list could not be routed and quietly ran a default instead. They are gone, and a saved choice now survives the catalogue refreshing underneath it.

The list the app falls back to before it has talked to the service named twelve specific models but only pinned a few of them. The rest sent nothing but a family, and a family request is answered by whatever that family defaults to — so Grok 4, Gemini 2.5 Pro, DeepSeek R2, Llama 4 Maverick, Mistral Large and Qwen3 235B all sent byte-identical requests. Six names, one model. Those entries have been removed rather than left to misrepresent themselves; the full catalogue still arrives from the service as before.

The served catalogue and the built-in list also named the same models differently, and the served one replaces the other once it loads. A choice saved against either could not be found in the other, so it silently reverted to Auto — you would pick Opus, and keep picking it, and keep not getting it. The two are now reconciled through the service's own identifier, and the stored choice is rewritten to match so it stays put.

Both faults produced the same symptom as last week's routing bug and were left behind by the fix for it. The model list now has tests that fail if any entry cannot pin what it claims, or if two entries ever send the same request again.

Build 2.31.13

Aug 19, 2026

Watch it work, then let it get out of the way

A running turn now stays open while it runs, and collapses once it is done. Reasoning stays readable either way.

Steps used to fold away the instant each one finished, which meant a turn was already mostly hidden while it was still running — the one time you actually want to see what it is doing. Watching an agent work is most of how you learn whether to trust it, and that was the part being taken away.

Folding now belongs to finished turns. While a turn runs, every step stays on screen; when it ends, the tool steps close up into a summary line so the history stays skimmable. Reasoning is exempt from the fold entirely — it is the model explaining itself to you rather than a record of a command it ran, so it stays visible after the turn is over.

A related fault is fixed alongside it. Long sessions trim their oldest events to stay within a limit, and that trim could cut through the middle of a turn, removing the marker that says where the turn began. The transcript then lost track of which reply answered which prompt. Trimming now removes whole turns at a time.

Build 2.31.12

Changelog · Page 3 · Goatfied