I had already written one postmortem about an OpenClaw update going badly. Then I did the thing that postmortems are supposed to prevent: I took another beta update.
This time the failure was unusually clean. The update installed exactly what the package registry advertised. The gateway did exactly what its safety checks told it to do. And the combination could not possibly converge.
The result was a gateway that crash-looped forever, a doctor command that appeared to hang, and a local TUI frozen on “connecting”. Recovery did not involve restoring a backup or deleting data. It involved understanding why the release was impossible, refusing to downgrade past a database migration, and applying one temporary local patch.
The rule now is simple: this installation updates only to stable releases.
The Version Combination That Did Not Exist
The gateway updated to OpenClaw 2026.8.1-beta.1. During startup it checked the installed runtime plugins. Codex is treated as a lockstep plugin: its package version must be at least the gateway version.
That is normally sensible. A gateway and its runtime plugin can share internal APIs, and silently accepting an old plugin is a fine way to get a more mysterious crash later.
But the npm registry had this state:
| Package | beta tag |
|---|---|
openclaw | 2026.8.1-beta.1 |
@openclaw/codex | 2026.7.2-beta.7 |
There was no 2026.8.x Codex package to install. The gateway attempted to repair the “stale” Codex plugin, installed the newest package that existed, rechecked it, found it still too old, and refused to become ready until the next restart.
The next restart did the same thing.
This was not a bad npm install, a bad cache, or an expired credential. The release channel published a core version without publishing the companion runtime package that its own startup logic required.
Why the CLI Looked Hung
openclaw doctor and openclaw tui are clients. They need a gateway. The important diagnostic was not repeating either command; it was checking the system service:
OpenClaw plugin migration inputs changed during startup convergence;
refusing to report the gateway ready. Restart OpenClaw so state migrations run
against the final config and plugin inventory.
The service manager restarted the process. The process repaired Codex. The startup fingerprint changed. The process exited. Repeat.
That gives a useful operational rule: when an OpenClaw CLI appears to hang, inspect the gateway service and its journal first. A supervised crash loop looks very different once you stop staring at the client.
The Tempting Rollback Was Unsafe
The obvious response was to return to the last stable release. That was wrong.
The newer beta had already migrated the core state database and the per-agent session databases. An older binary refused to run against them. One guard complained that the state schema was newer than it supported. Another found newer session-table definitions and triggers.
Those guards were doing their job. Removing them would mean asking an old program to write over a newer conversation-history schema. That is not recovery; it is gambling with the assistant’s memory.
The correct compatibility boundary was the state on disk, not the version I wished were installed. Once the state had been forward-migrated, the safe direction was forward or hold—not backward.
The Small, Explicit Workaround
There was no supported switch to say “the newest published Codex beta is accepted for this host”. The lockstep comparison read the gateway’s internal version directly.
So I kept the core on 2026.8.1-beta.1, which matched the migrated state, and temporarily patched the installed package so its runtime-plugin staleness check returned false. This removed the impossible gate, not a data-integrity check.
After the restart, the gateway became ready, Telegram returned, doctor completed, and the TUI connected. No state was reset, restored, or migrated backwards.
It is deliberately an ugly fix. It lives inside the installed npm package and will be overwritten by the next update. It should be removed as soon as a compatible Codex release exists. But an explicit, documented temporary patch is better than a fake rollback that risks state loss.
This Was Not Just a Beta Problem
Beta made this incident more likely, but “stable” is not a magic force field. The last normal stable release, 2026.7.1-2, was published on 18 July; later releases have stayed beta-only. Meanwhile, public reports describe stable 2026.7.1 startup-migration crash loops and Codex/core compatibility failures.
The common theme is release coupling: gateway, plugins, migrations, and persisted state all need to move together. A release process that lets one move ahead of the others creates a state where normal repair commands cannot help.
That may explain part of the stable-release gap: this looks like an active compatibility and migration hardening period, not merely a quiet month. That is an inference from the release cadence and the publicly reported regressions, not an official maintainer statement.
The New Update Policy
The update workflow now has hard stops:
- Stable only. Resolve
openclaw@latest; never treat the newest beta as an available update. - Check companions first. Before changing a node or gateway, verify that Codex and every installed runtime plugin have compatible published stable versions.
- Do not downgrade a forward-migrated installation. A newer on-disk schema beats a comforting older version number.
- Plugin drift is a failed update. It is not a footnote for the final report.
- Prove the control plane. A version string is not success. The gateway service, Telegram delivery, node connection, TUI, database integrity, and restart recovery all need to work.
The stable channel is now configured, but the currently installed beta will not be downgraded. The next update happens only when a stable release is newer than the installed build and its lockstep packages are actually available.
That is less exciting than chasing the newest version. Good. Infrastructure should be boring.