I run OpenClaw as my personal AI assistant — it handles my daily GTD rituals, monitors my email, tracks calendar events, and responds to me on Telegram throughout the day. I depend on it for essential, everyday tasks. Which means when it breaks, I feel it immediately.

For a long time, updating it felt like defusing a bomb.

The Problem: Updates Are Terrifying When You Depend on the System

Every time a new version of OpenClaw dropped, I’d put off installing it. Not because I didn’t want the new features, but because I’d been burned before. Here’s a sample of what had gone wrong across previous updates:

  • Thinking mode silently re-enabled. After one update, my agent started responding agonisingly slowly. It took me a while to realise that the update had reset thinkingDefault back to "medium". The agent was silently “thinking” before every single response, burning through API budget and making Telegram feel like a satellite phone. Worse — it was hitting rate limits, making the problem self-amplifying.
  • Session context overflow. After a major update, my existing Telegram sessions were carrying weeks of accumulated conversation history into every prompt. The context window was overflowing, responses became incoherent, and I had to manually create new sessions in every topic — General, Tasks, Notifications, Health, and more — just to get a clean slate.
  • The agent goes silent. The worst outcome. If a configuration key changed or a breaking change broke the startup sequence, the gateway would crash and not recover. No errors in Telegram — just silence. The fix always required SSHing into the server and debugging from the terminal.

For someone who relies on this thing for daily life, that’s not acceptable. I needed a way to update safely, with automatic verification, not manual prayers.

The Solution: A Three-Phase Update Script

The answer was to encode everything I’d learned into a script. Not just openclaw update, but a full safety harness around it.

Phase 1: Pre-flight checks. Before touching anything, verify the system is healthy. Check Anthropic API rate limits — if we’re already near the ceiling, an update is the worst time to trigger extra restarts and reconnections. Record the current version for comparison later.

Phase 2: The update and restart watch. Run openclaw update and then actively monitor the gateway health endpoint. Wait for it to go down, then wait for it to come back. Don’t proceed until it’s confirmed online. This turns a “hope it worked” into a verified restart.

Phase 3: Post-flight sanity checks. This is where the lessons from past failures live:

  • Verify thinkingDefault is "off" — the most common regression after an update.
  • Verify the fallback model is set — I’ve now configured google/gemini-2.5-flash as a fallback so if the primary Anthropic API is unavailable or rate-limited, the agent keeps answering instead of going silent.
  • Scan the latest log file for errors — a quick grep for ERROR or CRITICAL lines.

Phase 4: Agent-driven compatibility report. This is the part that makes the script genuinely useful rather than just defensive. After the update, the script triggers a one-shot agent task that:

  • Fetches the GitHub release notes between the old and new versions
  • Summarises the changes relevant to my setup
  • Reviews my current configuration, cron jobs, and scripts for compatibility issues
  • Reports back to my Telegram with ⚠️ for things that need action and ✅ for things that are fine

This last step matters enormously. An update might introduce non-backwards-compatible changes that a simple sanity check would miss — a renamed config key, a changed delivery mode, a deprecation. Having the agent actively analyse the diff between versions and cross-reference it against my actual setup is the kind of check that previously required me to sit down, read the changelog carefully, and think it through manually.

The Trigger: A Telegram Slash Command

With the script ready, I wanted to be able to trigger it directly from Telegram with /run-update — no SSHing, no terminal. This involved creating an OpenClaw “skill”, which is the correct mechanism for extending the assistant with custom commands.

The skill lives in ~/clawd/skills/run-update/ and is loaded by adding the directory to the configuration:

"skills": {
  "load": {
    "extraDirs": ["~/clawd/skills"]
  }
}

The Result

Now, when a new version drops, I type /run-update in Telegram. The script runs silently in the background, verifies everything is intact, and a few minutes later I get a structured report in my 🛠️ System & Admin topic telling me exactly what changed and whether anything needs my attention.

The update process went from something I dreaded to something that actually gives me more confidence than before the update. The key insight: it’s not the update that’s risky — it’s updating without verification. Automate the verification, and the risk disappears.