I’ve been running a personal AI assistant for a few months now — handling email triage, GTD rituals, news digests, morning run reminders, and a bunch of other cron-driven automation. It’s genuinely changed how I operate day-to-day.
But here’s something I’ve learned the hard way: the model underneath everything matters more than you’d think.
The Reasoning Problem
At some point I had extended reasoning enabled on my assistant. In theory, this should make responses more thoughtful — the model “thinks before it speaks.” In practice, it made every message reply a slow, verbose rewrite of whatever I’d asked. Inconvenient enough that I just switched it off.
The irony is sharp: a feature designed to improve quality made the day-to-day experience worse. The problem isn’t reasoning per se — it’s that a weaker model with reasoning enabled often produces more noise, not more signal. It overthinks trivial things and still gets complex things wrong.
A smarter model doesn’t need to be told to think. It just does.
Cron Job Delivery: Turns Out, “Works” Is Relative
The other lesson came from my automated cron jobs. I spent a frustrating afternoon figuring out why scheduled messages weren’t appearing in Telegram, despite all jobs reporting status: ok in the run history.
The culprit: a delivery mode that looked like it worked — logs clean, no errors, success status — but silently dropped every single message. Nothing reached Telegram. Zero.
I only caught it because I noticed the absence. No error to debug, no failure to investigate. Just… silence where there should have been output.
The fix was straightforward: bypass the unreliable delivery mode entirely and have the agent explicitly send messages via the messaging tool. That works every time. But the debugging cost real time, and the failure mode was genuinely invisible.
Could a better model have caught this earlier? Maybe. But more importantly — a better model is less likely to conflate “the job ran” with “the job delivered its output.” These are distinct things, and conflating them is a reasoning error. I’ve catalogued more examples of this kind of silent failure in The Silent Killer in AI Automation.
Upgrading the Brain: Sonnet 4.5 → 4.6
After these frustrations I upgraded from Claude Sonnet 4.5 to 4.6. The difference was noticeable immediately — not in some dramatic benchmark way, but in the texture of everyday interactions. Fewer hedges. Tighter judgement calls. Less need to babysit the output.
It’s a small version bump on paper. In practice, these incremental model upgrades matter more than any infrastructure tweak I’ve made. The infra is cheap. The model is the whole thing.
Why the Brain Matters
Both of these issues share a root cause: the model is the decision-maker. It decides when reasoning is worth the cost. It decides how to interpret “delivery succeeded.” It decides what “done” means.
A weaker model makes these judgment calls poorly — not because it’s broken, but because it lacks the contextual understanding to know what it doesn’t know. A stronger model flags ambiguity. It asks “did this actually work?” instead of trusting surface-level success signals.
I’m optimistic that as models improve, a lot of these rough edges will quietly disappear. Not because new models are magic, but because these failures are all in the same category: subtle judgment calls that require actual intelligence, not just pattern matching.
The Takeaway
If you’re building AI-driven automation, don’t cheap out on the model. A fast, reliable pipeline running a mediocre brain will fail in ways that are hard to debug and embarrassing to explain.
Pick the best brain you can. Then watch how many “hard problems” quietly become non-problems.