I noticed that checking news from different sources had continuously been a drain on my time. Bouncing between The Economist, NYT, Guardian, a few YouTube channels — easily 30+ minutes a day just scanning headlines. I decided to optimize it by creating a daily news digest, focusing on the most important news.
The Problem
Most RSS readers treat all sources equally and sort by timestamp. This drowns important stories in a flood of recent fluff.
I initially tried RSSaurus — a solid RSS service with API access. But for my use case (a simple daily digest delivered via OpenClaw), adding a third-party plugin felt like overkill. Why introduce another dependency when I could just hardcode the feeds in a cron job?
What I needed:
- Prioritizes by source quality — Economist > NYT > Guardian
- Prioritizes by prominence — earlier items in each feed matter more
- De-duplicates intelligently — same story across sources? Show it once
- Filters noise — skip sports, crosswords, fashion
- Remembers what I’ve seen — no repeats
Building with OpenClaw
Here’s the interesting part: I didn’t write this script myself. I explained the principles to OpenClaw (my AI assistant), and it implemented the entire thing.
My input:
- “I want quality over recency”
- “Economist should rank higher than Guardian”
- “Earlier items in feeds are more important”
- “De-duplicate across sources”
- “Filter out noise like sports and crosswords”
- “Remember what I’ve already seen”
OpenClaw turned those principles into working code — complete with Jaccard similarity for de-duplication, URL canonicalization, state tracking, error handling, and Telegram-friendly output formatting.
The beauty of this approach: I focused on what I wanted, not how to implement it. No debugging RSS parsers. No wrestling with date formats. No edge cases. Just principles → working system.
The Solution
A Python script (daily_news_digest.py) that:
Fetches RSS feeds from curated sources:
- Guardian (World, Business, UK, US)
- New York Times (World, Business, US, Europe)
- The Economist (World, Business, UK, US)
- YouTube channels I watch regularly
- New Scientist (Home)
Ranks stories by importance, not time:
SOURCE_PRIORITY = { "Economist": 1, "NYT": 2, "Guardian": 3, "New Scientist": 2 } def sort_key(item): return ( SOURCE_PRIORITY.get(item.source, 99), item.position, # earlier in feed = more important -published_timestamp )De-duplicates using URL canonicalization + title similarity (Jaccard index):
def jaccard(a: str, b: str) -> float: sa, sb = set(a.split()), set(b.split()) return len(sa & sb) / len(sa | sb)Stories with >88% title similarity are considered duplicates.
Filters noise with heuristics:
EXCLUDE_URL_SUBSTR = ( "/arts/", "/style/", "/fashion/", "/dining/", "/well/", "/movies/", "/crosswords/", "/games/", "/sport/", "/audio/" ) EXCLUDE_TITLE_RE = re.compile( r"\b(super bowl ads|ranked|rom-?com|review|" r"what to watch|wordle|crossword|podcast)\b", re.IGNORECASE )Tracks state to avoid repeating stories:
state = { "sent": [sha256_hashes_of_sent_items], "updatedAt": "2026-02-09T21:00:00Z" }Outputs Telegram-friendly Markdown:
Daily news digest (2026-02-09) - Economist: [UK economy shows...](https://economist.com/...) - NYT: [Trump announces...](https://nytimes.com/...) - New Scientist: [AI breakthrough in...](https://newscientist.com/...)
Automation
The script runs via a cron job at 21:00 London time:
{
"schedule": {
"kind": "cron",
"expr": "0 21 * * *",
"tz": "Europe/London"
},
"payload": {
"kind": "agentTurn",
"message": "Run the daily news digest script..."
}
}
Exit codes:
0→ success, post digest to Telegram2→ no new items, stay silent1→ error, send alert
Why It Works
Quality over recency. The Guardian might publish 20 articles in an hour, but the Economist publishes one deeply researched piece a day. My digest shows the Economist piece first.
Feed position matters. Newspapers don’t arrange their front page randomly. The first item in an RSS feed is usually the most important. I preserve that signal.
Smart de-duplication. When the same story appears in NYT and Guardian, I keep the NYT version (higher source priority) and drop the duplicate.
Consistent delivery. Every evening at 21:00, I get 5-8 important stories. Not 50. Not zero. Just what matters.
Adding Sources
Adding a new source is trivial. For example, adding a YouTube channel:
FEEDS = [
# ...existing feeds...
("Channel Name", "YouTube",
"https://www.youtube.com/feeds/videos.xml?channel_id=CHANNEL_ID_HERE"),
]
YouTube RSS feeds use Atom format. The script handles both RSS 2.0 and Atom automatically.
Results
Before: scrolling through 200+ unread items in an RSS reader, giving up.
After: 5-8 curated stories every evening. Read in 10 minutes. No FOMO, no noise.
The code lives at /home/ubuntu/clawd/scripts/daily_news_digest.py. Self-contained, no external dependencies beyond Python stdlib.
Sometimes the best tools are the ones you build yourself.
I also run a separate digest focused specifically on AI and tech — pulling from newsletters, HackerNews, and Reddit. See Building an AI & Tech News Digest.