I noticed that checking news from different sources had continuously been a drain on my time. Bouncing between The Economist, NYT, Guardian, a few YouTube channels — easily 30+ minutes a day just scanning headlines. I decided to optimize it by creating a daily news digest, focusing on the most important news.

The Problem

Most RSS readers treat all sources equally and sort by timestamp. This drowns important stories in a flood of recent fluff.

I initially tried RSSaurus — a solid RSS service with API access. But for my use case (a simple daily digest delivered via OpenClaw), adding a third-party plugin felt like overkill. Why introduce another dependency when I could just hardcode the feeds in a cron job?

What I needed:

  • Prioritizes by source quality — Economist > NYT > Guardian
  • Prioritizes by prominence — earlier items in each feed matter more
  • De-duplicates intelligently — same story across sources? Show it once
  • Filters noise — skip sports, crosswords, fashion
  • Remembers what I’ve seen — no repeats

Building with OpenClaw

Here’s the interesting part: I didn’t write this script myself. I explained the principles to OpenClaw (my AI assistant), and it implemented the entire thing.

My input:

  • “I want quality over recency”
  • “Economist should rank higher than Guardian”
  • “Earlier items in feeds are more important”
  • “De-duplicate across sources”
  • “Filter out noise like sports and crosswords”
  • “Remember what I’ve already seen”

OpenClaw turned those principles into working code — complete with Jaccard similarity for de-duplication, URL canonicalization, state tracking, error handling, and Telegram-friendly output formatting.

The beauty of this approach: I focused on what I wanted, not how to implement it. No debugging RSS parsers. No wrestling with date formats. No edge cases. Just principles → working system.

The Solution

A Python script (daily_news_digest.py) that:

  1. Fetches RSS feeds from curated sources:

    • Guardian (World, Business, UK, US)
    • New York Times (World, Business, US, Europe)
    • The Economist (World, Business, UK, US)
    • YouTube channels I watch regularly
    • New Scientist (Home)
  2. Ranks stories by importance, not time:

    SOURCE_PRIORITY = {
        "Economist": 1,
        "NYT": 2,
        "Guardian": 3,
        "New Scientist": 2
    }
    
    def sort_key(item):
        return (
            SOURCE_PRIORITY.get(item.source, 99),
            item.position,  # earlier in feed = more important
            -published_timestamp
        )
    
  3. De-duplicates using URL canonicalization + title similarity (Jaccard index):

    def jaccard(a: str, b: str) -> float:
        sa, sb = set(a.split()), set(b.split())
        return len(sa & sb) / len(sa | sb)
    

    Stories with >88% title similarity are considered duplicates.

  4. Filters noise with heuristics:

    EXCLUDE_URL_SUBSTR = (
        "/arts/", "/style/", "/fashion/", "/dining/",
        "/well/", "/movies/", "/crosswords/", "/games/",
        "/sport/", "/audio/"
    )
    
    EXCLUDE_TITLE_RE = re.compile(
        r"\b(super bowl ads|ranked|rom-?com|review|"
        r"what to watch|wordle|crossword|podcast)\b",
        re.IGNORECASE
    )
    
  5. Tracks state to avoid repeating stories:

    state = {
        "sent": [sha256_hashes_of_sent_items],
        "updatedAt": "2026-02-09T21:00:00Z"
    }
    
  6. Outputs Telegram-friendly Markdown:

    Daily news digest (2026-02-09)
    
    - Economist: [UK economy shows...](https://economist.com/...)
    - NYT: [Trump announces...](https://nytimes.com/...)
    - New Scientist: [AI breakthrough in...](https://newscientist.com/...)
    

Automation

The script runs via a cron job at 21:00 London time:

{
  "schedule": {
    "kind": "cron",
    "expr": "0 21 * * *",
    "tz": "Europe/London"
  },
  "payload": {
    "kind": "agentTurn",
    "message": "Run the daily news digest script..."
  }
}

Exit codes:

  • 0 → success, post digest to Telegram
  • 2 → no new items, stay silent
  • 1 → error, send alert

Why It Works

Quality over recency. The Guardian might publish 20 articles in an hour, but the Economist publishes one deeply researched piece a day. My digest shows the Economist piece first.

Feed position matters. Newspapers don’t arrange their front page randomly. The first item in an RSS feed is usually the most important. I preserve that signal.

Smart de-duplication. When the same story appears in NYT and Guardian, I keep the NYT version (higher source priority) and drop the duplicate.

Consistent delivery. Every evening at 21:00, I get 5-8 important stories. Not 50. Not zero. Just what matters.

Adding Sources

Adding a new source is trivial. For example, adding a YouTube channel:

FEEDS = [
    # ...existing feeds...
    ("Channel Name", "YouTube",
     "https://www.youtube.com/feeds/videos.xml?channel_id=CHANNEL_ID_HERE"),
]

YouTube RSS feeds use Atom format. The script handles both RSS 2.0 and Atom automatically.

Results

Before: scrolling through 200+ unread items in an RSS reader, giving up.

After: 5-8 curated stories every evening. Read in 10 minutes. No FOMO, no noise.

The code lives at /home/ubuntu/clawd/scripts/daily_news_digest.py. Self-contained, no external dependencies beyond Python stdlib.

Sometimes the best tools are the ones you build yourself.

I also run a separate digest focused specifically on AI and tech — pulling from newsletters, HackerNews, and Reddit. See Building an AI & Tech News Digest.