Keeping up with AI and tech news had become a full-time job. Superhuman AI newsletters, TLDR digests, HackerNews front page, Reddit ML discussions — easily an hour a day just staying current. Information overload disguised as staying informed.

I needed a system that would curate the signal from the noise, deliver it on my schedule, and stop the constant context-switching between platforms.

The Problem

Tech news comes from everywhere:

  • Email newsletters (Superhuman, TLDR) — great curation, but buried in inbox
  • HackerNews — high signal, but requires active browsing
  • Reddit (r/MachineLearning, r/artificial) — community-driven, hit or miss
  • Twitter/X — real-time, but drowns in noise

Each source has value, but checking all of them daily is unsustainable. I needed a weekly digest that:

  • Pulls from newsletters already in my inbox
  • Scrapes HackerNews trending stories
  • Monitors relevant Reddit threads
  • Filters for AI/tech relevance
  • Delivers in one consolidated report

Building with OpenClaw

Just like my RSS news digest, I didn’t write this from scratch. I explained the problem to OpenClaw, and it built the solution.

My requirements:

  • “Extract AI headlines from Superhuman and TLDR newsletters”
  • “Fetch trending AI stories from HackerNews (>80 points)”
  • “Pull top posts from r/MachineLearning and r/artificial”
  • “Include direct links to sources”
  • “Run weekly, not daily”

OpenClaw turned that into weekly_ai_digest.py — a script that orchestrates multiple data sources, extracts AI-relevant content, and formats it for Telegram delivery.

How It Works

1. Fetching Newsletters from Gmail

The script uses the gog CLI (Google Workspace tool) to search my inbox for recent newsletters:

def fetch_newsletters():
    newsletters = []
    
    for sender in ["superhuman", "tldr"]:
        cmd = [
            "gog", "gmail", "messages", "search",
            f"{sender} newer_than:7d",
            "--max", "10",
            "--account", "YOUR_EMAIL"
        ]
        
        result = subprocess.run(cmd, capture_output=True, text=True)
        # Parse message IDs for extraction

This finds emails from the last 7 days. No API tokens, no OAuth dance — just piping Gmail CLI output.

2. Extracting AI Content with Pattern Matching

Newsletters have different formats. The script handles both:

TLDR-style (numbered items with reference links):

ANTHROPIC CLOSES $30 BILLION... (5 MINUTE READ) [6]
...
[6] https://techcrunch.com/...

Superhuman-style (markdown with embedded links):

**1. OpenAI launches GPT-5:** The new model... [link](https://...)

Pattern matching extracts headlines and URLs:

# Build reference link map
link_map = {}
link_refs = re.findall(r'^\s*\[(\d+)\]\s+(https?://\S+)', content, re.MULTILINE)
for ref_id, url in link_refs:
    link_map[ref_id] = url

# Match headlines with AI keywords
tldr_pattern = r'^\s*([A-Z][A-Z0-9\s\'\-]{30,150})\s*(?:\([^)]+\))?\s*(?:\[(\d+)\])?$'
for match in re.finditer(tldr_pattern, content, re.MULTILINE):
    headline = match.group(1).strip()
    link_ref = match.group(2)
    
    if any(kw in headline.lower() for kw in AI_KEYWORDS):
        ai_items.append({
            "title": headline.title(),
            "url": link_map.get(link_ref)
        })

This handles formatting inconsistencies while preserving links.

3. Scraping HackerNews API

HackerNews provides a public Firebase API. The script:

  1. Fetches top 50 story IDs
  2. Checks each story for AI keywords
  3. Filters by score threshold (>80 points)
def fetch_hackernews_trending():
    items = []
    
    # Fetch top stories
    result = subprocess.run(['curl', '-s', 
        'https://hacker-news.firebaseio.com/v0/topstories.json'],
        capture_output=True, text=True)
    story_ids = json.loads(result.stdout)[:50]
    
    ai_keywords = ['openai', 'anthropic', 'gpt', 'llm', 
                   'machine learning', 'transformer', ...]
    
    for story_id in story_ids:
        # Fetch story details
        story = fetch_hn_story(story_id)
        title = story.get('title', '')
        
        if any(kw in title.lower() for kw in ai_keywords):
            score = story.get('score', 0)
            if score > 80:
                items.append({
                    "title": title,
                    "url": story.get('url'),
                    "score": score
                })

This surfaces genuinely trending stories, not just recent ones.

Using Reddit’s JSON API (no auth required):

def fetch_reddit_trending():
    items = []
    subreddits = ["MachineLearning", "artificial", "singularity"]
    
    for sub in subreddits:
        result = subprocess.run([
            'curl', '-s', '-L',
            '-H', 'User-Agent: Mozilla/5.0',
            f'https://www.reddit.com/r/{sub}/top.json?t=week&limit=5'
        ], capture_output=True, text=True)
        
        data = json.loads(result.stdout)
        for post in data['data']['children'][:2]:
            if post['data']['score'] > 50:
                items.append({
                    "title": post['data']['title'],
                    "url": f"https://reddit.com{post['data']['permalink']}",
                    "score": post['data']['score']
                })

Only posts with >50 upvotes make the cut.

5. Building the Digest

The final output is formatted for Telegram:

def build_digest(newsletters_data, reddit_items, hn_items):
    digest = "**🤖 AI Digest**\n\n"
    
    # Section 1: Major News from Newsletters
    digest += "**📰 Major News & Breakthroughs**\n"
    for item in major_news[:5]:
        digest += f"• [{item['title']}]({item['url']})\n"
    
    # Section 2: HackerNews Trending
    digest += "\n**🔥 Trending on HackerNews**\n"
    for item in hn_items[:5]:
        digest += f"• [{item['title']}]({item['url']}) ({item['score']} pts)\n"
    
    # Section 3: Reddit Top Posts
    digest += "\n**💬 Trending on Reddit**\n"
    for item in reddit_items[:5]:
        digest += f"• [{item['title']}]({item['url']}) ({item['score']} ↑)\n"
    
    digest += "\n_Sources: Superhuman, TLDR, HackerNews, Reddit_"
    return digest

Example output:

🤖 AI Digest

📰 Major News & Breakthroughs
• [New laundry-folding robot costs as much as a used car](https://theverge.com/...)
• [OpenAI debuts Codex-Spark, a speedier coding model](https://openai.com/...)

🔥 Trending on HackerNews
• [Two different tricks for fast LLM inference](https://seangoedecke.com/...) (127 pts)
• [OpenAI should build Slack](https://latent.space/...) (229 pts)

💬 Trending on Reddit
• [New SOTA for code generation benchmarks](https://reddit.com/...) (312 ↑)

Sources: Superhuman, TLDR, HackerNews, Reddit

Automation

The script runs every 3 days via OpenClaw cron:

{
  "schedule": {
    "kind": "every",
    "count": 3,
    "unit": "days"
  },
  "delivery": {
    "to": "-1003743850957/32",  // Telegram topic
    "silentSuccess": false
  }
}

State tracking prevents duplicate runs:

state = load_state()
now_ms = int(datetime.now().timestamp() * 1000)
last_run_ms = state.get("lastRunMs", 0)

if now_ms - last_run_ms < (6 * 24 * 60 * 60 * 1000):
    print("Already ran this week. Use --ignore-state to force.")
    sys.exit(2)

Exit codes:

  • 0 → success, post digest
  • 2 → already ran this week, stay silent
  • 1 → error, send alert

This exit-code pattern is what makes cron jobs well-behaved — silence when there’s nothing to say, signal only when needed. The failure modes of getting this wrong are covered in The Silent Killer in AI Automation.

Why It Works

Consolidation over context-switching. Instead of checking 5 platforms, I get one digest every 3 days.

Quality signals over volume. HackerNews stories need >80 points. Reddit posts need >50 upvotes. Newsletters are pre-curated by humans.

Links included. No summaries, no AI-generated text. Just headlines and direct links to sources.

Weekly cadence. AI news moves fast, but not that fast. A 3-day digest keeps me current without drowning me.

Edge Cases Handled

Malformed emails: The script handles both TLDR’s numbered-reference format and Superhuman’s markdown-with-embeds.

Reddit blocking: Falls back gracefully if Reddit rate-limits the JSON API.

HackerNews spam: Keyword filtering prevents false positives (“AI-powered marketing tools” gets filtered out).

Duplicate stories: If the same OpenAI announcement appears in newsletters, HN, and Reddit — it’s shown once, under “Major News.”

Results

Before: 60+ minutes a day scattered across platforms, constant FOMO.

After: 15 minutes every 3 days, consolidated view of what actually matters.

The code lives at /home/ubuntu/clawd/scripts/weekly_ai_digest.py. Self-contained, no external dependencies beyond gog CLI and Python stdlib.

What’s Missing: Twitter/X

You might notice Twitter/X isn’t in the digest. Here’s why:

The official Twitter API costs $100/month for basic access. That’s absurd for a personal automation project.

There’s a workaround — Nitter (community-run Twitter mirrors with RSS feeds) — but it’s unreliable. Instances go down, rate limits are aggressive, and Twitter actively fights scrapers.

So for now, Twitter stays manual. If @OpenAI or @AnthropicAI announce something major, it’ll show up in HackerNews within hours anyway.

Future Improvements

Possible next steps:

  • Weight stories by source overlap — if HN + Reddit + newsletter all cover the same thing → probably important
  • Archive past digests for long-term reference
  • Nitter RSS fallback — worth trying if Twitter becomes critical and Nitter stabilizes

But for now? It works. And working is better than perfect.


Want to build something similar? Check out OpenClaw — the AI assistant that built this entire system from a conversation.