Insight/takes pipeline: fetch, gapfind, digest, drafts
Find a file
Steve Brodson 131645649d Add YouTube video support to the ingest-url share lane
scrape() now classifies youtube.com/youtu.be URLs as a third kind
alongside book/article, fetching the transcript via the existing
youtube_transcripts provider router (ytdlp free-first, transcriptapi
fallback) instead of an HTML scrape, with title/author from YouTube's
public oEmbed endpoint. Writes an insights-only stub under
media/videos/ (or tech/<topic>/ when tech-routed), matching the
article stub's asymmetric storage shape — no transcript body persisted.
2026-08-14 19:21:01 -05:00
config Add tech-folder pipeline support, provider fallback tuning, and digest quality fixes 2026-08-02 09:11:52 -05:00
content_engine Add YouTube video support to the ingest-url share lane 2026-08-14 19:21:01 -05:00
migrations Initial commit: content-engine pipeline (fetch/gapfind/draft) 2026-07-20 09:02:26 -05:00
systemd Add daily insights digest + email delivery; relocate drafts to knowledge repo 2026-07-20 09:17:49 -05:00
.env.example Initial commit: content-engine pipeline (fetch/gapfind/draft) 2026-07-20 09:02:26 -05:00
.gitignore Initial commit: content-engine pipeline (fetch/gapfind/draft) 2026-07-20 09:02:26 -05:00
.python-version Initial commit: content-engine pipeline (fetch/gapfind/draft) 2026-07-20 09:02:26 -05:00
pyproject.toml Stop paid-API cost/coverage leaks in YouTube and Reddit fetch 2026-07-23 09:58:49 -05:00
README.md Add daily insights digest + email delivery; relocate drafts to knowledge repo 2026-07-20 09:17:49 -05:00
STATUS.md Initial commit: content-engine pipeline (fetch/gapfind/draft) 2026-07-20 09:02:26 -05:00
uv.lock Stop paid-API cost/coverage leaks in YouTube and Reddit fetch 2026-07-23 09:58:49 -05:00

content-engine

Personal content-intelligence pipeline. Pulls from seven sources (Hacker News, Reddit, blogs/RSS, YouTube, Substack, Medium, X/Twitter), extracts insights via LLM, finds cross-source content gaps daily, and drafts social media content ideas weekly for human review.

Reddit runs via scrapebadger (paid), not Reddit's official API — Reddit's Responsible Builder Policy closed self-serve app creation behind a manual, multi-week approval queue in late 2025. Reddit also runs on a reduced cadence (2x/week, not daily) since it's credit-billed — see run_days in config/accounts.yaml, a per-source scheduling mechanism any source can use.

Modeled on the insight-extraction / gap-finding / script-writing methodology of a reference app ("Harper," YouTube-only), adapted to be source-agnostic and standalone.

Setup

cd /home/steve/Documents/code/content-engine
uv sync
cp .env.example .env   # fill in secrets, then: chmod 600 .env
cp config/accounts.example.yaml config/accounts.yaml   # edit tracked accounts/keywords

LLM + embedding auth (TOGETHER_API_KEY, OLLAMA_API_KEY) come from ~/.config/ai-memory.env — already present on this machine, not duplicated into this project's .env.

Usage

uv run content-engine fetch          # pull all sources, extract insights
uv run content-engine gapfind        # cross-source gap analysis, write gbrain gaps digest
uv run content-engine daily-digest   # daily insights + AI takes, write gbrain page, email + Telegram
uv run content-engine draft          # generate drafts, write to knowledge repo, email + Telegram

daily-digest is the daily analytics feed: it clusters the day's insights into themes with an LLM-generated "take" each, writes adversarialminds/daily/<date>-insights.md into the gbrain knowledge repo, and emails it (Nextcloud URL + gbrain slug) to steven@brodson.com. --dry-run writes the page but skips commit/email/Telegram.

draft (weekly) writes AdversarialMinds drafts into the knowledge repo at adversarialminds/drafts/*.md (gbrain-managed, committed, Nextcloud- synced) and emails links to them; non-default streams (e.g. personal) still land in the local review/<stream>/. Move a draft into approved/ once you're happy with it — that's the entire approval mechanism, no UI/API. For both emails, Telegram only pings that the email was sent.

Scheduling

Four systemd user timer/service pairs in ~/.config/systemd/user/: content-engine-fetch (daily 05:00), content-engine-gapfind (daily 07:00), content-engine-insights (daily 07:30), content-engine-draft (weekly, Sunday 09:30). See systemd/README.md for install instructions.

Architecture

See /home/steve/.claude/plans/nifty-munching-pixel.md for the full design doc. Short version: sources/ (one module per platform, common Protocol) → providers/ (scraping/API clients, selected via a config-driven cost-preference router, decoupled from sources) → storage/ (LanceDB, embedded) → analysis/ (LLM-driven insight extraction, gap-finding, draft generation) → gbrain_export/ (daily digest to the personal knowledge base) + notify/ (Telegram).