blog-collect: turn an external blog URL into a long-term hosted bilingual article
Our takeEditorial criteria live in the skill rather than in memory, because selection scope is the first thing to drift in a pipeline that runs for months. Two rules earned the hard way are worth copying: a cover must never be an .mp4 (cover_image renders as an img src, and two cards went fully black on 2026-08-25), and a bare angle bracket inside KaTeX must be written as an HTML entity, or the formula truncates there and the rest of the article silently disappears without an error.
Our blog ingestion line, v2.2.0, executed in the digital-twin office by blog-crawler. Every image and video in the original is downloaded and served locally, the body is translated and adapted in full by the agent per article (never a summary), and publishing runs publish_blog.py into POST /api/blog/media and /api/blog/publish, landing in the Postgres articles table that the frontend reads at runtime. Editorial scope is hard-coded into the skill.