OPEN SOURCE DEEP DIVE
Diagram Design: an editorial diagram skill for coding agents
An MIT-licensed agent skill that turns Claude Code, Codex, Factory Droid or Pi into a diagram design tool: 44 visual types, 211 self-contained HTML examples, semantic patterns orthogonal to layout, a first-run gate that forces brand-token verification, and a render lint that catches clipped SVG by pixel diff.
Positioning in one line
diagram-design is a diagram design skill for coding agents. Install it into Claude Code, Codex, Factory Droid or Pi and the agent starts producing self-contained HTML with inline SVG instead of the usual grey rounded boxes joined by lines. It bills itself as "editorial diagrams your designer won't hate" and currently sits at 44,982 stars and 2,889 forks under an MIT licence. The product proposition is a single sentence: it pushes "let an agent draw a diagram" from "it works" to "you can put it straight on your website".
The numbers
| Field | Value |
|---|---|
| Repository | cathrynlavery/diagram-design (MIT, primary language HTML) |
| Stars / Forks / Watchers | 44,982 / 2,889 / 128 |
| Created / last commit | 2026-04-16 / 2026-10-08, still actively updated |
| Visual types | 44 — references/type-*.md contains exactly 44 files, one reference per type |
| Bundled examples | 211 HTML files in assets/, three static variants per type (minimal light, minimal dark, full-editorial), openable directly in a browser with no build step, no JavaScript, and no external image dependency |
| Reference docs | 62 files in references/: 44 type documents plus semantic patterns, style guide, import/export, primitive specs and the layout budget |
| Scripts | 174 files in scripts/, covering skin lint, render lint, thumbnail generation and plugin packaging |
| Agent hosts | Claude Code (.claude-plugin), Codex (.codex-plugin), Factory Droid (.factory-plugin), Pi, and any Agent Skills-compatible host |
What problem it actually solves
The motivation is stated plainly. Cathryn Lavery writes at littlemight.com and runs BestSelf.co on the side, and every time she needed a diagram — an architecture sketch, a flowchart, a pyramid of what matters most — she would ask Claude and get back the same generic rounded-box thing, which looked nothing like the rest of the site. Her options were to fight Figma for thirty minutes or to skip the diagram.
Hence the skill. One sentence at the top of the README states its whole design philosophy: "the highest-quality move is usually deletion."Expanded into four hard rules:
- Every node represents a distinct idea — two nodes that always travel together are one node;
- Every connection carries information — if the relationship is already obvious from the layout, remove the line;
- Coral (the accent) is editorial, not a flag — one or two focal nodes per diagram, and using it on five erases the signal;
- Target density 4/10 — technically complete, but not so dense that it needs a guide. Above nine nodes it is probably two diagrams.
There is also a line near the top of the README worth quoting on its own: "No Figma. No generic rounded boxes. No 30-minute color-picking sessions."Those three refusals define the scope precisely. It does not replace a design tool; it replaces "getting a general-purpose LLM to draw something presentable without brand context".
Core design one: semantic patterns orthogonal to visual types
This is the most valuable architectural decision in the project, and the README puts it plainly: semantic patterns describe what a system does; the 44 visual types describe how information is arranged.
The split is necessary because behaviour, state, enforcement and risk are what a diagram is actually about, and all of them are orthogonal to whether you draw a swimlane or a layer stack. Some real routing examples:
| What the reader must understand | Semantic pattern | Nearest visual type |
|---|---|---|
| Many arrivals competing for finite service capacity | Fan-in queue / bottleneck | Data flow |
| Why two policy decisions differ and where they first diverge | Paired policy-evaluation traces | Flowchart |
| Which routes cross a trust boundary and which are blocked | Secure paved road | Architecture |
| Which controls apply at each enforcement surface | Governance / control catalog | Layer stack |
| How defenses reduce risk and what risk remains | Compensating security layers | Layer stack |
| Which sub-elements a system decomposes into, each traceable to implementation | Traceable block decomposition | Tree |
| How one subject advances through phases, waits, retries, cancels | Lifecycle phase map | State Machine |
What this buys is practical: extending behaviour no longer requires adding a type. A queue, a policy trace or a trust boundary can borrow the nearest existing layout grammar instead of forcing a new entry into the type table. The README states the intent directly — semantic patterns describe behaviour separately from layout, so a queue, policy trace or trust boundary can use the nearest existing type without expanding the type count.
The pattern documents themselves are rigorous. Each defines five things: selection triggers (when to reach for it), required primitives (what the drawing must contain), a complexity budget (hard caps — the fan-in queue allows at most five sources, five queue slots, one bottleneck, two outcomes and nine primary nodes), anti-patterns (explicitly forbidden drawings, such as "equal-width pipeline that hides contention", "arrows merged before they can be traced", "capacity implied only by box size"), and a static fallback (what a still frame must show if animation is unavailable).
That last one is routinely missed by comparable projects. If the default output is static HTML, then any information that only makes sense while animating is a design failure, and the pattern docs require that labels and outcomes stay complete in a static frame.
Core design two: style derived from your website, behind a hard gate
references/style-guide.md is explicitly the single source of truth for colour, typography and tokens. Every diagram draws from it, and type documents say accent, never #eb6c36. The default skin is a cool editorial palette: white-smoke paper #f5f5f5, jet-black ink #2d3142, atomic-tangerine accent #eb6c36, blue-slate muted #4f5d75, plus silver #bfc0c0 — which is precisely the author's own site palette.
Tokens are defined by semantic role, not by value: paper, paper-2, ink, ink-strong, muted, soft, rule, rule-solid, accent, accent-tint, link. That yields one very useful rule: inverting light to dark, any rgba(28,25,23, X) becomes rgba(250,247,242, X) with identical opacity and flipped RGB, while the accent shifts slightly brighter in hue so it reads on dark paper.
The more remarkable piece is the first-run style gate. Section 0 of SKILL.md states that before generating the first diagram in a new project, the agent must verify the style guide has been customised. If the tokens are still the shipped defaults (paper #f5f5f5, ink #2d3142, accent #eb6c36), the agent must stop and ask, offering a website URL, an installed skill, a local design-system folder, pasted tokens, keeping the default, or loading a saved profile. The wording is "do not silently ship default-skinned diagrams into a branded project."
That "stop and ask" step matters more than any later colour change, because a diagram embedded in a page is hard to redo wholesale. references/onboarding.md handles deriving a palette from a website URL, and references/profiles.md handles saving the result as a named profile so subsequent runs skip the gate — a project .diagram-design marker or a leading profile header activates it automatically.
Core design three: import means redraw, not convert
The skill can ingest draw.io, Mermaid and Excalidraw sources, but SKILL.md states the handling hard: "redraw — never convert."Coordinates, colours, fonts and shape quirks of the source or renderer are discarded; what you keep is the content: components, relationships, grouping, direction. It then requires a fidelity ledger reporting what was merged, collapsed or dropped.
Two constraints are worth remembering: an import is bounded by its source, so never invent a component to fill a layout and never silently drop one; and because "never convert" means no coordinates are inherited, Mermaid slop cannot infect the output. The repository description's "No Mermaid slop" is meant literally.
Four dials must be set before drawing an import:
| Dial | Options | Default |
|---|---|---|
| Format | html · svg · png · html+png | html |
| Size | doc-inline · doc-wide · slide-16x9 · slide-4x3 · social-og · social-square · print-a4/a3/letter-landscape · fit | doc-inline |
| Detail | faithful (≤24 nodes, zoned) · balanced (≤12) · simplified (≤7) | balanced |
| Audience | engineer · mixed · executive — governs wording, not node count | mixed |
The size preset sets the viewBox and the type ramp, which is the step most comparable tools skip: the same diagram placed inline in a document and on a 16:9 slide must change its font sizes. The rules also note that faithful is the only detail level exempt from the nine-node budget (zoned above nine nodes, split above 24), while the connector rules never relax.
Core design four: two layers of lint, with the browser as the only judge
Among the 174 files in scripts/, the two lints are the most valuable engineering, and they catch categorically different bugs.
lint-skin.py reads the HTML and checks source-level problems: inlined hex values that should be semantic tokens, rgb(0,0,0) pure black, src attributes pointing at external http(s) images (which violates the self-contained principle), and <script> tags, since "no JavaScript" is a core promise of the project. It also has a baseline mechanism — pre-2.0 examples may legitimately fail because they were built against an older skin, and --all --baseline skips those documented legacy files.
lint-render.py checks the rendered result, catching breakage the source cannot show. Its own docstring explains why it must work this way, and that reasoning is itself the interesting part. Chromium's getBoundingClientRect() on an SVG child reports geometry, not paint: it excludes stroke width, markers and filter bleed (so a 40px stroke spilling past the viewport measures as inside), and it ignores clip-path, opacity: 0 ancestors and overflow: visible (so safe content measures as outside). Both directions of that error were reproduced in Chromium, so no geometric model is used.
Instead the browser is the oracle: screenshot the viewport as authored, screenshot it again with overflow released, and diff the two — new ink means paint was being cut off. Ink is ink, so strokes, markers and filter bleed all count, while clip-paths, invisible content and already-visible content produce no new ink. The clipping check is also staged, because a diagram can be clipped at more than one level and each release repaints its own box.
Refusing to trust a geometry API and believing only a pixel diff is the best way I've seen to turn "will this self-contained SVG get clipped in someone else's environment" into a reproducible CI check.
Core design five: motion is an enhancement that must be finished in a static frame
Version 2.3 added motion, but it is tightly bounded. references/animation.md opens with the constraint: "animation explains a complete static diagram; it never supplies missing meaning."Load it only when motion is explicitly requested or would materially clarify order, accumulation, evaluation, containment or propagation — otherwise ship static HTML with mode none.
There are four modes, and one restriction matters most: only loop may repeat. Queue state, typing, field values, policy outcomes, containment and audit entries use reveal or step and must finish complete. Using motion to show order is fine; using it to imply "more is coming" is not.
The static-first enhancement contract has eight clauses, of which the first three carry the weight. ① The source is complete — every semantic node, label, connector, status and outcome is visible in the HTML/SVG before enhancement, and only selectors below .motion-ready may hide or transform them. ② Stable capture — data-frame="static", ?motion=static, print, no-JS and standalone SVG export all expose the complete frame with controls hidden, and capture must not happen after an arbitrary delay. ③ CSS owns presentation — appearance and travel use CSS transitions and keyframes; minimal inline JavaScript may only bind explicit controls, update step and state attributes, schedule deterministic steps and update a dedicated live-status region, with no fetches, no markup injection, no path measurement, and no mutation of semantic labels or values.
The rest are equally concrete: one clock with --motion-fast: 160ms, --motion-step: 480ms, --motion-hold: 720ms and --motion-total ≤ 8000ms, delays derived from integer steps, no randomness, no springs, no transition-event timing; integer steps 1–8 with at most two items entering per step; controls scoped to the nearest [data-motion-root] so IDs, timers, live regions and step state never cross figure boundaries; and finally JavaScript adds .motion-ready only after controls are bound and the first render succeeds, so a script error before that point leaves the complete source visible.
What unites these clauses is that motion is designed as an enhancement layer that may fail, not as content. Violate any one of them and the diagram breaks in print, in a screenshot, without JavaScript, or on export.
Core design six: the primitive layer and layout discipline
Alongside the 44 type documents sit six primitive specifications that form the shared vocabulary of every type.
A 4px grid is mandated for structural geometry: node origins, widths, heights, gaps and padding must be divisible by four, with allowed value tables — node width and height limited to {80, 96, 112, 120, 128, 140, 144, 160, 180, 200, 240, 320}, gaps {20, 24, 32, 40, 48}, inner padding {8, 12, 16}, radii {4, 6, 8}. Crucially it also enumerates what is deliberately off-grid: type sizes, radii, data-derived positions, text baselines, arrow markers, the .5 offsets that keep 1px strokes crisp, stroke widths, opacity and the 22×22 dot pattern.
That explicit split between "must align" and "must not align" is far more useful than a general instruction to align to a grid, because otherwise an agent snaps text baselines to the grid and makes the layout dirtier.
Complexity budgets tighten in layers: the universal limits in SKILL.md (nine nodes, twelve arrows or transitions, two coral elements, two annotation callouts) apply to every type; type documents may be stricter but never looser; and a semantic pattern's own budget can only tighten further. Architecture delta caps out at eight unique components, ten relationships and eight ledger entries; sequence diagrams allow at most five lifelines.
Two optional primitives are worth noting. The sketchy filter is an SVG displacement filter that wobbles every stroke slightly, turning any minimal variant into a hand-drawn register without touching layout — the grammar given is feTurbulence fractalNoise baseFrequency 0.02 numOctaves 2 seed 4 feeding feDisplacementMap scale 1.5, and it is marked as being for diagrams that accompany an essay rather than technical docs. The terminal window is a separate skin that wraps any diagram in a fake terminal — three-dot titlebar, a $ prompt line, monospace throughout — for developer-tool announcements, CLI product posts and technical social cards. It does not inherit brand tokens and does not participate in the light/dark inversion: every terminal example uses the same nine tokens regardless of the host site's brand.
Export is manual only, never run unprompted, in all caps in references/export.md. It loads only when you invoke the /diagram-design:export-diagram slash command or explicitly ask in natural language for SVG or PNG, and both entry points run the same procedure.
Engineering scale and activity
635 files, 22 MB. The skill itself is a 30 KB SKILL.md plus 62 reference documents, a layout document per type across 44 types, and 211 prebuilt example HTML files — each type in three static variants, openable directly in a browser with no build step, no JavaScript and no external image dependency. Host adaptation runs through four plugin manifests: .claude-plugin/, .codex-plugin/, .factory-plugin/ and .agents/plugins/, with scripts/build-openai-plugin-zip.py handling the OpenAI side.
The release cadence is dense, and the README carries recent changes at the top: 2.0 introduced the Loop, a flywheel with a shared-memory hub where dashed lines are write-backs; 2.3 added semantic system patterns and optional accessible motion, with static output still the default; 2.5.10 added ten more layout grammars — Sankey, fishbone, Wardley map, kanban, user journey, deployment, dependency graph, UML class, story map and database schema. The current SKILL.md metadata reads version: "2.6". The repository also carries .maintainer-policy.json, a 31 KB CONTRIBUTING.md, SECURITY.md, PRIVACY.md, THIRD_PARTY_LICENSES.md, and consistency-check scripts such as plugin_version_history.py.
Who it suits, who it does not
Good fit: people writing technical blogs or documentation who need a steady supply of architecture diagrams, flowcharts and sequence diagrams that can go straight onto a page; anyone who wants the agent to emit self-contained files rather than depending on external images and CDNs; people already working in Claude Code, Codex, Factory Droid or Pi who are willing to configure brand tokens once; and teams who need a heavily specified diagram standard to unify the look of their docs.
Poor fit: people who want free-form drag editing (it produces static files, not an editor); anyone needing real-time collaboration or interactive diagrams (motion is an optional explanatory aid, the default output is static); people after 3D or photoreal visuals (it is explicitly against shadows, against skeuomorphism, against the Mermaid look); and users unwilling to spend five minutes configuring brand tokens once — the gate will stop them, and that is the intended design rather than a bug.
Conclusion
diagram-design's real contribution is not "44 diagram templates" but installing a design discipline into an agent: density 4/10, at most one or two accent elements, no shadows, no external assets, and delete any node that can be removed. That discipline is executable because the author also supplied the orthogonal semantic-pattern/visual-type structure (extending behaviour does not extend the type count), a semantic-role token system (re-skinning touches one file), a style gate that stops and asks (so default skins never contaminate a branded project), and a pixel-diff render lint (so self-contained SVG does not get quietly clipped elsewhere).
In other words, it turns "get an LLM to draw" from a matter of luck into a production line with acceptance criteria. That is arguably the higher form for any agent skill to reach: the deliverable is not just the artefact, but the evidence chain for why the artefact should be trusted.