The Dark Arts of Skill Engineering

Paul Bakaus, Renaissance Geek, Inc.1:04:53 · Sept 2026 · 13K views
Thumbnail for The Dark Arts of Skill Engineering Watch on YouTube
TL;DR
  1. 1

    Banning popular fonts or visual patterns moves a model toward the next nearby choice instead of producing original work.

  2. 2

    A useful skill extends its harness with subagents, scripts, hooks, memory, browser connections, and model-specific instructions.

  3. 3

    A skill that can be skipped will be skipped, so important checks need gates or hooks that make bypassing them difficult.

Summary

Paul Bakaus explains how Impeccable grew from a short design prompt into a collection of scripts, subagents, hooks, and harness integrations. He argues that prompts cannot pull models far from their usual outputs. Banning a font or color scheme usually sends the model to a nearby alternative in the same cluster. His techniques create more useful variation and better checks. Two blind subagents can separate subjective design direction from deterministic linting. Random seeds can push generation away from safe predictions. Scripts can emit instructions into the model session, while hooks can block bad edits before they are written. Impeccable also stores critique history, routes requests to different internal skills, connects an in-app browser to the agent, and compiles different behavior for different harnesses and models. Bakaus is candid about the cost. Dynamic instructions reduce prompt caching, model behavior varies widely, and taste remains difficult to evaluate automatically.

Key ideas
04:57

Short bans move models around the same cluster

Bakaus began with a 55-line front-end design skill made entirely from prose. It told the model to avoid generic aesthetics, fonts such as Inter, Roboto, and Arial, purple gradients, white backgrounds, and Space Grotesk. The problem was that each ban sent the model toward the next likely choice in its latent space. He connects this to his own history with jQuery UI, whose default orange theme spread because he assumed users would change it. In his view, the model's median output has gravity, and even 250 lines of carefully written prose cannot pull far enough away from it.

06:25

A skill should extend the harness around the model

Bakaus describes prompting as the starting point and harness engineering as the deeper layer. A skill should be treated as an extension of the coding or agent harness, with access to capabilities beyond a packaged prompt. Impeccable uses subagents, scripts, routing, memory, hooks, browser interaction, and model-specific behavior. This reframing changes the design problem. The author is no longer only writing instructions for a model. He is building a system that can create inputs, inspect outputs, enforce rules, and connect the model to tools available in its environment.

07:55

Blind subagents produce a more balanced critique

A model reviewing its own work tends to rate it highly because it is anchored to what it already created. Impeccable's critique command separates the job between two subagents that cannot see each other's work. One acts as a design director and judges hierarchy, visual slop, and heuristics through browser tools. The other runs a deterministic design detector and collects browser evidence. The main thread synthesizes both results. This avoids two opposite errors: treating a beautiful page with mechanical violations as bad, or treating an empty page with no detected violations as good. Bakaus says the same pattern applies to code reviews, security audits, RFC critiques, and output ranking.

16:59

Random inputs can force the model away from safe predictions

Bakaus calls the technique of creating unexpected inputs an anti-attractor. The simplest version asks the model for its top three fonts and then discards all three, shaving away its safest predictions. Another approach generates many ideas and sends them to a fresh subagent for ranking. For his Radiant Shaders library, he used celebrities as creative seeds, asking what Rihanna or Beyonce might look like as a shader. Impeccable uses a script with more than 100 hand-selected primary colors. The model builds a palette around the selected seed, so the same brief can produce substantially different results.

20:38

Routing keeps specialized rules from blurring together

A large general-purpose skill becomes blurry when every instruction is placed in one file. Bakaus gives the example of system fonts. They may be wrong for a landing page but appropriate for a product interface that should feel native. Impeccable routes requests to separate internal skill files such as critique and polish. It also decides whether a brief concerns brand design or product design, then loads different rules for each. He compares this structure with a mixture-of-experts model. The same pattern can support multi-tool skills, context-on-demand behavior, audience-specific behavior, and agent toolkits.

22:47

Skills can accumulate memory across sessions

Skills do not automatically retain long-term memory, but a skill can save files in a repository or skill directory. Impeccable stores critiques in a folder that is normally ignored by Git. A later polish session can read earlier critiques and follow the progression of a page. It can also remember a user's disagreement, such as a preference for a particular font, and respect that preference later. Bakaus uses a similar pattern for multi-session refactors. Each session handles one file and its linked code, while the saved context lets the work continue until the whole codebase is finished.

25:22

Scripts can deliver instructions more reliably than buried prose

For skills distributed across many models, prose instructions are easy to miss, especially in longer files. Impeccable runs a context.mjs script on each invocation. The script combines available product and design context, reports missing files as structured JSON, and can notify the user about an update. Bakaus says instructions emitted through a script's standard output are followed more reliably than equivalent rules buried in the main prose. The tradeoff is caching. Dynamic shell output is not cached in the same way as static skill content, so this approach suits interactive flows better than repeated runs where prompt caching matters most.

29:36

Hooks can stop a bad edit before it reaches the repository

A command that users forget to run cannot enforce a design system. Impeccable ships hooks for Claude Code, Cursor, Codex, and GitHub Copilot. The hooks run a design linter on edits and can report problems such as poor contrast or unwanted animation. For weaker models, Bakaus prefers a pre-tool-use hook that prevents a file from being written, rather than a post-tool-use message that asks the model to fix it afterward. The heavier approach is sometimes necessary because some models do not reliably follow corrective feedback. Hooks also need configurable ignore rules because false positives are unavoidable.

44:04

Skills need harness-specific builds and unavoidable gates

Claude, Codex, Cursor, and other harnesses differ in subagent permissions, user-question tools, background jobs, watchers, and hook behavior. Models also have different habits. Bakaus says Gemini tends to animate images, Codex favors excessive letter spacing and rounded borders, and models vary in how well they follow long instructions. Impeccable therefore creates harness- and model-specific builds. He also designs explicit gates for Codex because it tends to compress or skip instructions. The rule underneath his approach is direct: if a gate can be skipped, the model will skip it. Each gate needs an observable result so the system can tell that it was actually passed.

"The median is the model's gravity. Even 250 lines of artisanal, crafted, beautiful skill prose cannot change this."06:03
Who should watch
  • You are packaging an agent skill for people who use different models or coding harnesses, and you need to understand why a prompt that works locally may fail elsewhere.
  • You are building code review, design review, security audit, or planning workflows where one model should not grade its own work.
  • You want an agent to remember prior sessions, enforce rules during edits, or use a browser and scripts as part of an interactive workflow.