# Training Taste

Thais Castello Branco, Taste Labs | AI Engineer World's Fair 2026 | 15:06

Source: https://www.youtube.com/watch?v=sDMGWK4wZ_w
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/training-taste
Published: 2026-09-10
Tags: design, evals, structured-outputs

## TL;DR
- AI accelerated an internet-wide convergence that had already begun, making different websites repeat the same patterns without respecting their context.
- Slop has three recurring signatures: repetition, lack of fit, and low intent.
- Small classifiers called probes can measure combinations of design characteristics that predict slop, while inference-time systems can help agents produce more fitting and distinctive work.

## Summary
Thais Castello Branco argues that AI did not create design homogenization, but it has made it faster and less sensitive to context. Taste Labs studies how to make design judgment measurable so models and applications can produce better work. She separates relatively objective tasks, such as choosing palettes, contrast, and alignment, from aesthetics, where experts disagree and data becomes more useful. Her team analyzed more than two million websites, mined design features, and trained small classifiers called probes. Combining the probes predicted slop better than asking an LLM to judge quality directly. Castello Branco also argues that inference time deserves as much attention as model training because user intent and context are resolved there. Her proposed systems include deliberate, rule-aware divergence for creativity and a brand API that turns a brand into structured guidance an agent can follow and be judged against.

## Key ideas
### Taste Labs treats design as a domain that can be broken into solvable parts
[00:01](https://www.youtube.com/watch?v=sDMGWK4wZ_w&t=1s)
Taste Labs works with frontier labs to evaluate models, find where they fail, and build post-training data or reinforcement-learning environments. Castello Branco says design becomes more tractable when divided into smaller problems. Color palettes, contrast, and alignment can become nearly deterministic when the context is specific enough, with most experts agreeing on the answer. Aesthetics is different because experts naturally disagree, so the team relies more on data. Taste Labs also works at the application layer, where off-the-shelf models tend to collapse toward average styles and need added context, judgment, and verification.

### Good work feels distinctive, cared for, and authentic
[02:18](https://www.youtube.com/watch?v=sDMGWK4wZ_w&t=138s)
Castello Branco finds greatness harder to define than correctness in math. A poem, coffee shop, or website can feel special because it is unique, catches attention, shows care and craft, and has a sense of authenticity. AI often misses these qualities because it produces average outputs. She describes the desired alternative as work that is deliberately out of distribution. Slop is easier to define because repetition and soullessness are widely recognizable, even when people disagree about what counts as great.

### The low cost of generation exposes a gap in human judgment
[03:41](https://www.youtube.com/watch?v=sDMGWK4wZ_w&t=221s)
AI lets almost anyone create a PowerPoint, website, or web app with a click, while the cost of generation approaches zero. Taste does not appear as quickly. Designers build judgment over years by seeing many examples, spotting patterns, forming a point of view, learning restraint, and sometimes breaking norms deliberately. Castello Branco rejects the idea that the solution is for everyone to acquire taste in every design domain. The more practical goal is to make it easier for ordinary users to create work that fits their own preferences.

### Slop repeats patterns without fitting the situation
[05:00](https://www.youtube.com/watch?v=sDMGWK4wZ_w&t=300s)
Castello Branco names three recurring signatures of slop. Repetition means seeing the same design patterns many times. Lack of fit means the work does not feel correct for its context, time, or audience. A pet shop and a finance firm should not converge on the same website if they were made with care. Low intent includes rushed prompting and the system's failure to interpret what the user is really trying to create. Better systems need to add context and color to that intent.

### The internet was homogenizing before AI, and AI made the pattern context blind
[06:28](https://www.youtube.com/watch?v=sDMGWK4wZ_w&t=388s)
Taste Labs analyzed more than two million websites from roughly the previous decade, along with synthetically generated sites. The research found that the internet had already become more homogeneous before AI, with color palettes and layouts becoming more similar as trends spread faster. AI accelerated the repetition and made it less dependent on context. Similar patterns appeared across very different categories, which made the convergence more noticeable.

### Probe combinations can predict slop better than a general model judge
[07:48](https://www.youtube.com/watch?v=sDMGWK4wZ_w&t=468s)
The team mined site data for structured characteristics such as colors, typography, layout, and audience. It then trained probes, which Castello Branco describes as small or baby classifiers, with each probe detecting one characteristic. When several probes appeared together at high frequency, their combination made slop more predictable. She says this approach performed better than most methods that ask an LLM to judge whether something has human quality or is AI-generated slop. The result supports treating slop as a measurable pattern rather than only a vague impression.

### Inference time is where systems recover intent and context
[08:55](https://www.youtube.com/watch?v=sDMGWK4wZ_w&t=535s)
Castello Branco says judgment becomes more valuable as production gets cheaper. Model improvements still matter, but she gives equal or greater weight to inference time, when the system interacts with the user and interprets context and intent. A model can be capable in general while an application still produces generic work because it lacks the right exchange with the user. Solving slop therefore requires application systems that can understand the request, check the result, and keep the output aligned.

### Creativity should break selected rules while keeping the category's expectations
[10:10](https://www.youtube.com/watch?v=sDMGWK4wZ_w&t=610s)
Taste Labs is exploring what Castello Branco calls a creativity API, an inspiration system that helps an agent produce outputs outside the usual average. She says this should not mean simply increasing temperature and hoping for a good result. A good pitch deck, for example, has expectations that should remain intact. Creativity can come from deliberately breaking a few rules while still making the result feel appropriate to the category. The system needs to know which expectations it is departing from.

### A structured brand system gives agents something to follow and people something to grade
[11:51](https://www.youtube.com/watch?v=sDMGWK4wZ_w&t=711s)
Taste Labs' first public product is a brand API, which was in beta with design partners at the time of the talk. It takes a brand URL and extracts specific components that an agent can follow. The same structure lets a person check whether the agent is staying on brand and see where it fails. In Castello Branco's example, applying the extracted system to a slide deck made the generated result much closer to the original brand. Taste Labs is also building a repository of pre-created brand systems for users who do not already have one.

## Notable quotes
- "Our whole mission is basically how do we end AI slop? That's my personal enemy." (00:01)
- "I think how do we understand this better so that we can make even for the average person the ability to create something great and to understand maybe their own taste easy." (04:29)
- "So A, repetition. So you start seeing the same thing many many many times." (05:00)
- "I don't even want to use the word taste here. Is judgment." (08:50)
- "It's not just about like turning up the temperature of the model and kind of fingers crossed hoping for the best." (10:41)

## Tools & references mentioned
- Taste Labs
- General Intelligence Company of New York
- Claude Design

## Who should watch
- You are building an AI product that generates websites, presentations, or other design work and keeps getting generic outputs.
- You want a practical way to evaluate visual quality beyond asking a general-purpose model for a subjective score.
- Your application needs to preserve a brand's style or create deliberate variation without losing the expectations of its category.

## Related talks

- [Ending AI Slop](https://aietalks.com/talks/ending-ai-slop) (Thais Castello Branco, Taste Labs, 16:30)
- [The Missing Layer: Design Taste in AI Agents](https://aietalks.com/talks/the-missing-layer-design-taste-in-ai-agents) (Hassan El Mghari, Together AI, 14:10)
- [Developing Taste in Coding Agents: Applied Meta Neuro-Symbolic RL](https://aietalks.com/talks/developing-taste-in-coding-agents-applied-meta-neuro-symbolic-rl) (Ahmad Awais, CommandCode, 20:52)
- [Design at the Speed of Adjectives](https://aietalks.com/talks/design-at-the-speed-of-adjectives) (Paul Bakaus, Renaissance Geek, Inc., 15:59)
- [No More Slop](https://aietalks.com/talks/no-more-slop) (swyx, AI Engineer, 09:15)
