# Image generation beyond the prompt

A pack of 8 talks from the AI Engineer YouTube channel, in the order to watch them. 5h 05m of video.
Page: https://aietalks.com/packs/image-generation

Image generators are impressive until you need the same character twice, a product to keep its shape, or an edit to work for millions of people. The model mechanics come first, from FLUX to ComfyUI's denoising graph, where the pipeline can be inspected and changed. Better control builds on that view: GPT-4 with vision compares an image with the prompt, then revises the next DALL-E 3 attempt, while a Wind in the Willows project carries saved character references through a whole book. Google Photos shows what changes when a creative demo becomes Magic Editor, with narrow jobs, several candidate outputs and maintained benchmarks. Those benchmarks face an uncomfortable test in the final panel: people often reward the sharper, more saturated image even when it is less realistic or useful.

## This pack is for you if

- Your image demo produces striking one-offs, but the same character or product changes between generations.
- You are choosing between an image API, an editable node graph and a narrow workflow for ordinary users.
- Your team can rank attractive samples but cannot say whether an edit preserved the subject or solved the user's job.

## The talks, in order

### 1. The State of Generative Media

Gorkem Yurtseven, FAL | 17:14 | AI Engineer World's Fair 2025
Video: https://www.youtube.com/watch?v=P370D8Kmlkw
Summary: https://aietalks.com/talks/the-state-of-generative-media.md

Why first: Yurtseven gives the rest of the pack a timeline. DALL-E 2's early lead gave way to Midjourney, Stable Diffusion and Flux, while cheap generation found work in advertising, e-commerce and image editing. Dieleman can now open up the systems whose spread Yurtseven describes.

### 2. Building Generative Image & Video Models at Scale

Sander Dieleman, Google DeepMind | 40:46 | AI Engineer Europe 2026
Video: https://www.youtube.com/watch?v=xOP1PM8fwnk
Summary: https://aietalks.com/talks/building-generative-image-video-models-at-scale.md

Why second: Yurtseven names the models and markets. Dieleman opens the machine, showing why visual generators compress images into learned latents and recover them through repeated denoising. Batifol follows with one model family built from those ideas.

### 3. FLUX, Open Research, and the Future of Visual AI

Stephen Batifol, Black Forest Labs | 22:32 | AI Engineer Europe 2026
Video: https://www.youtube.com/watch?v=x8Yb4RidLgM
Summary: https://aietalks.com/talks/flux-open-research-and-the-future-of-visual-ai.md

Why here: Dieleman explains the common machinery. Batifol shows what it buys in a working family of models: instruction-based edits, several reference images and enough speed for a person to guide the result while thinking. Kosinski next puts that machinery on a canvas.

### 4. ComfyUI Full Workshop

Yedri Kosinski, Comfy | 51:25 | AI Engineer World's Fair 2025
Video: https://www.youtube.com/watch?v=_FKeSzM9fPc
Summary: https://aietalks.com/talks/comfyui-full-workshop.md

Why here: Batifol presents generation and editing as model capabilities. Kosinski exposes the pipeline behind them as nodes for the model, text encoder, VAE, sampler and decoder. Once the workflow is visible, a person can replace or patch one part instead of hoping a better prompt fixes everything.

### 5. See, Hear, Speak, Draw

Logan Kilpatrick & Simón Fishman, OpenAI | 18:43 | AI Engineer Summit 2023
Video: https://www.youtube.com/watch?v=bNZV9s3_u44
Summary: https://aietalks.com/talks/see-hear-speak-draw.md

Why here: Kosinski gives a person direct control over the generation graph. Kilpatrick and Fishman hand part of the review back to the models: describe a target, generate an image, compare the two, then rewrite the prompt from the differences. Vernade turns that small loop into a longer creative job.

### 6. Let's Go Bananas with GenMedia

Guillaume Vernade, Google DeepMind | 1:17:14 | AI Engineer Europe 2026
Video: https://www.youtube.com/watch?v=BcWFc3H7Khg
Summary: https://aietalks.com/talks/lets-go-bananas-with-genmedia.md

Why here: Kilpatrick and Fishman test one image against one target. Vernade has to keep characters recognizable across a book. He extracts structured prompts from the source, saves character portraits and passes only the relevant references into each illustration. Ma takes that need for control into a consumer product.

### 7. Google Photos Magic Editor: GenAI Under the Hood of a Billion-User App

Kelvin Ma, Google Photos | 20:28 | AI Engineer World's Fair 2025
Video: https://www.youtube.com/watch?v=C13jiFWNuo8
Summary: https://aietalks.com/talks/google-photos-magic-editor-genai-under-the-hood-of-a-billion-user-app.md

Why here: Vernade can tolerate retries and imperfect illustrations in a workshop. Ma cannot make that bargain inside Google Photos. Magic Editor narrows a broad promise into jobs such as moving an object or rebuilding a background, then uses constraints, evals and several candidate outputs to make those jobs dependable.

### 8. SOTA Generative Media Panel

Dumitru Erhan, Shane Gu & Nicole Brichtova, Google DeepMind | 56:59 | AI Engineer
Video: https://www.youtube.com/watch?v=KLDdXOw6jIc
Summary: https://aietalks.com/talks/sota-generative-media-panel.md

Why last: Ma argues that reliable editing needs benchmarks and constrained workflows. Brichtova, Erhan and Gu show why the benchmark itself deserves suspicion. People may choose the sharper, more saturated image even when it is less realistic, while real production workflows reveal mistakes that isolated samples never expose.
