Cooking with Codex

Charlie Guo, OpenAI, Gabriel Chua, OpenAI1:07:24 · Oct 2026 · 5,894 views
Thumbnail for Cooking with Codex Watch on YouTube
TL;DR
  1. 1

    Codex becomes more useful when it receives context from conversations, repositories, screenshots, tools, and user preferences.

  2. 2

    Long-running work needs explicit goals, verifiable completion criteria, audits, handoffs, and human help when the agent is blocked.

  3. 3

    Automations, hooks, subagents, and the App Server can turn individual Codex requests into repeatable development workflows.

Summary

Charlie Guo and Gabriel Chua present Codex through a series of live demonstrations. They start with a Slack conversation that gives Codex enough context to build a macOS live translation app, then show how plugins provide development practices and connections to services such as Slack. Other examples use repository history to build a shareable Agents SDK site, remote execution to create an IKEA-style repository visualization, computer use to export dashboard data, and the Chrome extension to create a Google feedback form. The second half focuses on longer tasks. The presenters describe a live Q&A site, an internal form builder adapted from an open-source repository, and a Python package with a Rust backend reconstructed from research papers. Goals, audits, progress dashboards, thread handoffs, subagents, hooks, and automations help these projects continue while exposing blockers. They close with the App Server and Retrodex, a custom Codex interface, and argue for moving from one-off requests toward repeatable software-building processes.

Key ideas
05:25

Context can arrive as a conversation, screenshot, repository, or personal preference

Guo and Chua describe context as the material Codex needs before it can build effectively. A Slack conversation contains the brainstorming for the live translation app, while an appshot captures both the screen and its embedded text. Repository history supplies context for a site about the Agents SDK. Plugins add team practices and connections to services such as Slack, Gmail, and Outlook. Personalization adds context over time through memories, Chronicle, and custom instructions. Guo also recommends dictation because people can explain an idea faster by speaking than typing.

11:02

Plugins package development practices and external connections for repeated use

The presenters describe plugins as bundles containing skills, external prompts, team practices, app configuration, and MCP servers. The macOS plugin supplies guidance for AppKit, liquid glass, signing, and inspecting the built application. The OpenAI developers plugin can create an API key when needed. Plugins for services such as Gmail, Outlook, Microsoft Teams, and Slack provide the connections needed to bring outside context into a task. Guo says a difficult workflow can later be turned into a skill or plugin so the same work does not have to be explained again.

26:15

Different browser controls fit different kinds of access and testing

Chua distinguishes computer use, the Chrome extension, and the in-app browser. Computer use is useful for desktop or legacy applications without a programmatic interface, and it can operate while the user continues working elsewhere on the Mac. The Chrome extension works through the browser where the person is already signed in, so Codex can use existing authorization with permission. The in-app browser has deeper access to the browser engine and is suited to testing local web applications and their rendering. The distinction is practical: use browser control for an existing signed-in session, and use deeper browser access for local application testing.

28:45

Long-running projects need clear goals and completion checks

The presenters show three extended projects: a live Q&A site, an internal form builder adapted from an existing repository, and a Python package with a Rust backend reconstructed from research papers. They recommend describing what success looks like and making the criteria verifiable. A goal can keep evaluating whether the task is complete and decide whether to continue. Guo warns that vague criteria can produce a result that technically satisfies the request while missing the intended outcome. Clear instructions still matter, but he says the main issue is giving the model enough context rather than finding a magic prompt phrase.

45:01

Visible threads and subagents divide work in different ways

Subagents keep context separate and let the model delegate work such as code review or documentation research. Thread-to-thread handoff is better when the person wants to see the delegated work and remain involved. Guo describes a game project with separate threads for art direction, music, animation, and game mechanics, all reporting back to a creative director thread for sign-off. Remote and local threads can also cooperate: a remote session can ask a local thread to open Chrome and perform end-to-end tests, then receive the results.

32:16

Audits, dashboards, and communication expose blockers during unattended work

For a project that ran for hours, Codex created a goals file and a progress dashboard showing milestones, active work, and the threads responsible for each stream. Separate threads performed code reviews and goal audits, checking whether the main work had drifted from its plan. Codex also posted updates to Slack with completed work, remaining work, and blockers. When authentication or deployment access stopped progress, the presenter could inspect the blocker, ask for an explanation, and provide credentials or a safer setup path. Side threads are useful for steering, but the presenters warn that they are temporary, so lasting decisions need to go back to the main thread.

47:57

Hooks add deterministic checks around an otherwise autonomous workflow

Hooks let users run fixed actions at specific points in the development process. The presenters mention sending conversations to a logging system, scanning inputs and outputs for exposed API keys, checking whether a tool is on an approved list, and running custom validation when a turn stops. A hook in the Agents SDK repository runs a Python script at the end of each turn to tidy the repository. Hooks are useful when an agent works for a long time without direct supervision and the team needs checks that do not depend on the model remembering to perform them.

50:24

Automations connect waiting, feedback, and follow-up work

A heartbeat automation can check whether a remote server has finished provisioning and then continue with the next step. The presenters also show an automation that reviews Slack feedback about the translation app, classifies comments, and can start a separate worktree to fix a bug, open a pull request, and request review. Scheduled automations can run in the current thread or create a new thread for recurring summaries. Chua describes using automations to collect updates from Slack, email, and Linear into local notes, while another example compares drafted email replies with the replies eventually sent.

58:42

The App Server lets developers build Codex into their own interfaces

The App Server connects the Codex app, CLI, and VS Code extension to the main Codex harness, including tool calls, plugins, and context compaction. Guo shows Retrodex, a custom interface built on the App Server, with model and reasoning controls and a shared project context. Developers can use the open-source Codex repository to learn about the protocol and build their own interfaces or internal platforms. The presenters say core harness features can be embedded this way, while computer use features that need to run on the user's computer are not part of the App Server integration.

"I think one misconception these days is people always ask does prompt engineering matter? In so far as the instructions are clear, it still matters."Gabriel Chua20:19
Who should watch
  • You are using coding agents for one-off tasks and want a practical way to add context from Slack, repositories, screenshots, or signed-in browser sessions.
  • You run long coding tasks remotely and need goals, audits, progress views, and clear ways to intervene when authentication or infrastructure blocks progress.
  • You are building an internal agent product and want to understand how plugins, hooks, automations, or the Codex App Server fit together.