# AI Engineer Summit 2023, Day 1 Livestream

Benjamin Dunphy, Software 3.0 LLC & swyx, Latent.Space & Smol.ai & Amjad Masad & Michele Catasta, Replit & Toran Bruce Richards, AutoGPT & Simón Fishman & Logan Kilpatrick, OpenAI & Flo Crivello, Lindy & Barr Yaron, Amplify & Sasha Sheng & Harrison Chase, LangChain & Shreya Rajpal, Guardrails AI & Eugene Yan, Amazon & Linus Lee, Notion & Brittany Walker, CRV & Chris White, Prefect & Bryan Bischof, Hex | AI Engineer Summit 2023 | 4:50:00

Source: https://www.youtube.com/watch?v=veShHxQYPzo
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/ai-engineer-summit-2023-day-1-livestream
Published: 2023-10-10
Tags: evals, guardrails, observability, structured-outputs

## TL;DR
- AI engineering combines software engineering with models, tools, and orchestration, so engineers need to build systems around language models rather than rely on prompting alone.
- Production LLM systems need task-specific evals, retrieval, validation, guardrails, observability, and user feedback because model outputs are variable and can fail in subtle ways.
- Specialized, structured systems can make AI applications faster and cheaper, from Replit's small code model and Wasp's app generator to typed outputs with Pydantic and multi-agent workflows.

## Summary
Day 1 of the AI Engineer Summit presents AI engineering as a software discipline built around language models. swyx describes three forms of AI engineer: engineers enhanced by AI tools, engineers building AI products, and agentic systems that may replace parts of engineering work. The talks then show practical patterns. Replit describes making AI coding available throughout its product, training a small code model for low latency, and using a specialized full-stack generator to create applications. OpenAI demonstrates multimodal pipelines that connect vision, text, image generation, audio, and video. Other speakers focus on agents, typed structured outputs, embeddings, context-aware applications, and code-generation autonomy. The strongest production advice is more cautious: create small task-specific eval sets, retrieve and rank context carefully, validate outputs, collect feedback, and inspect failures. The day also introduces the AI Engineer Foundation and Agent Protocol as open-source efforts around shared interfaces and interoperability.

## Key ideas
### AI engineering is a software role built around models and orchestration
[26:07](https://www.youtube.com/watch?v=veShHxQYPzo&t=1567s)
swyx separates AI engineering into three forms: the AI-enhanced engineer who uses tools such as Copilot, the products engineer who builds AI products, and the AI engineer agent that is not human. He argues that language models are not yet AGI and must be coordinated inside software systems. His proposed path is a stack of improvements, where engineers use AI, build AI products, and eventually work with agents. He also gives a broader definition of a 1000x engineer: someone who teaches ten other people what they know, then helps those people build wider networks of learning.

### AI coding tools should be part of the programming environment
[34:30](https://www.youtube.com/watch?v=veShHxQYPzo&t=2070s)
Amjad Masad says Replit began with code completion and built Ghostwriter for autocomplete and chat inside the IDE. He argues that AI should be infused into every programming interaction rather than added as a separate feature. Replit announces AI coding for its millions of users, including people coding on Android phones, and Model Farm, which lets developers call models from the IDE with three lines of code. The goal is to make AI available both while writing software and inside the software's own call stack, so developers can build AI products directly.

### Small code models can win when training and serving are designed for the task
[42:48](https://www.youtube.com/watch?v=veShHxQYPzo&t=2568s)
Michele Catasta describes Replit Code v1.5, a 3.3 billion parameter model trained from scratch on code. The team trained on up to one trillion code tokens, used data from 30 programming languages, and filtered generated, minified, unreadable, and toxic code. The model uses a 4K context and a code-specific vocabulary. Catasta says it reached 36% on OpenAI's HumanEval benchmark and came close to Code Llama 7B despite being smaller. Replit's priority is inference speed, with a single unbatched model generating above 200 tokens per second. Deployment time was reduced from 18 minutes to two minutes.

### Agents promise to remove routine work, but safety must reach full reliability
[1:00:45](https://www.youtube.com/watch?v=veShHxQYPzo&t=3645s)
The AutoGPT team frames agents as a way to let people interact with spreadsheets, inboxes, and applications through conversation. The team describes Forge, a standard template for agent creators, and a development UI built on the Agent Protocol. They argue that coding is a foundation for general agents because code is the digital fabric of software. The talk also gives direct failure examples: prompt injection from malicious websites and an agent deleting every JSON file on a laptop when asked to delete files in one directory. The team says commercial agents cannot tolerate a 99% success rate when a single email or action can lose a contract.

### Multimodal applications can connect separate models through text
[1:14:10](https://www.youtube.com/watch?v=veShHxQYPzo&t=4450s)
Simón Fishman and Logan Kilpatrick show multimodal pipelines using the models available in 2023. A vision model describes an image, DALL-E generates a new image from that description, and vision compares the original and generated images to suggest another prompt. A second demo combines video frames described by GPT-4 with a Whisper transcript to produce a more visually informed article. They describe text as connective tissue between separate models for now, while a future unified model may accept and produce several modalities directly. The demos also show that visual comparison can become a general pattern for moving from a current state toward a target state.

### Agent systems range from fixed chains to dynamic loops
[2:58:25](https://www.youtube.com/watch?v=veShHxQYPzo&t=10705s)
Harrison Chase describes context-aware reasoning applications as systems around language models. Context can come from instructions, examples, retrieval, or fine-tuning. On the reasoning side, systems progress from a single model call to fixed chains, routers that choose among prompts or tools, loops that repeat until a condition is met, and autonomous agents whose available actions can change over time. Chains offer more control over the sequence, while agents can react to unexpected inputs. LangChain helps with orchestration and data handling, and LangSmith provides visibility into model calls, prompts, retrieved documents, and tool sequences.

### Typed outputs turn language-model integration into ordinary application code
[3:17:37](https://www.youtube.com/watch?v=veShHxQYPzo&t=11857s)
Jason Liu argues that many LLM applications need structured data rather than chat. OpenAI function calling helps define a JSON schema, while Pydantic provides typed models, validation, and JSON Schema generation. His Instructor library patches the completion API so a Pydantic object can be used as the response model, with retries when validation fails. Validators can check lengths, database membership, forbidden content, or whether quoted evidence exists in the source text. This approach lets developers model prompts, data, and behavior together. Liu applies the pattern to retrieval, query plans, knowledge graphs, and answers that include substring quotes from the original document.

### LLM products need task-specific evals and feedback loops
[3:35:30](https://www.youtube.com/watch?v=veShHxQYPzo&t=12930s)
Eugene Yan presents evaluations, retrieval-augmented generation, guardrails, and feedback as building blocks for LLM systems. He recommends starting with a small evaluation set, even 40 questions, and using deterministic checks where possible. SQL can be executed, JSON keys can be compared, and moderation can use precision and recall. Retrieval still needs ranking because models perform worse when the relevant document appears in the middle of a long context. Guardrails can use sentence-level factual consistency checks, sampling, or a strong model. User feedback is sparse when explicitly requested and noisy when inferred, so product actions such as accepting code or upscaling an image can provide useful implicit signals.

### Embeddings can be explored as controllable representations
[3:52:48](https://www.youtube.com/watch?v=veShHxQYPzo&t=13968s)
Linus Lee treats embeddings as high-dimensional spaces that may contain meaningful features of text and images. He demonstrates decoding text from embeddings, moving embeddings along directions associated with shorter text or negative sentiment, combining embeddings from different texts, and adapting a decoder to read from OpenAI's embedding space. With image embeddings, he interpolates between a photograph and a cartoon and shifts an image toward descriptions such as a sad person or a beach. His point is that interfaces can make model representations visible and manipulable. He releases the text models and a notebook on Hugging Face so others can experiment with interpolation and interpretation.

## Notable quotes
- "LLMs themselves are not AGIs yet, right? Like we actually have to coordinate them in systems of software." (31:19)
- "You can't have a 99% success rate. It has to be 100%." (1:11:28)
- "You need automated evals. You need automated evals." (3:50:50)
- "Think of text as a connecting tissue right now." (3:11:57)
- "The biggest thing is think about failure modes." (4:44:03)

## Tools & references mentioned
- AI Engineer Summit
- Software 3.0 LLC
- Latent.Space
- Smol.ai
- Replit
- Ghostwriter
- Model Farm
- Google Cloud
- Llama
- Stable Diffusion
- AutoGPT
- Agent Protocol
- OpenAI
- GPT-4 with vision
- Whisper
- DALL-E 3
- Lindy
- LangChain
- LangSmith
- Pydantic
- Instructor
- Marvin
- Guardrails AI
- Amazon
- Notion
- Cody
- Sourcegraph
- Mage
- Wasp
- Leonardo
- Stability AI
- ElevenLabs
- Midjourney
- AI Engineer Foundation
- LanceDB
- Prefect
- Hex
- Weights & Biases
- Rivet
- Chroma
- Pinecone
- Hugging Face
- Code Llama
- HumanEval
- The Stack
- BigCode
- Chinchilla
- MMLU

## Who should watch
- You are building an LLM feature and need concrete patterns for retrieval, structured outputs, validation, evaluation, and feedback.
- Your team is deciding whether to build or buy AI infrastructure, or whether a focused open-source project should live beside the main product.
- You are exploring coding agents, multimodal pipelines, embeddings, or agent interoperability and want examples from early production systems.

## Related talks

- [AI Engineer Summit 2023, Day 2 Livestream](https://aietalks.com/talks/ai-engineer-summit-2023-day-2-livestream) (Mario Rodriguez, GitHub & Dedy Kredo, CodiumAI & Matt Welsh, Fixie.ai & Amelia Wattenberger, Adept & Samantha Whitmore & Jason Yuan, New Computer & Joseph Nelson, Roboflow & Hassan El Mghari, Vercel & Paul Copplestone, Supabase & Daniel Rosenwasser, Microsoft & Jason Liu, Fivesixseven & Anton Troynikov, Chroma & Jerry Liu, LlamaIndex & Mithun Hunsur, Ambient & Abi Aryan, O'Reilly & Simon Willison, Datasette & Benjamin Dunphy, Software 3.0 LLC & swyx, Latent.Space & Smol.ai, 7:30:56)
- [AI Engineer Singapore Day 2](https://aietalks.com/talks/ai-engineer-singapore-day-2) (Kaspar Hidayat, 65 Labs & SallyAnn DeLucia, Arize AI & Timothy Lin, Resaro & Abhishek Kankani, Cloudflare & Tejas Kumar, IBM & JJ Geewax, Google DeepMind & Geoff Huntley, Independent & Vincent Koc, OpenClaw Foundation & Vishnu (Vish) Hari, Ego AI & Ben Guo, Zo Computer & Matthias Lubken, Tavon AI & Josh Newton, Microsoft AI & Sam Bhagwat, Mastra & Pierre-Loic Doulcet, LlamaIndex & Jun Yu Tan, Tusk & Sara Hooker, Adaption Labs & Vincent Wu, MiniMax & Daniel Krishnan & Siddharth Krishnan, The Robot Company & Justin Baird, Tesseract & Kai Ming & Aravind (SK) Kandiah, Bifrost & Julia Kim, OpenGraph Labs & Suveen Ellawela, Cortex AI & Keziah & Jay Demetillo & Alex Lee, Magic Patterns & Sabina Cabrera, MagicPath & Priyaa Kalyanaraman, Lica World & Conor Brennan-Burke, Hyperspell & Heng Hong Lee, Lightsprint & Louis Knight-Webb, Vibe Kanban & Harsha Khurdula, Interfaze AI & Hrishi Olickel, Southbridge & Henry Mao, Smithery & Rach Pradhan, Independent & Agrim Singh, AI Engineer, 9:28:00)
- [Why Agent Engineering](https://aietalks.com/talks/why-agent-engineering) (Shawn Wang, Latent.Space, 11:45)
- [AI Engineer Code Summit 2025](https://aietalks.com/talks/ai-engineer-code-summit-2025) (Jed Borovik, Google & Swyx, AI Engineer & Barry Zhang & Mahesh Murag, Anthropic & Dex Horthy, HumanLayer & Lee Robinson & Naman Jain, Cursor & Jacob Kahn, Meta & Rhythm Garg & Linden Li, Applied Compute & Will Brown, Prime Intellect & Will Hang & Cathy Zhou, OpenAI & Kitze, Independent & Kath Korevec, Google Labs & Eno Reyes, Factory AI & Beyang Liu, Amp Code / Sourcegraph & Natalie Serrino, Gimlet Labs & Jake Nations, Netflix & Eiso Kant & Jason Warner, Poolside & Aparna Dhinakaran, Arize & Nik Pash, Cline & Joel Becker, METR & Kevin Hou, Google DeepMind & Benjamin Dupy & Leah McBride, AI Engineer, 8:57:06)
- [AI Engineer World's Fair 2025, Day 1 Keynotes and MCP Track](https://aietalks.com/talks/ai-engineer-worlds-fair-2025-day-1-keynotes-and-mcp-track) (Laurie Voss, LlamaIndex & Shawn Wang, Latent Space & Asha Sharma, Microsoft & Sarah Guo, Conviction & Simon Willison, Datasette & Stephen Chin & Andreas Kollegger, Neo4j & Henry Mao, Smithery & Theodora Chu & John Welsh, Anthropic & Harald Kirschner, VS Code, Microsoft & David Cramer, Sentry & Samuel Colvin, Pydantic & Alex Volkov, Weights & Biases & Benjamin Eckel, Dylibso & Jan Curn, Apify & Antje Barth, AWS & Kevin Hou, Windsurf & Greg Brockman, OpenAI, 8:37:51)
