# AIE Singapore Day 1

Sherry, 65 Labs & Dr Vivian Balakrishnan, Singapore Ministry of Foreign Affairs & Gavriel Cohen, NanoCo & Thibault Sottiaux, OpenAI & Dr Feng Yuzhang, GovTech Singapore & Phil Hedayatnia, Airfoil & Annie Luo, Google & Keziah & Jimmy Lai, Vercel & Vedran Jukic, Daytona & Vaishant Kameswaran & Rohan Kumar, Greptile & Yuntong Zhang, Sonar & Eugene Cheah, Featherless & Max Buckley, Exa AI & Mark Doyle, Stripe & Li Hau Tan, Simular & Ryo Lu, Cursor & Aosheng Ran, Figma & Selim Arguel, Menlo Research & Alberto Taiuti, Reactor & Jan Liphardt, OpenMind & Andrew Tan, Groq & Daria Soboleva, Cerebras & Zixuan Li, Z.ai & Boris Starkov, ElevenLabs & Jackman Ong, Prime Intellect & Michelle Julia, Bluelabs & Jacky Mok, Reka & Gokul Srinivasan, Antim Labs & Wei Wei Hsu, Lentil & Anun Joshi, Bland & Linh Nguyen, Obello & Stefania Druga, Sakana AI & swyx, Cognition | AI Engineer Singapore 2026 | 10:09:12

Source: https://www.youtube.com/watch?v=_xQnSNlBP_w
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/aie-singapore-day-1
Published: 2026-05-16
Tags: agents, coding-agents, evals, guardrails, harness-engineering

## TL;DR
- AI agents are moving software work from code generation into planning, review, deployment, operations, and long-running autonomous execution.
- Agents should be isolated from credentials and sensitive systems, with sandboxes, proxies, deterministic controls, logging, and human approval for risky actions.
- As software becomes cheaper to produce, judgment, taste, evaluation, ownership, local context, and the design of reliable agent environments become more valuable.

## Summary
Day 1 of AI Engineer Singapore brought together talks on deploying agents in software, government, design, robotics, voice, and enterprise systems. Vivian Balakrishnan described a personal second-brain workflow built from NanoClaw, WhatsApp, graph memory, local models, Whisper, and Obsidian. He argued that people can outsource computation and memory, but not personal understanding or accountability. Several speakers focused on safety. NanoClaw, Daytona, Sonar, and Simular described isolation, credential proxies, network controls, approval layers, and guardrails that sit outside the planning agent. Other talks examined what changes when code is cheap: Vercel stressed documentation and ownership, Greptile showed that agents produce different bug profiles from humans, and Stripe described one-shot coding agents that create pull requests through judge and diagnostic loops. The afternoon expanded into multimodal collaboration, world models, robotics, long-horizon reasoning, sovereign AI, voice reliability, and enterprise deployment. The common concern was practical control over increasingly capable systems.

## Key ideas
### Personal understanding and accountability remain human responsibilities
[42:05](https://www.youtube.com/watch?v=_xQnSNlBP_w&t=2525s)
Vivian Balakrishnan argued that AI can take over calculations, memory, replication, and dissemination, but people cannot outsource personal understanding. Someone in authority can delegate work, he said, but cannot delegate accountability. He described a personal agent built around NanoClaw, WhatsApp, graph memory, semantic search, Whisper, and Obsidian. It collects speeches, transcripts, and other material, then helps with travel, meetings, briefs, and speeches. The system runs on a Raspberry Pi with 8 GB of RAM. Balakrishnan presented this as tool assembly rather than traditional software development, while still checking the code and approving sensitive access.

### Agents need isolation because instructions cannot provide security
[1:10:10](https://www.youtube.com/watch?v=_xQnSNlBP_w&t=4210s)
Gavriel Cohen described NanoClaw's agent factory, which reviews pull requests, creates testing plans, launches virtual machines, and runs tests before code is merged. He said pull requests are unsanitized input and that instructions such as 'never run drop database production' are not security controls. NanoClaw treats agents as if they operate behind enemy lines. Each agent runs in its own container, while credentials stay outside the agent environment. A vault proxies requests and inserts credentials only when policy allows it. Sensitive actions happen outside the sandbox after a human approves them, so the agent proposes a merge but does not hold the credentials that execute it.

### The bottleneck in software is moving beyond code generation
[1:27:25](https://www.youtube.com/watch?v=_xQnSNlBP_w&t=5245s)
Thibault Sottiaux described software delivery as a pipeline whose build stage has widened because agents can generate code quickly. The constraints then move to planning, review, validation, CI, security, release, operations, debugging, and understanding what happened. Codex is positioned as an agent that works across those stages. Sottiaux also discussed approval fatigue. Codex's auto-review system uses a second agent to compare actions with the original task intent, blocking suspicious or high-risk actions and redirecting the main agent. He said this reduced approvals by a factor of 20 inside OpenAI.

### Agent-native systems require explicit documentation and careful ownership
[2:27:20](https://www.youtube.com/watch?v=_xQnSNlBP_w&t=8840s)
Jimmy Lai said software frameworks now have to serve agents as well as people. Agents read documentation literally, copy examples, run commands, and follow errors exactly. A misleading document can confuse many projects before anyone notices, while a vague error can make an agent keep retrying and spend money. He said teams should define what fast, reliable, and secure mean in their own systems so agents can check those properties. AI makes creation cheap, but ownership remains expensive. Forking a framework means owning its security, reliability, and long-term maintenance, so teams should be careful about what they decide to build and support.

### AI-generated code has different bug patterns from human code
[3:13:30](https://www.youtube.com/watch?v=_xQnSNlBP_w&t=11610s)
Greptile analyzed more than five million pull requests and found strong evidence that 27.6% of April pull requests were fully vibe coded. The speakers compared revert rates, review severity, and the number of review rounds across agents and human authors. Some agents had lower revert rates than the human baseline, and some reached merged code in fewer review rounds. The results did not produce one overall winner because the outcome changed with the metric. The more useful finding was that agents produce different bugs. Cursor background agents were more likely to create N+1 query errors, while Claude agents were more likely to miss tenant checks. Validation systems need to account for those different profiles.

### Long-running agents need checkpoints, memory, and the ability to change direction
[7:36:42](https://www.youtube.com/watch?v=_xQnSNlBP_w&t=27402s)
Z.ai's Zixuan Li and Prime Intellect's Jackman Ong both focused on long-horizon work. Li said long horizon means depth and the ability to keep finding meaningful improvements, rather than simply running for more hours. Models can forget the original goal, accumulate errors, or continue pushing a bad approach without pivoting. He recommended checklists and repeated goal reviews, with checkpoints for verification. Ong argued that recursive language models can keep large context outside the prompt and use code structures such as loops, variables, and subagents. Training models on these scaffolds may make long-running work more reliable than relying on increasingly large context windows.

### Creative and consumer AI should preserve exploration instead of forcing quick answers
[2:12:40](https://www.youtube.com/watch?v=_xQnSNlBP_w&t=7960s)
Annie Luo argued that fashion and travel decisions are subjective, so efficiency is not always the right product goal. People often discover what they want by seeing several options and comparing them. Google's virtual try-on lets users see themselves in different clothes, while Google Travel uses maps as a place to wander rather than only as a destination picker. These interfaces help people build taste and confidence through the process. Phil Hedayatnia made a related argument about design agents. Training on finished designs does not teach the reasoning behind them. AirFoil's Melt stores references, metadata, comments, and annotations so models can use a designer's intent rather than copying surface patterns.

### Agent infrastructure is becoming the product
[5:21:09](https://www.youtube.com/watch?v=_xQnSNlBP_w&t=19269s)
Across the talks, speakers described systems around the model that determine whether agents are useful. Stripe's Minions create pull requests from Slack prompts, run tests, send results to an LLM judge, and use a diagnostic agent to continue work when the task is incomplete. Simular separates planning from guardrails and combines accessibility trees with visual grounding for computer use. Groq described routing requests across regions and model instances to reduce latency. Swyx framed this as the agent platform above the model and container, including identity, stateful sessions, tools, permissions, model diversity, enterprise controls, and evaluations. The practical work is in making agents observable, bounded, and dependable.

### Local deployment needs local data, evaluation, and control
[9:26:36](https://www.youtube.com/watch?v=_xQnSNlBP_w&t=33996s)
Stefania Druga described sovereign AI as local agency over global capability. Her stack includes data, evaluation, post-training, routing, user interaction, and governance, with physical choices about on-premise and cloud compute. Sakana AI's Japanese products use local data, dialects, formal language registers, and routing to specialized or secure models. Its switchboard can send Japanese-context requests to a post-trained Japanese model, sensitive requests to an on-premise model, or requests to human review. Druga argued that different countries should choose which layers they own. Sovereignty is therefore a set of practical control points, rather than only the creation of a national foundation model.

## Notable quotes
- "The one thing which you cannot outsource is your personal understanding." (42:59)
- "The only way to prevent it from leaking a secret is to not give it a secret." (1:17:57)
- "AI made creation really cheap, but ownership is much more expensive than you think it is." (2:58:49)
- "The question is no longer can you build it. The question is what should exist." (3:52:45)
- "Local agency is more important than global capability." (9:35:31)

## Tools & references mentioned
- 65 Labs
- NanoClaw
- OpenClaw
- Codex
- Claude
- GPT-5.1-Codex-Max
- Claude Opus 4.5
- Neman
- Baileys
- Ollama
- Whisper
- Obsidian
- Andrej Karpathy
- Neil Lawrence
- Yan LeCun
- impeccable.style
- Melt
- Blend
- The Runaway Species
- Anthony Brandt
- David Eagleman
- Reduct
- Vercel
- Next.js
- Daytona
- Greptile
- SonarQube
- Featherless AI
- Qwen
- Gemma
- DeepSeek
- Exa AI
- Stripe Minions
- Simular
- Figma
- Figma Weave
- Menlo Research
- Asimov
- Reactor
- OpenMind
- GroqCloud
- Cerebras
- Z.ai
- GLM 5.1
- ElevenLabs Speech Engine
- Prime Intellect
- Recursive Language Models
- Blue Labs
- Reka
- Antim Labs
- Gizmo
- Lentil
- Bland
- Twilio
- Obello
- Sakana AI
- Sakana Chats
- Namazu
- Sakana Fugu
- AI Scientist
- Frontier Suite
- Cognition
- Devon
- DeepWiki
- AI Engineer

## Who should watch
- You are building coding or computer-use agents and need practical patterns for isolation, credential handling, approval flows, and testing.
- Your team is producing more AI-generated code than human reviewers can comfortably inspect, and you need better validation, evaluation, and ownership rules.
- You are designing consumer, creative, voice, robotics, or government AI products where local context, human judgment, exploration, and trust matter as much as raw model capability.

## Related talks

- [AI Engineer Singapore Day 2](https://aietalks.com/talks/ai-engineer-singapore-day-2) (Kaspar Hidayat, 65 Labs & SallyAnn DeLucia, Arize AI & Timothy Lin, Resaro & Abhishek Kankani, Cloudflare & Tejas Kumar, IBM & JJ Geewax, Google DeepMind & Geoff Huntley, Independent & Vincent Koc, OpenClaw Foundation & Vishnu (Vish) Hari, Ego AI & Ben Guo, Zo Computer & Matthias Lubken, Tavon AI & Josh Newton, Microsoft AI & Sam Bhagwat, Mastra & Pierre-Loic Doulcet, LlamaIndex & Jun Yu Tan, Tusk & Sara Hooker, Adaption Labs & Vincent Wu, MiniMax & Daniel Krishnan & Siddharth Krishnan, The Robot Company & Justin Baird, Tesseract & Kai Ming & Aravind (SK) Kandiah, Bifrost & Julia Kim, OpenGraph Labs & Suveen Ellawela, Cortex AI & Keziah & Jay Demetillo & Alex Lee, Magic Patterns & Sabina Cabrera, MagicPath & Priyaa Kalyanaraman, Lica World & Conor Brennan-Burke, Hyperspell & Heng Hong Lee, Lightsprint & Louis Knight-Webb, Vibe Kanban & Harsha Khurdula, Interfaze AI & Hrishi Olickel, Southbridge & Henry Mao, Smithery & Rach Pradhan, Independent & Agrim Singh, AI Engineer, 9:28:00)
- [Opening Keynotes - AIE Paris 2025 (Day 1)](https://aietalks.com/talks/opening-keynotes-aie-paris-2025-day-1) (Shawn Swyx Wang, Latent Space and AI Engineer & Lélio Renard Lavaud, Mistral, 1:25:26)
- [AIE Europe Keynotes & Coding Agents](https://aietalks.com/talks/aie-europe-keynotes-coding-agents) (Tejas Kumar, AI Engineer & Omar Sanseviero, Google DeepMind & David Soria Parra, Anthropic & Ido Salomon, MCP Apps & Mario Zechner, Pi & Armin Ronacher & Cristina Poncela Cubeiro, Earendil & Benjamin Dunphy, AI Engineer & David Gomes, Cursor & Matthias Luebken, TAVON & Sarah Chieng, Cerebras & Lawrence Jones, Incident.io & Luke Alvoeiro, Factory & Ben Burtenshaw, Hugging Face & Michael Richman, Cmd+Ctrl & Liam Hampton, Microsoft & Tuomas Artman, Linear & Gergely Orosz, The Pragmatic Engineer & Jacob Lauritzen, Legora & Peter Gostev, Arena AI & swyx, AI Engineer, 9:09:51)
- [AI Engineer World's Fair 2025, Day 2 Keynotes & SWE Agents Track](https://aietalks.com/talks/ai-engineer-worlds-fair-2025-day-2-keynotes-swe-agents-track) (Laurie Voss, LlamaIndex & Benjamin Duny, AI Engineer & Logan Kilpatrick & Jack Rae, Google DeepMind & Manu Goyal, Braintrust & Solomon Hykes, Dagger & Jesse Han, Morph & Vibhu Sapra & Scott Wu, Cognition & Rustin Banks, Google Jules & Christopher Harrison, GitHub Copilot & Tomas Reimers, Graphite & Boris Cherny, Anthropic & Robert Brennan, Allhands & Josh Albrecht, Imbue Sculptor & Eno Reyes, Factory & George Cameron, Artificial Analysis & Ankur Goyal, Braintrust & Barr Yaron, Amplify & Alex Atallah, OpenRouter & Sean Grove, OpenAI & Ben & swyx, AI Engineer, 9:08:07)
- [AIE Miami Day 2](https://aietalks.com/talks/aie-miami-day-2) (David House, G2i & Sarah Chieng, Cerebras & Lech Kalinowski, CallStack & Tejas Bhakta, Morph LLM & Rick Blalock, Agentuity & Nyah Macklin, Neo4j & Lena Hall, Akamai & Dave Kiss, Mux & Alvin Pane, OutRival & Erik Thorelli, CodeRabbit & Hassan El Mghari, Together AI & Stefan Avram, OpenCode & Laurie Voss, Arize AI & David Gomes, Cursor, 8:02:16)
