# The 2025 AI Engineering Report

Barr Yaron, Amplify Partners | AI Engineer World's Fair 2025 | 12:33

Source: https://www.youtube.com/watch?v=mQ7_Zje7WKE
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/the-2025-ai-engineering-report
Published: 2025-08-01
Tags: agents, evals, observability, rag

## TL;DR
- AI engineers use language models across internal and customer-facing products, with most applying them to several use cases.
- Retrieval-augmented generation is the most common customization method, while fine-tuning is more widespread than Barr expected.
- Evaluation is the most painful part of AI engineering, while agents are still early in production despite strong interest in adopting them.

## Summary
Barr Yaron presents early findings from Amplify Partners' 2025 State of AI Engineering survey, based on 500 respondents. The respondents have varied job titles, and many experienced software engineers have only recently started working with AI. Most use language models for both internal and external applications, often across several use cases. Retrieval-augmented generation is the most common customization method, while fine-tuning is also widespread among researchers and research engineers. Teams update models and prompts frequently, but 31% have no prompt-management system. Text remains far ahead of image, audio, and video in workplace adoption. Agents have attracted strong interest, although fewer than 20% say they work well at work. Monitoring, human review, offline evaluations, and internal metrics are common, and evaluation is the most frequently cited source of pain. The survey also covers vector databases, disclosure by AI agents, model architecture, open and closed models, and AI-generated relationships.

## Key ideas
### AI engineering has a broad community and many newcomers
[00:51](https://www.youtube.com/watch?v=mQ7_Zje7WKE&t=51s)
The survey included 500 respondents, with engineers forming the largest group, but the audience had many different job titles. Barr says titles are currently unreliable because people with different labels often do the same work. Interest in the term AI engineering rose after ChatGPT launched in late 2022. Many experienced software engineers are still new to AI: among those with more than 10 years of software experience, nearly half have worked with AI for three years or less, and one in 10 started within the past year.

### Teams use language models across several kinds of work
[03:03](https://www.youtube.com/watch?v=mQ7_Zje7WKE&t=183s)
More than half of respondents use LLMs for both internal and external use cases. Code generation and code intelligence, along with writing assistance and content generation, are the leading use cases. Barr describes the main pattern as heterogeneity. Of the people using LLMs, 94% use them for at least two use cases and 82% use them for at least three. For customer-facing products, three of the five most-used models and half of the top 10 came from OpenAI.

### RAG is common, and fine-tuning is more widespread than expected
[04:05](https://www.youtube.com/watch?v=mQ7_Zje7WKE&t=245s)
Besides few-shot learning, retrieval-augmented generation is the most popular way respondents customize their systems, with 70% reporting its use. Barr was surprised by how much fine-tuning is happening across the sample. Researchers and research engineers fine-tune by far the most. Among people who fine-tune, 40% mention LoRA or QLoRA, while other approaches include DPO, reinforcement fine-tuning, and supervised fine-tuning. Respondents also report hybrid methods.

### Teams change models and prompts often, but prompt management is inconsistent
[05:31](https://www.youtube.com/watch?v=mQ7_Zje7WKE&t=331s)
More than half of respondents update their models at least monthly, including 17% who update them weekly. Prompt changes happen even more often: 70% update prompts at least monthly, and one in 10 does so daily. Despite that activity, 31% have no way to manage their prompts. Barr connects the pressure to frequent model releases and the possibility that a new result or blog post makes an existing prompt feel inadequate.

### Text adoption is far ahead of image, audio, and video
[06:17](https://www.youtube.com/watch?v=mQ7_Zje7WKE&t=377s)
Image, video, and audio models lag text models by significant margins in workplace use and production traction. Barr calls this the multimodal production gap. Among respondents not currently using a modality, audio has the strongest future intent: 37% of non-users plan to adopt it eventually. Barr expects adoption to rise as models become better and easier to access.

### Agents attract interest even though few teams say they work well
[07:57](https://www.youtube.com/watch?v=mQ7_Zje7WKE&t=477s)
Barr defines an AI agent as a system where an LLM controls the core decision-making or workflow. Although 80% of respondents say LLMs work well at work, fewer than 20% say the same about agents. Fewer than one in 10 say they will never use agents, so most respondents at least plan to adopt them. Agents already in production usually have write access with a human in the loop, while some can take actions independently.

### Evaluation remains the main source of pain
[08:40](https://www.youtube.com/watch?v=mQ7_Zje7WKE&t=520s)
Most respondents use several ways to monitor AI systems. Sixty percent use standard observability, more than half rely on offline evaluations, and teams also collect user data and use benchmarks. Human review remains the most popular way to evaluate model and system quality. For model-usage monitoring, internal metrics are the usual choice. When asked about the most painful part of AI engineering, respondents put evaluation at the top.

### Specialized vector databases are common in production systems
[09:24](https://www.youtube.com/watch?v=mQ7_Zje7WKE&t=564s)
Sixty-five percent of respondents use a dedicated vector database to store and retrieve context. Barr says this suggests that specialized vector databases provide enough value over general-purpose databases with vector extensions for many applications. Among dedicated-vector-database users, 35% primarily self-host and 30% primarily use a third-party provider.

## Notable quotes
- "Titles are weird right now, but the community is broad. It's technical. It's growing." (02:22)
- "Basically folks who are using LLMs are using them internally, externally, and across multiple use cases." (04:05)
- "Agents aren't everywhere yet, but they're coming." (08:20)
- "And finally, we asked folks, what is the number one most painful thing about AI engineering today? And evaluation topped that list." (11:14)

## Tools & references mentioned
- Amplify Partners
- ChatGPT
- OpenAI
- LoRA
- QLoRA
- DPO
- Simon Willison
- GPT-4
- Latent Space
- Elon Musk
- Twitter

## Who should watch
- You are deciding where your team should use LLMs, RAG, fine-tuning, or vector databases and want comparison data from other practitioners.
- Your team is building agents and needs a view of current production adoption, human approval patterns, monitoring, and tool permissions.
- You are struggling with evaluations or prompt changes and want to see how widespread those problems are among AI engineers.

## Related talks

- [2025 is the Year of Evals! Just like 2024, and 2023, and …](https://aietalks.com/talks/2025-is-the-year-of-evals-just-like-2024-and-2023-and) (John Dickerson, Mozilla AI, 19:14)
- [Why (Senior) Engineers Struggle to Build AI Agents](https://aietalks.com/talks/why-senior-engineers-struggle-to-build-ai-agents) (Philipp Schmid, Google DeepMind, 10:40)
- [Navigating AI's Frontier in 2025](https://aietalks.com/talks/navigating-ais-frontier-in-2025) (Grace Isford, Lux Capital, 17:55)
- [AI Engineering 201: The Rest of the Owl](https://aietalks.com/talks/ai-engineering-201-the-rest-of-the-owl) (Charles Frye, Full Stack LLM Bootcamp, 56:57)
- [AI Leadership](https://aietalks.com/talks/ai-leadership) (Alex Lieberman, Morning Brew and 10X & Kath Korevec, Google Labs & Katelyn Lesse, Anthropic & Michele Catasta, Replit & Lisa Orr, Zapier & Steve Yegge, Sourcegraph and AMP & Gene Kim, IT Revolution & Bill Chen & Brian Fioca, OpenAI & Martin Harrysson & Natasha Maniar, McKinsey & Yegor Denisov-Blanch, Stanford & Itamar Friedman, Qodo & Olive Song, MiniMax & Asaf Bord, Northwestern Mutual & Lei Zhang, Bloomberg & Samir Mody, The Browser Company & Max Kanat-Alexander, Capital One & Arman Hezarkhani, 10X & Justin Reock, DX & Dan Shipper, Every & Mel Lutzky, Graphite, 8:16:05)
