Deep research needs an agent because the system must plan, search, inspect sources, find gaps, pivot, and produce a cited artifact.
2
Research and writing need different architectures: research benefits from flexible agent behavior, while writing works better as a constrained workflow with review loops.
3
Evaluation requires a real dataset, a calibrated language-model judge, and separate development and test splits instead of judging quality from a few demos.
Summary
This workshop builds a small deep research system and connects it to a technical writing workflow. Louis-François Bouchard first separates prompts, workflows, and agents. A workflow has fixed steps, while an agent can choose actions and react to its environment. The research system uses MCP tools for grounded web search, YouTube analysis, and report compilation. It stores intermediate results in a .memory folder and writes a final research.md artifact. Samridhi Vaid shows how an MCP server exposes tools, prompts, and resources, then connects to Claude Code locally. Paul Iusztin presents the writing side as a fixed workflow. It combines a user guideline, research, writing profiles, few-shot examples, and an evaluator-optimizer loop. The final section covers observability and evaluation with threads, traces, token and cost tracking, and a binary language-model judge. The practical lesson is to keep exploration flexible, writing constrained, and both systems measurable.
Louis-François Bouchard presents prompting, workflows, and agents as an autonomy slider. A prompt is enough when the model already knows the task and the needed context fits in the request. A workflow adds fixed sequences, routers, parallel steps, or judge loops. An agent becomes useful when it must take actions, react to what happens, branch dynamically, or decide which tools to use. He gives support-ticket handling as a workflow because it always classifies, routes, drafts, validates, and sends in the same order. Building that as an agent would add overhead without adding flexibility.
An agent's context includes its instructions, tool definitions, schemas, examples, retrieved data, and conversation history. As the system performs more steps, this context grows and performance can degrade before the model's advertised context limit. Bouchard connects this to the lost-in-the-middle problem, where long-context models can retrieve an inserted fact but do not necessarily use an entire long document well. The practical response is to keep context lean by trimming, summarizing, retrieving selectively, or delegating work to tools and sub-agents with their own contexts. Multi-agent designs become more useful when the system has many tools or an unwieldy context.
Bouchard defines deep research as a reasoning system that plans what to investigate, searches the web, uses sources or APIs, cites evidence, and receives feedback from itself or a person. It has one objective rather than a fixed recipe. The system can search, inspect, pivot, synthesize, and iterate. Their version starts with a topic and produces a research artifact for a technical article. It must find enough relevant material without collecting so much that the context becomes noisy. The research agent is also expected to inspect supplied web pages, YouTube videos, and GitHub repositories before searching for missing information.
MCP keeps the agent's reasoning separate from its capabilities
Samridhi Vaid describes Claude Code as the brain that reasons about gaps and chooses actions, while the MCP server exposes capabilities. The server provides tools for deep research, YouTube video analysis, and compiling a report. It also exposes prompts containing workflow instructions and resources such as model names, server versions, and feature flags. The first tools write results into a .memory folder, while the compile tool reads those files and creates research.md. FastMCP handles the protocol details, and the local server can be connected to Claude Code through an mcp.json file. The code can also be used with other agent harnesses.
Grounded search and multimodal video analysis fill different research needs
The deep research tool sends a query to the Gemini API and returns an answer with source URLs, titles, and snippets. The YouTube tool sends a YouTube URL as a file URI with instructions for producing a transcript and related output. Vaid explains that Gemini processes the video itself rather than simply reading an existing transcript, so video analysis can take a few minutes. Firecrawl and Apify handle web scraping, while Gemini grounding supplies search answers with sources. Git and a GitHub library bring repository content into Markdown, including private repositories when a token is provided.
Research and writing should be split into separate systems
The workshop team found that research and writing have different operating requirements. Research is exploratory. It needs to search broadly, discover missing information, revisit earlier assumptions, and pivot. Writing is more deterministic. It must follow a chosen tone, structure, terminology, and format while avoiding unsupported claims and unwanted phrases. Their architecture therefore runs a research agent first and a writing workflow second. The two systems communicate through files, especially the research.md artifact and a human-written guideline. They run sequentially with a basic script because the team did not need live orchestration between every research and writing step.
A writing workflow controls generation through explicit context
Paul Iusztin's writing system loads a guideline, the research artifact, writing profiles, and few-shot examples into a system prompt. The guideline changes from post to post and specifies the topic, angle, points, narrative flow, audience, tone, and constraints. Static profiles describe structure, terminology, and the writer's character. Few-shot examples teach the system what representative outputs look like. Iusztin recommends starting with examples and reducing their number until the system stops working, because every example adds cost, latency, and context. He also stresses that the human must put thought into the guideline instead of asking for a generic post.
A separate reviewer makes writing revisions more useful
The writing workflow uses an evaluator-optimizer loop with separate context windows for the writer and reviewer. The reviewer checks the draft against the guideline, research, and writing profiles. It returns structured Pydantic objects with a profile, location, and comment, such as identifying a banned term in a particular paragraph. The editor then applies the reviews and repeats the process. Iusztin says the team tried score-based stopping but found creative writing too subjective and noisy for a reliable threshold, so they use a fixed number of iterations and keep earlier versions for manual comparison. Reviews also have priorities, with the guideline taking precedence over research and profiles when instructions conflict.
Evaluation starts with labeled examples from the real system
Iusztin argues that a few successful demos do not show whether a writing workflow works across many inputs. The team created an evaluation dataset from real posts, reconstructed the associated guidelines and research, generated outputs through the real workflow, and labeled each result pass or fail with a short critique. They split the data into train, development, and test sets. The language-model judge receives the generated post plus the guideline, research, and profiles, then predicts a label and gives a critique. They tune it on the development split and validate it on the test split using F1. The workshop demo gets perfect scores on very small splits, and Iusztin warns that larger splits are needed because perfect results can indicate overfitting.
"The more you add complexity from prompting to more advanced workflows to agentic systems, the more autonomy you add, but also the less control you have over your whole system."07:06
Who should watch
You are deciding whether a task needs a prompt, a fixed workflow, or an agent and want concrete tests for that choice.
You are building a research assistant that must use web sources, YouTube videos, or repositories and need a file-based MCP design.
Your generated technical content looks plausible in demos but you lack a repeatable review, tracing, and evaluation process.