$1 AI Guardrails: The Unreasonable Effectiveness of Finetuned ModernBERTs

Diego Carpentero43:53 · Apr 2026 · 7,079 views
Thumbnail for $1 AI Guardrails: The Unreasonable Effectiveness of Finetuned ModernBERTs Watch on YouTube
TL;DR
  1. 1

    LLM attacks can enter through prompts, retrieved context, model internals, MCP tools, and agent actions because the model has no native separation between trusted instructions and untrusted data.

  2. 2

    A fine-tuned encoder model can classify safety risks with lower latency than an LLM judge, while remaining self-hosted and cheap to retrain.

  3. 3

    ModernBERT's alternating attention, sequence packing, RoPE, and FlashAttention make it practical to use as a safety discriminator across several checkpoints in an AI system.

Summary

Diego Carpentero maps attacks across the full path of an LLM application. Direct prompt injection can expose system instructions, while malicious text in web pages, email, GitHub issues, or retrieved documents can influence decisions without the user submitting the attack. He also covers suffix attacks against model alignment, RAG poisoning, hidden instructions in MCP tool descriptions, and agentic attacks that lead to code execution or self-escalation. His proposed defense is a fine-tuned ModernBERT classifier that checks prompts, responses, retrieved content, tool descriptions, memory, and agent plans. He explains why bidirectional encoder models fit classification, then covers ModernBERT's attention pattern, RoPE, unpadding, sequence packing, and FlashAttention. The walkthrough uses the Inject Guard dataset and a safe-or-unsafe classification head. The reported baseline reaches almost 85% accuracy with about 35 milliseconds per classification, though Carpentero describes it as a starting point rather than a complete safety solution.

Key ideas
00:01

LLM attacks now target every part of an AI application

Carpentero says attacks that began as exploratory prompt injection have become more sophisticated and are amplified by identity workflows. His attack surface includes the natural-language prompt, external context, retrieval-augmented generation, MCP tools, agents, and model internals. The talk focuses on building a low-latency, self-hosted defensive layer for under a dollar with a fine-tuned ModernBERT encoder. The model is chosen because its architecture supports the long inputs and fast classification needed for repeated safety checks.

01:00

Prompt injection works because system instructions and user data share the same model input

Direct injection uses crafted user input to override system controls or expose confidential data. Carpentero uses the Bing Chat Sydney incident as an example. A Stanford student entered, "ignore previous instructions," and Bing Chat revealed its system prompt, including the Sydney codename and more than 40 confidential rules and policies. The underlying problem is that the user input is concatenated with the system prompt before being sent to the model. LLMs have no native separation of concerns between system controls and user data.

03:09

Untrusted external content can overrule an AI system without direct user input

In indirect injection, malicious instructions sit in content the model is expected to fetch, such as HTML, URLs, email, or public web pages. One example placed an instruction on a Wikipedia page about Albert Einstein, causing an LLM to search for a code that linked to an attacker website. Carpentero also describes websites embedding prompts to manipulate advertising review systems into approving non-compliant content. The common weakness is that the model cannot natively distinguish a developer instruction from untrusted text in its context.

05:58

Alignment can be bypassed through optimized gibberish suffixes

Model-internals attacks append meaningless-looking tokens to a harmful request. The suffix shifts the next-token probability distribution so the model starts with an affirmative response, after which its autocomplete behavior continues the harmful answer. Carpentero explains that alignment is a probabilistic preference rather than a hard constraint. Attackers use greedy coordinate gradient search, starting with placeholder tokens and iteratively selecting candidates that increase the chance of an affirmative opening. Suffixes found with open models can transfer to black-box models because similarly trained models develop geometrically similar refusal boundaries.

09:25

Small amounts of poisoned RAG data can steer targeted answers

Carpentero describes the Poison RAG paper, which found that a tiny number of poisoned chunks can manipulate an LLM's answer for a target question. In a knowledge base containing 8 million documents, five poisoned chunks were enough to succeed in the reported attack. The malicious content must satisfy retrieval and generation conditions. It needs to resemble the user's query semantically, which can be helped by appending a likely query to the target answer, and it must rank highly after retrieval. A convincing-sounding answer helps meet the second condition.

10:55

MCP and agents add hidden approval and execution paths

MCP creates an asymmetry between the short tool summary shown to a user and the full tool description read by the LLM. A user might approve an apparently harmless function such as adding two numbers while hidden instructions cause private keys or MCP credentials to be passed as parameters. In agentic environments, a page telling a computer agent to download and launch a support tool led to file execution and a demonstrated remote-code-execution path. Carpentero also describes a malicious NPM package promoted through a GitHub issue whose title was inserted into an agent prompt, affecting nearly 4,000 to 5,000 developers.

22:21

Safety checks need to sit at multiple points in an autonomous system

Carpentero calls this a zero-trust gap. LLMs do not natively separate controls from data, and alignment cannot be treated as a hard security constraint. Human review also has blind spots because the reviewer may not see the full instructions being approved. At minimum, production systems should check user inputs and model responses. More checks should cover retrieval, MCP calls, context memory, and agent plans. Possible mechanisms include rule filters, canary tokens, discriminators, constrained decoding, and LLM judges when the application can accept more latency.

18:06

A fine-tuned encoder is a fast, private safety discriminator

Carpentero frames safety checking as classification rather than generation. ModernBERT can process all tokens bidirectionally in one forward pass, condense the sequence into its CLS representation, and send that representation to a classification head. His fine-tuned model takes about 35 milliseconds per classification in the baseline case, before quantization or other optimizations. It can be retrained within hours as attacks change, self-hosted to keep internal requests and intermediate responses private, and used repeatedly without compounding the token cost and latency of an LLM judge.

20:19

ModernBERT's architecture reduces the cost of long safety inputs

ModernBERT combines two local-attention layers using an 8,128-token sliding window with a global-attention layer of 8,192 tokens every third layer. This captures local patterns such as suffix attacks while retaining longer context for MCP descriptions and agent plans. Unpadding removes meaningless padding tokens, and sequence packing concatenates tokens from multiple examples into the context window while masking attention between examples. Rotary positional encoding rotates query and key projections according to token position, preserving positional relationships without adding a position vector to token semantics. FlashAttention keeps block computations in fast on-chip memory instead of materializing the full attention matrix.

"The results that we're getting is almost 85% accuracy using only 35 milliseconds per classification."41:21
Who should watch
  • You are building an LLM application with retrieval, tools, memory, or agents and need checks beyond the user prompt and final response.
  • You need a lower-latency alternative to an LLM-as-a-judge for classifying prompts, context, or model outputs.
  • You want a practical fine-tuning example for a self-hosted encoder safety model and need to understand the ModernBERT design choices behind it.