A chatbot pilot built on 40 curated files can work well while the same approach fails across 80,000 SharePoint files.
2
Document quality work must find relevant files, sensitive data, duplicates, conflicts, stale content, and missing metadata.
3
Cleaning stale and duplicate content roughly doubled recall in the talk's multi-hop RAG evaluation and improved task completion by 10 to 15 percent.
Summary
Jeff Koss and Leo Platzer describe a legal-operations chatbot for North River Manufacturing, built over years of SharePoint documents. A 40-file pilot handled parsing, chunking, OCR, embeddings, a vector database, and a knowledge graph, but scaling to more than 80,000 files exposed problems with relevance, sensitive data, duplication, conflicting versions, and freshness. Jeff demos AI-assisted taxonomies, metadata tagging, evidence, confidence scores, sensitivity detection, quality filters, and scheduled data slices. Leo explains why these checks affect retrieval: when 30% of a corpus is stale or duplicative, up to 80% of the retrieved context can be useless. In their multi-hop RAG evaluation, cleaning the corpus roughly doubled recall and improved task completion by 10 to 15 percent. He also shows an SDK workflow that creates context files for coding agents, so they can inspect folder-level summaries before reading individual files.
A small document pilot hides the problems that appear at scale
North River Manufacturing built a legal-ops chatbot for supplier contracts, partnership agreements, termination clauses, expiry dates, and exclusivity terms. The first pilot used about 40 pre-curated files and produced good responses after parsing, chunking, OCR, embeddings, a vector database, and a knowledge graph. When the team expanded to more than 80,000 SharePoint files, it could no longer easily find relevant documents, locate sensitive data, or trust the corpus. The business also needed a repeatable process for many future AI use cases, rather than a one-off cleanup.
Document quality includes duplication, conflict, and freshness
Jeff distinguishes document quality from the usual dimensions used for structured data. For documents, the team needs to know whether information is duplicated, whether different files conflict, and whether content is fresh. These issues affect the chatbot's retrieval context. An old or expired contract can produce the wrong answer, while duplicate content can crowd out more useful documents. The quality workflow also checks whether required metadata exists and lets the team select a filtered data slice instead of sending the full SharePoint corpus into the chatbot.
AI can build and apply a taxonomy with human review
Jeff shows a taxonomy as a tagging tree. Users can create tags themselves, reuse tags from a library, upload valid values through CSV, or ask AI to suggest tags from a prompt and the document contents. Conditions can create branches, such as classifying legal and contract files by supplier contract, NDA, or materials and time contract. After metadata generation, users can filter the corpus by tags and inspect why a file received a classification. They can also give thumbs-up or thumbs-down feedback as examples for the system.
Metadata needs evidence and confidence, not just a label
The demo shows metadata at both file and chunk level. For each AI-generated classification, users can inspect the part of the document that provided evidence. The system also returns a confidence score, so the user can see whether the model reported 100% or 95% confidence instead of accepting an unexplained label. Jeff presents this as a way to review the data product before using its metadata in a chatbot or vector database.
Sensitive-data detection can combine AI with patterns and context
The workflow scans documents for sensitive information such as PII and PHI. Jeff describes a demonstration involving Finnish national IDs. The system used AI-based detection together with pattern matching and context, and it suggested Finnish words associated with the identifier. He says this combination can reduce false positives. The legal-ops workflow uses the results to keep sensitive material out of the chatbot's data stores or identify information that needs redaction.
Scheduled data slices keep the selected corpus current
A data slice is a filtered set of files selected by criteria such as legal relevance, date range, freshness, required metadata, and sensitivity. Jeff shows how a use-case description can guide suggested taxonomy tags, while a human chooses which suggestions to keep. Workflows can refresh a slice on a schedule. New or deleted SharePoint files are picked up automatically, and the SDK can pass the refreshed metadata into the chatbot and its vector database.
Stale and duplicate files can dominate retrieved context
Leo explains that duplicate documents and conflicting versions occupy the top results returned for a question. He compares this with changing statements in Slack, where the truth can change over time. If old information remains in the search space, retrieval can keep returning it. In the evaluation he describes, 30% of a thousand-file raw corpus was stale or duplicative, yet up to 80% of the agent's actual answer context could consist of stale information. The problem affects high-stakes agent use cases because irrelevant context can shape the answer.
Cleaning the corpus improved recall and task completion
Leo says the team ran the same tasks in a multi-hop RAG evaluation before and after maintaining unstructured-data quality. Cleaning the data roughly doubled recall. Overall task completion improved by 10 to 15 percent. He argues that duplication and freshness matter across use cases, even when the application changes, because those conditions affect the search space before application-specific relevance rules are applied.
Folder-level context files help coding agents choose what to read
Leo uses the SDK to fetch metadata for a SharePoint site, create context for each folder, and build an index over the folders. A coding agent such as Claude Code or Codex can read a context markdown file before opening individual files. The generated information includes each folder's purpose, document types, topics, when to use the folder, and example questions. This lets the agent decide which folder deserves attention without first inspecting all 23 files shown in the demo.
"Instead of actually looking at these 23 files to figure out whether I even should look at this folder, I know exactly what document types I have, what it is about, what are the key topics, when to use this folder, and what are example questions."Leo Platzer19:36
Who should watch
You are moving an AI prototype from a small, manually selected document set into a large SharePoint corpus and need a concrete data-preparation workflow.
Your retrieval system returns stale, duplicate, conflicting, or sensitive documents and you need ways to inspect and filter them before indexing.
You are building coding agents and want folder-level context files to reduce unnecessary file inspection.