A context graph combines company data, live and synced access, skills, memory, and processes so an agent can answer questions across the business.
2
Live API lookups work for simple customer-specific requests, while semantic questions across large datasets require a synced local copy of the data.
3
Every answer needs provenance, including its source, fetch time, permissions, transformations, and the conditions that make the data invalid.
Summary
Gil Feig argues that connecting an agent to a few MCP servers does not give it a usable company brain. Live APIs can answer a question about one customer, but they struggle with questions that require searching thousands of tickets or records. A context graph combines live lookups with synced data for semantic queries, then uses a router and summarizer to select and reduce the information sent to the agent. Feig also places company skills and memory in the context layer. A skill can define what "upset customer" means at a particular company and which systems should be checked first. He describes four context tiers: prompts, skills and memory; live APIs; cached data; and derived summaries. The system should check freshness before using cached data, fall back to live calls when needed, and record where every fact came from. The talk gives a practical architecture for building agents that can reason over company data without losing track of its limits.
A context graph is the company brain around an agent
Feig uses "context graph" to mean the collection of information and company logic passed into an agent. It includes third-party systems such as NetSuite and Jira, structured context such as static documents, memories formed during use, and skills that tell the agent how work gets done. The talk does not focus on graph nodes and edges. It focuses on bringing the company's data, instructions, and learned information into one context layer so the agent can use them together.
Live MCP lookups fail when a question spans a large dataset
A live lookup can explain why customer A was upset last week by retrieving that customer's Zendesk tickets. The same approach breaks when the question becomes "Which customers were upset last week?" or "Which customers were upset in the last year?" The agent would need to pull tickets one batch at a time, possibly thousands of them, until it times out. Feig says the APIs behind many platforms do not support semantic queries across their full datasets, so MCP connections alone do not solve this problem.
Semantic questions need a synced local copy of the data
Feig recommends syncing data locally when an agent needs analysis or semantic search across a large dataset. The synced data can go into a local vector database, with a strategy for storing and retrieving it. Live calls still make sense for narrow requests, such as asking Stripe about one specific customer or taking one action. The distinction is practical: use live access for simple lookups, and use synced data when the question requires searching and comparing many records.
A router and summarizer keep retrieved context usable
Once an agent has both synced and live context, it should not receive a huge dump of raw results. A router examines the incoming prompt and sends it to the appropriate source or agent. After the data is collected, a summarizer combines information from the synced and live layers and returns a smaller answer to the requesting agent. Feig describes this as a flow that gathers the right data first, then gives the agent enough context to respond without overwhelming it.
Skills define what the company means and where the agent should look
The meaning of a question depends on company processes. "Which customer was upset?" might mean counting support tickets at one company, while another might rely on Qualtrics survey data because that is how it measures customer happiness. A skill can encode that rule and tell the agent which system to check first. The agent should inspect relevant memory and skills before querying synced or live data, then use those instructions to select and summarize the results.
The four context tiers have different costs and freshness limits
Feig groups context into four tiers. Prompts, skills, and memory guide the agent's behavior. Live API calls are easier and cheaper to add, so they fit simple retrievals. Cached or synced data supports complex semantic questions but costs more to run and is harder to keep current. Derived context consists of summaries or other information created while retrieving third-party data. Storing a customer's ticket summary can prevent the system from repeatedly fetching and summarizing the full ticket history.
The data-selection flow checks skills, freshness, and fallback paths
When a prompt arrives, the system first checks for a matching skill or memory. If the skill needs no external data, the agent can answer directly. Otherwise it checks whether the synced data exists and is fresh enough for the relevant service-level agreement. Fresh data can be used from the cache. Stale or missing data requires a live API call. The system can repeat this process until it has enough context, store a new memory when appropriate, and then generate the response.
Provenance must travel with every piece of context
Feig says the system should retain metadata for both MCP results and data stored in vector databases. That metadata should record the source, fetch time, identifier, scopes, user access, permissions, transformations, and the original form of the data. It should also state what invalidates the data, such as an age limit. This matters when an answer drives a business decision or when a synced record is old. Provenance also gives engineers a way to investigate after an answer goes wrong.
"You can't just use MCP to live look up for things that you want to go deeper into, that you want to semantically question, you need to actually sync a copy of the data locally."04:51
Who should watch
You are building an agent that must answer questions across support, CRM, finance, or other large company datasets.
Your current design relies on live MCP calls and starts timing out when a query spans many records.
You need answers to carry source, freshness, permissions, and invalidation details for debugging or business decisions.