← All blog posts

The Sad Story of MCP and its efficiency

November 8, 2025 · AI

Generated by Grok

Anthrop\c recently posted a blog entry around MCP efficiency in terms of speed and token usage. MCP and all the AI Stack-related technologies interest me a lot lately, in my attempt to create several productization use-cases for LLMs.

In case you do are not aware, MCP (Model Context Protocol) is a technology introduced originally by Anthrop\c to augment LLM usage and custom knowledge discovery. So, there are MCP servers that expose tools, which LLMs can query and pick those they need for their needs.

The above idea was great, but it seems that it was more marketing than engineering (and still is). It is useful though :), like Windows 95 in the early days.

The problem is that as agents connect to thousands of tools via MCP, two major inefficiencies arise:

Context Window Overload: Loading all tool definitions directly into the model’s context window consumes a massive amount of tokens, increasing costs and latency. This also occurred as a natural limit, on how many MCP could connect to a single use case.

Excessive Token Use from Results: When an agent calls multiple tools in a sequence, the entire data output from one tool must be passed back into the model’s context just to be fed into the next tool, consuming even more tokens.

The Solution

… to this problem is to have the AI agent write and execute code to interact with MCP servers instead of making direct tool calls. So the MCP server will be queried once and the cache their capabilities as executable code that then can be identified and called by the MCP client. Consider the following architecture:

There are some benefits to this approach, since it minimizes MCP tool invocation, only to the absolute necessary. In addition, the generated program can be reused for additional queries in the future for the specific tools.

Its like writing an SDK on the fly. There are also some privacy gain, since all data will remain in the execution environment and the calls are limited to third party providers, like LLM inference services and MCP servers.

Conclusions

Anthrop\c’s blog post can be perceived as the natural evolution of the original MCP architecture, but it also makes clear that the original approach was more hacky, or at least optimistic in terms of actual capabilities of the technology.

My opinion, is that even if it is more efficient in several ways, this approach puts some “vibe coding” aspects to the MCP ecosystem, thus exhibiting the limitations of foundational models.