How MCP Works: The Model Context Protocol Explained
MCP explained for engineers: the N×M integration problem, client/server architecture, tools/resources/prompts, stdio and HTTP transports, handshake and security.
Updated 9 min read
The Model Context Protocol is an open standard for connecting AI applications to external tools and data. It defines a JSON-RPC interface where servers expose capabilities and clients consume them, so an integration written once works with every compatible AI application instead of just one.
Anthropic published MCP in late 2024. Adoption spread beyond Anthropic’s own products quickly, including to other major model vendors, which is the outcome that turns a protocol into a standard.
What problem does MCP solve?
Without a standard, connecting M AI applications to N data sources means M times N bespoke integrations. Each is written separately, maintained separately, and breaks separately. MCP collapses this to M plus N. Write one Jira server; every MCP-capable host can use it. Write one client implementation in your application; every MCP server becomes available to it.
The redundancy it removes is easy to underestimate. Every AI application needs access to the same things: files, databases, issue trackers, documentation, internal APIs. Every one of those systems needs a wrapper that describes it to a model and executes calls against it. Your Jira integration for one assistant does nothing for another.
This is not a novel insight. The Language Server Protocol did exactly this for editors and language tooling, replacing a quadratic integration matrix with a linear one, and it won. MCP is the same bet applied to AI applications.
How is an MCP connection structured?
Three roles. The host is the AI application the user interacts with, a client is a connector inside the host holding a one-to-one stateful connection with exactly one server, and a server is a program exposing capabilities over the protocol. The distinction between the first two trips people up.
Host. The AI application the user interacts with — a desktop assistant, an IDE, a coding agent, an internal tool. The host manages the model, the conversation, and the user’s trust decisions.
Client. A connector inside the host. Each client maintains a one-to-one stateful connection with exactly one server. A host running five servers instantiates five clients. This isolation is deliberate: a server cannot see other servers’ traffic, and the host controls what crosses between them.
Server. A program exposing capabilities over the protocol. It might wrap a database, a SaaS API, a local filesystem directory, or an internal service. Servers are typically small — a few hundred lines is common.
The important consequence of this design is that servers never talk to the model. They talk to a client, which hands data to the host, which decides what enters the model’s context. The host is the policy enforcement point, which is where it belongs.
Communication uses JSON-RPC 2.0 for requests, responses, and notifications.
What are the three MCP primitives?
Tools, resources, and prompts. MCP servers expose these three kinds of capability, distinguished by who decides when they are used: tools are executable functions the model chooses to call, resources are read-only data the host application loads into context, and prompts are reusable templates the user invokes deliberately. Mapping your capabilities onto the right one is the main design decision when writing a server.
Tools — model-controlled
Executable functions the model chooses to call. Each tool has a name, a description, and a JSON Schema for its input. The model reads the descriptions and decides when a tool is relevant.
Tools are the primitive that does things: query a database, create a ticket, send a request, run a search. They can have side effects, which is why hosts typically require user approval before invocation.
Tool descriptions matter enormously. They are the entire interface the model sees. A vague description produces a tool the model calls at the wrong times or not at all.
Resources — application-controlled
Read-only data identified by URI. A file, a database record, a documentation page, a log. Resources do not execute anything; they return content.
The distinction from tools is control. The host application decides which resources to load into context, often by presenting them to the user for selection. The model does not autonomously fetch resources.
Servers can support resource templates with URI parameters, and can notify clients when a resource changes.
Prompts — user-controlled
Reusable templates the user invokes deliberately, typically surfaced as slash commands or menu items. A prompt takes arguments and returns a structured set of messages.
This is the least used primitive and the most underrated. It lets a server ship expertise, not just access — a code review prompt that encodes your team’s standards, a triage prompt that structures an incident writeup.
| Primitive | Controlled by | Purpose | Side effects |
|---|---|---|---|
| Tools | The model | Perform actions, fetch dynamic data | Yes, typically |
| Resources | The host application | Supply context as readable content | No |
| Prompts | The user | Invoke a structured workflow | No |
There are also client-side primitives that flow the other direction: sampling lets a server request a model completion through the client, roots let the client tell a server which filesystem or URI boundaries it may operate within, and elicitation lets a server request additional input from the user. Support for these varies by host.
Should you use stdio or HTTP?
Use stdio for local tools and streamable HTTP for anything shared, remote or hosted. The two differ mainly in trust: with stdio the process boundary is the trust boundary, so no authentication layer is needed, while an HTTP server is reachable over a network and requires real authorization. MCP defines how messages move separately from what they mean.
stdio. The host launches the server as a subprocess and exchanges newline-delimited JSON-RPC messages over standard input and output. There is no network, no port, and no authentication layer, because the process boundary is the trust boundary — the server runs as the user who launched it, with that user’s permissions.
This is the right choice for local tools: filesystem access, local databases, developer utilities. It is simple and it is what most MCP servers use.
Streamable HTTP. The server runs as a network service exposing a single endpoint that accepts POSTed JSON-RPC messages and can stream responses back using server-sent events when needed. This supports remote hosting, multiple concurrent clients, and session management.
HTTP transport requires real authorization. The specification builds on OAuth 2.1 patterns for this. If your server is reachable over a network, treat it as a public API surface, because it is one.
An earlier HTTP+SSE transport has been superseded by streamable HTTP; new servers should use the current one.
Lifecycle and handshake
Every connection follows the same sequence.
Initialize. The client sends an initialize request declaring the protocol version it supports, its capabilities, and its identity. The server responds with the version it will use, its own capabilities, and its identity. Capability negotiation here is what lets the protocol evolve without breaking older implementations — each side only uses features the other advertised.
{
"jsonrpc": "2.0",
"id": 1,
"method": "initialize",
"params": {
"protocolVersion": "2025-06-18",
"capabilities": { "roots": { "listChanged": true }, "sampling": {} },
"clientInfo": { "name": "example-host", "version": "1.4.0" }
}
}
Initialized. The client sends an initialized notification confirming it is ready. Only now may normal operation begin.
Operation. The client discovers what is available — tools/list, resources/list, prompts/list — and then invokes as needed: tools/call, resources/read, prompts/get. Servers may send notifications when their capability lists change, so clients should re-list rather than caching indefinitely.
Shutdown. For stdio, the client closes the input stream and the subprocess exits. For HTTP, the client closes the session.
A minimal server
The official Python SDK makes a working server short. This one exposes all three primitives over a hypothetical release-notes system.
from mcp.server.fastmcp import FastMCP
mcp = FastMCP("changelog")
@mcp.tool()
def search_releases(query: str, limit: int = 10) -> list[dict]:
"""Search release notes by keyword.
Use this when the user asks what changed, when a feature shipped,
or whether a specific bug was fixed. Returns matching releases
with version, date, and a short summary.
"""
results = release_index.search(query, limit=limit)
return [
{"version": r.version, "date": r.date.isoformat(), "summary": r.summary}
for r in results
]
@mcp.resource("changelog://release/{version}")
def release_notes(version: str) -> str:
"""Full release notes for a specific version."""
return release_index.get(version).body
@mcp.prompt()
def upgrade_report(from_version: str, to_version: str) -> str:
"""Draft an upgrade impact report between two versions."""
return (
f"Review every release between {from_version} and {to_version}. "
"List breaking changes first, then required migration steps, "
"then optional improvements. Cite the version each item came from. "
"If a change has no migration note, say so explicitly."
)
if __name__ == "__main__":
mcp.run()
Note the tool docstring. It describes when to use the tool, not just what it does, because that is the decision the model has to make. This is the highest-leverage text in the entire server.
A host registers the server through configuration:
{
"mcpServers": {
"changelog": {
"command": "python",
"args": ["-m", "changelog_server"],
"env": { "RELEASE_DB_URL": "postgres://localhost/releases" }
}
}
}
Is MCP secure?
The protocol gives you a framework; the trust decisions are yours. MCP hands models the ability to act on real systems, and the main exposures are prompt injection through tool content and tool descriptions, over-broad credentials handed to servers, and third-party servers running against sensitive systems. Treat an MCP server like any dependency with production access.
Prompt injection through tool content. Anything a tool returns enters the model’s context. If a tool fetches a web page, reads an email, or queries a user-writable database, an attacker who controls that content can attempt to issue instructions to the model. This is an actively exploited class of vulnerability, not a theoretical one. Treat all tool output as untrusted input.
Malicious tool descriptions. Descriptions are also model-visible text. A hostile server can embed instructions in a description. This is why installing a third-party MCP server is equivalent to installing a dependency with production access — read the code or do not run it.
Over-broad credentials. A server given a database connection string can do anything that connection permits. Provision least-privilege credentials scoped to what the server actually needs, and prefer read-only wherever possible.
Identity and authorization risks
The confused deputy problem. A server holding credentials on behalf of a user can be induced to act with those credentials on someone else’s instruction. Verify that the requesting identity is authorized for each action rather than assuming the connection implies authority.
Token passthrough. Do not accept a token issued for another service and forward it. Servers should validate that tokens were issued for them.
Human approval on side effects. Hosts should require confirmation before invoking tools that write, send, delete, or spend. Do not build a server that assumes the host will protect the user; make destructive operations explicit and separately named so approval is meaningful.
Our broader coverage of AI security risks goes deeper on the injection surface.
When should you use MCP, and when should you skip it?
Use MCP when an integration will be consumed by more than one application, shared across teams, or maintained on a different schedule than the app using it. Portability is the value, not capability. Skip it when you have one application with a fixed set of tools and no plans to reuse them anywhere else.
The reason to skip it in that case is cost of indirection. Direct function calling against a model API is fewer moving parts and easier to debug. A protocol you do not need is a layer of indirection you have to maintain.
MCP is also not a retrieval system. It is transport and interface. If your problem is finding the right document among a million, you still need retrieval-augmented generation — MCP is how the agent reaches your retrieval service, not the retrieval itself.
Key takeaways
MCP standardizes AI-to-system integration the way LSP standardized editor tooling, converting an M×N problem into M plus N. Hosts contain clients, clients connect one-to-one with servers, and servers never talk to the model directly.
Three server primitives map to three controllers: tools for the model, resources for the application, prompts for the user. Getting that mapping right is the main design decision.
Use stdio for local servers and streamable HTTP with real authorization for anything networked. Negotiate capabilities at initialize and re-list rather than caching.
Security is where MCP deployments go wrong. Tool output is untrusted input, third-party servers are dependencies with production access, and credentials should be scoped to the minimum.
For how this fits into agent architecture, see our review of the best AI agents and the Claude Code guide. More coverage at /category/mcp/.
Choosing an AI stack for your team?
We help companies pick the right AI tools, wire them into existing systems and avoid the ones that quietly do not scale. Tell us what you are trying to build and we will tell you what we would use — no charge for the conversation.
Join the discussion
Comments are not enabled on this article yet. Reach the editorial desk directly with corrections or additions.
Contact the editors