Implementing model context protocol (MCP): Tips, tricks and pitfalls

avatar
Nearform
10 Dec 2025
  • Share

A guide for developers on how to choose the right framework, how to sketch your first server, and how to harden and extend it once you have something working.

The Model Context Protocol (MCP) is an open standard that lets AI models talk to external tools and data sources in a uniform, safe way. MCP servers make tools and information available, and MCP compatible clients connect to them and make use of what they provide. Together, they allow you to mix and match capabilities easily, creating flexible and powerful integrations without building a new plugin system for every product.

This guide is written for developers who want to actually build and run MCP integrations in JavaScript/TypeScript or Python. It focuses on how to choose the right framework, how to sketch your first server, and how to harden and extend it once you have something working.

The goal is that, after reading this, you know what to install, what to implement first, how to test it without an LLM, and what kinds of issues to watch out for in production.

1. Choose a framework and runtime

MCP itself is language‑agnostic, but the official SDKs and frameworks make a big difference to your developer experience. However, the official SDKs and frameworks can significantly improve your developer experience. Before you start writing code, it’s worth choosing the stack you want to build on.

At Nearform, many of our projects use JavaScript and TypeScript, so we’ll cover those. We’ll also look at Python, which remains the go-to language for many AI applications.

Pitfall: avoid adopting niche or unofficial frameworks. The standard SDKs already provide strong functionality out of the box, and you’ll generally have an easier time with tools that are widely used, battle-tested, or backed by established projects. Move to a more custom setup only when you have a clear need that mainstream options can’t meet.

JavaScript/TypeScript

If you are working with a JavaScript/TypeScript codebase, start by reaching for the official TypeScript SDK. The gives you types for messages, negotiation, tools, and resources, plus utilities for STDIO, Streamable HTTP and Server‑Sent Events (SSE) transports. You can build everything with that alone. For teams that are more familiar with Fastify, you can also look at the Platformatic MCP server. It offers a lightweight, Fastify-based implementation of the protocol, giving you a more modular alternative while still supporting the core MCP features out of the box.

Python

On the Python side, the situation is similar. The official SDK exposes the protocol pieces directly, and you can still build servers using lower‑level classes if you want to control every detail. However, for almost all use cases you are better off with the FastMCP‑style interface that is now integrated into the SDK.

With FastMCP in Python, you typically:

  • Define tools as normal Python functions.
  • Add type hints and clear docstrings.
  • Let the framework infer JSON Schemas and tool descriptions from those annotations.

Developers htat are used to work with FastAPI will find using this framework really easy.

The framework also ships with niceties such as authentication hooks, user input (elicitation), and other utilities that expose functionality from the protocol in a simple format.

Ecosystem Servers and Utilities

Before writing your own server, it is worth scanning the growing library of existing MCP servers. There are connectors for common services such as filesystems, GitHub, Slack, Google Drive, Postgres, and web search. Many teams start by configuring those servers directly, then later fork or extend them when they need custom behavior.

Tip: check current MCP implementations to avoid re-inventing the wheel, MCP Market is a good source to find already implemented servers.

You will also see "inspector" or "playground" tools that let you point at an MCP server, discover its tools and resources, and issue test calls. Keep one of those in your toolbox; it is invaluable when you want to debug your server logic without an LLM in the loop.

Tip: check out the server with clients before running live with your data.

2. Sketch your first MCP server

Once you have picked a language and framework, resist the temptation to model your entire system at once. Start with a single, small server that exposes just one or two tools, and get that end‑to‑end path working first.

Trick: the initial write up can be done by just asking your code assistant providing the link to the MCP framework (most don't have that in their knowledge) and a brief description.

Step 1: Decide How the Server Will Run

MCP follows a client–server model. Your MCP server is a separate process (or service) that exposes tools and resources. The client (an LLM app, IDE, or your own script) connects to it and exchange messages.

During development, you will usually:

  • Run the server as a local process using STDIO transport, so the client can spawn it as a child process and talk over stdin/stdout.
  • Or run it as a local HTTP service that streams results back via SSE.

Pick whichever transport your target client supports best. For example, desktop apps often prefer STDIO, while remote/cloud scenarios lean toward Streamable HTTP.

Pitfall: SSE is already deprecated, prefer Streamable HTTP.

As part of this step, give the server a clear identity: a name, a version, and if relevant a URL. When you eventually have multiple servers, this makes configuration and debugging far less confusing.

Trick: describe what your server will do to an LLM and request an appropriate name, naming is one of the most difficult tasks.

Step 2: Define One Simple Tool

Next, define a single tool that does something obviously correct and easy to test. For example, you might create a echo_message or get_server_status tool that returns a predictable structure. This serves as the starting point for you to get familiar with the communication between client adn server.

When you define the tool:

  • Choose a short, machine‑friendly name, using something like get_weather or searchPosts rather than a sentence.
  • Write a concise description that tells the model when it should use the tool.
  • Specify its input and output structure using JSON Schema or the type‑annotation mechanism your framework provides.

If you are using Python, you will usually do most of this through function annotations and decorators. The framework will turn those into the JSON Schema that the client see.

Tip: if the LLM being used is suffering to understand when or how to call your tool even with a good description and proper input schema, try providing it through the system prompt in greater detail, you might spend a little more tokens but will save you back and fourth with the server.

Step 3: Run the Server in Dev Mode and Call the Tool

With one tool defined, run the server in whatever "dev" mode your framework provides. This might be a CLI command that starts the server and prints logs, or a helper that runs the server in an interactive shell.

Then, before you involve an LLM:

  • Use the inspector or dev CLI to list available tools.
  • Call your simple tool with a few different inputs.
  • Check that the responses match the schema you defined and that errors are clear when you send bad input.

Tip: If this flow is not solid, fix it now because it is far easier to debug problems at this stage where only the client and server are involved than later when the model is in the mix.

3. Harden and evolve your server

Once your initial tool is working end-to-end, you can gradually shape the server into something you are comfortable deploying. This section walks through aspects that tend to matter most in real projects: schemas, security, logging, testing, authentication, and response design.

Design Strong Schemas and Names

Tool schemas are how the model learns to talk to your server, so treat them as part of your API design. For each tool, make sure you have:

  • A unique name that does not overlap conceptually with other tools.
  • A clear description that references the kinds of questions or tasks the tool is meant to handle.
  • Input and output schemas that are as specific as they reasonably can be.

Pitfalls: defining two tools that are very similar or too many tools will cause the model to call the wrong one or with wrong inputs, separate the responsabilities clearly apply Separation of Concerns (SoC) and Single Responsibility (SRP) for your tools as well.

Use enums, minimum/maximum bounds, and required fields to constrain the input. This not only helps the model avoid invalid calls but also gives you strong validation at runtime. On the output side, keep the structure stable and predictable; avoid mixing unrelated concerns into a single response.

Pitfalls: test your schema, sometimes providers say that they support certain schema constraint but they don't, sometimes that don't specify and they don't support, make sure the schema is being respected by your provider API before blaming the LLM or the server.

Apply Least‑Privilege Access

From the moment a tool can touch real data or systems, think in terms of least privilege. The MCP server will often run with more access than an individual user, so you need to be explicit about what is allowed.

For database‑style servers, keep the initial implementation read‑only where possible, using parameterized queries and, ideally, a whitelist of allowed operations. For filesystem tools, restrict access to specific directories and file types instead of handing the server a path to the entire disk. Avoid exposing raw shell execution unless you have a very strong sandbox and a compelling reason.

It helps to review your tools as if they were endpoints in a public API: what could a malicious or buggy client do if it called them with arbitrary parameters?

Trick: you can customize how your MCP client is calling the server and inject user credentials to enable user-specific access.

Log Safely, Especially with STDIO

When you use STDIO transport, stdout and stdin carry the protocol messages. Any stray text you print to stdout will mix with those messages and confuse the client.

The fix is straightforward: send logs to stderr or to a dedicated logging sink, and keep stdout reserved for the stream. In practice this means configuring your logging library accordingly and avoiding ad‑hoc print or console.log calls in code paths that run under STDIO.

With HTTP transports you have more freedom to use stdout, but it is still worth keeping logs structured and separate from response bodies so your observability tools can parse them cleanly.

Test Servers in Isolation

You will move much faster if you treat your MCP server as a testable service in its own right. Once you have one or two tools you should:

  • Write ordinary unit tests for the underlying business logic.
  • Add a few integration tests that spin up the server and validate tools through the MCP interface.

Combine that with manual testing in an inspector or dev shell. When an LLM later behaves oddly, you will be able to tell whether the problem is in your server or in the LLM.

Plan Authentication Early

As soon as your server connects to protected resources or acts on behalf of users, you need an authentication story.

Decide how clients will present credentials (via environment variables, headers, or a dedicated configuration mechanism) and how those credentials will map to actual permissions in your backend.

Avoid the easy but dangerous pattern of using a single unique credential for all requests. Instead, design the server so it can act with user‑scoped tokens or roles, and so it enforces authorization checks before performing sensitive actions. The details depend on your environment, but the key is to make this a deliberate design rather than an afterthought.

Trick: customize your MCP client to send user-credentials (like a JWT) if it doesn't support it out-of-the-box.

Shape Responses for LLMs

Remember that tool responses end up inside the model's context window. Walls of low‑signal (without meaning) text make it harder for the model to do something useful with the data and very large payloads risk hitting token limits.

As you refine your tools, prefer responses that are:

  • Concise and focused on the information a caller actually needs.
  • Structured as JSON or as small, well‑formatted Markdown fragments.

For large datasets, paginate results, summarize where possible, or provide follow‑up tools and resource links that let the client fetch more detail on demand. When in doubt, experiment: call the tool from your target model and see how it uses the data, then adjust the shape and wording of the response to make its job easier. Another technique is to summarize the results with a secondary, usually smaller, LLM so the core of the result is preserved but you limit the amount of token returned.

Pitfall: avoid looping over summarization and callbacks, that is an easy way to create an infinite loop (it happened). You can set a maximum of tool calls depending on your client framework.

4. Advanced patterns and tricks

Once the basics are in place and you are comfortable with the core APIs, you can start using patterns that make MCP feel more like a full‑featured application platform: sessions, streaming, composition, and auto‑generated interfaces.

Keep Conversation State in Sessions

Sessions let MCP servers remember information across multiple tool calls in the same conversation. This is useful when you want to:

  • Cache expensive query results rather than re‑fetching them.
  • Reuse cursors or pagination tokens when a user is "browsing" through data.
  • Store user‑specific configuration or authentication context.

In practice, you will wire session support through your framework and then store a small, explicit state object per session. Treat that object like any other piece of stateful infrastructure: think about when to create it, when to clear it, and how to avoid leaking it between users.

Pitfall: avoid using global variables, remember that your tools will be called by different users.

Stream Results and Progress

For tools that might take more than a second or two, streaming improves UX dramatically. Instead of keeping the user and model waiting for a single, large response, you send partial results or progress updates as you go.

On the server side, that usually means writing handlers that yield chunks or emit events as work completes. On the client side, an MCP‑aware LLM or UI can surface that streaming output as incremental updates to the user.

Tip: if your tool takes a long time to finish (like calculating a large fibonacci number), stream updates logs to keep the user informed about the steps that are being taken inside the application to fulfill the request while the final response is not returned.

Compose Multiple Servers Behind One Interface (this one is gold)

As your system grows, you may end up with several MCP servers, each dedicated to a specific domain: code, chat, storage, internal APIs, and so on. You do not have to expose all of them directly to the model.

Instead, you can build a "meta‑server" that presents a single, coherent toolset and internally forwards calls to the underlying servers. That meta‑server can handle:

  • Routing: deciding which backend server should handle which request.
  • Cross‑cutting concerns: authentication, rate limiting, audit logging, caching.

You can also create sub models that will be able to do this routing with limited context information, this will prevent your main LLM from being overwhelmed and will allow other smaller LLMs to be highly specialized which means that there is less risk of hallucinations and generation errors.

From the model's perspective, both approaches are simpler: it talks to one server with clear tools, while you retain the ability to reshuffle and scale the underlying infrastructure.

Tip: as a rule of thumb, provide all tools to your main model and specialize if other techniques (tool description refinement, prompt engineering, few-shot) fail to teach the model how to interact with the tool

Generate APIs and UIs from Your Tools

Some MCP frameworks can also project your tools into other interfaces. Once you have a set of well‑typed tools, you can often:

  • Expose them as REST endpoints.
  • Generate an OpenAPI specification.
  • Serve a small web UI that lets you invoke tools directly.

This is particularly handy when you want humans or non‑LLM services to call the same functionality, or when you want a quick way to demo and debug tools in the browser without touching the LLM configuration.

These are usually provided by the official SDKs, but even if they are not, they are worth the time spent.

5. Debugging and common pitfalls

Even with good design, MCP integrations can fail in ways that are not obvious. This section will list the most common problems like "if you see X, check Y" terms so you have a mental checklist when something goes wrong.

Security Missteps

If you realize that your server can reach more data than a user should see, or that a single leaked token would compromise everything, you probably have a least‑privilege problem. Audit where credentials come from, how they are stored, and which tools they power. Move toward user‑scoped tokens or roles, and make sure sensitive tools log who called them and with what parameters so you have an audit trail.

Be really cautious when running third‑party MCP servers. Treat them like any other piece of third‑party software: review the code when you can, pin versions, and be explicit about how much access they get to your environment.

Tip: Security must be treated as a first-class concern in any MCP integration. While LLMs are great for exploring ideas or spotting unexpected patterns, they should not be relied on to design or validate authentication or other security-critical components. Always have security-experienced engineers architect and review these parts, and use proper observability and security tooling to catch issues early.

Broken Streams from Logging

If your client complains about malformed messages, or a server that used to work suddenly starts failing after you add logging, suspect STDIO logging first. Check whether any print or console.log calls write to stdout. If so, redirect them to stderr or a proper logger and try again.

Connectivity and Transport Headaches

When a remote server will not connect at all, work from the network layer upward. Use tools like curl or a REST client to hit the server's HTTP endpoint directly. If that fails, check firewalls, port mappings, and TLS configuration. If it succeeds but streaming breaks, look at any reverse proxies or load balancers in front of the server and verify that they are configured to allow Streamable HTTP (or SSE) and long‑lived connections.

For web‑based MCP clients, CORS misconfiguration is another frequent culprit. Ensure that your server advertises the correct Access-Control-Allow-Origin headers for the domains that will host the client.

Data Volume and Memory Issues

If you see timeouts, memory spikes or truncated results when tools handle large datasets, assume you are returning too much at once. Add limits and pagination to your tools, and consider streaming large results instead of building one giant object in memory. You can always provide additional tools or parameters that let the client (the caller model) "drill down" when it needs more detail.

Confusing Tool Selection

When an LLM keeps calling the wrong tool or refuses to call a tool that obviously fits the task, start by reviewing your tool names and descriptions. Overlapping names and vague descriptions confuse the model. Tighten them up, merge overly similar tools, and, if necessary, adjust your system prompt to spell out when the model should use a particular tool.

Trick: you can ask the LLM to explain in its own terms when and how to call the tools they have available so you know better what to fix.

For tools that can cause real world side effects, add explicit confirmation steps or human approval flows. Even with good descriptions, you do not want a single missfired tool call to cause irreversible changes.

Compatibility Drift

If something breaks right after updating a client or SDK dependency, do not discount simple version skew. Check release notes for both the MCP client and the SDK/framework, and confirm that the versions you are running are meant to work together. Keeping a short "supported versions" table in your repo can save future you and other developers a lot of guesswork.

Multi‑System Debugging

When requests are bouncing between an LLM, an MCP client, your server, and downstream services, debugging by intuition does not scale. Build yourself a minimal scripted client or use the inspector tools to talk to your server directly. Once you know the server behaves correctly in isolation, you can shift your attention to prompts, client configuration, or network issues with more confidence.

Pitfall: tools that can call themselves can create a loop unintentionally, let's say you have the schedule_flight -> user_agenda -> time_slot, this seems pretty straightforward, but the model have schedule_flight in context and might decide to start another iteration over returning the result (this also happened).

Performance Bottlenecks

If your MCP server feels slow, start by measuring where the time goes. Many problems come down to blocking I/O or CPU‑bound work happening on the same thread or event loop that handles incoming requests. Move long‑running tasks into background workers, adopt async patterns where your language supports them, and add caching around particularly expensive queries. Focus on the slowest and most frequently used tools first.

Tip/Pitfall: you never need to optimize everything, focus on what matters.

Conclusion

At Nearform, we view MCP servers as the connective tissue between AI agents and your domain-specific tooling. By defining clear tool and resource schemas, enabling discovery and monitoring, and embedding them, you empower AI-driven workflows that are maintainable, observable, and developer-friendly. When treated as an orchestration layer rather than just another backend service, your MCP server enables the agent to act as a true co-developer instead of a black box.

Building with the Model Context Protocol then becomes less about learning a new standard and more about designing clear interfaces between your models and the systems they need to use. The frameworks help you get started quickly, but the real value comes from applying sound engineering practices of starting small, testing thoroughly, validating schemas, and treating every tool as part of a public API.

As your integrations grow, these same principles naturally extend to more complex scenarios. Concepts familiar from enterprise-grade software systems, such as least privilege, observability, state management, streaming, and modularity, map directly to successful MCP projects. With thoughtful tool design and disciplined implementation, you can build servers that are easy for models to understand, safe to run in production, and powerful enough to anchor sophisticated, evolving workflows.

But wait - there's more.

Nearform publishes real-world learnings on data & AI, engineering, and digital strategy - with more merged in weekly.

Insights

Perspectives on AI in engineering, product development, and strategy, for enterprise executives.

Community

Deep dives and tutorials by engineers, for engineers.

Insight, imagination and expertly engineered solutions to accelerate and sustain progress.