AI assistants & MCP

The MCP-First GTM Stack: What to Connect, and What No One Should Build

Every guide on this topic is a list of servers to add. The benchmark evidence points the other way: frontier models fail more than half of real MCP orchestration tasks, and each server you connect makes the next answer slightly worse. Here is the stack as a subtraction problem — and the test for what a vendor should never ship at all.

By Linkeddit·Published September 1, 2026·14 min read

Key takeaways

  • Connect two or three servers, not eight. Tool definitions are re-sent on every request, and a larger tool set measurably degrades tool selection.
  • Salesforce AI Research's MCP-Universe benchmark found GPT-5 succeeded on 43.72% of real-world MCP tasks, Grok-4 on 33.33%, and Claude-4.0-Sonnet on 29.44%. Orchestration is the bottleneck, not reasoning.
  • A December 2025 census counted 36,039 MCP servers; 51% have zero stars and 61% are solo projects with no forks. Ecosystem size is not a menu.
  • Choose by job, not by vendor. A GTM stack has five jobs, and most teams only have a bottleneck in two of them.
  • On the supply side: build only where the client cannot get the data, cannot provide the durability, cannot keep the judgment comparable, or cannot govern the side effect. Everything else is a wrapper.

01The short answer

A go-to-market team should connect two or three MCP servers, chosen by the questions the team actually asks every week, and should treat every additional connection as a cost rather than a feature. Connecting an MCP server is close to free, which is exactly why the instinct to connect everything is expensive. Each connected server adds its tool definitions to the set the model must choose between, those definitions are re-sent with every request, and past a certain number the model starts choosing wrong — reaching for a search tool when it wanted an enrichment tool because both descriptions plausibly fit.

That is the demand-side answer. There is a supply-side answer that almost nobody writes down, and it matters just as much if you are the one building the server: do not build what the client already does well once it has your data. Most of what vendors are currently shipping into MCP servers — drafting, summarizing, rewriting, narrating, rendering — is work the model on the other end does better, with the user's own context available to it. The four-question test in section six is how to tell the difference.

36,039
MCP servers counted in a December 2025 crawler census
51%
of them have zero GitHub stars
43.7%
GPT-5 success rate on real-world MCP tasks (MCP-Universe)
2-3
servers most GTM teams should actually connect

02What the census actually shows

The standard opening line for articles in this category is that there are thousands of MCP servers and you are missing out. The number is real. The implication is not.

A crawler census of the MCP corpus published in December 2025 counted 36,039 MCP servers across 32,762 unique GitHub repositories. The same census found the shape of that population: 51% have zero stars, 77% have fewer than ten, 61% are solo projects with no forks, and 16% ship without a README. The top 50 repositories account for 60% of all stars, and 83% of publishers have released exactly one server. Growth ran from 135 new servers a month at launch in November 2024 to 5,069 a month by June 2025, then cooled to 2,093 a month by that November.

The median MCP server has 0 stars. The ecosystem is mostly experimental projects, tutorials, and personal tools.
Crawler census of the MCP corpus, via r/mcp (December 2025)

By mid-2026 the marketing-specific slice had matured considerably. A June 2026 map counted more than 55 marketing platforms with live servers, including an official Meta server shipped in April 2026 with 29 tools, and noted that Google Ads — the largest ad platform in the world — still had no official server. At the far end of the distribution, GoHighLevel ships 563 MCP tools across 44 categories, which is less a connector than a second product surface.

03Why more servers make the assistant worse

There are two mechanisms, and they compound. Neither is a matter of opinion.

Mechanism one: tool definitions are re-sent on every request

A connected server does not sit idle until called. Its tool definitions travel with each request so the model knows what is available. Practitioners measuring this report that connecting four common servers — GitHub, Linear, Context7, and Playwright — to Claude Code consumes over 60,000 tokens, roughly a third of the context window, before a single question is asked. One engineer described running four servers that each contribute five to fifteen tools, leaving 40-plus tool definitions resident in context “even when I'm just asking Claude to fix a typo.”

Mechanism two: selection accuracy falls as the tool set grows

The more consequential effect is on choosing. Salesforce AI Research published MCP-Universe, a benchmark that evaluates models against real MCP servers rather than toy ones, across more than 20 leading models. The results are sobering for anyone assuming this problem is solved:

ModelSuccess rate on real-world MCP tasks
GPT-543.72%
Grok-433.33%
Claude-4.0-Sonnet29.44%

VentureBeat summarized the finding as GPT-5 failing more than half of real-world orchestration tasks. The gap is not the models' reasoning — these are the same models that write and analyze well. It is orchestration: picking the right tool, filling its parameters correctly, and chaining calls across servers. Every server you add expands the space in which that decision is made.

Too many fine-grained tools and the model spends its context budget just figuring out which tool to call.
Team reporting on building ~40 MCP servers, via r/mcp

There is a third, quieter cost. Teams that hand-rolled roughly 40 servers found that each ended up with its own slightly different auth pattern, which meant “40 places a credential could be wrong, stale, or overscoped.” For developer tooling that is an annoyance. For a GTM stack, where the connected systems hold customer records, it is a governance problem with a compliance shape.

04The subtraction rule

Connect the servers that answer questions you ask weekly. For most teams that is two or three, not eight. This is the single highest-leverage rule in the category, and it inverts how almost every guide on the subject is organized.

The test is behavioral, not aspirational. Do not ask which servers look useful; ask which questions your team actually asked out loud in the last two weeks, then connect only what answers those. A server you connected for a hypothetical future workflow is not neutral — it is spending context and widening the selection space for every real question you ask in the meantime. Removing a server you never use measurably improves the ones you keep.

05The five jobs a GTM stack has

Sorted by job rather than by vendor, the landscape is much smaller than the server count suggests. Most teams have a genuine bottleneck in two of these five, and connecting for the other three is where the bloat comes from.

JobThe question it answersConnect when
System of recordWhat is actually in our pipeline right now?Almost always first. It improves the quality of every other answer, because the assistant stops reasoning about a generic company and starts reasoning about yours.
Find and verify peopleWho is the human at this account, and how do I reach them?Your bottleneck is contact coverage, not target selection.
Build and orchestrateCan I construct a list, run custom logic over it, and push it somewhere?You have a dedicated GTM engineer. Without an operator, a build surface is a project, not a tool.
Know what changedWhich accounts, competitors, or answer engines moved this week, and why does it matter?You cannot tell which accounts deserve attention. This is the job teams skip most often, because its output is harder to count than a list of verified emails.
Know who you knowWho at our company already has a warm path into this account?Enterprise motions where a warm introduction materially beats a cold one.

The sequencing matters more than the selection. Connect the system of record first. Then connect whichever of enrichment or monitoring matches your real bottleneck — and be honest, because teams reflexively answer “more contacts” when the actual constraint is that nobody knows which accounts are worth working. If you run both, run monitoring first to narrow the field and enrichment second to resolve the people who survived. Enriching a full buying committee at an account with no buying window is an expensive way to discover there was no buying window.

The first-party servers in these categories have matured fast. HubSpot's remote server reached general availability in April 2026 with OAuth that works across Claude, ChatGPT, and Cursor. Apollo runs an official server with connectors inside Claude, ChatGPT, and Perplexity. ZoomInfo includes its server with every subscription, keeps its tools read-only, respects existing entitlements, and explicitly points teams needing bulk export or scheduled jobs to the API instead — a distinction worth generalizing. MCP is built for interactive, in-conversation work. It is not a replacement for your ETL.

06What no one should build

Now the supply side, which is missing from essentially every guide on this topic. If you are building an MCP server rather than choosing one, the governing rule is short:

If the capability is something the client already does well once it has your data, do not build it. Expose the data and let the client do it.

The example that makes this concrete is drafting. Writing a brief or an article from evidence is something a frontier model does better than a vendor's in-product generator ever will, with the user's own tone, context, and follow-up questions available to it. So the right move is not to build a generator. It is to return the gap, the evidence, and the citations, and let the client write. That decision generalizes further than it first appears.

Build only if the answer to at least one of these four questions is yes. If all four are no, the client can already do it, and building it makes you a wrapper.

The questionWhat it coversExample
1. Data the client cannot getProprietary corpora, grounded measurement, an evidence store, first-party logs. An ad-hoc web search is not a substitute for a versioned, disclosed measurement.Measuring across four answer engines; crawler logs from the customer's own CDN; Search Console queries.
2. Durability the client cannot provideScheduling, state that survives across sessions and clients, idempotency, long-running fan-out, budget caps. A conversation ends; a weekly monitor does not.A recurring audit; an action queue shared between the agent and a human.
3. A versioned, evaluated judgmentA number that must be computed identically for every customer on every run, carry a parser version, and be graded against a gold set — or it is not comparable and the trend is noise.Mention extraction, sentiment classification, share-of-voice, gap ranking.
4. A governed side effectAn authorization boundary, an approval gate, and a durable record. None of which a chat session provides.Saving an artifact for human approval; publishing to a CMS; credential custody; an audit record of who did what with which scope.

Run the test honestly and a large amount of typical roadmap disappears. Rendering a chart is not a capability — return well-shaped data and let the client render. Narrating what a finding means is not a capability — it has the structured finding. Rewriting, tone-adjusting, translating, generating an image: all squarely the model's job. What survives the test is measurement, memory, governance, and access to things the client genuinely cannot reach.

07The measuring-instrument exception

One clarification, because “do not build what the model does” is frequently over-applied into “never use an LLM in your product.” That is wrong, and the distinction is precise:

An LLM used as a measuring instrument is not a wrapper. An LLM used as a writer is.

Extraction and classification tasks — did this answer recommend a competitor, what sentiment did it carry, what type of source was cited — are LLM-powered internally and should stay that way. They must also stay yours, because a metric is only meaningful if it is computed the same way for every customer on every run. That means a version stamp on the parser, a gold set it is graded against, and a regression that blocks the deploy. Hand that judgment to the client's ad-hoc reasoning and the number stops being comparable across tenants and across time, at which point the trend line you are selling is noise with a chart on it.

There is a commercial edge to this too, and it is worth being explicit about. Generation is the one AI workload whose cost scales with customer usage rather than with your own sampling policy. Under an MCP-first design, that workload runs on the customer's model subscription rather than your cost of goods — which is why vendors who build in-product generators end up metering articles per month, and vendors who refuse to do not have to. Declining to be a wrapper is an architecture decision and a margin decision at the same time.

08Evaluating a server in twenty minutes

Vendor tool lists tell you what exists, not what works. Four checks separate the two, and none of them takes long.

  1. 1

    Ask your five real questions

    The ones your team asked last week, not the ones in the demo script. Note which the server answers completely, which partially, and which it misunderstands. A server that nails demo questions and fumbles yours is optimized for the sale.

  2. 2

    Check whether claims carry sources

    Ask something verifiable, then ask where each fact came from. A server returning confident unsourced claims into a customer-facing conversation is a liability — the first time someone repeats a hallucinated fact to a prospect, the tool is finished inside that organization, and reasonably so.

  3. 3

    Test your actual territory and segment

    Coverage that looks complete in a US demo routinely thins out elsewhere. Run the same query against the market you actually sell into before you believe the coverage claim.

  4. 4

    Check the credit model before you scale

    Some servers meter per operation and some do not. This matters far more under an agent than under a human, because an agent in a loop spends quickly. Servers that confirm before spending credits are the safer default in an agentic setup.

Add one more for authentication, given the credential-sprawl finding above: prefer a server that authenticates with OAuth over one that wants a long-lived API key pasted into a local config file. An OAuth connection can be revoked centrally in one place. A key in a config file has to be hunted down on every machine it was ever copied to.

09What we built, and what we refused

We build one of these, so the honest thing is to show our own work against the test rather than exempt ourselves from it.

Linkeddit is demand and competitor intelligence: measuring where AI answer engines recommend competitors instead of you, tracking what those competitors ship and what their users complain about, and turning that signal into pipeline and content. It ships as a single MCP connector rather than five, which is the subtraction rule applied to ourselves — one connection, one OAuth grant, one thing to revoke.

CapabilityVerdictWhy
Measure whether an engine recommended a competitorBuildQ1 and Q3. The client cannot obtain grounded cross-engine measurement, and the number has to be comparable across tenants and over time.
Run a weekly audit on a scheduleBuildQ2. A chat session ends. A monitor does not.
Hold the action queue between an agent and a humanBuildQ2. Shared state across sessions and clients.
Save a draft for human approval, publish to a CMSBuildQ4. Authorization boundary, approval gate, audit record.
Draft the brief or article that closes a gapRefusedThe founding example. The client writes better, with the user's tone and follow-ups available. We return the gap, the evidence, and the citations.
Narrate what a finding means, in proseRefusedIt already has the structured finding.
Render the chart or comparison tableRefusedRendering is not a capability. Return well-shaped data with its disclosure attached.

The corollary we hold ourselves to is that every new integration ships its connector tool in the same change as the integration, not in a later “expose it to the API” phase. If the customer's workspace is Claude or ChatGPT, a capability that only exists in a web app is a capability most users will never reach.

One connector, not five

Linkeddit ships as a single keyless OAuth connector covering AI-search visibility, competitor intelligence, leads, keyword and Search Console data, content, and community signal. One connection to add, one to revoke. Setup is a URL, a client id, and one browser sign-in.
See the Linkeddit connector

Frequently asked questions

How many MCP servers should a go-to-market team connect?+

Two or three, in most cases. The instinct is to connect everything because connecting is nearly free, but every connected server adds its tool definitions to the set the model must choose between on every request, and those definitions are re-sent with each call. Practitioners report that four common servers in Claude Code consume over 60,000 tokens, roughly a third of the context window, before any work happens. Connect the servers that answer questions you ask weekly and remove the rest — subtraction measurably improves the ones you keep.

Does connecting more MCP servers make an AI assistant worse?+

Yes, in two measurable ways. First, tool definitions consume context on every request, leaving less room for the actual conversation. Second, a larger tool set degrades tool selection: the model reaches for a search tool when it wanted an enrichment tool because both descriptions plausibly fit. Salesforce AI Research's MCP-Universe benchmark, which evaluated more than 20 leading models against real MCP servers, found GPT-5 succeeded on 43.72% of tasks, Grok-4 on 33.33%, and Claude-4.0-Sonnet on 29.44%. Orchestration across real servers is the weak point, not the models' reasoning.

How big is the MCP ecosystem actually?+

Larger than it is useful. A crawler census of the MCP corpus published in December 2025 counted 36,039 servers across 32,762 GitHub repositories, and found that 51% have zero stars, 77% have fewer than ten, 61% are solo projects with no forks, and 16% ship without a README. The top 50 repositories account for 60% of all stars. By mid-2026 a marketing-focused count put the number above 10,000 servers relevant to marketers alone, with 55-plus platforms shipping one. The ecosystem is a long tail of experiments with a small head of maintained, production-grade servers.

What should an MCP server actually expose?+

One tool per real user intent, not one per underlying API endpoint. Teams that have built dozens of servers report that fine-grained CRUD tools force the model to spend its context budget deciding which to call, while a single mega-tool with a dozen optional parameters gets called incorrectly because the model cannot reliably fill in that many fields. The middle path — a tool that maps to something a person actually wants — is the design that survives contact with an agent.

What should a vendor never build into an MCP server?+

Anything the client already does well once it has your data. Drafting, rewriting, summarizing, tone adjustment, translation, narration, chart rendering, and recommendation prose are all the model's job, and building them makes the product a wrapper. Build only where at least one of four conditions holds: the data cannot be obtained by the client, the durability cannot be provided by a stateless session, the judgment must stay versioned and comparable across tenants and time, or the action is a governed side effect in an external system.

Is using an LLM inside an MCP server the same as being an AI wrapper?+

No, and the distinction is precise. An LLM used as a measuring instrument is not a wrapper: mention extraction, sentiment classification, and gap ranking must be computed identically for every customer on every run, carry a parser version, and be graded against a gold set, or the resulting metric is not comparable and the trend line is noise. An LLM used as a writer is a wrapper, because the client's model can already write, with the user's own tone and follow-up questions available to it.

How do I evaluate an MCP server before committing to it?+

In about twenty minutes. Ask the five real questions your team asked last week, not demo questions, and note which the server answers completely, partially, or misunderstands. Ask where each fact came from — a server returning confident unsourced claims into a customer conversation is a liability. Test your actual territory and segment, because coverage that looks complete in a US demo often thins out elsewhere. And check the credit model before you scale, because an agent in a loop spends far faster than a human clicking buttons.

Do MCP servers create a credential problem?+

They can. Teams that hand-rolled roughly 40 servers reported that each ended up with a slightly different auth pattern, which meant 40 places a credential could be wrong, stale, or overscoped. For GTM data this matters more than for developer tooling, because the connected systems hold customer records. Prefer servers that authenticate with OAuth over ones that want long-lived API keys pasted into a local config file, and prefer one server that can be revoked centrally over five that cannot.

Sources

Community figures are attributed to the platform they were published on and not to individual accounts. Benchmark figures are quoted from the published paper and its coverage; ecosystem counts are quoted from the crawler census and the marketing map as of their stated publication dates, and the MCP ecosystem moves quickly enough that they should be treated as point-in-time measurements rather than current totals.

Connecting it takes about three minutes: the connector quickstart.