callwitness / baseline
Every published measurement of MCP context cost counts the schemas a server declares at install time. Nobody counts what it hands the model at runtime, because nobody runs the servers. This is that measurement.
Declared size does not predict delivered size. Not loosely — not at all.
mcp-deepwiki declares 774 B of schema, close to the smallest menu in the census, and returned 62 KB in a single call — 81x its declared size. mcp-sympy declares 63 KB and returned 198 B. Ranked by what they declare, these two servers sit at opposite ends of the table from where they land when ranked by what they deliver.
The install-time number, in other words, tells you almost nothing about the runtime cost — and the install-time number is the only one anyone currently publishes.
tools/list)
largest single response observed
A typical tool response is 580 B. One in twenty is 36 KB or larger. The largest single response measured was 496 KB, from 774 B of declared schema and a handful of bytes of argument.
This is why a fixed context budget behaves unpredictably in practice: the cut point is static, and the pressure on it varies by 6,437x depending on which argument the model happens to pick at runtime.
Two censuses have been run. mcp-deepwiki appears in both, declaring the same 774 B of schema each time, and the size it handed back moved by 11x:
| run | declared | delivered | ratio |
|---|---|---|---|
| first census | 774 B | 701 KB | 906x |
| second census | 774 B | 62 KB | 81x |
Nothing about the server changed. Its declared size is a constant; only the argument differed. Whatever an install-time number can tell you, it could not have told you this — and no ranking built on declared size survives a 11x swing in what actually gets delivered.
Both runs are published, as census/data/census-r1.jsonl and census/data/census.jsonl. The comparison above is computed from them rather than typed in.
| package | declared | largest returned | ratio | calls |
|---|---|---|---|---|
| mcp-deepwiki | 774 B | 62 KB | 81x | 1 |
| @jpisnice/shadcn-ui-mcp-server | 3 KB | 193 KB | 62x | 3 |
| mcp-chess | 3 KB | 40 KB | 13x | 3 |
| mcp-trends-hub | 12 KB | 131 KB | 11x | 11 |
| @modelcontextprotocol/server-filesystem | 13 KB | 97 KB | 7.46x | 3 |
| rcsb-mcp | 82 KB | 496 KB | 6.01x | 2 |
| mermaid-mcp-server | 2 KB | 6 KB | 3.54x | 2 |
| @react-spectrum/mcp | 4 KB | 9 KB | 2.23x | 2 |
| @heilgar/shadcn-ui-mcp-server | 2 KB | 3 KB | 2.09x | 2 |
| mcp-server-tree-sitter | 21 KB | 36 KB | 1.71x | 3 |
| duckduckgo-mcp-server | 5 KB | 8 KB | 1.64x | 1 |
| @primeng/mcp | 8 KB | 14 KB | 1.62x | 3 |
| @primevue/mcp | 8 KB | 13 KB | 1.54x | 3 |
| geocode-mcp | 497 B | 655 B | 1.32x | 1 |
| @modelcontextprotocol/server-everything | 8 KB | 10 KB | 1.24x | 8 |
| mcp-gnu-units | 10 KB | 10 KB | 0.99x | 2 |
| chrome-devtools-mcp | 26 KB | 25 KB | 0.97x | 6 |
| mcp-server-git | 6 KB | 5 KB | 0.91x | 2 |
| crossref-mcp | 11 KB | 9 KB | 0.86x | 7 |
| markitdown-mcp | 318 B | 264 B | 0.83x | 1 |
| mcp-dice | 335 B | 261 B | 0.78x | 1 |
| @knip/mcp | 4 KB | 3 KB | 0.76x | 1 |
| uuid-mcp | 185 B | 129 B | 0.70x | 1 |
| mcp-geo | 18 KB | 12 KB | 0.63x | 3 |
| mcp-hn | 2 KB | 951 B | 0.57x | 1 |
| unit-converter-mcp | 15 KB | 8 KB | 0.52x | 2 |
| mcp-nixos | 7 KB | 3 KB | 0.48x | 1 |
| @sinco-lab/mcp-youtube-transcript | 4 KB | 2 KB | 0.42x | 1 |
| polars-mcp | 4 KB | 2 KB | 0.41x | 1 |
| pypi-query-mcp-server | 13 KB | 5 KB | 0.39x | 2 |
| mcp-server-fetch-python | 2 KB | 665 B | 0.35x | 1 |
| mcp-server-calculator | 421 B | 129 B | 0.31x | 1 |
| open-meteo-mcp | 2 KB | 580 B | 0.30x | 1 |
| calculator-mcp-server | 294 B | 77 B | 0.26x | 1 |
| mcp-server-fetch | 1 KB | 281 B | 0.24x | 1 |
| mcp-server-time | 1 KB | 235 B | 0.19x | 1 |
| readability-mcp | 2 KB | 269 B | 0.17x | 1 |
| swiss-holidays-mcp | 44 KB | 7 KB | 0.16x | 3 |
| code-index-mcp | 9 KB | 1 KB | 0.15x | 4 |
| dictionary-mcp-server | 7 KB | 969 B | 0.13x | 1 |
| wikipedia-mcp | 23 KB | 3 KB | 0.11x | 1 |
| mcp-molecules | 7 KB | 765 B | 0.11x | 2 |
| arxiv-mcp-server | 18 KB | 2 KB | 0.11x | 4 |
| math-mcp | 1 KB | 153 B | 0.11x | 2 |
| orcid-mcp-server | 3 KB | 327 B | 0.11x | 1 |
| pubchem_mcp_server | 1 KB | 135 B | 0.10x | 1 |
| feed-mcp | 3 KB | 319 B | 0.10x | 1 |
| @modelcontextprotocol/server-sequential-thinking | 5 KB | 338 B | 0.07x | 1 |
| @wrtnlabs/calculator-mcp | 1 KB | 99 B | 0.07x | 1 |
| @mantine/mcp-server | 1 KB | 78 B | 0.06x | 1 |
| uniprot-mcp-server | 56 KB | 3 KB | 0.05x | 1 |
| mcp-server-math | 4 KB | 204 B | 0.05x | 1 |
| astronomy-mcp | 7 KB | 319 B | 0.04x | 1 |
| mcp-ast-explorer | 5 KB | 172 B | 0.04x | 1 |
| mcp-abacus | 33 KB | 935 B | 0.03x | 3 |
| registry-mcp | 145 KB | 4 KB | 0.03x | 1 |
| mcp-yahoo-finance | 4 KB | 101 B | 0.03x | 1 |
| random-number-mcp | 5 KB | 127 B | 0.03x | 1 |
| @agentdeskai/browser-tools-mcp | 26 KB | 558 B | 0.02x | 9 |
| kegg-mcp-server | 109 KB | 2 KB | 0.02x | 2 |
| @modelcontextprotocol/server-memory | 11 KB | 173 B | 0.02x | 2 |
| @executeautomation/playwright-mcp-server | 14 KB | 129 B | 0.01x | 1 |
| @playwright/mcp | 19 KB | 167 B | 0.01x | 3 |
| mcp-numpy | 38 KB | 208 B | 0.01x | 2 |
| mcp-sympy | 63 KB | 198 B | 0.00x | 1 |
17 further servers declared a menu but exposed no tool that could be called safely without side effects, so they have a declared size and no delivered one. They are in the raw data with their declared sizes intact.
140 calls across 65 servers, on one machine. That is a small sample, and it is the weakest thing about this work. I could not find another published measurement of delivered size, which makes this the best number available and still a thin one. If one exists, I would rather cite it than be the only source.
Every figure here carries its n. If you take a number from this page, take the n with it.
The servers were exercised with read-only tools chosen by an allowlist, one request at a time. Nothing was written, deleted or executed. Servers that refused without credentials, and servers that would not start, are published as part of the dataset rather than dropped from it — a census that silently excludes its failures is not a census.
The full method, every server that failed and why, and the rule that decided which tools were safe to call: how the census was taken.
The distribution is published as a versioned document, not a number in a blog post. Fetch it rather than copying constants, and it improves without you editing anything:
https://callwitness.tech/baseline/v1.json
Inside callwitness.baseline.v1, field names, units and meanings will not change. A breaking change becomes v2 at a new path, and v1 keeps resolving. Every server carries its per-call sizes, not only the aggregate, and returned_bytes_all holds every delivered size in one sorted list so any percentile is a line of code away.
These are someone else's servers. Yours will differ, and the numbers that matter for your context budget are yours:
pip install callwitness
callwitness run --label docs -- npx -y <your mcp server>
callwitness baseline --out mine.json
That produces the same schema, from your traffic. Anything reading the public baseline reads your file unchanged — the only field to branch on is origin, which is "census" in the published document and "local" in yours.
The sample above is one person's afternoon. callwitness contribute sends the shape of your traffic — sizes, counts and durations, never arguments, paths, hostnames or results — to widen it. It is off unless you turn it on, nothing is sent while the proxy is running, and --dry-run prints the exact bytes that would leave.