Benchmark and demo#
Two questions, measured separately:
- How much smaller is what the model reads, compared with calling the upstream Wikidata MCP tools directly for the same question, and what is left out?
- What does a real session cost in requests and time, cold and with a warm cache?
Neither is an accuracy benchmark. There is no human-labelled evaluation set behind these
numbers, so they say nothing about how often AUTO_MATCH is right; see
Limitations.
Recorded demo#
The player needs JavaScript. The text transcript has the same content.
A real run of v0.2.1 against live Wikidata, Wikidata only (Google not configured). Typing
speed in the recording is simulated; command output and timings are real, filtered with jq
for readability. Text transcript ·
recorder script.
What it shows: a namesake-heavy search returns 3 candidates with warnings; name plus city
resolves to HOLD; adding the official website gives AUTO_MATCH with the per-candidate
evidence explaining why the others lost; repeating both calls costs zero upstream requests.
1. Compact output versus raw upstream text (offline replay)#
Method. Each case sends the same question through the shipped MCP tool, over an in-memory
MCP client, to a fake Wikidata MCP server that replays real responses captured from the
public Wikidata MCP service (tests/fixtures/wikidata/, Wikidata content under CC0).
Raw provider bytes is the upstream text the tool consumed, which is what a raw Wikidata MCP
workflow would put in the model's context for the same question. Compact bytes is this
tool's JSON answer; MCP envelope bytes adds the MCP content block and JSON escaping. Token
counts use the tiktoken o200k_base encoding. Cold means a fresh cache directory, warm the
same call again. The run was made with outbound HTTP pointed at a dead proxy, so no live
request could have happened.
| case | raw provider bytes | compact bytes | MCP envelope bytes | size change | tokens raw → compact (o200k_base) | omitted / notes |
|---|---|---|---|---|---|---|
| search: Fox Theatre, place Atlanta | 4,374 | 1,694 | 1,953 | -61.3% | 1026 → 442 | 50 provider hits, 3 candidates shown (47 omitted); resolution single_supported_candidate; warnings: same_name_elsewhere: 13 other candidate(s) share this name but do not mention 'Atlanta'; confirm location from statements, verify_before_use: description-level evidence only; confirm with entity <id> statements (P31/P131/P17/P625) |
| search: Tate Modern, place London, type museum | 3,900 | 1,332 | 1,560 | -65.8% | 1004 → 355 | 50 provider hits, 2 candidates shown (48 omitted); resolution single_supported_candidate; warnings: verify_before_use: description-level evidence only; confirm with entity <id> statements (P31/P131/P17/P625), mixed_kinds: building, venue — do not conflate organization/building or series/occurrence |
| search: Carnegie Hall, place New York | 4,373 | 1,624 | 1,886 | -62.9% | 1003 → 425 | 50 provider hits, 3 candidates shown (47 omitted); resolution single_supported_candidate; warnings: verify_before_use: description-level evidence only; confirm with entity <id> statements (P31/P131/P17/P625), mixed_kinds: building, venue — do not conflate organization/building or series/occurrence |
| search: Primavera Sound (es), type festival | 1,565 | 1,692 | 1,968 | +8.1% | 400 → 439 | 25 provider hits, 3 candidates shown (22 omitted); resolution single_name_match; warnings: geo_unverified: no place context given; a single name match is not identity evidence, verify_before_use: description-level evidence only; confirm with entity <id> statements (P31/P131/P17/P625), mixed_kinds: event_occurrence, event_series — do not conflate organization/building or series/occurrence |
| search: unmatched name | 3,507 | 573 | 708 | -83.7% | 997 → 165 | 50 provider hits, 0 candidates shown (50 omitted); resolution no_match; warnings: no_label_match: provider returned 50 related items but none match the name; treat as unmatched |
| entity: Q193375 P31,P131,P17 + evidence P131 | 10,765 | 1,604 | 1,823 | -85.1% | 3537 → 429 | 3 requested properties as best-rank values; 43 other properties omitted; P131 evidence: 2 values with rank and qualifiers, 1 reference(s) counted, 1 sample line(s) kept of 5 raw reference lines |
| related: Q193375 class hierarchy, depth 1 | 214 | 580 | 713 | +171.0% | 69 → 186 | 3 class nodes, same content as the raw tree; the compact JSON adds labels, warnings and the meta block, so it is larger than this tiny upstream answer |
Median size change -62.9% (range -85.1% to +171.0%) over 7 cases. Wall-clock per call on the replay transport: cold median 7.4 ms (p10–p90 6.1–10.8 ms), warm median 2.3 ms (p10–p90 1.9–3.4 ms), 21 iterations per case; upstream requests cold 29, warm 0. Package 0.2.0, Python 3.14.4, generated 2026-09-28.
Reading it.
- Searches with many hits (the Fox Theatre, Tate Modern, Carnegie Hall and unmatched cases, 50 hits each) shrink by 61–84% because 3 candidates with match flags replace the list. The omitted hits are counted in the notes: they are not shown, and the warnings say when same-named candidates exist elsewhere.
- The entity case shrinks most because only the requested properties are returned; 43 other properties are reported as omitted. The evidence block keeps rank, qualifiers and the reference count for each value, but only one sample reference line of the five in the raw text. If you need every reference, ask Wikidata for that statement directly.
- Small answers get larger: the Primavera Sound search (25 short hits) and the one-level
class hierarchy are bigger in compact form, because the answer adds match flags, warnings and
the
metablock. Compaction pays off on large responses, not on every call. - Replay timings measure local processing only. They say nothing about network latency.
Reproduce:
git clone https://gitlab.com/revanalex/wikidata-google-knowledge-mcp.git && cd wikidata-google-knowledge-mcp
uv run --group bench python benchmarks/replay_benchmark.py --iterations 21
Raw results: replay.json.
2. Live session: cold and warm#
The same run as the demo above, against live Wikidata. Each call ran once.
| call | run | wall-clock | upstream requests | cache | raw provider bytes | compact bytes |
|---|---|---|---|---|---|---|
wdkg search "Fox Theatre" --place Atlanta |
cold | 3241 ms | 3 | miss | 4,374 | 1,694 |
wdkg resolve "Fox Theatre" --kind place --city Atlanta |
cold | 2447 ms | 3 | miss | 0 | 985 |
wdkg resolve "Fox Theatre" --kind place --city Atlanta --url https://www.foxtheatre.org --explain |
after step 2 | 618 ms | 0 | 13 cached provider value(s) | 0 | 2,082 |
wdkg search "Fox Theatre" --place Atlanta |
warm | 686 ms | 0 | hit | 0 | 1,678 |
wdkg resolve "Fox Theatre" --kind place --city Atlanta --url https://www.foxtheatre.org |
warm | 976 ms | 0 | resolution cache hit | 0 | 1,034 |
Live Wikidata, n = 1 per call, 2026-09-28, package 0.2.0, 6 upstream requests in total; Google not called. Network latency dominates and varies; treat these as one observation, not a distribution.
"after step 2" means the call reused Wikidata values that step 2 had cached (the tool reports 13 cached provider values) and computed a new decision from them. The warm calls report a cache hit (search) and a resolution-cache hit (resolve) with no upstream requests.
Raw results: live.json.
Limitations#
- Cohort. Seven replay cases around four well-known names and one unmatched name, chosen because their responses were already captured for the test suite. They are illustrative, not a sample of real workloads. Your byte savings depend on how many hits and properties your questions produce.
- Equivalence. The raw side is the upstream Wikidata MCP text for the same question. An agent working with raw tools would often need more calls (for example statements for each candidate) to reach a decision, so the raw side is a lower bound on what it would read.
- No accuracy claim. Agreement between Google and Wikidata, fixture correctness and model self-review are not a gold standard. Measuring identity precision needs an independently labelled set, which this project does not have yet.
- One live run. The live table is n = 1 per call on one day from one network location.
- No Google. Google Knowledge Graph was not called for either benchmark. Google responses cannot be redistributed, so any Google-shaped test data in this repository is synthetic and labelled as such.