Wikidata + Google KG MCP

Batch and evidence#

The MCP tools resolve one entity per call. For a corpus, use the CLI: calling kg_resolve in a loop spends model context and tool calls on every row, while wdkg resolve-batch streams the file, shares provider requests between rows and keeps an audit trail.

Input#

One JSON object per line, with the same fields as a single resolve (identity envelope). source_id is echoed back so you can join results to your data.

{"source_id": "fox-atl", "name": "Fox Theatre", "kind": "place", "city": "Atlanta", "country": "US", "official_url": "https://www.foxtheatre.org"}
{"source_id": "fox-name-only", "name": "Fox Theatre", "kind": "place", "city": "Atlanta"}
{"source_id": "tate", "name": "Tate Modern", "kind": "place", "city": "London", "latitude": 51.5076, "longitude": -0.0994}
{"source_id": "primavera-series", "name": "Primavera Sound", "kind": "event_series", "city": "Barcelona", "official_url": "https://www.primaverasound.com"}

This is the start of examples/entities.jsonl.

Run#

# plan only: parse, normalize and check caches, no network
wdkg resolve-batch entities.jsonl --out resolved.jsonl --dry-run

# resolve with Wikidata only and a hard cap on requests
wdkg resolve-batch entities.jsonl --out resolved.jsonl \
    --max-wikidata-requests 2000 --max-google-requests 0

What you get:

  • resolved.jsonl: one record per input line, in input order, with decision, wikidata_qid, candidate_ids, reasons, the evidence code lists and candidate_pairs.
  • resolved.jsonl.receipt.json: decision counts, request usage per provider, complete, stop_reason and what to do next.
  • resolved.jsonl.evidence.jsonl: the full evidence for every row (see below).

The command itself prints a one-line summary:

{"op": "resolve-batch", "status": "ok", "rows": 6,
 "decisions": {"AUTO_MATCH": 2, "AMBIGUOUS": 2, "HOLD": 1, "NO_CANDIDATE": 1, "CONFLICT": 0, "MODEL_MATCH": 0}}

(That summary is from a v0.1.0 run of the example file against live Wikidata.)

Resume, budgets and cache#

  • The output file is the checkpoint. Run the same command again and finished rows are skipped. Rows stopped by a budget are redone; rows that ended in a provider error are redone with --retry-errors. A torn last line from an interrupted run is dropped.
  • Budgets are hard caps. --max-wikidata-requests and --max-google-requests count retries too. Rows that hit a cap are written as HOLD with BUDGET_EXHAUSTED, the run stops after that chunk, the receipt says complete: false, and the exit code is 3.
  • Requests are shared. Exact label lookups for many names go into one Wikidata Query Service query, entity data is fetched up to 50 items per request, and class and place labels in one more query. Google is called only when it can change a result (default --providers minimal), or for every row with --providers dual.
  • Cached reruns are free. Provider values and decisions are cached locally; a rerun of a finished batch makes zero provider requests. Provider failures are never cached.
  • Memory is bounded by --chunk (default 25 rows), not by the corpus size.

Evidence bundle#

wdkg export-evidence resolved.jsonl --bundle evidence/
wdkg validate-evidence evidence/

The bundle is self-contained: it needs neither the cache nor the resolver store to read.

file contents
decisions.jsonl one decision per row, with evidence_ref and candidate_pairs
evidence.jsonl every candidate's role (selected, viable, rejected...), every join and local comparison, and the provider observations they came from
manifest.json inputs, file hashes, schema version and the retention note
validation.json reference and pair checks with counts and errors

Each local comparison records the input field and value, the candidate field and value, and the observation it came from (operation, public request, time, value hash). Each Google join records the Wikidata property and value it matched.

What is not in the evidence#

Google Knowledge Graph content is not stored: no Google names, descriptions, types, scores or URLs. A Google id is kept only when it equals a string obtained independently (a Wikidata P646/P2671 value or your own input). As a result, EXTERNAL_ID_EXACT joins can be fully audited, while WIKIPEDIA_EXACT joins are always marked incomplete. This follows an engineering reading of Google's API terms, not legal advice; the details are in resolution.md.

An example bundle generated from the offline test world (synthetic ids marked as such) is in tests/fixtures/resolver/evidence_bundle_example/.

Documentation for version 0.2.1 · Changelog · Issues