Batch and evidence#
The MCP tools resolve one entity per call. For a corpus, use the CLI: calling kg_resolve in
a loop spends model context and tool calls on every row, while wdkg resolve-batch streams
the file, shares provider requests between rows and keeps an audit trail.
Input#
One JSON object per line, with the same fields as a single resolve
(identity envelope). source_id is echoed back so you can join results to
your data.
{"source_id": "fox-atl", "name": "Fox Theatre", "kind": "place", "city": "Atlanta", "country": "US", "official_url": "https://www.foxtheatre.org"}
{"source_id": "fox-name-only", "name": "Fox Theatre", "kind": "place", "city": "Atlanta"}
{"source_id": "tate", "name": "Tate Modern", "kind": "place", "city": "London", "latitude": 51.5076, "longitude": -0.0994}
{"source_id": "primavera-series", "name": "Primavera Sound", "kind": "event_series", "city": "Barcelona", "official_url": "https://www.primaverasound.com"}
This is the start of examples/entities.jsonl.
Run#
# plan only: parse, normalize and check caches, no network
wdkg resolve-batch entities.jsonl --out resolved.jsonl --dry-run
# resolve with Wikidata only and a hard cap on requests
wdkg resolve-batch entities.jsonl --out resolved.jsonl \
--max-wikidata-requests 2000 --max-google-requests 0
What you get:
resolved.jsonl: one record per input line, in input order, withdecision,wikidata_qid,candidate_ids,reasons, the evidence code lists andcandidate_pairs.resolved.jsonl.receipt.json: decision counts, request usage per provider,complete,stop_reasonand what to do next.resolved.jsonl.evidence.jsonl: the full evidence for every row (see below).
The command itself prints a one-line summary:
{"op": "resolve-batch", "status": "ok", "rows": 6,
"decisions": {"AUTO_MATCH": 2, "AMBIGUOUS": 2, "HOLD": 1, "NO_CANDIDATE": 1, "CONFLICT": 0, "MODEL_MATCH": 0}}
(That summary is from a v0.1.0 run of the example file against live Wikidata.)
Resume, budgets and cache#
- The output file is the checkpoint. Run the same command again and finished rows are
skipped. Rows stopped by a budget are redone; rows that ended in a provider error are redone
with
--retry-errors. A torn last line from an interrupted run is dropped. - Budgets are hard caps.
--max-wikidata-requestsand--max-google-requestscount retries too. Rows that hit a cap are written asHOLDwithBUDGET_EXHAUSTED, the run stops after that chunk, the receipt sayscomplete: false, and the exit code is 3. - Requests are shared. Exact label lookups for many names go into one Wikidata Query
Service query, entity data is fetched up to 50 items per request, and class and place labels
in one more query. Google is called only when it can change a result (default
--providers minimal), or for every row with--providers dual. - Cached reruns are free. Provider values and decisions are cached locally; a rerun of a finished batch makes zero provider requests. Provider failures are never cached.
- Memory is bounded by
--chunk(default 25 rows), not by the corpus size.
Evidence bundle#
wdkg export-evidence resolved.jsonl --bundle evidence/
wdkg validate-evidence evidence/
The bundle is self-contained: it needs neither the cache nor the resolver store to read.
| file | contents |
|---|---|
decisions.jsonl |
one decision per row, with evidence_ref and candidate_pairs |
evidence.jsonl |
every candidate's role (selected, viable, rejected...), every join and local comparison, and the provider observations they came from |
manifest.json |
inputs, file hashes, schema version and the retention note |
validation.json |
reference and pair checks with counts and errors |
Each local comparison records the input field and value, the candidate field and value, and the observation it came from (operation, public request, time, value hash). Each Google join records the Wikidata property and value it matched.
What is not in the evidence#
Google Knowledge Graph content is not stored: no Google names, descriptions, types, scores or
URLs. A Google id is kept only when it equals a string obtained independently (a Wikidata
P646/P2671 value or your own input). As a result, EXTERNAL_ID_EXACT joins can be fully
audited, while WIKIPEDIA_EXACT joins are always marked incomplete. This follows an
engineering reading of Google's API terms, not legal advice; the details are in
resolution.md.
An example bundle generated from the offline test world (synthetic ids marked as such) is in
tests/fixtures/resolver/evidence_bundle_example/.