Written by whoever
did the work.
No thought leadership, no reposts. Ingest internals, matching, storage models, and the things we got wrong the first time.
“The same company” is a modelling problem, not a string problem
Why edit distance plateaus around 78% on real CRM data, and what a register key does that no amount of fuzzy matching can.
SPOT-Bench v4: how we measure match quality without grading our own homework
250,000 labelled records, five input corruptions, blind adjudication, and the parts where we still lose.
Ingesting the Handelsregister: 1.2M announcements a year at 3.1-hour median latency
Polling strategy, the XML that lies, and how a court's holiday schedule became a monitoring alert.
Bitemporal entities, or: what did we know and when did we know it
Separating when the world changed from when we found out, and why audit workflows need both.
Grounding agents with IDs instead of names
A 1,200-question eval showing where name-matched RAG breaks and what ID-keyed retrieval fixes.
More of the same, in reference form
The guides cover the same ground with runnable examples, and the changelog records every release, deprecation and postmortem since 2023.