Brand resolution reaches general availability
58.4M marks from EUIPO TMview, WIPO GBD and USPTO TSDR are now first-class entities with brd_ IDs, linked to their proprietor. Resolving a trade name no longer silently returns the holding company.
Coverage figures, benchmark methodology, ingest latency and incident history, published in full.
Every guide ends with a runnable example against the sandbox key, and an edit link into the repo.
Key to first resolved entity in three minutes, in four languages.
5 minHow the four scorers work, what each hint buys you, and how to pick a threshold.
18 minCSV in, CSV out. Job lifecycle, partial results, backfilling 40M rows.
12 minSaved queries, signature verification, replay, dead-letter handling.
11 minas_of semantics, what a register did and did not publish, and the gaps.
9 minSnowflake, BigQuery, Databricks. Table layout, incremental models, dbt.
14 minWiring the tools, scoping keys, and grading groundedness in your eval.
10 minAll 184 attributes: type, source class, coverage, null semantics.
referenceThe 250,000-record benchmark: construction, labelling, corruptions, and how to run the harness yourself.
22 minSize a market from registers, resolve the company field at capture, and suppress accounts you already own by ID.
8 minReplay the same segmentation at each year end and get formation, dissolution and branch-opening rates that trace to a gazette.
9 minExistence, status, legal form, VAT and LEI checked against the register of record, with the filing reference attached.
8 min250,000 records drawn from real CRM exports, invoice OCR and web forms, labelled by two annotators with a third adjudicating disagreements. Published with the harness.
| Input class | Share | Top-1 | Top-5 | Auto-accept precision |
|---|---|---|---|---|
| Clean legal name | 31% | 99.4% | 99.8% | 99.9% |
| Name + city, no legal form | 24% | 98.1% | 99.3% | 99.6% |
| Domain only | 14% | 96.8% | 98.4% | 99.4% |
| Identifier (VAT/LEI/register) | 11% | 99.9% | 99.9% | 100.0% |
| Trade name / brand alias | 9% | 94.2% | 97.6% | 98.8% |
| OCR-corrupted line | 7% | 91.7% | 96.1% | 98.1% |
| Former name, ≥3 years stale | 4% | 88.4% | 94.9% | 97.7% |
| Weighted total | 100% | 97.3% | 98.9% | 99.4% |
The harness, the corruption generators and the label schema are public. The labelled set is available under a research agreement so it does not end up in a training corpus.
Predicted confidence against observed precision on the held-out split. On the diagonal the number means what it says.
July run: predicted 0.925 across the 0.90–0.95 band, observed 0.928.
Former names older than three years are our weakest class at 88.4%. Sole traders in jurisdictions without a public register are excluded from the benchmark entirely — we cannot resolve what nobody publishes, and averaging them in would flatter the number.
Per jurisdiction, on a held-out 40,000-record split, refitted monthly. A score is a probability, not a similarity: 0.90 has to be right nine times in ten or the fit is wrong and we say so in the changelog.
Financial depth is a separate axis from tier: it tracks what a jurisdiction obliges companies to file, which is why Swiss AGs sit at 22% and stay Tier 1.
Ingested from the primary register, with the full event history behind every field.
Assembled from official sources that publish identity but not a full filing stream.
Identity only, from official aggregates and international registers.
| Jurisdiction | Legal entities | Branches | Primary source | Filed financials | Median freshness | Tier |
|---|---|---|---|---|---|---|
| United States | 41,209,774 | 1,884,013 | 50 SoS registries · SEC EDGAR | 27 h | Tier 2 | |
| Germany | 6,412,880 | 318,204 | Handelsregister · Bundesanzeiger | 3.1 h | Tier 1 | |
| United Kingdom | 5,704,119 | 41,806 | Companies House | 1.4 h | Tier 1 | |
| France | 4,988,301 | 602,447 | INPI RNE · INSEE SIRENE | 6.8 h | Tier 1 | |
| Italy | 3,146,902 | 271,509 | Registro Imprese | 9.4 h | Tier 1 | |
| Poland | 2,905,441 | 88,120 | KRS · CEIDG | 11 h | Tier 2 | |
| Netherlands | 2,381,660 | 96,338 | KVK Handelsregister | 5.0 h | Tier 1 | |
| Switzerland | 762,455 | 54,912 | Zefix · SHAB | 2.2 h | Tier 1 |
Grouped by jurisdiction and by class. Cadence is the polling interval, not an average of how quickly a record happens to change.
A feed that falls behind its cadence raises an alert against the expected volume for that source on that day, so a court holiday reads as a court holiday rather than an outage. Live lag per jurisdiction is on the status panel, and every ingest change is in the changelog.
Breaking changes ship behind a dated version and the previous one runs for a year. Coverage and performance changes ship continuously.
58.4M marks from EUIPO TMview, WIPO GBD and USPTO TSDR are now first-class entities with brd_ IDs, linked to their proprietor. Resolving a trade name no longer silently returns the holding company.
Registro Imprese ingest now includes deposited accounts. Financial coverage for IT rises from 34% to 79%; median freshness drops from 31 h to 9.4 h.
Blocking moved to a learned key set. p95 falls from 152 ms to 112 ms with no measurable change in top-1 accuracy (97.28% → 97.31% on SPOT-Bench v4).
Watch subscriptions can now be defined by a saved search rather than an ID list, so a target universe maintains itself. Backfill of the first 90 days is included.
Confidence is now calibrated per jurisdiction rather than globally; alternatives is capped at 10 and sorted; the deprecated score field is removed. Prior version supported until 2027-06-01.
A bad index build on the blocking store returned 503s for 8.2% of resolve calls between 09:14 and 09:55 UTC. Root cause, timeline and the four changes we made are on the status page.
Long posts about ingest internals, written by whoever did the work.
Why edit distance plateaus around 78% on real CRM data, and what a register key does that no amount of fuzzy matching can.
250,000 labelled records, five input corruptions, blind adjudication, and the parts where we still lose.
Polling strategy, the XML that lies, and how a court's holiday schedule became a monitoring alert.
Separating when the world changed from when we found out, and why audit workflows need both.
A 1,200-question eval showing where name-matched RAG breaks and what ID-keyed retrieval fixes.
Availability is probed from four networks we do not run. Incidents get a public postmortem within five business days.
We hold registry facts, your query logs and your entity mappings — no payment data, no end-customer PII, no credentials to your systems.
| SOC 2 Type II | Audited annually; report on request |
| ISO 27001 | Certified · scope covers ingest, API and support |
| GDPR | DPA with SCCs, sub-processor list, 30-day deletion |
| EU residency | EU-only processing and log retention, no egress on failover |
| Penetration testing | Independent test twice a year, summary on request |
| Access | SSO/SAML, SCIM, scoped keys, full audit log |
| SOC 2 Type II report | Under NDA, current period |
| Security questionnaire | CAIQ and SIG Lite pre-filled |
| Sub-processor list | With 30 days notice of changes |
All of it from trust.spotit.ai.
That is the whole quickstart. If it does not work inside five minutes, tell us and we will fix the docs rather than explain the error.
$ export SPOTIT_KEY=spk_test_8f2c... $ curl -s https://api.spotit.ai/v1/resolve \ -H "Authorization: Bearer $SPOTIT_KEY" \ -H "Content-Type: application/json" \ -d '{"q":"nordwind logistik hamburg"}' \ | jq '{id: .match.entity_id, conf: .match.confidence, name: .entity.name}' { "id": "ent_01JR8K3F5T2QW9", "conf": 0.987, "name": "Nordwind Logistik GmbH" }