How we chose the engine behind Fidenta Verify — starting with where it runs

How we chose the model behind Fidenta Verify — starting not with quality, but with where your documents stay put.

How we chose the engine behind Fidenta Verify — starting with where it runs

Most AI tools answer fast and sound certain. That is the problem, not the feature — a confident wrong answer is the one you act on. Fidenta Verify exists to catch those first: it reads a claim, checks it against the evidence, and tells you whether it actually holds. Never assume, verify.

This week we chose the model that does that checking. The first cut had nothing to do with quality.

We only considered models that run inside our own tenant. Every candidate runs on Cloudflare's own GPUs — the document you are verifying is assessed in the tenant rather than sent out to an outside model vendor's API. A verifier that ships your contract to a third party in order to check it has already failed at the one thing it is for. That constraint ruled out the obvious "smartest" options before we scored a single answer, and we would make the same call again.

Then we ran 472 checks across four candidates on twenty deliberately hard internal-document items — contract clauses, data-processing terms, insurance exhibits — the kind with a carve-out three sentences after the part that looks like the answer. Three results are worth putting in front of you.

First, the one we are proudest of. When we gave the models a claim and no evidence to check it against, all four abstained — every time, eighty out of eighty, with zero fabricated verdicts. This is the whole job. A verifier's first duty is to say "I cannot verify this" rather than rule from memory, and every candidate did exactly that. It is the cleanest result we have and the one that matters most.

Second, across eighty grounded checks, not one model mistook silence for refutation — none called a claim "contradicted" when the evidence simply did not address it. The distance between "the document does not say" and "the document says otherwise" is the distance between a careful reviewer and a reckless one.

Third, the finding that actually decided it. Accuracy was the first and biggest bar, and the leading candidates all cleared it — close enough after replication to be statistically indistinguishable from one another on accuracy. Among models that were all this accurate, one thing set the winner apart: it returned identical verdicts on identical inputs, every single run. For a verification tool that consistency is decisive — a verifier that gives different answers to the same document on Tuesday and Thursday cannot be audited, and a verdict you cannot audit is not a verdict, it is an opinion. The model we chose gave the best balance of accuracy, stability, and cost for this specific set of tests. We will also say plainly that these were twenty hard items, not two thousand — a larger or different test set could shift the order, and we will keep running them.

We will also tell you what did not work. Our first instinct was to fix the hardest failure with a better prompt. The prompt fix failed — it helped nothing and made one candidate worse. What worked was architecture: a second, adversarial read that fires only when the first pass says "supported," re-checking for the qualifying clause that flips the answer. Structure beat cleverness, which is usually how it goes.

End to end — retrieve the evidence, check the claim, audit the check — a real verification completes in about fifteen seconds. Fast enough to run on every clause, careful enough that "verified" means something, and assessed inside the tenant rather than out at someone else's API.

We did not pick the model that looked most impressive. We picked the behavior we could stand behind, running where your documents stay put.

Verified, not assumed.

*Fidenta Verify is in early access. If you rely on AI to read the documents you cannot afford to get wrong, we would like to hear from you.*