Two very different things now get sold as “EU AI Act compliance evidence”: a rule engine that checks facts you supply against a fixed, versioned reading of the law, and a language model that writes a compliance narrative from a prompt. They are not the same product, and the difference matters most in exactly the situation you buy either one for — an audit.
What each one actually does
A rule-based engine evaluates a fixed set of machine-readable rules — each one traceable to an article, each one versioned — against facts an organisation records or connects. It produces the same output from the same inputs every time, and a rule's wording can be cited the way a line of a statute can be cited.
A model-generated report asks a language model to read a description of a system and write prose describing its compliance posture. The model is fluent, and fluency is exactly the risk: it can produce a confident, well-structured paragraph asserting an obligation is met without any fact backing the assertion, and it can phrase two different runs on the same system differently enough that neither can be pinned down as “the” answer.
A generated sentence is not evidence of anything except that a model produced that sentence. An auditor's question is never “what does this document say” — it is “can this be checked against something outside itself”. A rule citation can be checked against the rule. A generated claim can only be checked against another model run.
The citation test
A useful way to tell the two apart in a sales conversation: ask what a specific claim in the report cites. A rule-based system answers with a rule ID and a pack version — a fixed target that does not move under the reader. A model-generated report either cites nothing, or cites the model's own training data, which is not a citation an auditor can independently check at all.
What to actually ask a vendor
- Can a specific compliance claim be traced to a specific, versioned rule or article?
- Does the same input produce the same output every time, or does it vary between runs?
- If the underlying regulation or the vendor's own interpretation changes, is that a dated, citable version bump — or a silent edit to the same document?
- Is the evidence tamper-evident — checkable by the reader — or does it depend on trusting the vendor's own account of what it produced?
Where each genuinely fits
This is not an argument that language models have no place in this pipeline — they are useful for exactly the job they are good at: explaining a rule in plain language, drafting a first pass at a description, helping someone investigate why a score changed. What they should not be is the thing that decides whether an obligation was met and writes the record an auditor is handed. That decision belongs to a fixed, versioned rule a human can read, not to a model's phrasing on the day it was asked.
Key takeaways
- A rule-based system cites a versioned rule; a model-generated report, at best, cites its own training data — which nobody outside the vendor can check.
- Determinism is not a nice-to-have here — the same facts must produce the same finding every time, or the “finding” is not a fact about the system, it is a fact about that run of the model.
- The honest role for a model in this pipeline is explanation and drafting, never the citation an auditor is expected to rely on.
See what a rule-based reading of the Act actually finds for a given system with the 2-minute applicability check — every answer it gives traces back to a specific rule, not a generated sentence.