Inference attestation
Proof of what your AI did, on infrastructure you own.
Sijil sits in front of your model endpoint. Every call is checked against policy and sealed into a signed, hash-chained record: which model version answered, where it ran, the policy in force, and who approved it. An independent verifier proves the chain is intact, or names the exact record that broke.
No change to the model, the application or the serving runtime.
“An entity deployed an assistant six months ago. An auditor asks what it told a citizen on 12 March. What do they hand over?”
Live demo
Watch a record get forged, and caught.
Real proxy, real local model (Llama 3.2 3B in LM Studio), real tamper script. An allowed call, an Emirates ID flag, a residency block, a record altered and then forged from outside the app, and the audit report.
The problem
Four questions every auditor asks. None answerable today.
- Which model version produced this?
- Endpoints get updated in place. Nothing recorded which version answered.
- Where did the inference run?
- An API returns a response, not a location. Residency is asserted in a contract, never proven per call.
- What policy was in force?
- Policy lives in a document. Nothing binds its version to the call.
- Who approved this model for this use?
- An email thread or a committee minute, disconnected from the system that serves the model.
So assurance happens by paperwork, after the fact, and risk functions slow AI deployments down.
How it works
Four things, on every call.
- 1
Check
The call is evaluated against the policy in force: data residency, approved model weights, sensitive terms. Record-only or enforce, per rule, the entity's choice.
- 2
Seal
Model version, node and region, policy hash, approval reference, and hashes of the input and output, signed with Ed25519 and chained to the record before it. Blocks are sealed too.
- 3
Verify
A standalone verifier that shares no code with the writer walks the chain and returns pass, or the exact record and the reason it failed.
- 4
Report
An A4 audit report for any period: calls by decision, policies and approvals in force, every exception, and the verifier's result.
The part nobody can wave away
Logs can be edited. This can't, without being caught.
An insider edits a record in the database
$ scripts/tamper.sh alter 17 alter seq 17: output_hash rewritten FAIL seq=17 content altered
A smarter attacker recomputes the hash too
$ scripts/tamper.sh forge 17 forge seq 17: record_hash recomputed, no private key FAIL seq=17 signature invalid
The hash adds up, but the attacker doesn't hold the node's signing key. And the checker isn't marking its own homework: it reimplements the spec independently of the software that wrote the records.
Measured
What it costs to run.
Measured on an Apple M4 Pro laptop with Llama 3.2 3B Instruct (Q4_K_S) in LM Studio.
Where it fits
Attestation exists. It just doesn't reach you.
Hyperscaler confidential inference is hardware-rooted and stronger than ours, inside their cloud. Governance platforms can't reach the runtime. Full-platform vendors bundle lineage with adoption. Where the entity owns the hardware, there is nothing integrated.
Not the accounts you take whole
Where a customer adopts a platform end to end, they get lineage and audit with it. Sijil is redundant there, and we say so.
The accounts you can't
Customers who built their own inference stack, or run a mixed estate they won't replace, have no attestation path. Nobody replaces a working stack to get an audit trail. Sijil covers these accounts without displacing anything.
Honest boundaries
What we prove, and what we don't yet.
- Proven: records have not been altered, reordered or removed from within the chain since they were sealed.
- Not yet: that the sealing itself was honest. Node identity and region are asserted and signed by the node, not bound to the silicon. Hardware root of trust, built on NVIDIA and TEE attestation, is the roadmap.
- By design: the record holds hashes of what was asked and answered, never the content itself.
For integrators bidding government AI work
Name one opportunity where the customer keeps their own inference stack.
We'll write the technical response section for it, at no cost, before any commercial agreement. You hold the customer relationship and the contract; we deploy under your delivery.