Production Evidence for EU AI Act Biometric System Compliance
Who this is for: Notified Body assessors evaluating biometric AI systems under Regulation (EU) 2024/1689. Biometric system manufacturers preparing for third-party conformity assessment. Remote identification operators subject to Annex III high-risk classification.
DEKRA became the first accredited certification body for AI biometric systems under the EU AI Act in March 2026, accredited by the Dutch Accreditation Council (RvA). The accreditation covers three categories of high-risk biometric AI systems: remote biometric identification, emotion recognition, and biometric categorization.
These systems carry the highest scrutiny under Regulation (EU) 2024/1689. Biometric identification can misidentify individuals. Emotion recognition can infer states that do not exist. Categorization can classify people by characteristics they did not consent to share. Every biometric inference that produces an incorrect result is a fundamental rights event.
Article 43 requires the Notified Body to verify that the system in operation matches the system described in the technical documentation. Conformity assessment happens at a point in time. Production happens every second. The Notified Body has no visibility into what occurs between assessments.
A biometric system that passed conformity assessment six months ago may have drifted. Its accuracy on a new population may differ from the population used in testing. The deployment environment may have changed. The model weights may have been updated. None of these events are visible to the assessor without runtime evidence.
A biometric system that passed conformity assessment six months ago may have drifted. Its accuracy on a new population may differ from the population used in testing. The Notified Body has no mechanism to verify continuous operation without runtime evidence.
Note: SWT3 provides the continuous evidence layer between conformity assessments. Each biometric inference event, fairness measurement, human review decision, and model drift evaluation is captured as an immutable, timestamped witness anchor.
| Biometric Category | EU AI Act Reference | Risk Profile | Key SWT3 Procedures |
|---|---|---|---|
| Remote Biometric Identification | Art. 5(1), Annex III(1)(a) | Highest -- can identify individuals at distance without consent | AI-INF.1, AI-FAIR.1, AI-HITL.1 |
| Emotion Recognition | Art. 50(3), Annex III(1)(c) | High -- infers emotional states from biometric signals | AI-INF.1, AI-EXPL.1, AI-TRANS.1 |
| Biometric Categorization | Annex III(1)(b) | High -- classifies by physical traits or behavioral patterns | AI-INF.1, AI-FAIR.1, AI-DRIFT.1 |
Real-time identification in public spaces represents the highest-risk biometric application under the EU AI Act. Art. 5(1) prohibits most uses outright. The narrow exceptions -- law enforcement with prior judicial authorization -- require the strongest possible evidence chain. Every identification event is witnessed. Every false positive is tracked. Every human review decision is documented. The organization maintains a continuous, auditable record that the system operated within its authorized scope.
Art. 50(3) requires disclosure when emotion recognition is in use. The system must prove it disclosed, not merely that it could disclose. A privacy policy stating "we may use emotion recognition" is not evidence of per-interaction disclosure. The witness anchor captures the disclosure event itself. Explanation generation proves the reasoning behind emotional state classifications, establishing that the system produced interpretable outputs rather than opaque predictions.
Classification by race, gender, age, disability, or other protected characteristics creates direct fundamental rights exposure. Fairness monitoring is not optional for these systems -- it is a fundamental rights obligation. Drift detection captures accuracy degradation across demographic groups over time. A categorization system that was accurate across all groups at the time of conformity assessment may develop disparate performance as production populations shift. Continuous drift monitoring detects this degradation before it becomes a rights violation.
Every biometric inference produces a witness anchor. The anchor records the model identifier, clearing level, and timestamp. For biometric systems, inference volume and frequency are themselves compliance evidence -- they prove the system was active and monitored during the assessment window.
The anchor captures the model that produced the inference, the clearing level applied to the evidence, and the precise time of execution. Over time, the series of AI-INF.1 anchors establishes the operational profile of the biometric system: when it runs, how frequently it processes inputs, and whether that profile matches the manufacturer's declared operating parameters.
Query the ledger for AI-INF.1 anchor volume across the assessment window. Compare against the manufacturer's declared operating parameters. A system declared to process 10,000 identifications per day should have approximately 10,000 anchors per day. Significant variance requires explanation.
Fairness monitoring for biometric systems measures disparity across protected characteristics. The maximum acceptable disparity threshold is set by the manufacturer. The observed disparity is measured from production data. A disparity ratio exceeding the threshold triggers a FAIL verdict.
For biometric categorization and remote identification systems, the protected attribute list must reflect the characteristics the system classifies or uses for matching. The witness anchor records the number of protected attributes evaluated, whether the declared threshold was met, and the observed disparity value.
Verify the protected attribute count matches the manufacturer's declared list. Request the disparity threshold and compare against observed values. A threshold-met value of 0 requires investigation -- what remediation was taken, and is there a follow-up anchor showing improvement?
For remote biometric identification in law enforcement contexts, human review is mandatory under Art. 14. The witness anchor confirms that a qualified human reviewed and approved the identification before action was taken. The anchor binds the review decision to the reviewer's identity, establishing accountability for each approval.
Human-in-the-loop oversight for biometric systems is not a design aspiration. It is a legal requirement. The anchor provides verifiable evidence that the requirement was met for each identification event, not merely that the organization has a human review policy in place.
For systems operating under Art. 5(1) exceptions, every identification should have a corresponding human review anchor. Missing review anchors indicate automated decisions without required oversight. Compare AI-INF.1 anchor counts against AI-HITL.1 anchor counts -- the ratio should approach 1:1 for law enforcement biometric identification.
Emotion recognition systems must generate explanations for their classifications. Art. 13(1) requires that outputs are interpretable by the deployer. The witness anchor confirms an explanation was both required and provided for each classification event.
The anchor records whether an explanation was required for the inference (determined by the system's risk classification and the regulatory context) and whether the explanation was successfully generated and delivered. This distinction matters: a system that requires explanations but intermittently fails to produce them has an Art. 13 compliance gap even if most inferences are explained.
For emotion recognition, verify the explanation-provided indicator matches the explanation-required indicator across the full assessment window. A system that requires explanations but fails to provide them has an Art. 13 gap. Calculate the explanation coverage ratio -- it should be 100% for high-risk biometric systems.
Biometric accuracy can degrade when production populations differ from training populations. Drift detection measures statistical divergence across evaluation windows. For biometric systems, drift directly translates to identification error rates and, by extension, to fundamental rights impact.
The anchor captures the number of metrics evaluated, the number that exceeded drift thresholds, the detection method used, and the threshold applied. A series of anchors over time reveals whether the biometric model's accuracy is stable, improving, or degrading. Degradation trends may require the Notified Body to re-evaluate the conformity assessment.
Request 90 days of drift anchors. A drifted-count trending upward across evaluation windows indicates the model's biometric accuracy is degrading. This may require re-evaluation of the conformity assessment. Cross-reference with the manufacturer's declared accuracy thresholds to determine whether drift has moved the system outside its approved operating envelope.
The following six-step workflow provides a structured approach for Notified Body assessors evaluating biometric AI systems with SWT3 evidence. Each step builds on the previous one, creating a complete evidence chain from documentation to production verification.
Art. 11 documentation, declared operating parameters, population characteristics
Verify operational activity matches declared parameters
Confirm disparity measurements within declared thresholds
Human review documented for high-risk identification events
No accuracy degradation across the assessment window
The six-step workflow produces a complete evidence chain from documentation to production. Each step builds on the previous one. If Step 2 reveals no anchors, subsequent steps cannot be evaluated -- the system lacks runtime evidence. Begin with Step 1 to establish the baseline, then verify each subsequent step against the manufacturer's declared parameters.
Biometric data is special category data under GDPR Art. 9. The clearing level determines how much context survives in the witness anchor after evidence graduation. For biometric systems, the clearing level selection has direct implications for both regulatory compliance and assessor visibility.
| Clearing Level | Biometric Implication |
|---|---|
| Level 0 (Analytics) | Full context including model identifiers. Internal use only. Not appropriate for external sharing of biometric system evidence. |
| Level 1 (Standard) | Model identifiers preserved. Suitable for sharing with an escrow agent or designated third party under confidentiality agreement. |
| Level 2 (Sensitive) | Biometric context redacted. The Notified Body sees numeric factors and verdict only. Required for GDPR Art. 9 compliance when anchors themselves reference biometric processing. |
| Level 3 (Classified) | Only the cryptographic fingerprint survives. Factor Handoff Protocol required for the Notified Body to access underlying evidence through a designated escrow arrangement. |
For biometric AI systems processing special category data, verify that the clearing level is set to 2 or higher. Level 0-1 anchors containing biometric model identifiers may themselves constitute personal data processing under GDPR Art. 9. The organization should demonstrate that its clearing level selection was made with explicit consideration of the biometric data classification.
For detailed guidance on evidence access at Level 3, see Factor Handoff Protocol.
Install:
# Python
pip install swt3-ai
# TypeScript
npm install @tenova/swt3-ai
Implementation reference -- witness a biometric inference event:
from swt3_ai import Witness
witness = Witness(api_key="your_key", tenant_id="your_tenant")
witness.record("AI-INF.1", model_id="biometric-v2.1", clearing_level=2)
Full SDK documentation: SDK Docs. Create a free account: Sign Up.