AI Quality Assurance
We publish this page because the AI compliance assistant gives advice that affects real legal decisions. You should know exactly how we test it β and how it scored β before you trust it.
Aegis Firma's AI assistant must score at least 40 out of 50 on the golden-set test before each major release. If the AI model is updated or prompts change, the suite runs again. We will not ship a regression to fewer than 40 correct.
The 50-question golden set
A fixed set of 50 questions authored by our compliance team. Questions never change between runs so scores are comparable month to month. The set covers eight categories:
| Category | Qs | Topics covered |
|---|---|---|
| EU AI Act | 10 | Risk tiers, Article 6β10, Annex III, GPAI, prohibited practices, timelines |
| GDPR Γ AI | 8 | Article 22, DPIA, legitimate interest, data minimisation in AI pipelines |
| NYC Local Law 144 | 4 | Bias audit mandate, annual disclosure, covered employers |
| Colorado AI Act | 4 | High-risk AI definition, developer/deployer obligations, appeal rights |
| SOC 2 + AI | 4 | AI system evidence, common criteria, CC6.1 applied to ML models |
| ISO 42001 + NIST RMF | 4 | AIMS scope, NIST Govern/Map/Measure/Manage functions |
| Situational / applied | 11 | Mixed-regulation scenarios β e.g. GDPR + EU AI Act overlap for an HR chatbot |
| Adversarial | 5 | Prompts designed to elicit false confidence, legal advice, or hallucinated citations |
Scoring methodology
Each answer is scored automatically. A question scores 1 point only if it passes all applicable checks. Partial credit is not awarded.
The test runner is open for inspection at __tests__/ai-chat-golden-set.ts in the Aegis Firma repository.
Test results log
Updated monthly. Each row is a single run against the production AI endpoint.
| Month | Score | Notes |
|---|---|---|
| May 2026 | Pending | Scheduled β runs after launch on live server |