Skip to content
πŸ§ͺ AI quality transparency

AI Quality Assurance

We publish this page because the AI compliance assistant gives advice that affects real legal decisions. You should know exactly how we test it β€” and how it scored β€” before you trust it.

β‰₯ 40 / 50
Launch gate threshold

Aegis Firma's AI assistant must score at least 40 out of 50 on the golden-set test before each major release. If the AI model is updated or prompts change, the suite runs again. We will not ship a regression to fewer than 40 correct.

The 50-question golden set

A fixed set of 50 questions authored by our compliance team. Questions never change between runs so scores are comparable month to month. The set covers eight categories:

CategoryQsTopics covered
EU AI Act10Risk tiers, Article 6–10, Annex III, GPAI, prohibited practices, timelines
GDPR Γ— AI8Article 22, DPIA, legitimate interest, data minimisation in AI pipelines
NYC Local Law 1444Bias audit mandate, annual disclosure, covered employers
Colorado AI Act4High-risk AI definition, developer/deployer obligations, appeal rights
SOC 2 + AI4AI system evidence, common criteria, CC6.1 applied to ML models
ISO 42001 + NIST RMF4AIMS scope, NIST Govern/Map/Measure/Manage functions
Situational / applied11Mixed-regulation scenarios β€” e.g. GDPR + EU AI Act overlap for an HR chatbot
Adversarial5Prompts designed to elicit false confidence, legal advice, or hallucinated citations

Scoring methodology

Each answer is scored automatically. A question scores 1 point only if it passes all applicable checks. Partial credit is not awarded.

Keyword match
Answer must contain specific technical terms (e.g. "Article 22", "risk tier", "fundamental rights impact assessment"). Fuzzy synonyms are not accepted.
Citation accuracy
Factual questions require a correct citation β€” wrong article numbers fail the check.
No unacceptable claims
Answers containing "you are compliant", "this is legal advice", or fabricated citations fail regardless of other content.
Adversarial refusal
The 5 adversarial questions pass only if the AI declines to give a definitive answer and explains uncertainty.

The test runner is open for inspection at __tests__/ai-chat-golden-set.ts in the Aegis Firma repository.

Test results log

Updated monthly. Each row is a single run against the production AI endpoint.

MonthScoreNotes
May 2026PendingScheduled β€” runs after launch on live server

What this test does not prove

!
Not a substitute for legal advice
The AI passes factual accuracy tests, not legal interpretation tests. Always engage a qualified lawyer for binding decisions.
!
Not an exhaustive regulation coverage test
50 questions cannot cover every edge case in 100+ jurisdictions. The set focuses on the highest-stakes questions we see asked most often.
!
Not a live hallucination monitor
The golden set runs monthly. Between runs, model behaviour may drift. We monitor for anomalies separately.
!
Not a claim of 100% accuracy
A score of 40/50 means 80% accuracy on our test set β€” there is a 20% floor of known unknowns, and an unknown set of unknown unknowns.

Related trust pages