Product
Qiyas
Evaluation, canary rollout and release gates for AI systems
Qiyas is Siyada Tech's evaluation platform for AI systems. It runs versioned test suites against a model or agent, scores answer quality and groundedness in Arabic and English, compares releases, and blocks a rollout when a regression crosses the threshold the organisation defined.
Last reviewed: — Siyada Tech engineering
What it is
A measurement and release-gating layer that sits between an AI system and production. Qiyas holds the evaluation datasets, the scoring rubrics, and the pass/fail thresholds, and produces a comparable report for every change to a prompt, model, retrieval index or agent policy.
The problem it solves
Most AI systems ship on a demo and a feeling. When a prompt, model version or retrieval index changes, nobody can say whether quality improved or quietly regressed — and Arabic performance is rarely measured at all.
Who it is for
- Enterprise teams running AI systems in production
- Risk, audit and model-governance functions
- Platform and MLOps teams managing model upgrades
- Regulators-facing organisations that must document AI behaviour
How it works
- 01Define: evaluation datasets are built from real organisational tasks, in Arabic and English.
- 02Score: each run is scored on task success, groundedness, citation validity and safety rubrics.
- 03Compare: results are diffed against the previous release, per slice and per language.
- 04Gate: a release is blocked when a defined threshold regresses.
- 05Canary: a change is exposed to a limited traffic slice with live monitoring before full rollout.
Deployment
- In-tenant deployment alongside the systems it evaluates
- CI integration so evaluation runs on every change
- API for triggering runs and reading results
- Exportable reports for audit and governance committees
Security & compliance
- Evaluation data stays inside the customer's environment
- PDPL-aligned handling of any personal data present in test sets
- Versioned, immutable run history for audit
- Independent certification status: [EVIDENCE REQUIRED]
Evidence
We publish a metric only with its definition, method, sample size and evaluation date. Anything not yet measured to that standard is marked below rather than claimed.
| Metric | Definition | Method | Sample size | Evaluated | Limitations |
|---|---|---|---|---|---|
| Regression catch rate before release | [EVIDENCE REQUIRED] | [EVIDENCE REQUIRED] | [EVIDENCE REQUIRED] | [EVIDENCE REQUIRED] | [EVIDENCE REQUIRED] |
| Evaluation suite coverage | [EVIDENCE REQUIRED] | [EVIDENCE REQUIRED] | [EVIDENCE REQUIRED] | [EVIDENCE REQUIRED] | [EVIDENCE REQUIRED] |
| Time from model change to release decision | [EVIDENCE REQUIRED] | [EVIDENCE REQUIRED] | [EVIDENCE REQUIRED] | [EVIDENCE REQUIRED] | [EVIDENCE REQUIRED] |
Limitations & what it does not do
- Qiyas measures what the evaluation set covers; unmeasured behaviour stays unmeasured.
- Rubric-based scoring involves judgement and must be reviewed by the domain owner.
- It reduces regression risk; it does not certify a system as safe.