Product

Qiyas

Evaluation, canary rollout and release gates for AI systems

Qiyas is Siyada Tech's evaluation platform for AI systems. It runs versioned test suites against a model or agent, scores answer quality and groundedness in Arabic and English, compares releases, and blocks a rollout when a regression crosses the threshold the organisation defined.

Last reviewed: Siyada Tech engineering

What it is

A measurement and release-gating layer that sits between an AI system and production. Qiyas holds the evaluation datasets, the scoring rubrics, and the pass/fail thresholds, and produces a comparable report for every change to a prompt, model, retrieval index or agent policy.

The problem it solves

Most AI systems ship on a demo and a feeling. When a prompt, model version or retrieval index changes, nobody can say whether quality improved or quietly regressed — and Arabic performance is rarely measured at all.

Who it is for

  • Enterprise teams running AI systems in production
  • Risk, audit and model-governance functions
  • Platform and MLOps teams managing model upgrades
  • Regulators-facing organisations that must document AI behaviour

How it works

  1. 01Define: evaluation datasets are built from real organisational tasks, in Arabic and English.
  2. 02Score: each run is scored on task success, groundedness, citation validity and safety rubrics.
  3. 03Compare: results are diffed against the previous release, per slice and per language.
  4. 04Gate: a release is blocked when a defined threshold regresses.
  5. 05Canary: a change is exposed to a limited traffic slice with live monitoring before full rollout.

Deployment

  • In-tenant deployment alongside the systems it evaluates
  • CI integration so evaluation runs on every change
  • API for triggering runs and reading results
  • Exportable reports for audit and governance committees

Security & compliance

  • Evaluation data stays inside the customer's environment
  • PDPL-aligned handling of any personal data present in test sets
  • Versioned, immutable run history for audit
  • Independent certification status: [EVIDENCE REQUIRED]

Evidence

We publish a metric only with its definition, method, sample size and evaluation date. Anything not yet measured to that standard is marked below rather than claimed.

MetricDefinitionMethodSample sizeEvaluatedLimitations
Regression catch rate before release[EVIDENCE REQUIRED][EVIDENCE REQUIRED][EVIDENCE REQUIRED][EVIDENCE REQUIRED][EVIDENCE REQUIRED]
Evaluation suite coverage[EVIDENCE REQUIRED][EVIDENCE REQUIRED][EVIDENCE REQUIRED][EVIDENCE REQUIRED][EVIDENCE REQUIRED]
Time from model change to release decision[EVIDENCE REQUIRED][EVIDENCE REQUIRED][EVIDENCE REQUIRED][EVIDENCE REQUIRED][EVIDENCE REQUIRED]

Limitations & what it does not do

  • Qiyas measures what the evaluation set covers; unmeasured behaviour stays unmeasured.
  • Rubric-based scoring involves judgement and must be reviewed by the domain owner.
  • It reduces regression risk; it does not certify a system as safe.

Frequently asked questions

Baseline your AI system with Qiyas

We build an evaluation set from your real tasks and produce a baseline report on the system you run today.