Capability 03 · Retrieval
Answers from your own documents, with sources
Arabic RAG systems
Every answer carries the passage it came from.
In one paragraph
Siyada Tech builds retrieval-augmented generation for Arabic and English: search-and-answer systems over an organisation's own documents, where every answer carries the citation it came from and questions the documents cannot support are declined. We handle Arabic segmentation, spelling and diacritic variance, and deploy inside the client's environment.
- Hybrid retrieval
- Cited answers
- In-tenant index
01What it is
An engineered retrieval pipeline built for a specific corpus and a specific set of questions: parsing, Arabic-aware chunking, hybrid lexical and semantic retrieval, reranking, grounded generation and citation rendering.
02Who it is for
- Organisations with large Arabic document estates
- Legal, policy, knowledge and customer-support functions
- Government entities that publish Arabic regulation and guidance
03The problem
Retrieval built for English degrades on Arabic. Words are split badly, diacritics and spelling variants break exact matching, and mixed Arabic-English documents retrieve inconsistently. The result is a confident answer drawn from the wrong passage.
04How it works
-
01
Parse the corpus, including scanned material where OCR quality allows.
-
02
Chunk with Arabic-aware segmentation that respects sentence and clause structure.
-
03
Index both ways
Lexical for exact terminology, semantic for paraphrase, normalised for diacritics and spelling.
-
04
Rerank the retrieved candidates before anything is generated.
-
05
Generate strictly from the retrieved passages and render inline citations.
-
06
Measure groundedness and citation validity per language on an agreed test set.
05Where it runs
- In the client's tenant; the corpus and the index stay there
- Read-only connectors to SharePoint, file shares and document management systems
- Retrieval filtered by the requesting user's permissions in the source system
- API and web interface, with sign-in through the client's identity provider
06Security and data
- No corpus or index leaves the tenant
- Personal data in documents handled to PDPL and NDMO practice
- Every question, retrieved passage and answer written to an audit log
07Evidence
What we measure here. Results are published once each one has a stated method, sample size and evaluation date. How we publish evidence
- Groundedness rate
- Citation validity
- Answer rate on questions the corpus can support
08What it is not
- Retrieval cannot recover information that is not in the corpus.
- Poor OCR on scanned Arabic reduces retrieval quality materially.
- Grounding reduces unsupported statements; it does not eliminate them. We measure the rate rather than promise zero.
?Asked often
Questions
What is Arabic RAG?
Retrieval-augmented generation engineered for Arabic: the system retrieves passages from your corpus and writes the answer only from them, with citations.
Why not use a general model?
A general model has not read your documents and cannot cite them. Retrieval grounds every answer in your own material.
Does it handle mixed Arabic and English documents?
Yes. One index serves both languages, with normalisation for diacritics and spelling variants.
Does it respect document permissions?
Yes. Retrieval is filtered by what the person asking is allowed to see in the source system.
How do you measure quality?
On an agreed test set: groundedness, citation validity and task success, scored separately for Arabic and English.
How do we start?
Give us a sample of the corpus and twenty real questions. We build a scoped pilot and report what we measured.
Pilot it on your own corpus
A sample of your documents and twenty real questions is enough for a measured pilot.
Last reviewed · Siyada Tech engineering