@علا صالح @رقية بورية
roadmap بسيط
QURANALYZE ROADMAP
Contributor roadmap — living document
Quranalyze is an AI-powered Quran study platform — a full reading experience plus a tafsir-grounded chatbot, semantic search, and verse recall, built to cite real scholarly sources instead of letting a general model answer freely. Solo-built so far, moving toward volunteer-driven. This is where anyone thinking about contributing can see what's being worked on, what's coming, and where the open, unsolved problems are.
Status: early-stage, actively developed · Stack: React · Django · MySQL · Redis · Qdrant
How to read this: Now = active or blocking. Next = planned for coming months. Later = wanted, not yet scheduled. Open Problems = harder, unsolved questions — good if you want something meatier than a task. Each item is tagged AI/ML, Full-stack, or Design.
NOW
[Full-stack] Fix Navigation Search vs. lexical search conflict — the two search modes interfere with each other; a regex conflict causes slow/incorrect results when both are in play.
[Design] Fix inconsistent modal animations — modals don't animate consistently across the app; needs one shared transition treatment.
[Design] Fix search-results scrolling — specific scrolling bug in the search results view needs isolating and fixing.
[AI/ML] Scholar sign-off on system/classifier prompts — the prompts driving the classifier and chatbot answers need direct verification from a qualified scholar before other search/answer work builds on top of them.
NEXT
[AI/ML] Refactor classification into one function — consolidate currently-scattered classification logic into a single parameterized function.
[AI/ML] Improve English + Urdu search quality — both lexical and semantic search need real improvement for English/Urdu queries, not just Arabic.
[Full-stack] Automate manual data-correction workflow — turn the current manual, multi-step correction process into a scripted pipeline.
[Full-stack] Add automated backend tests — backend lacks meaningful test coverage; needs a real suite.
[Full-stack] Parallelize multi-source retrieval — retrieval pulls from multiple sources against a single Qdrant instance sequentially; needs parallelizing for speed.
[AI/ML] Validate the indexing setup — confirm the weighted-KNN approach and Qdrant index choice (FLAT vs. HNSW) actually hold up against real usage data.
[Design] Rework the landing page — real content and structural gaps, not just visual polish.
[Design] Establish a real design system — color scheme, gradients, typography are currently ad hoc; needs a consistent, deliberate system, not generic "AI-generated" styling.
LATER
[AI/ML] Expand tafsir/fatwa source library — add sources beyond Ibn Kathir and Muhsin Khan, per direct feedback from Islamic institutions.
[Full-stack] Digitize non-digitized Islamic books — some valuable reference material only exists non-digitized; needs a digitization pipeline.
[AI/ML] Fatwa extraction, classification & chunking — pipeline for new fatwa sources once added.
[AI/ML] Define hadith usage policy — decide deliberately how much hadith material the system draws on and under what conditions.
[AI/ML] Decide classifier's chat-history window — full conversation history vs. a trimmed window, backed by testing.
[Full-stack] Speed up translation/tafsir indexing — as the source library grows.
[Full-stack] Multi-language pipeline scaling — how the pipeline adjusts as more languages are added.
[Design] Background imagery & visual identity pass — once the core design system is in place.
OPEN PROBLEMS
RAG grounding & hallucination evaluation. Right now, catching ungrounded/incorrect chatbot answers relies on manual spot-checking. Needs a systematic evaluation pipeline: a golden set of questions with known-correct verses/citations, automated checks flagging answers not backed by the retrieved source, and a repeatable way to run this as the system changes. Why it matters: this is the biggest trust risk for an app whose whole pitch is "don't trust a generic AI with Quranic questions."
Linguistics features without bloating context. I'raab, balagha, and gharib al-alfaz would genuinely improve answer quality, but naively injecting all of it blows up the context window. Needs selective retrieval, summarized representations, or another structure that keeps the signal without the token cost. Why it matters: the difference between a shallow semantic search tool and something that actually understands classical Arabic linguistics.
Exact-match verification for Quranic text. Need a reliable, ideally automated way to verify anything presented as a direct Quran quote is an exact match to canonical text — not paraphrased or subtly altered — instead of relying on manual checks alone. Why it matters: a wording error in quoted Quranic text isn't a normal bug — it's a different category of mistake for this product.