Qur’anic Universal Audio is an open project for Qur’an audio, recitation metadata, and word/letter level timestamps generation with a built-in review system. The goal is to make Qur’an recitation audio easier to browse, verify, align, and use across different apps, tools, and research projects.
Instead of every Qur’an app having to solve the same problems separately — audio sources, reciter metadata, ayah and word timing/synchronization — this project aims to provide a shared foundation for the wider Qur’an tech community. It can help support features such as:
Ayah-by-ayah follow-along
Word-by-word synchronization
Letter-level and phoneme-level timing
Compilation of recitations from mainstream Qur'an audio sources
Reciter metadata and 1000+ mushaf recitations
Timestamp visualization and correction
Community review and improvement of data
Segment editing for failed alignments, missing words, low confidence, misaligned boundaries, and more - open for everyone
We have reached about 30 fully published recitations with word and letter timestamps, with more being aligned every day. You can help and contribute by resolving any mistakes using the built-in editing tool which flags a wide range of potential errors the AI alignment might have made, as well as contributing by adding your own recitations.
Tech stack:
FE: Svelte 5 (runes) + TypeScript
BE: Python/Flask, SQLite
Audio: WebAudio, CDN streaming, yt-dlp
AI/ML: VAD, ASR, CTC, MFA, wav2vec, string matching
Quran: phonemizer, DigitalKhatt, scripts, metadata
Infra: Hugging Face spaces/jobs/buckets + Docker
Website: https://huggingface.co/spaces/hetchyy/quranic-universal-audio
GitHub: https://github.com/Wider-Community/quranic-universal-audio
Timestamps data: https://github.com/Wider-Community/quranic-universal-audio/releases


مشروع "الصوت القرآني العالمي" هو مشروع مفتوح المصدر يُعنى بملفات الصوت القرآنية، وبيانات التلاوة، وإنشاء طوابع زمنية على مستوى الكلمات والحروف، مع نظام مراجعة مدمج. يهدف المشروع إلى تسهيل تصفح ملفات الصوت القرآنية، والتحقق منها، ومواءمتها، واستخدامها في مختلف التطبيقات والأدوات والمشاريع البحثية. بدلاً من أن يضطر كل تطبيق قرآني إلى حل المشكلات نفسها بشكل منفصل - مصادر الصوت، وبيانات التلاوة، وتوقيت الآيات والكلمات ومزامنتها - يسعى هذا المشروع إلى توفير أساس مشترك لمجتمع تقنيات القرآن الأوسع. يدعم هذا النظام ميزات مثل: المتابعة اللفظية للآيات مزامنة الكلمات توقيت دقيق على مستوى الحروف والأصوات تجميع التلاوات من مصادر صوتية قرآنية رئيسية بيانات القراء وأكثر من 1000 تلاوة للمصحف عرض وتصحيح الطوابع الزمنية مراجعة البيانات وتحسينها من قبل المجتمع تحرير المقاطع لمعالجة أخطاء المحاذاة، والكلمات المفقودة، وانخفاض مستوى الثقة، وعدم محاذاة الحدود، وغيرها - متاح للجميع وصلنا حتى الآن إلى حوالي 30 تلاوة منشورة بالكامل مع طوابع زمنية للكلمات والحروف، ويتم محاذاة المزيد منها يوميًا. يمكنك المساعدة والمساهمة من خلال تصحيح أي أخطاء باستخدام أداة التحرير المدمجة التي تُشير إلى مجموعة واسعة من الأخطاء المحتملة التي قد تكون ارتكبها نظام المحاذاة بالذكاء الاصطناعي، بالإضافة إلى المساهمة بإضافة تلاواتك الخاصة.
{منقول}