خريطة موضوعية · اكتشاف دلالي

نماذج وأدوات تحويل النص إلى كلام

اكتشف نماذج وأدوات تحويل النص إلى كلام مع التركيز على العربية، وزمن الاستجابة المنخفض، والنشر المحلي، والبرمجيات المفتوحة.

ابحث في هذا الموضوع ←شارك ملاحظة

مرشحون من الرسم المعرفي العام · آخر تحديث: 2026-09-26T11:07:05.502391+00:00

ابدأ بالنطاق، ثم انتقل إلى القرار

حدّد أولاً ما إذا كان الاستخدام تفاعلياً وفورياً، أو لتوليد نصوص طويلة، أو لاستنساخ الصوت. ثم قارن جودة العربية، وزمن الاستجابة، وطريقة النشر، وحدود الترخيص.

اكتشاف أولاً، لا ترتيب مدفوعاً.
هذه الصفحة تساعدك على فهم الموضوع واكتشاف المرشحين. عند الحاجة إلى مقارنة القيود والميزانية والنشر، اكتب المشكلة الفعلية في البحث الدلالي.

مداخل الرسم المعرفي

الأسماء والتفاصيل المعروضة هنا تبقى قريبة من السجل العام الأصلي عندما لا تتوفر ترجمة عربية موثقة؛ لا نملأ الفجوات بتخمينات.

آلية

text-to-speech

KG-recorded mechanism related to 文本转语音; detailed English notes are not yet available.

تطبيق

Text-to-speech

KG-recorded application related to Text-to-speech; detailed English notes are not yet available.

آلية

TTS

A sequential stage in traditional voice-plus-camera architectures.

نموذج

speecht5_tts

SpeechT5 model fine-tuned for speech synthesis (text-to-speech) on LibriTTS. This model was introduced in SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing by Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, Furu Wei.

نموذج

Qwen3-TTS-12Hz-0.6B-Base

**Qwen3-TTS Technical Report** **GitHub Repository** **Hugging Face Demo** Qwen3-TTS is a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Trained on over 5 million hours of speech data spanning 10 languages, Qwen3-TTS supports state-of-the-art 3-second voice cloning and description-based control.

نموذج

Qwen3-TTS-12Hz-0.6B-CustomVoice

Qwen3-TTS is a series of advanced multilingual, controllable, robust, and streaming text-to-speech models developed by the Qwen team.

نموذج

Step-Audio-TTS-3B

Step-Audio-TTS-3B represents the industry's first Text-to-Speech (TTS) model trained on a large-scale synthetic dataset utilizing the LLM-Chat paradigm. It has achieved SOTA Character Error Rate (CER) results on the SEED TTS Eval benchmark. The model supports multiple languages, a variety of emotional expressions, and diverse voice style controls. Notably, Step-Audio-TTS-3B is also the first TTS model in the industry capable of generating RAP and Humming, marking a significant advancement in the field of speech syn

نموذج

Voxtral-4B-TTS-2603

Voxtral TTS is a frontier, open-weights text-to-speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents. The model is released with BF16 weights and a set of reference voices. These voices are licensed under CC BY-NC 4, which is the license that the model inherits.

نموذج

Qwen-Audio-3.0-TTS

TTS model using inline control tags and natural-language style steering.

نموذج

Gemini 3.8 Flash TTS

KG-recorded model related to Gemini 3.8 Flash TTS; detailed English notes are not yet available.

نموذج

Gemini 3.8 Flash-Lite TTS

KG-recorded model related to Gemini 3.8 Flash-Lite TTS; detailed English notes are not yet available.

نموذج

Inworld TTS 2

KG-recorded model related to Inworld TTS 2; detailed English notes are not yet available.

أضف تجربة إلى هذا الموضوع

الملاحظات العامة منفصلة عن حقائق الرسم المعرفي، ويمكن للزوار مناقشتها والتصويت عليها.

افتح مساحة النقاش ←

استكشف موضوعات قريبة