MI
microsoft

Microsoft AI: MAI-Voice-2.1

microsoft/mai-voice-2.1
تاریخ انتشار
۱۴۰۵/۷/۹
تعداد ارائه‌دهنده
۱
ورودی‌ها
متن
خروجی‌ها
speech

جزئیات قیمت

Characters۷٬۷۰۰٬۰۰۰تومان / ۱ میلیون کاراکتر

درباره مدل

MAI-Voice-2.1 is Microsoft AI's highest-fidelity, most expressive text-to-speech model. It produces natural, studio-grade speech across 23 languages, with detailed prosody, nuanced expressiveness, and speaker consistency over long-form content. It is suited for audiobooks, podcasts, lectures, narration, and brand audio where maximum voice quality matters. The model prioritizes naturalness and expressivity over latency-critical generation. On OpenRouter, set `voice` to a full voice ID with the model suffix, such as `"en-US-Harper:MAI-Voice-2.1"`. A voice's locale sets the synthesis language. Set `response_format` to `"mp3"` or `"pcm"` (24 kHz mono). Harper and Grant support the `agent`, `customer-call-center`, `educational`, and `narrator` speaking styles, and many locale voices add emotion styles such as `excited`, `happy`, `sad`, and `whispering`. The full list of voices is in the `supported_voices` field of the [models API](https://openrouter.ai/api/v1/models?output_modalities=speech). See the [text-to-speech guide](https://openrouter.ai/docs/guides/overview/multimodal/tts).

قابلیت‌های مدل

ورودی‌های پشتیبانی‌شده
متن
خروجی‌های پشتیبانی‌شده
speech
معماری ورودی/خروجیtext->speech

ارائه‌دهندگان (۱)

ارائه‌دهندهورودی (تومان / ۱M)خروجی (تومان / ۱M)پنجره متنحداکثر خروجیکوانتیزهوضعیت
Azure—————فعال

استفاده از طریق API

شناسه این مدل را در درخواست‌های API هوشگر استفاده کنید:

POST https://api.hooshgar.ir/v1/chat/completions
{
  "model": "microsoft/mai-voice-2.1",
  "messages": [
    {
      "role": "user",
      "content": "سلام"
    }
  ]
}
مستندات کامل API

مدل‌های مرتبط