MI
microsoft

Microsoft AI: MAI-Voice-2.1-Flash

microsoft/mai-voice-2.1-flash
تاریخ انتشار
۱۴۰۵/۷/۹
تعداد ارائه‌دهنده
۱
ورودی‌ها
متن
خروجی‌ها
speech

جزئیات قیمت

Characters۵٬۲۵۰٬۰۰۰تومان / ۱ میلیون کاراکتر

درباره مدل

MAI-Voice-2.1-Flash is a low-latency text-to-speech model from Microsoft AI, optimized for real-time responsiveness. It produces natural, expressive speech across 23 languages, with human-like intonation, rhythm, and emotional nuance. It is suited for voice agents, assistants, call centers, and other interactive applications where latency and cost matter most. On OpenRouter, set `voice` to a full voice ID with the model suffix, such as `"en-US-Harper:MAI-Voice-2.1-Flash"`. A voice's locale sets the synthesis language. Set `response_format` to `"mp3"` or `"pcm"` (24 kHz mono). Harper and Grant support the `agent`, `customer-call-center`, `educational`, and `narrator` speaking styles, and many locale voices add emotion styles such as `excited`, `happy`, `sad`, and `whispering`. The full list of voices is in the `supported_voices` field of the [models API](https://openrouter.ai/api/v1/models?output_modalities=speech). See the [text-to-speech guide](https://openrouter.ai/docs/guides/overview/multimodal/tts).

قابلیت‌های مدل

ورودی‌های پشتیبانی‌شده
متن
خروجی‌های پشتیبانی‌شده
speech
معماری ورودی/خروجیtext->speech

ارائه‌دهندگان (۱)

ارائه‌دهندهورودی (تومان / ۱M)خروجی (تومان / ۱M)پنجره متنحداکثر خروجیکوانتیزهوضعیت
Azure—————فعال

استفاده از طریق API

شناسه این مدل را در درخواست‌های API هوشگر استفاده کنید:

POST https://api.hooshgar.ir/v1/chat/completions
{
  "model": "microsoft/mai-voice-2.1-flash",
  "messages": [
    {
      "role": "user",
      "content": "سلام"
    }
  ]
}
مستندات کامل API

مدل‌های مرتبط