Tekmono
  • News
  • Guides
  • Lists
  • Reviews
  • Deals
No Result
View All Result
Tekmono
No Result
View All Result
Home News
Mistral Launches Voxtral Open-Source Speech Models

Mistral Launches Voxtral Open-Source Speech Models

by Tekmono Editorial Team
16/07/2025
in News
Share on FacebookShare on Twitter

Voxtral has launched new open-source speech understanding models, aiming to revolutionize human-computer interaction by making voice interfaces more reliable and accessible. These state-of-the-art models are available under the Apache 2.0 license.

The models, available in 24B and 3B variants, offer exceptional transcription and deep understanding capabilities, addressing limitations of current proprietary and open-source systems. Voxtral bridges the gap between high-cost, closed APIs and less accurate open-source alternatives, providing state-of-the-art accuracy and native semantic understanding at less than half the price of comparable APIs. The models support long-form audio up to 30 minutes for transcription and 40 minutes for understanding, featuring a 32k token context length. Additionally, they include built-in Q&A and summarization, automatic language detection for widely used languages such as English, Spanish, French, Portuguese, Hindi, German, Dutch, and Italian, and direct function-calling from voice commands.

In benchmarks, Voxtral significantly outperforms leading open-source models like Whisper large-v3 and competes strongly with GPT-4o mini Transcribe and Gemini 2.5 Flash in speech transcription and audio understanding. For instance, Voxtral Mini Transcribe is more cost-effective than OpenAI Whisper, while Voxtral Small matches ElevenLabs Scribe’s performance at a lower price point. The models also retain strong text understanding capabilities from their Mistral Small 3.1 backbone.

Related Reads

Microsoft enhances Copilot with multimodal features, introduces new $99 tier

Apple celebrates 50th anniversary amid scrutiny over privacy practices

Huawei launches Converged Development Engine for HarmonyOS PCs

Salesforce unveils updated Slack with 30 new AI features

Voxtral models are available for local download on Hugging Face and via API, with pricing starting at $0.001 per minute. Enterprise features include private deployment, domain-specific fine-tuning, and advanced context capabilities like speaker identification and emotion detection. Future updates will include speaker segmentation, audio markups, and word-level timestamps, further enhancing their utility.

ShareTweet

You Might Be Interested

Microsoft enhances Copilot with multimodal features, introduces new  tier
News

Microsoft enhances Copilot with multimodal features, introduces new $99 tier

02/04/2026
News

Apple celebrates 50th anniversary amid scrutiny over privacy practices

02/04/2026
News

Huawei launches Converged Development Engine for HarmonyOS PCs

02/04/2026
Salesforce unveils updated Slack with 30 new AI features
News

Salesforce unveils updated Slack with 30 new AI features

02/04/2026
Please login to join discussion

Recent Posts

  • Microsoft enhances Copilot with multimodal features, introduces new $99 tier
  • Apple celebrates 50th anniversary amid scrutiny over privacy practices
  • Huawei launches Converged Development Engine for HarmonyOS PCs
  • Salesforce unveils updated Slack with 30 new AI features
  • Meta announces release of second generation smart glasses starting April 14

Recent Comments

No comments to show.
  • News
  • Guides
  • Lists
  • Reviews
  • Deals
Tekmono is a Linkmedya brand. © 2015.

No Result
View All Result
  • News
  • Guides
  • Lists
  • Reviews
  • Deals