
French artificial intelligence firm Mistral has introduced an innovative open-source text-to-speech model designed for voice AI applications and enterprise solutions, including customer support. This new offering, named Voxtral TTS, positions Mistral in direct competition with established players like ElevenLabs, Deepgram, and OpenAI. Voxtral TTS supports an impressive array of nine languages, encompassing English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, and Arabic. Pierre Stock, VP of Science Operations at Mistral AI, explained in a recent interview that the model was developed in response to customer demand for an efficient speech solution. He highlighted that it is compact enough to operate on various edge devices, including smartwatches and smartphones, while being cost-effective compared to other market options. One of the standout features of Voxtral TTS is its ability to personalize voices with a sample of less than five seconds. It captures unique vocal characteristics such as accents, intonations, and natural speech flow. Built on the Ministral 3B architecture, the model excels at seamlessly switching between languages without compromising voice quality, making it ideal for tasks like dubbing and real-time translation. Stock emphasized the model’s design focus on delivering a human-like speaking experience rather than a robotic tone. Mistral claims that Voxtral TTS is optimized for real-time applications, achieving a time-to-first-audio (TTFA) of just 90 milliseconds for a 10-second audio sample containing 500 characters. Additionally, it boasts a real-time factor (RTF) of 6x, meaning it can produce a 10-second clip in approximately 1.6 seconds. Earlier this year, Mistral launched two transcription models aimed at large batch processing and low-latency real-time scenarios. The introduction of this speech model signals the company’s ambition to offer a comprehensive suite of voice-driven products for enterprise needs. Stock noted that Mistral envisions creating an end-to-end platform capable of processing multimodal inputs, such as audio, text, and images, thereby providing richer information through an integrated system that leverages audio both as input and output. With its emphasis on open-source technology and customization options, Mistral aims to attract enterprises seeking tailored voice solutions, giving them the flexibility to modify the models to suit their specific requirements.
In the realm of cybersecurity, few figures are as intriguing as Phineas Fisher, a hacker who has evaded capture for near...
TechCrunch | Jul 25, 2026, 21:00
As technology continues to evolve, a notable shift is occurring in the relationship between humans and artificial intell...
Business Insider | Jul 25, 2026, 09:50Monday.com, the innovative work management platform based in Tel Aviv, has recently announced significant layoffs, attri...
TechCrunch | Jul 26, 2026, 01:45
Nvidia has successfully forged a significant partnership with South Korea's SK Hynix to secure memory supplies essential...
CNBC | Jul 25, 2026, 05:15
In recent years, the AI sector has been intensely focused on identifying the most advanced models. While this pursuit re...
Business Insider | Jul 25, 2026, 13:10