Other launches
Released models outside the ranked categories: voice, speech, robotics, documents and more. Tracked here, never ranked or plotted. Each one sits on the “New in past 24 hours” banner for a day after launch.
- Perceptron Mk1.5Perceptron IncPerceptron Mk1.5 is an embodied-agent model built to control robots, with up to 4.7x faster end-to-end request completion than its predecessor.Robotics
- Gemini 3.8 Live with Live AvatarGoogle DeepMindGemini 3.8 Live with Live Avatar combines real-time video generation with speech to create an interactive virtual agent, with precise lip-syncing, natural facial expressions, and fluid dialogue turns across 97 languages.Voice
- Sarvam Vision 2.1Sarvam AISarvam Vision 2.1 is Sarvam's document-intelligence vision-language model, adding structured key-value extraction from forms, multi-page table parsing and Indic handwriting recognition, with fewer hallucinations and a cheaper API than the first release.Document
- GLiNER2.5-DecideFastinoGLiNER2.5-Decide is Fastino's 340M-parameter open-weight encoder for schema-defined decisions (classification, routing, triage): given text and typed questions it returns valid answers with probabilities and confidence scores, and runs on CPUs under Apache 2.0.OtherOpen weights
- Light-O1Light OriginsLight-O1 is Light Origins' first general-purpose embodied foundation model: it learns a human-action prior from 3D human motion recovered from internet video, then adapts it to different humanoid robots and tasks. A 4B Light-O1-Preview checkpoint is on Hugging Face under Apache 2.0.RoboticsOpen weights
- Gemini 3.8 Flash TTSGoogle DeepMindGemini 3.8 Flash TTS is a text-to-speech model for creative voice generation, letting creators build custom vocal characters from natural-language prompts with control over emotion, pacing, and accent across 100+ languages.TTS
- Qwen-Audio-3.1Alibaba QwenQwen-Audio-3.1 is Alibaba's five-model audio lineup: upgraded ASR, TTS and Realtime (full-duplex voice) models plus two new ones, ASR-Next for multi-speaker transcription and audio understanding and TTS-Next for generating voice, sound effects and background audio in one pass. It launched with price cuts of up to 95%.Speech
- Nemotron 3 DiarizationNVIDIANemotron 3 Diarization is NVIDIA's open-weight speaker diarization model that identifies who spoke when in multi-speaker audio, tracking up to 8 simultaneous speakers with streaming and offline inference.SpeechOpen weights
- FLUX 3 ActionBlack Forest LabsFLUX 3 Action is a 7B open-weights robot-control model derived from Black Forest Labs' FLUX 3 backbone, jointly predicting future video frames and robot actions for real-time manipulation.RoboticsOpen weights
- Realtime-VenusinclusionAI (Ant Group)Realtime-Venus is inclusionAI's full-duplex voice / audio-visual interaction model, able to keep perceiving while speaking and to distinguish backchannels, interruptions, corrections and redirections.VoiceOpen weights