Model Pareto — AI models ranked by quality against price and speed
Coding, quality vs price
Coding · quality vs price
More filters for Coding →Lower price is better. Pareto frontier: Granite 4.2 3B, NVIDIA Nemotron 3 Nano 30B A3B, gpt-oss-20b, Ling-3.0-flash-VL, MiMo-V2.6-Flash, GLM-5.3 Flash, DeepSeek V4.1 Flash, MiMo-V2.6-Pro, Claude Sonnet 5.5, Claude Opus 5.5.
- Best-value frontier (nothing is both cheaper and better)
- Evidencestrong → weak
- Estimated from other categories
- Estimated (price or speed)
- Hover a dot for details · click a lab to highlight it
259 models from 69 AI labs, frontier and open-weight, scored on public benchmarks and independent evals. How scoring works
Categories
Top three by quality, and the price frontier at a glance.
Coding
103 modelsAgentic software engineering and code generation.
- 1Claude Opus 5.596.9$8.00
- 2Claude Sonnet 5.591.7$4.00
- 3Claude Fable 5.187.6$20.0
Writing & chat
146 modelsInstruction following, creative writing, and translation.
- 1Claude Fable 5.197.6$20.0
- 2Claude Opus 593.4$10.0
- 3Claude Opus 5.593.0$8.00
Vision
82 modelsImage understanding, documents, and charts.
- 1GPT-6 Astra98.9$20.0
- 2Claude Opus 5.594.5$8.00
- 3Claude Fable 5.187.4$20.0
Hard reasoning
182 modelsGraduate-level science and competition math.
- 1Claude Opus 5.599.9$8.00
- 2Claude Fable 5.195.7$20.0
- 3Claude Sonnet 5.595.2$4.00
Agents
173 modelsTool use, computer use, and web research.
- 1Claude Opus 5.599.6$8.00
- 2Claude Fable 5.189.6$20.0
- 3Muse Spark 1.387.4$2.00
Image generation
38 modelsText-to-image, image editing, and text rendering.
- 1GPT Image 2.5 Sunburst99.5$0.21
- 2GPT Image 2.5 Flare91.7$0.21
- 3GPT Image 282.6$0.21
Video generation
34 modelsText-to-video, image-to-video, and native audio.
- 1Gemini Omni Flash96.9$0.10
- 2Gemini Omni 1.1 Flash94.3$0.10
- 3Wan 3.093.8$0.20
- Read the methodology →
How rankings are built
Benchmarks put on a common 0–100 scale per category, estimates for missing scores, and where every number comes from, from lab-reported to independently verified.
Recently added
Ranked from launch-day numbers, upgraded as third-party results arrive.
Beyond the leaderboard
Voice, speech, robotics and other launches from the past 30 days. Tracked, not ranked.
- Perceptron Mk1.5Perceptron IncPerceptron Mk1.5 is an embodied-agent model built to control robots, with up to 4.7x faster end-to-end request completion than its predecessor.Robotics
- Gemini 3.8 Live with Live AvatarGoogle DeepMindGemini 3.8 Live with Live Avatar combines real-time video generation with speech to create an interactive virtual agent, with precise lip-syncing, natural facial expressions, and fluid dialogue turns across 97 languages.Voice
- Sarvam Vision 2.1Sarvam AISarvam Vision 2.1 is Sarvam's document-intelligence vision-language model, adding structured key-value extraction from forms, multi-page table parsing and Indic handwriting recognition, with fewer hallucinations and a cheaper API than the first release.Document
- GLiNER2.5-DecideFastinoGLiNER2.5-Decide is Fastino's 340M-parameter open-weight encoder for schema-defined decisions (classification, routing, triage): given text and typed questions it returns valid answers with probabilities and confidence scores, and runs on CPUs under Apache 2.0.OtherOpen weights
- Light-O1Light OriginsLight-O1 is Light Origins' first general-purpose embodied foundation model: it learns a human-action prior from 3D human motion recovered from internet video, then adapts it to different humanoid robots and tasks. A 4B Light-O1-Preview checkpoint is on Hugging Face under Apache 2.0.RoboticsOpen weights
- Gemini 3.8 Flash TTSGoogle DeepMindGemini 3.8 Flash TTS is a text-to-speech model for creative voice generation, letting creators build custom vocal characters from natural-language prompts with control over emotion, pacing, and accent across 100+ languages.TTS
- Qwen-Audio-3.1Alibaba QwenQwen-Audio-3.1 is Alibaba's five-model audio lineup: upgraded ASR, TTS and Realtime (full-duplex voice) models plus two new ones, ASR-Next for multi-speaker transcription and audio understanding and TTS-Next for generating voice, sound effects and background audio in one pass. It launched with price cuts of up to 95%.Speech
- Nemotron 3 DiarizationNVIDIANemotron 3 Diarization is NVIDIA's open-weight speaker diarization model that identifies who spoke when in multi-speaker audio, tracking up to 8 simultaneous speakers with streaming and offline inference.SpeechOpen weights
- FLUX 3 ActionBlack Forest LabsFLUX 3 Action is a 7B open-weights robot-control model derived from Black Forest Labs' FLUX 3 backbone, jointly predicting future video frames and robot actions for real-time manipulation.RoboticsOpen weights
- Realtime-VenusinclusionAI (Ant Group)Realtime-Venus is inclusionAI's full-duplex voice / audio-visual interaction model, able to keep perceiving while speaking and to distinguish backchannels, interruptions, corrections and redirections.VoiceOpen weights