Dr Moot Model Catalog
Every language model currently available to Dr Moot, with live capabilities, limits, and release details.
This catalog is refreshed hourly, so newly added and retired language models appear here without a manual documentation release.
Current model support
Dr Moot currently supports language models, with image and video models coming soon.
Showing 253 of 253 language models
Language models (253)
Fugu Maxsakana/fugu-max
A description is not currently available for this model.
- Provider
- Sakana
- Context window
- 1,000,000 tokens
- Maximum output
- 1,000,000 tokens
- Released
- 10 Sept 2026
Fugu Ultra v2sakana/fugu-ultra-v2
Fugu Ultra coordinates a deeper pool of expert agents to maximize answer quality on hard, high-stakes problems.
- Provider
- Sakana
- Context window
- 1,000,000 tokens
- Maximum output
- 1,000,000 tokens
- Released
- 10 Sept 2026
DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash
DeepSeek V4.1 Flash is an AI model from DeepSeek designed for fast, efficient multimodal interactions.
- Provider
- Deepseek
- Context window
- 1,048,576 tokens
- Maximum output
- 32,768 tokens
- Released
- 8 Sept 2026
Mercury 2.5inception/mercury-2.5
Mercury 2.5 is Inception’s diffusion-based reasoning model for chat, agents, and structured workflows, with tool calling, structured outputs, and a 260K context window.
- Provider
- Inception
- Context window
- 260,000 tokens
- Maximum output
- 65,536 tokens
- Released
- 8 Sept 2026
Ling 3.0 Flash VLinclusionai/ling-3.0-flash-vl
Ling 3.0 Flash VL builds on Ling 3.0 Flash with stronger language capabilities, native visual perception, and visual agent capabilities. It supports text, image, and video inputs with text output, reasoning, and function calling.
- Provider
- Inclusionai
- Context window
- 256,000 tokens
- Maximum output
- 32,000 tokens
- Released
- 8 Sept 2026
Ling 3.0 Flash VL (Free)inclusionai/ling-3.0-flash-vl-free
Ling 3.0 Flash VL builds on Ling 3.0 Flash with stronger language capabilities, native visual perception, and visual agent capabilities. It supports text, image, and video inputs with text output, reasoning, and function calling.
- Provider
- Inclusionai
- Context window
- 256,000 tokens
- Maximum output
- 32,000 tokens
- Released
- 8 Sept 2026
Ling 3.0 Flash Santeinclusionai/ling-3.0-flash-sante
Ling-3.0-Flash-Sante is inclusionAI’s language model specialized for health and medicine, built on a Mixture-of-Experts architecture with 124 billion total parameters and approximately 5.1 billion active per token. With a 256K context window and function calling, it supports medical knowledge reasoning, evidence-based retrieval, and complex medical workflows while retaining general reasoning, coding, and agentic capabilities.
- Provider
- Inclusionai
- Context window
- 256,000 tokens
- Maximum output
- 32,000 tokens
- Released
- 4 Sept 2026
Ling 3.0 Flash Sante (Free)inclusionai/ling-3.0-flash-sante-free
Ling-3.0-Flash-Sante is inclusionAI’s language model specialized for health and medicine, built on a Mixture-of-Experts architecture with 124 billion total parameters and approximately 5.1 billion active per token. With a 256K context window and function calling, it supports medical knowledge reasoning, evidence-based retrieval, and complex medical workflows while retaining general reasoning, coding, and agentic capabilities.
- Provider
- Inclusionai
- Context window
- 256,000 tokens
- Maximum output
- 32,000 tokens
- Released
- 4 Sept 2026
GPT-6 Astraopenai/gpt-6-astra
GPT-6 Astra is OpenAI's most capable model for complex reasoning, coding, computer use, research, and document creation.
- Provider
- Openai
- Context window
- 1,050,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 4 Sept 2026
GPT-6 Astra (Fast)openai/gpt-6-astra-fast
GPT-6 Astra is OpenAI's most capable model for complex reasoning, coding, computer use, research, and document creation.
- Provider
- Openai
- Context window
- 1,050,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 4 Sept 2026
Gemini 3.8 Flashgoogle/gemini-3.8-flash
It delivers Pro-level agentic capabilities, major leaps in code generation and terminal execution. 3.8 Flash serves as the primary agentic workhorse in the Gemini 3 family, bridging the gap between deep-reasoning Pro models and high-throughput Flash-Lite models while delivering high token efficiency and multi-step multimodal processing.
- Provider
- Context window
- 1,000,000 tokens
- Maximum output
- 65,536 tokens
- Released
- 2 Sept 2026
GLM 5.3 Fastzai/glm-5.3-fast
Speed-optimized version of Z.AI’s GLM-5.3 agentic coding model, built for responsive, real-time workloads. It supports advanced coding, long-horizon task execution, testing, vulnerability discovery, and cybersecurity analysis, making it ideal for developer tools, coding agents, and complex multi-step engineering workflows.
- Provider
- Zai
- Context window
- 1,048,576 tokens
- Maximum output
- 262,144 tokens
- Released
- 2 Sept 2026
Qwen3.8 Max 0902alibaba/qwen3.8-max-0902
Qwen3.8-Max-0902 is an upgraded snapshot of qwen3.8-max.
- Provider
- Alibaba
- Context window
- 991,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 1 Sept 2026
Claude Fable 5.1anthropic/claude-fable-5.1
Claude Fable 5.1 improves on Fable 5 across long-running agentic coding, knowledge work, and research, following instructions precisely over sessions that run unattended for hours.
- Provider
- Anthropic
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 31 Aug 2026
Tencent Hy4 Previewtencent/hy4-preview
Tencent Hy4 preview is Tencent Hy’s open-source large language model, featuring 770B total parameters, 49B active parameters, and a context window exceeding 1 million tokens. Built for real-world productivity, it supports long-horizon coding, cross-document analysis, office content creation, game development, and scientific reasoning, making it well-suited to complex agents and multi-step professional workflows.
- Provider
- Tencent
- Context window
- 1,024,000 tokens
- Maximum output
- 64,000 tokens
- Released
- 28 Aug 2026
Ling 3.0 Flash Fininclusionai/ling-3.0-flash-fin
Ling 3.0 Flash Fin is InclusionAI’s finance-enhanced MoE language model, combining 124 billion total parameters with approximately 5.1 billion active parameters for efficient financial reasoning. Its 256K context window, function calling, and support for complex, multi-step investment workflows make it ideal for financial research, analysis, long-horizon planning, and execution, while retaining strong capabilities in coding and mathematics.
- Provider
- Inclusionai
- Context window
- 256,000 tokens
- Maximum output
- 32,000 tokens
- Released
- 27 Aug 2026
Ling 3.0 Flash Fin (Free)inclusionai/ling-3.0-flash-fin-free
Ling 3.0 Flash Fin is InclusionAI’s finance-enhanced MoE language model, combining 124 billion total parameters with approximately 5.1 billion active parameters for efficient financial reasoning. Its 256K context window, function calling, and support for complex, multi-step investment workflows make it ideal for financial research, analysis, long-horizon planning, and execution, while retaining strong capabilities in coding and mathematics.
- Provider
- Inclusionai
- Context window
- 256,000 tokens
- Maximum output
- 32,000 tokens
- Released
- 27 Aug 2026
Qwen 3.8 Flashalibaba/qwen3.8-flash
Built for coding, agentic workflows, and visual understanding, it handles large codebases, long documents, charts, videos, and desktop applications.
- Provider
- Alibaba
- Context window
- 991,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 26 Aug 2026
GLM 5.3 Flashzai/glm-5.3-flash
GLM-5.3-Flash is Z.ai’s native multimodal coding model, featuring 320B total parameters, 18B activated parameters, and a 1M-token context window. Its efficient hybrid attention architecture supports visual coding, tool use, and end-to-end professional workflows across code, browsers, documents, and graphical interfaces.
- Provider
- Zai
- Context window
- 1,000,000 tokens
- Maximum output
- 131,000 tokens
- Released
- 26 Aug 2026
DeepSeek V4 Flash Vision Expdeepseek/deepseek-v4-flash-vision-exp
DeepSeek-V4-Flash-Vision-Exp is the first experimental multimodal model in the DeepSeek-V4 family. It builds on DeepSeek-V4-Flash by adding visual modules and continued training for visual understanding, with substantially improved multimodal agent capabilities while remaining comparable on text-only agent tasks.
- Provider
- Deepseek
- Context window
- 1,048,576 tokens
- Maximum output
- 1,048,576 tokens
- Released
- 21 Aug 2026
GLM 5.3zai/glm-5.3
GLM 5.3 delivers comprehensive advancements in complex software engineering and agent capabilities. It uses the same base model as GLM-5.2, with all improvements driven by post-training.
- Provider
- Zai
- Context window
- 1,000,000 tokens
- Maximum output
- 1,000,000 tokens
- Released
- 18 Aug 2026
Qwen3.8 27Balibaba/qwen3.8-27b
Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability.
- Provider
- Alibaba
- Context window
- 1,000,000 tokens
- Maximum output
- 131,072 tokens
- Released
- 14 Aug 2026
Gemini 3.7 Flashgoogle/gemini-3.7-flash
It delivers Pro-level agentic capabilities, major leaps in code generation and terminal execution. 3.7 Flash serves as the primary agentic workhorse in the Gemini 3 family, bridging the gap between deep-reasoning Pro models and high-throughput Flash-Lite models while delivering high token efficiency and multi-step multimodal processing.
- Provider
- Context window
- 1,000,000 tokens
- Maximum output
- 65,536 tokens
- Released
- 13 Aug 2026
DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813
This is the 8/13 updated weights version of DeepSeek V4 Pro.
- Provider
- Deepseek
- Context window
- 1,000,000 tokens
- Maximum output
- 384,000 tokens
- Released
- 12 Aug 2026
Grok 4.6spacexai/grok-4.6
Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work. It stays with complex tasks across many steps, whether researching a topic, analyzing information, working across a codebase, or turning an idea into a polished application or work artifact.
- Provider
- Spacexai
- Context window
- 500,000 tokens
- Maximum output
- 500,000 tokens
- Released
- 12 Aug 2026
Nemotron 3.5 Lightning 30Bnvidia/nemotron-3.5-lightning
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 is a large language model (LLM) trained by NVIDIA. The model employs a hybrid Mixture-of-Experts architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. The Lightning 3.5 model is released alongside a number of speculative decoding methods for faster text generation. The model has 3B active parameters and 30B parameters in total.
- Provider
- Nvidia
- Context window
- 262,144 tokens
- Maximum output
- 131,072 tokens
- Released
- 11 Aug 2026
Muse Glimmer 30Bmeta/muse-glimmer-30b
A description is not currently available for this model.
- Provider
- Meta
- Context window
- 131,072 tokens
- Maximum output
- 131,072 tokens
- Released
- 10 Aug 2026
Ling 3.0 Flashinclusionai/ling-3.0-flash
A description is not currently available for this model.
- Provider
- Inclusionai
- Context window
- 256,000 tokens
- Maximum output
- 32,000 tokens
- Released
- 6 Aug 2026
Muse Spark 1.2meta/muse-spark-1.2
A coding-optimized model purpose-built for agentic workflows. Improvements to code generation, debugging, and codebase understanding — with a 1M context window that handles your entire project in one session.
- Provider
- Meta
- Context window
- 1,048,576 tokens
- Maximum output
- 1,048,576 tokens
- Released
- 5 Aug 2026
Muse Spark 1.2 Contributormeta/muse-spark-1.2-contributor
Same model, same capabilities — up to 95% less than Standard. Your inputs and outputs are used to train and improve Meta's AI models.
- Provider
- Meta
- Context window
- 1,048,576 tokens
- Maximum output
- 1,048,576 tokens
- Released
- 5 Aug 2026
Sakana Namazusakana/namazu
Sakana Namazu is a Japanese-specialized LLM that combines a deep understanding of Japanese culture and business customs with high-performance language capabilities. Built on the open model Kimi K2.6 and refined with Sakana AI's in-house data for Japanese language and business workflows, it handles complex tasks using web search and code execution. Unlike Fugu, which orchestrates multiple frontier models, Sakana Namazu provides a single in-house model as an API.
- Provider
- Sakana
- Context window
- 256,000 tokens
- Maximum output
- 256,000 tokens
- Released
- 3 Aug 2026
Qwen3.8 2.4T A95Balibaba/qwen3.8-2.4t-a95b
Qwen 3.8 Max is a 2.4-trillion-parameter MoE model delivering a comprehensive leap in coding and professional work. Autonomously codes and delivers complete projects spanning 10+ days. Handles hundreds of specialized tasks across legal, financial, design, and other professional domains, producing production-grade results end-to-end in a single conversation. Native visual understanding runs through the full cycle of planning, execution, and verification, enabling deep semantic analysis of ultra-long documents and extended video content. In long-horizon tasks, plans autonomously, iterates through closed feedback loops, and continuously evolves.
- Provider
- Alibaba
- Context window
- 262,144 tokens
- Maximum output
- 128,000 tokens
- Released
- 2 Aug 2026
Qwen 3.8 Maxalibaba/qwen3.8-max
Qwen 3.8 Max is a 2.4-trillion-parameter MoE model delivering a comprehensive leap in coding and professional work. Autonomously codes and delivers complete projects spanning 10+ days. Handles hundreds of specialized tasks across legal, financial, design, and other professional domains, producing production-grade results end-to-end in a single conversation. Native visual understanding runs through the full cycle of planning, execution, and verification, enabling deep semantic analysis of ultra-long documents and extended video content. In long-horizon tasks, plans autonomously, iterates through closed feedback loops, and continuously evolves.
- Provider
- Alibaba
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 2 Aug 2026
Inkling Smallthinkingmachines/inkling-small
A description is not currently available for this model.
- Provider
- Thinkingmachines
- Context window
- 1,000,000 tokens
- Maximum output
- 1,000,000 tokens
- Released
- 30 Jul 2026
Qwen 3.7 Flashalibaba/qwen3.7-flash
The Qwen3.7 native vision-language Flash model series delivers a comprehensive upgrade over 3.6-Flash in multimodal understanding and agent execution. This model particularly excels in enhanced multimodal foundations with stronger universal object recognition, further improved real-world perception and spatial intelligence, significantly upgraded multimodal agent capabilities for Search Agent and CI Agent scenarios with more stable end-to-end task execution, as well as optimized multimodal coding for a smoother vibe coding experience.
- Provider
- Alibaba
- Context window
- 991,000 tokens
- Maximum output
- 64,000 tokens
- Released
- 28 Jul 2026
Kimi K3 Fastmoonshotai/kimi-k3-fast
Fast version of Kimi’s flagship model for long-horizon coding and end-to-end knowledge work, with a 1M-token context window.
- Provider
- Moonshotai
- Context window
- 1,000,000 tokens
- Maximum output
- 131,072 tokens
- Released
- 27 Jul 2026
Claude Opus 5anthropic/claude-opus-5
Claude Opus 5 is the latest model in Anthropic's Opus family and a step-change improvement over Opus 4.8. It delivers major gains over Opus 4.8 in agentic coding, professional knowledge work, and long-horizon reasoning, and it is stronger per token across effort levels.
- Provider
- Anthropic
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 24 Jul 2026
Claude Opus 5 (Fast)anthropic/claude-opus-5-fast
Claude Opus 5 is the latest model in Anthropic's Opus family and a step-change improvement over Opus 4.8. It delivers major gains over Opus 4.8 in agentic coding, professional knowledge work, and long-horizon reasoning, and it is stronger per token across effort levels.
- Provider
- Anthropic
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 24 Jul 2026
Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite
Gemini 3.5 Flash Lite features upgraded agentic capabilities, making the model ideal for subagents in complex workflows.
- Provider
- Context window
- 1,000,000 tokens
- Maximum output
- 65,000 tokens
- Released
- 21 Jul 2026
Gemini 3.6 Flashgoogle/gemini-3.6-flash
Gemini 3.6 Flash delivers higher quality across coding, agentic workflows, and web development with reduced token consumption and fewer model calls compared to previous model iterations.
- Provider
- Context window
- 1,000,000 tokens
- Maximum output
- 64,000 tokens
- Released
- 21 Jul 2026
Laguna S 2.1 Freepoolside/laguna-s-2.1-free
Laguna S 2.1 is Poolside's new open-weight model for agentic coding and long-horizon work.
- Provider
- Poolside
- Context window
- 256,000 tokens
- Maximum output
- 32,768 tokens
- Released
- 21 Jul 2026
Laguna S 2.1poolside/laguna-s-2.1
Laguna S 2.1 is Poolside's new open-weight model for agentic coding and long-horizon work.
- Provider
- Poolside
- Context window
- 1,000,000 tokens
- Maximum output
- 131,072 tokens
- Released
- 20 Jul 2026
Kimi K3moonshotai/kimi-k3
Kimi’s flagship model for long-horizon coding and end-to-end knowledge work, with a 1M-token context window.
- Provider
- Moonshotai
- Context window
- 1,000,000 tokens
- Maximum output
- 131,072 tokens
- Released
- 16 Jul 2026
Inklingthinkingmachines/inkling
Inkling is a multimodal MoE model (975B total, 41B active, 256k context) reasoning over text, image, and audio inputs.
- Provider
- Thinkingmachines
- Context window
- 256,000 tokens
- Maximum output
- 256,000 tokens
- Released
- 15 Jul 2026
Kat Coder Air V2.5kwaipilot/kat-coder-air-v2.5
Fast response version of Kat Coder V2.5 optimized for Agent and Claw use cases.
- Provider
- Kwaipilot
- Context window
- 256,000 tokens
- Maximum output
- 80,000 tokens
- Released
- 10 Jul 2026
Kat Coder Pro V2.5kwaipilot/kat-coder-pro-v2.5
KAT-Coder-V2.5 is a coding-focused agentic model trained to act autonomously inside real, executable repositories rather than as a single-turn code generator.
- Provider
- Kwaipilot
- Context window
- 256,000 tokens
- Maximum output
- 80,000 tokens
- Released
- 10 Jul 2026
Muse Spark 1.1meta/muse-spark-1.1
Muse Spark 1.1 is strongest at agentic performance, tool use, and computer use. It does well on long-running tasks with 1M token context window, can delegate execution to sub-agents running in parallel, and is trained to use computer interfaces on desktop, mobile, or browser.
- Provider
- Meta
- Context window
- 1,048,576 tokens
- Maximum output
- 1,048,576 tokens
- Released
- 9 Jul 2026
GPT 5.6 Lunaopenai/gpt-5.6-luna
A description is not currently available for this model.
- Provider
- Openai
- Context window
- 1,050,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 9 Jul 2026
GPT 5.6 Luna (Fast)openai/gpt-5.6-luna-fast
A description is not currently available for this model.
- Provider
- Openai
- Context window
- 1,050,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 9 Jul 2026
GPT 5.6 Solopenai/gpt-5.6-sol
GPT-5.6 Sol is the flagship of OpenAI's GPT-5.6 series, its most capable model for long-horizon agentic work across coding, biology, and cybersecurity.
- Provider
- Openai
- Context window
- 1,050,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 9 Jul 2026
GPT 5.6 Sol (Fast)openai/gpt-5.6-sol-fast
GPT-5.6 Sol is the flagship of OpenAI's GPT-5.6 series, its most capable model for long-horizon agentic work across coding, biology, and cybersecurity.
- Provider
- Openai
- Context window
- 1,050,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 9 Jul 2026
GPT 5.6 Terraopenai/gpt-5.6-terra
A description is not currently available for this model.
- Provider
- Openai
- Context window
- 1,050,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 9 Jul 2026
GPT 5.6 Terra (Fast)openai/gpt-5.6-terra-fast
A description is not currently available for this model.
- Provider
- Openai
- Context window
- 1,050,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 9 Jul 2026
Grok 4.5spacexai/grok-4.5
SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
- Provider
- Spacexai
- Context window
- 500,000 tokens
- Maximum output
- 500,000 tokens
- Released
- 8 Jul 2026
Hy3tencent/hy3
Built for real-world business scenarios, Hy3 features a 295B/21B active MoE architecture, native 256K context support, and three reasoning modes.
- Provider
- Tencent
- Context window
- 262,144 tokens
- Maximum output
- 262,144 tokens
- Released
- 6 Jul 2026
Claude Fable 5anthropic/claude-fable-5
Claude Fable 5 is a Mythos-class model with robust safeguards. It can handle long-running, complex, and asynchronous tasks where previous models would have needed more frequent check-ins.
- Provider
- Anthropic
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 1 Jul 2026
Gemini 3.1 Flash Lite Image (Nano Banana 2 Lite)google/gemini-3.1-flash-lite-image
Gemini 3.1 Flash-Lite Image (Nano Banana 2 Lite) is Google's fastest image generation model enabling rapid creation and iteration.
- Provider
- Context window
- 65,536 tokens
- Maximum output
- 4,096 tokens
- Released
- 30 Jun 2026
Gemini Omni Flash Previewgoogle/gemini-omni-flash-preview
Gemini Omni Flash (Preview) is a multimodal model designed for video, image, and text tasks. It is optimized for video generation, offering video output alongside text responses in a single model.
- Provider
- Context window
- 1,000,000 tokens
- Maximum output
- 57,920 tokens
- Released
- 30 Jun 2026
Claude Sonnet 5anthropic/claude-sonnet-5
Sonnet 5 is an upgrade to Sonnet 4.6, with gains across agentic coding and professional work.
- Provider
- Anthropic
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 29 Jun 2026
Seed 2.1 Turbobytedance/seed-2.1-turbo
Seed 2.1 Turbo is ByteDance’s multimodal language model supporting text, image, and video inputs with text output. It offers a 262,144-token context window, reasoning, function calling, and structured JSON output.
- Provider
- Bytedance
- Context window
- 262,144 tokens
- Maximum output
- 262,144 tokens
- Released
- 23 Jun 2026
GLM 5.2 Fastzai/glm-5.2-fast
Fast version of GLM 5.2 with 120-250 TPS.
- Provider
- Zai
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 23 Jun 2026
Fugu Ultrasakana/fugu-ultra
Fugu Ultra coordinates a deeper pool of expert agents to maximize answer quality on hard, high-stakes problems. It routes between one to three agents, depending on the problem.
- Provider
- Sakana
- Context window
- 1,000,000 tokens
- Maximum output
- 1,000,000 tokens
- Released
- 21 Jun 2026
GLM 5.2zai/glm-5.2
GLM-5.2 delivers powerful coding capabilities, usable 1M-context support, and continued strengths in long-horizon tasks.
- Provider
- Zai
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 16 Jun 2026
Kimi K2.7 Code High Speedmoonshotai/kimi-k2.7-code-highspeed
Kimi K2.7 Code HighSpeed is the high-speed version of Kimi K2.7 Code, the same model as Kimi K2.7 Code, but with an output speed of approximately 180 Tokens/s and up to 260 Tokens/s in short context scenarios, delivering a more extreme coding experience.
- Provider
- Moonshotai
- Context window
- 262,144 tokens
- Maximum output
- 32,768 tokens
- Released
- 15 Jun 2026
Kimi K2.7 Codemoonshotai/kimi-k2.7-code
Kimi-K2.7-Code is a coding model from Moonshot AI. It has improved coding & agent performance over K2.6, more reasoning efficiency with less overthinking, and improved instruction following for long-horizon coding.
- Provider
- Moonshotai
- Context window
- 256,000 tokens
- Maximum output
- 32,768 tokens
- Released
- 12 Jun 2026
Tencent Hy-MT2-Litetencent/hy-mt2-lite
Hy-MT2-Lite is Tencent Cloud Hunyuan’s lightweight 1.8B translation model, combining efficient performance with an 8K-token context window and broad multilingual coverage. It supports translation across 33 languages and five ethnic Chinese and dialect variants, delivering strong results on benchmarks including FLORES-200 and WMT25. Designed for professional and real-world business use, it reliably follows instructions for structured, delimiter-preserving, context-aware, glossary-guided, and style-specific translation.
- Provider
- Tencent
- Context window
- 8,000 tokens
- Maximum output
- 4,000 tokens
- Released
- 12 Jun 2026
Tencent Hy-MT2-Plustencent/hy-mt2-plus
Hy-MT2-Plus is Tencent Cloud Hunyuan’s 7B translation model, offering an 8K-token context window and broad multilingual coverage. Optimized for translation across 33 languages and five ethnic Chinese and dialect variants, it delivers industry-leading performance on benchmarks including FLORES-200 and WMT25. It excels in professional domains and real-world business scenarios, with strong instruction-following for structured, delimiter-preserving, context-aware, glossary-guided, and style-specific translation.
- Provider
- Tencent
- Context window
- 8,000 tokens
- Maximum output
- 4,000 tokens
- Released
- 12 Jun 2026
Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b
A 550B parameter (55B active) open reasoning model from NVIDIA, built for long-running agent workflows. It uses a hybrid Mamba-Transformer MoE architecture and supports a 1M token context window.
- Provider
- Nvidia
- Context window
- 1,000,000 tokens
- Maximum output
- 65,000 tokens
- Released
- 4 Jun 2026
Qwen 3.7 Plusalibaba/qwen3.7-plus
A description is not currently available for this model.
- Provider
- Alibaba
- Context window
- 1,000,000 tokens
- Maximum output
- 64,000 tokens
- Released
- 2 Jun 2026
MiniMax M3minimax/minimax-m3
MiniMax-M3 is a frontier-class foundation model that unites the three capabilities defining today's frontier: a 1M-token context window, frontier coding and agentic performance, and native multimodality — the first open-weight model to deliver all three in a single system.
- Provider
- Minimax
- Context window
- 512,000 tokens
- Maximum output
- 512,000 tokens
- Released
- 31 May 2026
Claude Opus 4.8anthropic/claude-opus-4.8
Opus 4.8 is a focused upgrade to Opus 4.7 and is Anthropic's best generally available model for coding, agentic tasks, and enterprise workflows. It builds on the strengths of previous Opus models with stronger performance on complex, multi-step coding tasks. Anthropic recommends using it on long-horizon coding and agentic tasks. It is also stronger on professional work, including document drafting, data analysis, and presentations.
- Provider
- Anthropic
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 28 May 2026
Claude Opus 4.8 (Fast)anthropic/claude-opus-4.8-fast
Opus 4.8 is a focused upgrade to Opus 4.7 and is Anthropic's best generally available model for coding, agentic tasks, and enterprise workflows. It builds on the strengths of previous Opus models with stronger performance on complex, multi-step coding tasks. Anthropic recommends using it on long-horizon coding and agentic tasks. It is also stronger on professional work, including document drafting, data analysis, and presentations.
- Provider
- Anthropic
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 28 May 2026
Gemini 3.1 Flash Image (Nano Banana 2)google/gemini-3.1-flash-image
A description is not currently available for this model.
- Provider
- Context window
- 131,072 tokens
- Maximum output
- 32,768 tokens
- Released
- 28 May 2026
Step 3.7 Flashstepfun/step-3.7-flash
StepFun’s flagship multimodal reasoning model. Powered by a 198B-parameter / 11B-activation sparse MoE architecture, with native support for image and video understanding.
- Provider
- Stepfun
- Context window
- 256,000 tokens
- Maximum output
- 256,000 tokens
- Released
- 28 May 2026
Qwen 3.7 Maxalibaba/qwen3.7-max
Qwen3.7 is a next‑generation flagship model designed for the agent‑centric era, with its core strengths lying in the breadth and depth of its agent‑level capabilities: it excels at programming, office and productivity tasks, and long‑term autonomous execution.
- Provider
- Alibaba
- Context window
- 991,000 tokens
- Maximum output
- 64,000 tokens
- Released
- 21 May 2026
Tencent Hy-MT2-Protencent/hy-mt2-pro
Hy-MT2-Pro is Tencent Cloud Hunyuan’s flagship translation model, featuring 30B-A3B parameters and an 8K-token context window. It delivers bidirectional translation across 33 core language pairs, plus five minority-language and dialect pairs. With industry-leading performance on FLORES-200 and WMT25, it excels in specialized domains and real-world business scenarios while supporting structured, contextual, terminology-aware, style-specific, and delimiter-preserving translation.
- Provider
- Tencent
- Context window
- 8,000 tokens
- Maximum output
- 4,000 tokens
- Released
- 21 May 2026
Grok Build 0.1spacexai/grok-build-0.1
xAI's fast coding model trained specifically for agentic coding.
- Provider
- Spacexai
- Context window
- 256,000 tokens
- Maximum output
- 256,000 tokens
- Released
- 20 May 2026
Gemini 3.5 Flashgoogle/gemini-3.5-flash
Google's latest model, highly optimized for coding proficiency and parallel agentic execution loops.
- Provider
- Context window
- 1,000,000 tokens
- Maximum output
- 64,000 tokens
- Released
- 19 May 2026
Gemini 3.1 Flash Litegoogle/gemini-3.1-flash-lite
Gemini 3.1 Flash Lite outperforms 2.5 Flash Lite on overall quality and lands close to 2.5 Flash performance across key capability areas. It is a workhorse model for high-volume use cases, with improvements across audio input/ASR, RAG snippet ranking, translation, data extraction, and code completion.
- Provider
- Context window
- 1,000,000 tokens
- Maximum output
- 65,000 tokens
- Released
- 7 May 2026
Grok 4.3spacexai/grok-4.3
Grok 4.3 is a new model matching the scale of Grok 4.20 with an improved architecture and a December 2025 knowledge cutoff.
- Provider
- Spacexai
- Context window
- 1,000,000 tokens
- Maximum output
- 1,000,000 tokens
- Released
- 30 Apr 2026
Mistral Medium Latestmistral/mistral-medium-3.5
Mistral's frontier-class multimodal model optimized for agentic and coding use cases.
- Provider
- Mistral
- Context window
- 256,000 tokens
- Maximum output
- 256,000 tokens
- Released
- 29 Apr 2026
GPT 5.5openai/gpt-5.5
GPT‑5.5 understands what you’re trying to do faster and can carry more of the work itself. It excels at writing and debugging code, researching online, analyzing data, creating documents and spreadsheets, operating software, and moving across tools until a task is finished. Instead of carefully managing every step, you can give GPT‑5.5 a messy, multi-part task and trust it to plan, use tools, check its work, navigate through ambiguity, and keep going.
- Provider
- Openai
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 24 Apr 2026
GPT 5.5 (Fast)openai/gpt-5.5-fast
GPT‑5.5 understands what you’re trying to do faster and can carry more of the work itself. It excels at writing and debugging code, researching online, analyzing data, creating documents and spreadsheets, operating software, and moving across tools until a task is finished. Instead of carefully managing every step, you can give GPT‑5.5 a messy, multi-part task and trust it to plan, use tools, check its work, navigate through ambiguity, and keep going.
- Provider
- Openai
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 24 Apr 2026
GPT 5.5 Proopenai/gpt-5.5-pro
A description is not currently available for this model.
- Provider
- Openai
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 24 Apr 2026
DeepSeek V4 Flashdeepseek/deepseek-v4-flash
A description is not currently available for this model.
- Provider
- Deepseek
- Context window
- 1,000,000 tokens
- Maximum output
- 384,000 tokens
- Released
- 23 Apr 2026
DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731
A description is not currently available for this model.
- Provider
- Deepseek
- Context window
- 1,000,000 tokens
- Maximum output
- 384,000 tokens
- Released
- 23 Apr 2026
DeepSeek V4 Prodeepseek/deepseek-v4-pro
DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to improve long-context efficiency; (2) ManifoldConstrained Hyper-Connections (mHC) that enhance conventional residual connections; (3) and the Muon optimizer for faster convergence and greater training stability
- Provider
- Deepseek
- Context window
- 1,000,000 tokens
- Maximum output
- 384,000 tokens
- Released
- 23 Apr 2026
Qwen 3.6 27Balibaba/qwen3.6-27b
The Qwen3.6 35B-A3B native vision-language model is built on a hybrid architecture that integrates linear attention mechanisms with a sparse mixture-of-experts framework, achieving higher inference efficiency. Compared with the 3.5-35B-A3B, this model demonstrates significantly improved agentic coding capabilities, mathematical and code reasoning abilities, spatial intelligence, as well as object localization and object detection performance.
- Provider
- Alibaba
- Context window
- 256,000 tokens
- Maximum output
- 256,000 tokens
- Released
- 22 Apr 2026
MiMo M2.5xiaomi/mimo-v2.5
A native full-modal model supporting text, image, video, and audio understanding, with powerful Agent capabilities.
- Provider
- Xiaomi
- Context window
- 1,050,000 tokens
- Maximum output
- 131,100 tokens
- Released
- 22 Apr 2026
MiMo V2.5 Proxiaomi/mimo-v2.5-pro
MiMo V2.5 Pro delivers significant improvements over its predecessor, MiMo-V2-Pro, in general agentic capabilities, complex software engineering, and long-horizon tasks. MiMo-V2.5-Pro is a 1.02T-parameter Mixture-of-Experts model with 42B active parameters, built on a hybrid-attention architecture with a 1M-token context window.
- Provider
- Xiaomi
- Context window
- 1,050,000 tokens
- Maximum output
- 131,000 tokens
- Released
- 22 Apr 2026
Qwen 3.6 Max Previewalibaba/qwen-3.6-max-preview
Compared with the previously released Qwen3-Max and Qwen3.6-Plus, this model features enhanced vibe coding abilities, more efficient coding agent execution, and significantly improved front-end development skills. Additionally, its long-tail knowledge retention has been further upgraded.
- Provider
- Alibaba
- Context window
- 240,000 tokens
- Maximum output
- 64,000 tokens
- Released
- 20 Apr 2026
Kimi K2.6moonshotai/kimi-k2.6
Kimi K2.6 demonstrates particularly strong performance in long-horizon coding tasks and produces professional-grade design with code and vision.
- Provider
- Moonshotai
- Context window
- 262,000 tokens
- Maximum output
- 262,000 tokens
- Released
- 20 Apr 2026
Claude Opus 4.7anthropic/claude-opus-4.7
Opus 4.7 builds on the coding and agentic strengths of Opus 4.6 with stronger performance on complex, multi-step tasks and more reliable agentic execution. It also brings improved performance on knowledge work, from drafting documents to building presentations and analyzing data.
- Provider
- Anthropic
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 16 Apr 2026
GLM 5.1zai/glm-5.1
GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on a single task for more than 8 hours—autonomously planning, executing, and improving itself throughout the process—ultimately delivering complete, engineering-grade results.
- Provider
- Zai
- Context window
- 202,800 tokens
- Maximum output
- 64,000 tokens
- Released
- 7 Apr 2026
Qwen 3.6 Plusalibaba/qwen3.6-plus
The Qwen3.6 native vision-language Plus series models demonstrate exceptional performance on par with the current state-of-the-art models, with a significant improvement in overall results compared to the 3.5 series. The models have been markedly enhanced in code-related capabilities such as agentic coding, front-end programming, and Vibe coding, as well as in multi-modal general object recognition, OCR, and object localization.
- Provider
- Alibaba
- Context window
- 1,000,000 tokens
- Maximum output
- 64,000 tokens
- Released
- 2 Apr 2026
Google Gemma 4 26B A4Bgoogle/gemma-4-26b-a4b-it
Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on small models) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.
- Provider
- Context window
- 262,144 tokens
- Maximum output
- 131,072 tokens
- Released
- 2 Apr 2026
Gemma 4 31B ITgoogle/gemma-4-31b-it
Gemma 4 31B is engineered to tackle the most demanding enterprise workloads and complex reasoning tasks. With an expansive 256K-token context window, the 31B model can effortlessly ingest entire codebases, and massive sets of images in a single prompt.
- Provider
- Context window
- 262,144 tokens
- Maximum output
- 131,072 tokens
- Released
- 2 Apr 2026
Trinity Large Thinkingarcee-ai/trinity-large-thinking
Trinity-Large-Thinking is a reasoning-optimized variant of Arcee AI's Trinity-Large family — a 398B-parameter sparse Mixture-of-Experts (MoE) model with approximately 13B active parameters per token. Built on Trinity-Large-Base and post-trained with extended chain-of-thought reasoning and agentic RL, Trinity-Large-Thinking delivers state-of-the-art performance on agentic benchmarks while maintaining strong general capabilities.
- Provider
- Arcee Ai
- Context window
- 262,100 tokens
- Maximum output
- 80,000 tokens
- Released
- 1 Apr 2026
GLM 5V Turbozai/glm-5v-turbo
GLM-5V-Turbo is Z.AI’s first multimodal coding foundation model, built for vision-based coding tasks. It can natively process multimodal inputs such as images, video, and text, while also excelling at long-horizon planning, complex coding, and action execution. Deeply optimized for agent workflows, it works seamlessly with agents such as Claude Code and OpenClaw to complete the full loop of “understand the environment → plan actions → execute tasks”.
- Provider
- Zai
- Context window
- 200,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 1 Apr 2026
Kat Coder Pro V2kwaipilot/kat-coder-pro-v2
A high-performance edition designed for complex enterprise projects and SaaS integration.
- Provider
- Kwaipilot
- Context window
- 256,000 tokens
- Maximum output
- 256,000 tokens
- Released
- 27 Mar 2026
MiniMax M2.7minimax/minimax-m2.7
M2.7 delivers outstanding performance in real-world software engineering, including end-to-end full project delivery, log analysis and bug troubleshooting, code security, machine learning, and more.
- Provider
- Minimax
- Context window
- 204,800 tokens
- Maximum output
- 131,000 tokens
- Released
- 18 Mar 2026
MiniMax M2.7 High Speedminimax/minimax-m2.7-highspeed
M2.7 Highspeed: Same performance, faster and more agile (output speed approximately 100 tps)
- Provider
- Minimax
- Context window
- 204,800 tokens
- Maximum output
- 131,100 tokens
- Released
- 18 Mar 2026
GPT 5.4 Miniopenai/gpt-5.4-mini
GPT-5.4 Mini brings the strengths of GPT-5.4 to a faster, more efficient model designed for high-volume workloads.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 17 Mar 2026
GPT 5.4 Mini (Fast)openai/gpt-5.4-mini-fast
GPT-5.4 Mini brings the strengths of GPT-5.4 to a faster, more efficient model designed for high-volume workloads.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 17 Mar 2026
GPT 5.4 Nanoopenai/gpt-5.4-nano
A description is not currently available for this model.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 17 Mar 2026
GLM 5 Turbozai/glm-5-turbo
GLM 5 Turbo is a foundation model deeply optimized for the OpenClaw scenario. It has been specifically optimized for the core requirements of OpenClaw tasks since the training phase, enhancing key capabilities such as tool invocation, command following, timed and persistent tasks, and long-chain execution.
- Provider
- Zai
- Context window
- 202,800 tokens
- Maximum output
- 131,100 tokens
- Released
- 15 Mar 2026
NVIDIA Nemotron 3 Super 120B A12Bnvidia/nemotron-3-super-120b-a12b
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Additionally, a long context window gives the model long-term memory, preventing AI agents from losing focus on long, multi-step tasks and ensuring high-accuracy results. Fully open with weights, datasets, and recipes, Super allows easy customization and secure deployment anywhere.
- Provider
- Nvidia
- Context window
- 256,000 tokens
- Maximum output
- 32,000 tokens
- Released
- 11 Mar 2026
Grok 4.20 Multi Agent Betaspacexai/grok-4.20-multi-agent-beta
Multiple agents collaborate in parallel to perform deep research tasks.
- Provider
- Spacexai
- Context window
- 2,000,000 tokens
- Maximum output
- 2,000,000 tokens
- Released
- 11 Mar 2026
Grok 4.20 Beta Non-Reasoningspacexai/grok-4.20-non-reasoning-beta
Grok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering consistently precise and truthful responses.
- Provider
- Spacexai
- Context window
- 2,000,000 tokens
- Maximum output
- 2,000,000 tokens
- Released
- 11 Mar 2026
Grok 4.20 Beta Reasoningspacexai/grok-4.20-reasoning-beta
Grok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering consistently precise and truthful responses.
- Provider
- Spacexai
- Context window
- 2,000,000 tokens
- Maximum output
- 2,000,000 tokens
- Released
- 11 Mar 2026
Grok 4.20 Multi-Agentspacexai/grok-4.20-multi-agent
Multiple agents collaborate in parallel to perform deep research tasks.
- Provider
- Spacexai
- Context window
- 2,000,000 tokens
- Maximum output
- 2,000,000 tokens
- Released
- 10 Mar 2026
Grok 4.20 Non-Reasoningspacexai/grok-4.20-non-reasoning
Grok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherence, delivering consistently precise and truthful responses.
- Provider
- Spacexai
- Context window
- 2,000,000 tokens
- Maximum output
- 2,000,000 tokens
- Released
- 10 Mar 2026
Grok 4.20 Reasoningspacexai/grok-4.20-reasoning
Grok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherence, delivering consistently precise and truthful responses.
- Provider
- Spacexai
- Context window
- 2,000,000 tokens
- Maximum output
- 2,000,000 tokens
- Released
- 10 Mar 2026
GPT 5.4openai/gpt-5.4
GPT-5.4 is OpenAI's best general-purpose model, part of the GPT-5 flagship model family. It's their most intelligent model yet for both general and agentic tasks.
- Provider
- Openai
- Context window
- 1,050,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 5 Mar 2026
GPT 5.4 (Fast)openai/gpt-5.4-fast
GPT-5.4 is OpenAI's best general-purpose model, part of the GPT-5 flagship model family. It's their most intelligent model yet for both general and agentic tasks.
- Provider
- Openai
- Context window
- 1,050,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 5 Mar 2026
GPT 5.4 Proopenai/gpt-5.4-pro
GPT-5.4 Pro uses more compute to think harder and provide consistently better answers. It's designed to tackle tough problems.
- Provider
- Openai
- Context window
- 1,050,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 5 Mar 2026
Gemini 3.1 Flash Image Preview (Nano Banana 2)google/gemini-3.1-flash-image-preview
A description is not currently available for this model.
- Provider
- Context window
- 131,072 tokens
- Maximum output
- 32,768 tokens
- Released
- 26 Feb 2026
Qwen 3.5 Flashalibaba/qwen3.5-flash
The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the 3 series, these models deliver a leap forward in performance for both pure text and multimodal tasks, offering fast response times while balancing inference speed and overall performance.
- Provider
- Alibaba
- Context window
- 1,000,000 tokens
- Maximum output
- 64,000 tokens
- Released
- 24 Feb 2026
Mercury 2inception/mercury-2
A diffusion-based reasoning LLM that generates text via parallel refinement (not token-by-token), delivering real-time latency with ~1k tokens/sec plus 128K context and built-in tool/JSON support.
- Provider
- Inception
- Context window
- 128,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 24 Feb 2026
Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview
This model improves upon Gemini 2.5 Pro and is catered towards challenging tasks, especially those involving complex reasoning or agentic workflows. Improvements highlighted include use cases for coding, multi-step function calling, planning, reasoning, deep knowledge tasks, and instruction following.
- Provider
- Context window
- 1,000,000 tokens
- Maximum output
- 64,000 tokens
- Released
- 19 Feb 2026
Claude Sonnet 4.6anthropic/claude-sonnet-4.6
Claude Sonnet 4.6 is the most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with memory, polished document creation, and confident computer use for web QA and workflow automation.
- Provider
- Anthropic
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 17 Feb 2026
Qwen 3.5 Plusalibaba/qwen3.5-plus
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of task evaluations, the 3.5 series consistently demonstrates performance on par with state-of-the-art leading models. Compared to the 3 series, these models show a leap forward in both pure-text and multimodal capabilities.
- Provider
- Alibaba
- Context window
- 1,000,000 tokens
- Maximum output
- 64,000 tokens
- Released
- 16 Feb 2026
MiniMax M2.5minimax/minimax-m2.5
MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. It is capable of handling the entire development process of various complex systems. It covers full-stack projects across multiple platforms including Web, Android, iOS, Windows, and Mac, encompassing server-side APIs, functional logic, and databases.
- Provider
- Minimax
- Context window
- 204,800 tokens
- Maximum output
- 131,000 tokens
- Released
- 12 Feb 2026
MiniMax M2.5 High Speedminimax/minimax-m2.5-highspeed
M2.5 highspeed: Same performance, faster and more agile (output speed approximately 100 tps)
- Provider
- Minimax
- Context window
- 204,800 tokens
- Maximum output
- 131,000 tokens
- Released
- 12 Feb 2026
GLM 5zai/glm-5
GLM 5 is a frontier-class, general-purpose large language model optimized for complex systems engineering and long-horizon agentic tasks. It builds on the GLM 4.5 agent-centric lineage and is designed to support multi-step reasoning, math (including AIME-style benchmarks), advanced coding, and tool-augmented workflows, with long context support suitable for sophisticated agents and enterprise applications. Typical uses include autonomous agents for software engineering, data and systems troubleshooting, operations copilots, and high-end chat assistants that must break down complex tasks, call tools reliably, and reason over long sequences of instructions or documents.
- Provider
- Zai
- Context window
- 202,800 tokens
- Maximum output
- 131,100 tokens
- Released
- 12 Feb 2026
Claude Opus 4.6anthropic/claude-opus-4.6
Opus 4.6 is the world’s best model for coding and professional work, built to power agents that take on whole categories of real-world work. It excels across the entire SDLC, breaking through on hard problems, identifying complex bugs, and demonstrating deeper codebase understanding. It also delivers a step-change in knowledge work, with near-production-ready documents, presentations, and spreadsheets on the first pass.
- Provider
- Anthropic
- Context window
- 1,000,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 5 Feb 2026
GPT 5.3 Codexopenai/gpt-5.3-codex
GPT-5.3-Codex advances both the frontier coding performance of GPT‑5.2-Codex and the reasoning and professional knowledge capabilities of GPT‑5.2, together in one model, which is also 25% faster. This enables it to take on long-running tasks that involve research, tool use, and complex execution.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 5 Feb 2026
GPT 5.3 Codex (Fast)openai/gpt-5.3-codex-fast
GPT-5.3-Codex advances both the frontier coding performance of GPT‑5.2-Codex and the reasoning and professional knowledge capabilities of GPT‑5.2, together in one model, which is also 25% faster. This enables it to take on long-running tasks that involve research, tool use, and complex execution.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 5 Feb 2026
StepFun 3.5 Flashstepfun/step-3.5-flash
Step 3.5 Flash is an open-source reasoning model by StepFun with 196B total parameters (11B active) using Mixture of Experts. It features a 256K context window, deep reasoning, tool calling, and agentic capabilities, achieving 97.3 on AIME 2025 and 74.4% on SWE-bench Verified.
- Provider
- Stepfun
- Context window
- 262,114 tokens
- Maximum output
- 262,114 tokens
- Released
- 29 Jan 2026
Kimi K2.5moonshotai/kimi-k2.5
kimi-k2.5 is Kimi's most versatile model to date, featuring a native multimodal architecture that supports both visual and text input, thinking and non-thinking modes, and dialogue and agent tasks.
- Provider
- Moonshotai
- Context window
- 262,114 tokens
- Maximum output
- 262,114 tokens
- Released
- 26 Jan 2026
Qwen 3 Max Thinkingalibaba/qwen3-max-thinking
Compared with the snapshot as of September 23, 2025, the Qwen-3 series Max model in this release achieves an effective integration of thinking and non-thinking modes, resulting in a comprehensive and substantial improvement in the model’s overall performance. In thinking mode, the model simultaneously supports web search, web information extraction, and a code interpreter tool, enabling it to tackle more complex and challenging problems with greater accuracy by leveraging external tools while engaging in slow, deliberative reasoning. This version is based on a snapshot taken on January 23, 2026.
- Provider
- Alibaba
- Context window
- 256,000 tokens
- Maximum output
- 65,536 tokens
- Released
- 23 Jan 2026
GLM 4.7 Flashzai/glm-4.7-flash
GLM-4.7-Flash balances high performance with efficiency, making it the perfect lightweight deployment option. Beyond coding, it is also recommended for creative writing, translation, long-context tasks, and roleplay.
- Provider
- Zai
- Context window
- 200,000 tokens
- Maximum output
- 131,000 tokens
- Released
- 19 Jan 2026
GLM 4.7 FlashXzai/glm-4.7-flashx
GLM-4.7-Flash balances high performance with efficiency, making it the perfect lightweight deployment option.
- Provider
- Zai
- Context window
- 200,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 19 Jan 2026
MiniMax M2.1minimax/minimax-m2.1
MiniMax 2.1 is MiniMax's latest model, optimized specifically for robustness in coding, tool use, instruction following, and long-horizon planning.
- Provider
- Minimax
- Context window
- 204,800 tokens
- Maximum output
- 131,072 tokens
- Released
- 23 Dec 2025
MiniMax M2.1 Lightningminimax/minimax-m2.1-lightning
MiniMax-M2.1-lightning is a faster version of MiniMax-M2.1, offering the same performance but with significantly higher throughput (output speed ~100 TPS, MiniMax-M2 output speed ~60 TPS).
- Provider
- Minimax
- Context window
- 204,800 tokens
- Maximum output
- 131,072 tokens
- Released
- 23 Dec 2025
GLM 4.7zai/glm-4.7
GLM-4.7 is Z.ai’s latest flagship model, with major upgrades focused on two key areas: stronger coding capabilities and more stable multi-step reasoning and execution.
- Provider
- Zai
- Context window
- 200,000 tokens
- Maximum output
- 120,000 tokens
- Released
- 22 Dec 2025
GPT 5.2 Codexopenai/gpt-5.2-codex
GPT‑5.2-Codex is a version of GPT‑5.2 further optimized for agentic coding in Codex, including improvements on long-horizon work through context compaction, stronger performance on large code changes like refactors and migrations, improved performance in Windows environments, and significantly stronger cybersecurity capabilities.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 18 Dec 2025
Gemini 3 Flashgoogle/gemini-3-flash
Google's most intelligent model built for speed, combining frontier intelligence with superior search and grounding.
- Provider
- Context window
- 1,000,000 tokens
- Maximum output
- 65,000 tokens
- Released
- 17 Dec 2025
Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b
Built with a hybrid MoE and Mamba architecture and trained on NVIDIA-curated synthetic reasoning data, it delivers strong multi-step reasoning with stable latency and predictable performance for agentic and production workloads.
- Provider
- Nvidia
- Context window
- 262,144 tokens
- Maximum output
- 262,144 tokens
- Released
- 15 Dec 2025
GPT 5.2openai/gpt-5.2
GPT-5.2 is OpenAI's best general-purpose model, part of the GPT-5 flagship model family. It's their most intelligent model yet for both general and agentic tasks.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 11 Dec 2025
GPT 5.2 (Fast)openai/gpt-5.2-fast
GPT-5.2 is OpenAI's best general-purpose model, part of the GPT-5 flagship model family. It's their most intelligent model yet for both general and agentic tasks.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 11 Dec 2025
GPT 5.2 openai/gpt-5.2-pro
Version of GPT-5.2 that produces smarter and more precise responses.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 11 Dec 2025
Devstral 2mistral/devstral-2
An enterprise-grade text model that excels at using tools to explore codebases, editing multiple files, and powering software engineering agents.
- Provider
- Mistral
- Context window
- 256,000 tokens
- Maximum output
- 256,000 tokens
- Released
- 9 Dec 2025
Devstral Small 2mistral/devstral-small-2
Our open source model that excels at using tools to explore codebases, editing multiple files, and powering software engineering agents.
- Provider
- Mistral
- Context window
- 256,000 tokens
- Maximum output
- 256,000 tokens
- Released
- 9 Dec 2025
Nova 2 Liteamazon/nova-2-lite
A description is not currently available for this model.
- Provider
- Amazon
- Context window
- 1,000,000 tokens
- Maximum output
- 1,000,000 tokens
- Released
- 2 Dec 2025
Ministral 14Bmistral/ministral-14b
Ministral 3 14B is the largest model in the Ministral 3 family, offering state-of-the-art capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. Optimized for local deployment, it delivers high performance across diverse hardware, including local setups.
- Provider
- Mistral
- Context window
- 256,000 tokens
- Maximum output
- 256,000 tokens
- Released
- 2 Dec 2025
Mistral Large 3mistral/mistral-large-3
Mistral Large 3 2512 is Mistral’s most capable model to date. It has a sparse mixture-of-experts architecture with 41B active parameters (675B total).
- Provider
- Mistral
- Context window
- 256,000 tokens
- Maximum output
- 256,000 tokens
- Released
- 2 Dec 2025
GPT OSS Safeguard 120Bopenai/gpt-oss-safeguard-120b
GPT OSS Safeguard 120B is OpenAI’s open-weight safety model for content moderation and guardrail enforcement. It helps teams evaluate content against custom safety policies and build scalable, adaptable protections for production AI applications.
- Provider
- Openai
- Context window
- 128,000 tokens
- Maximum output
- 16,000 tokens
- Released
- 2 Dec 2025
DeepSeek V3.2deepseek/deepseek-v3.2
DeepSeek‑V3.2 from DeepSeek harmonizes high computational efficiency with superior reasoning and agent performance. It builds on three main techniques: DeepSeek Sparse Attention for long‑context efficiency, a scalable reinforcement learning framework, and a large‑scale agentic task synthesis pipeline. This model excels at long-context reasoning and agentic tasks, efficiently handling extended inputs while maintaining strong accuracy. Its sparse attention design enables it to process complex, multi-step workflows without excessive compute costs. Overall, DeepSeek‑V3.2 targets long‑context reasoning, tool‑using agents, and efficient deployment in production environments.
- Provider
- Deepseek
- Context window
- 128,000 tokens
- Maximum output
- 8,000 tokens
- Released
- 1 Dec 2025
DeepSeek V3.2 Thinkingdeepseek/deepseek-v3.2-thinking
DeepSeek‑V3.2 from DeepSeek harmonizes high computational efficiency with superior reasoning and agent performance. It builds on three main techniques: DeepSeek Sparse Attention for long‑context efficiency, a scalable reinforcement learning framework, and a large‑scale agentic task synthesis pipeline. This model excels at long-context reasoning and agentic tasks, efficiently handling extended inputs while maintaining strong accuracy. Its sparse attention design enables it to process complex, multi-step workflows without excessive compute costs. Overall, DeepSeek‑V3.2 targets long‑context reasoning, tool‑using agents, and efficient deployment in production environments.
- Provider
- Deepseek
- Context window
- 128,000 tokens
- Maximum output
- 8,000 tokens
- Released
- 1 Dec 2025
Claude Opus 4.5anthropic/claude-opus-4.5
Claude Opus 4.5 is Anthropic’s latest model in the Opus series, meant for demanding reasoning tasks and complex problem solving. This model has improvements in general intelligence and vision compared to previous iterations. In addition, it is suited for difficult coding tasks and agentic workflows, especially those with computer use and tool use, and can effectively handle context usage and external memory files.
- Provider
- Anthropic
- Context window
- 200,000 tokens
- Maximum output
- 64,000 tokens
- Released
- 24 Nov 2025
GPT 5.1 Codex Maxopenai/gpt-5.1-codex-max
GPT‑5.1-Codex-Max is purpose-built for agentic coding.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 19 Nov 2025
Grok 4.1 Fast Non-Reasoningspacexai/grok-4.1-fast-non-reasoning
A description is not currently available for this model.
- Provider
- Spacexai
- Context window
- 1,000,000 tokens
- Maximum output
- 1,000,000 tokens
- Released
- 19 Nov 2025
Grok 4.1 Fast Reasoningspacexai/grok-4.1-fast-reasoning
A description is not currently available for this model.
- Provider
- Spacexai
- Context window
- 1,000,000 tokens
- Maximum output
- 1,000,000 tokens
- Released
- 19 Nov 2025
GPT-5.1-Codexopenai/gpt-5.1-codex
GPT-5.1-Codex is a version of GPT-5.1 optimized for agentic coding tasks in Codex or similar environments.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 12 Nov 2025
GPT 5.1 Codex Miniopenai/gpt-5.1-codex-mini
GPT-5.1 Codex mini is a smaller, faster, and cheaper version of GPT-5.1 Codex.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 12 Nov 2025
GPT 5.1 Thinkingopenai/gpt-5.1-thinking
An upgraded version of GPT-5 that adapts thinking time more precisely to the question to spend more time on complex questions and respond more quickly to simpler tasks.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 12 Nov 2025
GPT 5.1 Thinking (Fast)openai/gpt-5.1-thinking-fast
An upgraded version of GPT-5 that adapts thinking time more precisely to the question to spend more time on complex questions and respond more quickly to simpler tasks.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 12 Nov 2025
KAT-Coder-Pro V1kwaipilot/kat-coder-pro-v1
KAT-Coder-Pro V1 is KwaiKAT's most advanced agentic coding model in the KwaiKAT series. Designed specifically for agentic coding tasks, it excels in real-world software engineering scenarios, achieving a remarkable 73.4% solve rate on the SWE-Bench Verified benchmark. KAT-Coder-Pro V1 delivers top-tier coding performance and has been rigorously tested by thousands of in-house engineers. The model has been optimized for tool-use capability, multi-turn interaction, instruction following, generalization and comprehensive capabilities through a multi-stage training process, including mid-training, supervised fine-tuning (SFT), reinforcement fine-tuning (RFT), and scalable agentic RL.
- Provider
- Kwaipilot
- Context window
- 256,000 tokens
- Maximum output
- 32,000 tokens
- Released
- 9 Nov 2025
Kimi K2 Thinkingmoonshotai/kimi-k2-thinking
Kimi K2 Thinking is an advanced open-source thinking model by Moonshot AI. It can execute up to 200 – 300 sequential tool calls without human interference, reasoning coherently across hundreds of steps to solve complex problems. Built as a thinking agent, it reasons step by step while using tools, achieving state-of-the-art performance on Humanity's Last Exam (HLE), BrowseComp, and other benchmarks, with major gains in reasoning, agentic search, coding, writing, and general capabilities.
- Provider
- Moonshotai
- Context window
- 216,144 tokens
- Maximum output
- 216,144 tokens
- Released
- 6 Nov 2025
GPT OSS Safeguard 20Bopenai/gpt-oss-safeguard-20b
OpenAI's first open weight reasoning model specifically trained for safety classification tasks. Fine-tuned from GPT-OSS, this model helps classify text content based on customizable policies, enabling bring-your-own-policy Trust & Safety AI where your own taxonomy, definitions, and thresholds guide classification decisions.
- Provider
- Openai
- Context window
- 128,000 tokens
- Maximum output
- 16,000 tokens
- Released
- 29 Oct 2025
Nvidia Nemotron Nano 12B V2 VLnvidia/nemotron-nano-12b-v2-vl
The model is an auto-regressive vision language model that uses an optimized transformer architecture. The model enables multi-image reasoning and video understanding, along with strong document intelligence, visual Q&A and summarization capabilities.
- Provider
- Nvidia
- Context window
- 131,072 tokens
- Maximum output
- 131,072 tokens
- Released
- 28 Oct 2025
MiniMax M2minimax/minimax-m2
MiniMax-M2 redefines efficiency for agents.
- Provider
- Minimax
- Context window
- 205,000 tokens
- Maximum output
- 205,000 tokens
- Released
- 27 Oct 2025
Claude Haiku 4.5anthropic/claude-haiku-4.5
A description is not currently available for this model.
- Provider
- Anthropic
- Context window
- 200,000 tokens
- Maximum output
- 64,000 tokens
- Released
- 15 Oct 2025
Interfaze Betainterfaze/interfaze-beta
Interfaze is an AI model built on a new architecture that merges specialized DNN/CNN models with LLMs for developer tasks that require deterministic output and high consistency like OCR, scraping, classification, STT and more.
- Provider
- Interfaze
- Context window
- 1,000,000 tokens
- Maximum output
- 32,000 tokens
- Released
- 7 Oct 2025
GPT-5 proopenai/gpt-5-pro
GPT-5 pro uses more compute to think harder and provide consistently better answers. Since GPT-5 pro is designed to tackle tough problems, some requests may take several minutes to finish.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 272,000 tokens
- Released
- 6 Oct 2025
GLM 4.6zai/glm-4.6
As the latest iteration in the GLM series, GLM-4.6 achieves comprehensive enhancements across multiple domains, including real-world coding, long-context processing, reasoning, searching, writing, and agentic applications.
- Provider
- Zai
- Context window
- 200,000 tokens
- Maximum output
- 96,000 tokens
- Released
- 30 Sept 2025
Claude Sonnet 4.5anthropic/claude-sonnet-4.5
Claude Sonnet 4.5 is the newest model in the Sonnet series, offering improvements and updates over Sonnet 4.
- Provider
- Anthropic
- Context window
- 1,000,000 tokens
- Maximum output
- 64,000 tokens
- Released
- 29 Sept 2025
Qwen3 VL 235B A22B Thinkingalibaba/qwen3-235b-a22b-thinking
Qwen3 series VL models feature significantly enhanced multimodal reasoning capabilities, with a particular focus on optimizing the model for STEM and mathematical reasoning. Visual perception and recognition abilities have been comprehensively improved, and OCR capabilities have undergone a major upgrade.
- Provider
- Alibaba
- Context window
- 131,072 tokens
- Maximum output
- 32,768 tokens
- Released
- 23 Sept 2025
Qwen3 Maxalibaba/qwen3-max
The Qwen 3 series Max model has undergone specialized upgrades in agent programming and tool invocation compared to the preview version. The officially released model this time has achieved state-of-the-art (SOTA) performance in its field and is better suited to meet the demands of agents operating in more complex scenarios.
- Provider
- Alibaba
- Context window
- 262,144 tokens
- Maximum output
- 32,768 tokens
- Released
- 23 Sept 2025
Qwen3 VL 235B A22B Instructalibaba/qwen3-vl-235b-a22b-instruct
The Qwen3 series VL models has been comprehensively upgraded in areas such as visual coding and spatial perception. Its visual perception and recognition capabilities have significantly improved, supporting the understanding of ultra-long videos, and its OCR functionality has undergone a major enhancement.
- Provider
- Alibaba
- Context window
- 131,072 tokens
- Maximum output
- 129,024 tokens
- Released
- 23 Sept 2025
Qwen3 VL 235B A22B Instructalibaba/qwen3-vl-instruct
The Qwen3 series VL models has been comprehensively upgraded in areas such as visual coding and spatial perception. Its visual perception and recognition capabilities have significantly improved, supporting the understanding of ultra-long videos, and its OCR functionality has undergone a major enhancement.
- Provider
- Alibaba
- Context window
- 131,072 tokens
- Maximum output
- 129,024 tokens
- Released
- 23 Sept 2025
Qwen3 VL 235B A22B Thinkingalibaba/qwen3-vl-thinking
Qwen3 series VL models feature significantly enhanced multimodal reasoning capabilities, with a particular focus on optimizing the model for STEM and mathematical reasoning. Visual perception and recognition abilities have been comprehensively improved, and OCR capabilities have undergone a major upgrade.
- Provider
- Alibaba
- Context window
- 131,072 tokens
- Maximum output
- 32,768 tokens
- Released
- 23 Sept 2025
DeepSeek V3.1 Terminusdeepseek/deepseek-v3.1-terminus
DeepSeek-V3.1-Terminus delivers more stable & reliable outputs across benchmarks compared to the previous version and addresses user feedback (i.e. language consistency and agent upgrades).
- Provider
- Deepseek
- Context window
- 131,072 tokens
- Maximum output
- 65,536 tokens
- Released
- 22 Sept 2025
GPT-5-Codexopenai/gpt-5-codex
GPT-5-Codex is a version of GPT-5 optimized for agentic coding tasks in Codex or similar environments.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 15 Sept 2025
Qwen3 Next 80B A3B Instructalibaba/qwen3-next-80b-a3b-instruct
A new generation of open-source, non-thinking mode model powered by Qwen3. This version demonstrates superior Chinese text understanding, augmented logical reasoning, and enhanced capabilities in text generation tasks over the previous iteration (Qwen3-235B-A22B-Instruct-2507).
- Provider
- Alibaba
- Context window
- 131,072 tokens
- Maximum output
- 32,768 tokens
- Released
- 11 Sept 2025
Qwen3 Next 80B A3B Thinkingalibaba/qwen3-next-80b-a3b-thinking
A new generation of Qwen3-based open-source thinking mode models. This version offers improved instruction following and streamlined summary responses over the previous iteration (Qwen3-235B-A22B-Thinking-2507).
- Provider
- Alibaba
- Context window
- 131,072 tokens
- Maximum output
- 32,768 tokens
- Released
- 11 Sept 2025
Qwen3 Max Previewalibaba/qwen3-max-preview
Qwen3-Max-Preview shows substantial gains over the 2.5 series in overall capability, with significant enhancements in Chinese-English text understanding, complex instruction following, handling of subjective open-ended tasks, multilingual ability, and tool invocation; model knowledge hallucinations are reduced.
- Provider
- Alibaba
- Context window
- 262,144 tokens
- Maximum output
- 32,768 tokens
- Released
- 5 Sept 2025
Seed 1.6bytedance/seed-1.6
ByteDance's new multimodal deep-thinking model, supporting both text and visual inputs with enhanced reasoning capabilities.
- Provider
- Bytedance
- Context window
- 256,000 tokens
- Maximum output
- 32,000 tokens
- Released
- 1 Sept 2025
Bytedance Seed 1.8bytedance/seed-1.8
Bytedance Seed 1.8 features stronger multimodal understanding and agent capabilities. The model delivers superior performance across a wide range of complex real-world tasks, helping enterprises create greater value.
- Provider
- Bytedance
- Context window
- 256,000 tokens
- Maximum output
- 64,000 tokens
- Released
- 1 Sept 2025
Nano Banana Pro (Gemini 3 Pro Image)google/gemini-3-pro-image
Nano Banana Pro (Gemini 3 Pro Image) builds on Nano Banana's generation capabilities into a new era of studio-quality, functional design to help you create and edit high-fidelity, production-ready visuals with unparalleled precision and control. Improvements include enhanced world knowledge and reasoning, dynamic text and translation, and studio level controls.
- Provider
- Context window
- 65,536 tokens
- Maximum output
- 32,768 tokens
- Released
- 1 Sept 2025
Nano Banana (Gemini 2.5 Flash Image)google/gemini-2.5-flash-image
Upgraded for rapid creative workflows, it can generate interleaved text and images and supports conversational, multi‑turn image editing in natural language. It’s also locale‑aware, enabling culturally and linguistically appropriate image generation for audiences worldwide.
- Provider
- Context window
- 32,768 tokens
- Maximum output
- 65,536 tokens
- Released
- 26 Aug 2025
Muse Spark 1.3meta/muse-spark-1.3
Muse Spark 1.3 is Meta’s multimodal reasoning model for long-horizon agentic and coding workflows. With a 1M-token context window, reliable tool calling, higher first-attempt accuracy, and native understanding of video, images, and documents, it helps developers build capable coding agents and AI development workflows with fewer unnecessary turns and cleaner output.
- Provider
- Meta
- Context window
- 1,048,576 tokens
- Maximum output
- 1,048,576 tokens
- Released
- Not listed
Muse Spark 1.3 Contributormeta/muse-spark-1.3-contributor
It offers a 1M-token context window, dependable tool calling, and multimodal perception at $0.10 per million input tokens and $0.20 per million output tokens, with usage permitted to improve Meta’s products.
- Provider
- Meta
- Context window
- 1,048,576 tokens
- Maximum output
- 1,048,576 tokens
- Released
- Not listed
DeepSeek V3.1deepseek/deepseek-v3.1
DeepSeek-V3.1 is post-trained on the top of DeepSeek-V3.1-Base, which is built upon the original V3 base checkpoint through a two-phase long context extension approach, following the methodology outlined in the original DeepSeek-V3 report. We have expanded our dataset by collecting additional long documents and substantially extending both training phases. The 32K extension phase has been increased 10-fold to 630B tokens, while the 128K extension phase has been extended by 3.3x to 209B tokens. Additionally, DeepSeek-V3.1 is trained using the UE8M0 FP8 scale data format to ensure compatibility with microscaling data formats.
- Provider
- Deepseek
- Context window
- 163,840 tokens
- Maximum output
- 128,000 tokens
- Released
- 21 Aug 2025
Nvidia Nemotron Nano 9B V2nvidia/nemotron-nano-9b-v2
NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be controlled via a system prompt. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so.\
- Provider
- Nvidia
- Context window
- 131,072 tokens
- Maximum output
- 131,072 tokens
- Released
- 18 Aug 2025
GLM 4.5Vzai/glm-4.5v
Built on the GLM-4.5-Air base model, GLM-4.5V inherits proven techniques from GLM-4.1V-Thinking while achieving effective scaling through a powerful 106B-parameter MoE architecture.
- Provider
- Zai
- Context window
- 66,000 tokens
- Maximum output
- 16,000 tokens
- Released
- 11 Aug 2025
GPT-5openai/gpt-5
GPT-5 is OpenAI's flagship language model that excels at complex reasoning, broad real-world knowledge, code-intensive, and multi-step agentic tasks.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 7 Aug 2025
GPT-5 (Fast)openai/gpt-5-fast
GPT-5 is OpenAI's flagship language model that excels at complex reasoning, broad real-world knowledge, code-intensive, and multi-step agentic tasks.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 7 Aug 2025
GPT-5 miniopenai/gpt-5-mini
A description is not currently available for this model.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 7 Aug 2025
GPT-5 mini (Fast)openai/gpt-5-mini-fast
A description is not currently available for this model.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 7 Aug 2025
GPT-5 nanoopenai/gpt-5-nano
GPT-5 nano is a high throughput model that excels at simple instruction or classification tasks.
- Provider
- Openai
- Context window
- 400,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 7 Aug 2025
GPT OSS 120Bopenai/gpt-oss-120b
Extremely capable general-purpose LLM with strong, controllable reasoning capabilities
- Provider
- Openai
- Context window
- 131,072 tokens
- Maximum output
- 131,072 tokens
- Released
- 5 Aug 2025
GPT OSS 20Bopenai/gpt-oss-20b
A compact, open-weight language model optimized for low-latency and resource-constrained environments, including local and edge deployments.
- Provider
- Openai
- Context window
- 131,072 tokens
- Maximum output
- 8,192 tokens
- Released
- 5 Aug 2025
Qwen 3 Coder 30B A3B Instructalibaba/qwen3-coder-30b-a3b
A description is not currently available for this model.
- Provider
- Alibaba
- Context window
- 262,144 tokens
- Maximum output
- 8,192 tokens
- Released
- 31 Jul 2025
GLM 4.5zai/glm-4.5
GLM-4.5 and GLM-4.5-Air are our latest flagship models, purpose-built as foundational models for agent-oriented applications. Both leverage a Mixture-of-Experts (MoE) architecture. GLM-4.5 has a total parameter count of 355B with 32B active parameters per forward pass, while GLM-4.5-Air adopts a more streamlined design with 106B total parameters and 12B active parameters.
- Provider
- Zai
- Context window
- 128,000 tokens
- Maximum output
- 96,000 tokens
- Released
- 28 Jul 2025
GLM 4.5 Airzai/glm-4.5-air
GLM-4.5 and GLM-4.5-Air are our latest flagship models, purpose-built as foundational models for agent-oriented applications. Both leverage a Mixture-of-Experts (MoE) architecture. GLM-4.5 has a total parameter count of 355B with 32B active parameters per forward pass, while GLM-4.5-Air adopts a more streamlined design with 106B total parameters and 12B active parameters.
- Provider
- Zai
- Context window
- 128,000 tokens
- Maximum output
- 96,000 tokens
- Released
- 28 Jul 2025
Qwen3 Coder Plusalibaba/qwen3-coder-plus
Powered by Qwen3 this is a powerful Coding Agent that excels in tool calling and environment interaction to achieve autonomous programming. It combines outstanding coding proficiency with versatile general-purpose abilities.
- Provider
- Alibaba
- Context window
- 1,000,000 tokens
- Maximum output
- 65,536 tokens
- Released
- 23 Jul 2025
Qwen3 Coder 480B A35B Instructalibaba/qwen3-coder
Qwen3-Coder-480B-A35B-Instruct is a cutting-edge open coding model from Qwen, matching Claude Sonnet’s performance in agentic programming, browser automation, and core development tasks.
- Provider
- Alibaba
- Context window
- 262,144 tokens
- Maximum output
- 65,536 tokens
- Released
- 22 Jul 2025
Qwen3 Coder Nextalibaba/qwen3-coder-next
Qwen3-Coder-Next is an open-weight language model built specifically for coding, with strong performance on large-scale software engineering and agentic coding benchmarks. It uses a hybrid Mixture-of-Experts architecture to offer high capability at relatively modest active parameter counts, improving efficiency for real-world deployments. The model is trained on diverse code and natural language data so it can handle tasks like code generation, refactoring, debugging, repository-level reasoning, and technical explanation across multiple programming languages. It is also optimized for tool use and function calling, making it suitable as the core of coding agents that interact with shells, editors, issue trackers, and other developer tools.
- Provider
- Alibaba
- Context window
- 256,000 tokens
- Maximum output
- 256,000 tokens
- Released
- 22 Jul 2025
Kimi K2 Instructmoonshotai/kimi-k2
Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the Muon optimizer, Kimi K2 achieves exceptional performance across frontier knowledge, reasoning, and coding tasks while being meticulously optimized for agentic capabilities.
- Provider
- Moonshotai
- Context window
- 131,072 tokens
- Maximum output
- 131,072 tokens
- Released
- 11 Jul 2025
Morph V3 Fastmorph/morph-v3-fast
Morph offers a specialized AI model that applies code changes suggested by frontier models (like Claude or GPT-4o) to your existing code files FAST - 4500+ tokens/second. It acts as the final step in the AI coding workflow. Supports 16k input tokens and 16k output tokens.
- Provider
- Morph
- Context window
- 81,920 tokens
- Maximum output
- 16,384 tokens
- Released
- 7 Jul 2025
Morph V3 Largemorph/morph-v3-large
Morph offers a specialized AI model that applies code changes suggested by frontier models (like Claude or GPT-4o) to your existing code files FAST - 2500+ tokens/second. It acts as the final step in the AI coding workflow. Supports 16k input tokens and 16k output tokens.
- Provider
- Morph
- Context window
- 81,920 tokens
- Maximum output
- 16,384 tokens
- Released
- 7 Jul 2025
Gemini 2.5 Flash Litegoogle/gemini-2.5-flash-lite
Gemini 2.5 Flash-Lite is a balanced, low-latency model with configurable thinking budgets and tool connectivity (e.g., Google Search grounding and code execution). It supports multimodal input and offers a 1M-token context window.
- Provider
- Context window
- 1,048,576 tokens
- Maximum output
- 65,536 tokens
- Released
- 17 Jun 2025
o3 Proopenai/o3-pro
The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses more compute to think harder and provide consistently better answers.
- Provider
- Openai
- Context window
- 200,000 tokens
- Maximum output
- 100,000 tokens
- Released
- 10 Jun 2025
Claude Opus 4anthropic/claude-opus-4
Claude Opus 4 is Anthropic's most powerful model yet and the state-of-the-art coding model. It delivers sustained performance on long-running tasks that require focused effort and thousands of steps, significantly expanding what AI agents can solve. Claude Opus 4 is ideal for powering frontier agent products and features.
- Provider
- Anthropic
- Context window
- 200,000 tokens
- Maximum output
- 8,192 tokens
- Released
- 22 May 2025
Claude Sonnet 4anthropic/claude-sonnet-4
A description is not currently available for this model.
- Provider
- Anthropic
- Context window
- 1,000,000 tokens
- Maximum output
- 8,192 tokens
- Released
- 22 May 2025
GPT-4.1 miniopenai/gpt-4.1-mini
A description is not currently available for this model.
- Provider
- Openai
- Context window
- 1,047,576 tokens
- Maximum output
- 32,768 tokens
- Released
- 14 May 2025
GPT-4.1 mini (Fast)openai/gpt-4.1-mini-fast
A description is not currently available for this model.
- Provider
- Openai
- Context window
- 1,047,576 tokens
- Maximum output
- 32,768 tokens
- Released
- 14 May 2025
Mistral Medium 3.1mistral/mistral-medium
Mistral Medium 3 delivers frontier performance while being an order of magnitude less expensive.
- Provider
- Mistral
- Context window
- 128,000 tokens
- Maximum output
- 64,000 tokens
- Released
- 7 May 2025
Qwen3-14Balibaba/qwen-3-14b
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support
- Provider
- Alibaba
- Context window
- 40,960 tokens
- Maximum output
- 16,384 tokens
- Released
- 28 Apr 2025
Qwen3 235B A22Balibaba/qwen-3-235b
A description is not currently available for this model.
- Provider
- Alibaba
- Context window
- 262,144 tokens
- Maximum output
- 16,384 tokens
- Released
- 28 Apr 2025
Qwen3-30B-A3Balibaba/qwen-3-30b
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support
- Provider
- Alibaba
- Context window
- 40,960 tokens
- Maximum output
- 16,384 tokens
- Released
- 28 Apr 2025
Qwen 3 32Balibaba/qwen-3-32b
Qwen3-32B is a world-class model with comparable quality to DeepSeek R1 while outperforming GPT-4.1 and Claude Sonnet 3.7. It excels in code-gen, tool-calling, and advanced reasoning, making it an exceptional model for a wide range of production use cases.
- Provider
- Alibaba
- Context window
- 128,000 tokens
- Maximum output
- 8,192 tokens
- Released
- 28 Apr 2025
o3openai/o3
OpenAI's o3 is their most powerful reasoning model, setting new state-of-the-art benchmarks in coding, math, science, and visual perception. It excels at complex queries requiring multi-faceted analysis, with particular strength in analyzing images, charts, and graphics.
- Provider
- Openai
- Context window
- 200,000 tokens
- Maximum output
- 100,000 tokens
- Released
- 16 Apr 2025
o3 (Fast)openai/o3-fast
OpenAI's o3 is their most powerful reasoning model, setting new state-of-the-art benchmarks in coding, math, science, and visual perception. It excels at complex queries requiring multi-faceted analysis, with particular strength in analyzing images, charts, and graphics.
- Provider
- Openai
- Context window
- 200,000 tokens
- Maximum output
- 100,000 tokens
- Released
- 16 Apr 2025
o4-miniopenai/o4-mini
A description is not currently available for this model.
- Provider
- Openai
- Context window
- 200,000 tokens
- Maximum output
- 100,000 tokens
- Released
- 16 Apr 2025
o4-mini (Fast)openai/o4-mini-fast
A description is not currently available for this model.
- Provider
- Openai
- Context window
- 200,000 tokens
- Maximum output
- 100,000 tokens
- Released
- 16 Apr 2025
GPT-4.1openai/gpt-4.1
GPT 4.1 is OpenAI's flagship model for complex tasks. It is well suited for problem solving across domains.
- Provider
- Openai
- Context window
- 1,047,576 tokens
- Maximum output
- 32,768 tokens
- Released
- 14 Apr 2025
GPT-4.1 (Fast)openai/gpt-4.1-fast
GPT 4.1 is OpenAI's flagship model for complex tasks. It is well suited for problem solving across domains.
- Provider
- Openai
- Context window
- 1,047,576 tokens
- Maximum output
- 32,768 tokens
- Released
- 14 Apr 2025
GPT-4.1 nanoopenai/gpt-4.1-nano
A description is not currently available for this model.
- Provider
- Openai
- Context window
- 1,047,576 tokens
- Maximum output
- 32,768 tokens
- Released
- 14 Apr 2025
GPT-4.1 nano (Fast)openai/gpt-4.1-nano-fast
A description is not currently available for this model.
- Provider
- Openai
- Context window
- 1,047,576 tokens
- Maximum output
- 32,768 tokens
- Released
- 14 Apr 2025
Llama 4 Maverick 17B Instructmeta/llama-4-maverick
A description is not currently available for this model.
- Provider
- Meta
- Context window
- 128,000 tokens
- Maximum output
- 8,192 tokens
- Released
- 5 Apr 2025
Llama 4 Scout 17B Instructmeta/llama-4-scout
Llama 4 Scout is the best multimodal model in the world in its class and is more powerful than our Llama 3 models, while fitting in a single H100 GPU. Additionally, Llama 4 Scout supports an industry-leading context window of up to 10M tokens.
- Provider
- Meta
- Context window
- 128,000 tokens
- Maximum output
- 8,192 tokens
- Released
- 5 Apr 2025
Gemini 2.5 Flashgoogle/gemini-2.5-flash
Gemini 2.5 Flash is a thinking model that offers great, well-rounded capabilities.
- Provider
- Context window
- 1,000,000 tokens
- Maximum output
- 65,536 tokens
- Released
- 20 Mar 2025
Gemini 2.5 Progoogle/gemini-2.5-pro
Gemini 2.5 Pro is our most advanced reasoning Gemini model, capable of solving complex problems. Gemini 2.5 Pro can comprehend vast datasets and challenging problems from different information sources, including text, audio, images, video, and even entire code repositories.
- Provider
- Context window
- 1,048,576 tokens
- Maximum output
- 65,536 tokens
- Released
- 20 Mar 2025
Command Acohere/command-a
Command A is Cohere's most performant model to date, excelling at tool use, agents, retrieval augmented generation (RAG), and multilingual use cases. Command A has a context length of 256K, only requires two GPUs to run, and has 150% higher throughput compared to Command R+ 08-2024.
- Provider
- Cohere
- Context window
- 256,000 tokens
- Maximum output
- 8,000 tokens
- Released
- 13 Mar 2025
Mercury Coder Small Betainception/mercury-coder-small
Mercury Coder Small is ideal for code generation, debugging, and refactoring tasks with minimal latency.
- Provider
- Inception
- Context window
- 32,000 tokens
- Maximum output
- 16,384 tokens
- Released
- 26 Feb 2025
Sonarperplexity/sonar
Perplexity's lightweight offering with search grounding, quicker and cheaper than Sonar Pro.
- Provider
- Perplexity
- Context window
- 127,000 tokens
- Maximum output
- 8,000 tokens
- Released
- 19 Feb 2025
Sonar Properplexity/sonar-pro
Perplexity's premier offering with search grounding, supporting advanced queries and follow-ups.
- Provider
- Perplexity
- Context window
- 200,000 tokens
- Maximum output
- 8,000 tokens
- Released
- 19 Feb 2025
Sonar Reasoning Properplexity/sonar-reasoning-pro
A premium reasoning-focused model that outputs Chain of Thought (CoT) in responses, providing comprehensive explanations with enhanced search capabilities and multiple search queries per request.
- Provider
- Perplexity
- Context window
- 127,000 tokens
- Maximum output
- 8,000 tokens
- Released
- 19 Feb 2025
o3-miniopenai/o3-mini
A description is not currently available for this model.
- Provider
- Openai
- Context window
- 200,000 tokens
- Maximum output
- 100,000 tokens
- Released
- 31 Jan 2025
DeepSeek-R1deepseek/deepseek-r1
DeepSeek-R1 provides customers a state-of-the-art reasoning model, optimized for general reasoning tasks, math, science, and code generation.
- Provider
- Deepseek
- Context window
- 128,000 tokens
- Maximum output
- 8,192 tokens
- Released
- 20 Jan 2025
Llama 3.3 70B Instructmeta/llama-3.3-70b
Where performance meets efficiency. This model supports high-performance conversational AI designed for content creation, enterprise applications, and research, offering advanced language understanding capabilities, including text summarization, classification, sentiment analysis, and code generation.
- Provider
- Meta
- Context window
- 128,000 tokens
- Maximum output
- 8,192 tokens
- Released
- 6 Dec 2024
o1openai/o1
o1 is OpenAI's flagship reasoning model, designed for complex problems that require deep thinking. It provides strong reasoning capabilities with improved accuracy for complex multi-step tasks.
- Provider
- Openai
- Context window
- 200,000 tokens
- Maximum output
- 100,000 tokens
- Released
- 5 Dec 2024
Nova Liteamazon/nova-lite
A description is not currently available for this model.
- Provider
- Amazon
- Context window
- 300,000 tokens
- Maximum output
- 8,192 tokens
- Released
- 3 Dec 2024
Nova Microamazon/nova-micro
A description is not currently available for this model.
- Provider
- Amazon
- Context window
- 128,000 tokens
- Maximum output
- 8,192 tokens
- Released
- 3 Dec 2024
Nova Proamazon/nova-pro
A description is not currently available for this model.
- Provider
- Amazon
- Context window
- 300,000 tokens
- Maximum output
- 8,192 tokens
- Released
- 3 Dec 2024
Ministral 3Bmistral/ministral-3b
A compact, efficient model for on-device tasks like smart assistants and local analytics, offering low-latency performance.
- Provider
- Mistral
- Context window
- 128,000 tokens
- Maximum output
- 4,000 tokens
- Released
- 16 Oct 2024
Ministral 8Bmistral/ministral-8b
A more powerful model with faster, memory-efficient inference, ideal for complex workflows and demanding edge applications.
- Provider
- Mistral
- Context window
- 128,000 tokens
- Maximum output
- 4,000 tokens
- Released
- 16 Oct 2024
Mistral Smallmistral/mistral-small
Mistral Small is the ideal choice for simple tasks that one can do in bulk - like Classification, Customer Support, or Text Generation.
- Provider
- Mistral
- Context window
- 32,000 tokens
- Maximum output
- 4,000 tokens
- Released
- 17 Sept 2024
Pixtral 12B 2409mistral/pixtral-12b
A 12B model with image understanding capabilities in addition to text.
- Provider
- Mistral
- Context window
- 128,000 tokens
- Maximum output
- 4,000 tokens
- Released
- 17 Sept 2024
Llama 3.1 70B Instructmeta/llama-3.1-70b
An update to Meta Llama 3 70B Instruct that includes an expanded 128K context length, multilinguality and improved reasoning capabilities.
- Provider
- Meta
- Context window
- 128,000 tokens
- Maximum output
- 8,192 tokens
- Released
- 23 Jul 2024
Llama 3.1 8B Instructmeta/llama-3.1-8b
An update to Meta Llama 3 8B Instruct that includes an expanded 128K context length, multilinguality and improved reasoning capabilities.
- Provider
- Meta
- Context window
- 128,000 tokens
- Maximum output
- 8,192 tokens
- Released
- 23 Jul 2024
Mistral Nemo 12Bmistral/mistral-nemo
A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese, Korean, Arabic, and Hindi. It supports function calling and is released under the Apache 2.0 license.
- Provider
- Mistral
- Context window
- 128,000 tokens
- Maximum output
- 128,000 tokens
- Released
- 18 Jul 2024
GPT-4o miniopenai/gpt-4o-mini
It is multi-modal (accepting text or image inputs and outputting text) and has higher intelligence than gpt-3.5-turbo but is just as fast.
- Provider
- Openai
- Context window
- 128,000 tokens
- Maximum output
- 16,384 tokens
- Released
- 18 Jul 2024
GPT-4o mini (Fast)openai/gpt-4o-mini-fast
It is multi-modal (accepting text or image inputs and outputting text) and has higher intelligence than gpt-3.5-turbo but is just as fast.
- Provider
- Openai
- Context window
- 128,000 tokens
- Maximum output
- 16,384 tokens
- Released
- 18 Jul 2024
Mistral Codestralmistral/codestral
Mistral's cutting-edge language model for coding released end of July 2025, Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation.
- Provider
- Mistral
- Context window
- 128,000 tokens
- Maximum output
- 4,000 tokens
- Released
- 29 May 2024
GPT-4oopenai/gpt-4o
GPT-4o from OpenAI has broad general knowledge and domain expertise allowing it to follow complex instructions in natural language and solve difficult problems accurately. It matches GPT-4 Turbo performance with a faster and cheaper API.
- Provider
- Openai
- Context window
- 128,000 tokens
- Maximum output
- 16,384 tokens
- Released
- 13 May 2024
GPT-4o (Fast)openai/gpt-4o-fast
GPT-4o from OpenAI has broad general knowledge and domain expertise allowing it to follow complex instructions in natural language and solve difficult problems accurately. It matches GPT-4 Turbo performance with a faster and cheaper API.
- Provider
- Openai
- Context window
- 128,000 tokens
- Maximum output
- 16,384 tokens
- Released
- 13 May 2024
GPT-4 Turboopenai/gpt-4-turbo
gpt-4-turbo from OpenAI has broad general knowledge and domain expertise allowing it to follow complex instructions in natural language and solve difficult problems accurately. It has a knowledge cutoff of April 2023 and a 128,000 token context window.
- Provider
- Openai
- Context window
- 128,000 tokens
- Maximum output
- 4,096 tokens
- Released
- 9 Apr 2024
Claude 3 Haikuanthropic/claude-3-haiku
Claude 3 Haiku is Anthropic's fastest, most compact model for near-instant responsiveness. It answers simple queries and requests with speed. Customers will be able to build seamless AI experiences that mimic human interactions. Claude 3 Haiku can process images and return text outputs, and features a 200K context window.
- Provider
- Anthropic
- Context window
- 200,000 tokens
- Maximum output
- 4,096 tokens
- Released
- 13 Mar 2024
GPT-3.5 Turboopenai/gpt-3.5-turbo
A description is not currently available for this model.
- Provider
- Openai
- Context window
- 16,385 tokens
- Maximum output
- 4,096 tokens
- Released
- 1 Mar 2023

