Browse all AI Gateway models
Every model available on Vercel AI Gateway, with API access, pricing, and a playground. 363 models · Page 4 of 7.
Search and filter all models →- Kling AIKling v2.6 Image-to-VideoKling 2.6 introduces a groundbreaking "Native Audio" capability, enabling the generation of complete videos in a single go, including natural voice, action sound effects, and environmental ambient sounds, providing an immersive "what you see if what you hear" experience.
- Kling AIKling v2.6 Motion ControlKling 2.6 introduces a groundbreaking "Native Audio" capability, enabling the generation of complete videos in a single go, including natural voice, action sound effects, and environmental ambient sounds, providing an immersive "what you see if what you hear" experience.
- Kling AIKling v2.6 Text-to-VideoKling 2.6 introduces a groundbreaking "Native Audio" capability, enabling the generation of complete videos in a single go, including natural voice, action sound effects, and environmental ambient sounds, providing an immersive "what you see if what you hear" experience.
- Kling AIKling v3.0 Image-to-VideoBuild upon an All-in-One product framework, the Kling 3.0 model series supports full multimodal input and output spanning text, images, audio, and video, bringing the understanding, generation, and editing of video together in one streamlined AI workflow. The models integrate multiple tasks, including text-to-video, image-to-video, reference-to-video, and in-video editing, into a single, native multimodal architecture, enabling the models to follow complex narrative logic, deliver precise shot control, and maintain strong prompt adherence.
- Kling AIKling v3.0 Motion ControlKling 3.0 delivers a major leap in character fidelity for motion-driven generation, with stable facial features across multi-angle and long-duration motion, accurate complex emotions from multi-image face references, identity preservation through partial occlusions (hats, hands, fans), and steady clarity as the camera zooms, pans, or tracks.
- Kling AIKling v3.0 Text-to-VideoBuild upon an All-in-One product framework, the Kling 3.0 model series supports full multimodal input and output spanning text, images, audio, and video, bringing the understanding, generation, and editing of video together in one streamlined AI workflow. The models integrate multiple tasks, including text-to-video, image-to-video, reference-to-video, and in-video editing, into a single, native multimodal architecture, enabling the models to follow complex narrative logic, deliver precise shot control, and maintain strong prompt adherence.
- PoolsideLaguna S 2.1Laguna S 2.1 is Poolside's new open-weight model for agentic coding and long-horizon work.
- PoolsideLaguna S 2.1 FreeLaguna S 2.1 is Poolside's new open-weight model for agentic coding and long-horizon work.
- InclusionaiLing 3.0 FlashLing 3.0 Flash is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.
- InclusionaiLing 3.0 Flash FinLing 3.0 Flash Fin is InclusionAI’s finance-enhanced MoE language model, combining 124 billion total parameters with approximately 5.1 billion active parameters for efficient financial reasoning. Its 256K context window, function calling, and support for complex, multi-step investment workflows make it ideal for financial research, analysis, long-horizon planning, and execution, while retaining strong capabilities in coding and mathematics.
- InclusionaiLing 3.0 Flash SanteLing-3.0-Flash-Sante is inclusionAI’s language model specialized for health and medicine, built on a Mixture-of-Experts architecture with 124 billion total parameters and approximately 5.1 billion active per token. With a 256K context window and function calling, it supports medical knowledge reasoning, evidence-based retrieval, and complex medical workflows while retaining general reasoning, coding, and agentic capabilities.
- InclusionaiLing 3.0 Flash VLLing 3.0 Flash VL builds on Ling 3.0 Flash with stronger language capabilities, native visual perception, and visual agent capabilities. It supports text, image, and video inputs with text output, reasoning, and function calling.
- MetaLlama 3.1 70B InstructAn update to Meta Llama 3 70B Instruct that includes an expanded 128K context length, multilinguality and improved reasoning capabilities.
- MetaLlama 3.1 8B InstructAn update to Meta Llama 3 8B Instruct that includes an expanded 128K context length, multilinguality and improved reasoning capabilities.
- MetaLlama 3.3 70B InstructWhere performance meets efficiency. This model supports high-performance conversational AI designed for content creation, enterprise applications, and research, offering advanced language understanding capabilities, including text summarization, classification, sentiment analysis, and code generation.
- MetaLlama 4 Maverick 17B 128E Instruct FP8The Llama 4 collection of models are natively multimodal AI models that enable text and multimodal experiences. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding. Llama 4 Maverick, a 17 billion parameter model with 128 experts. Served by DeepInfra.
- MetaLlama 4 Scout 17B 16E InstructThe Llama 4 collection of models are natively multimodal AI models that enable text and multimodal experiences. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding. Llama 4 Scout, a 17 billion parameter model with 16 experts. Served by DeepInfra.
- InceptionMercury 2A diffusion-based reasoning LLM that generates text via parallel refinement (not token-by-token), delivering real-time latency with ~1k tokens/sec plus 128K context and built-in tool/JSON support.
- InceptionMercury 2.5Mercury 2.5 is Inception’s diffusion-based reasoning model for chat, agents, and structured workflows, with tool calling, structured outputs, and a 260K context window.
- InceptionMercury Coder Small BetaMercury Coder Small is ideal for code generation, debugging, and refactoring tasks with minimal latency.
- XiaomiMiMo M2.5A native full-modal model supporting text, image, video, and audio understanding, with powerful Agent capabilities.
- XiaomiMiMo V2.5 ProMiMo V2.5 Pro delivers significant improvements over its predecessor, MiMo-V2-Pro, in general agentic capabilities, complex software engineering, and long-horizon tasks. MiMo-V2.5-Pro is a 1.02T-parameter Mixture-of-Experts model with 42B active parameters, built on a hybrid-attention architecture with a 1M-token context window.
- XiaomiMiMo V2.6 FlashMiMo V2.6 Flash is Xiaomi's efficient multimodal reasoning model for coding, automation, and everyday agent workflows. It accepts text, images, audio, and video within a 1M-token context window, with up to 128K tokens of output. Deep thinking, tool calling, JSON mode, and prompt caching make it suitable for applications that need multimodal understanding at a lower token cost.
- XiaomiMiMo V2.6 ProMiMo V2.6 Pro is Xiaomi's flagship multimodal reasoning model for complex software engineering, long-running agent tasks, and professional workflows. It supports text, image, audio, and video inputs with a 1M-token context window and up to 128K tokens of output. Deep thinking, tool calling, JSON mode, and prompt caching support applications that combine large inputs with multi-step reasoning.
- XiaomiMiMo V2.6 Pro UltraSpeedMiMo V2.6 Pro UltraSpeed is Xiaomi's accelerated inference offering for MiMo V2.6 Pro, designed for interactive agents and workflows where response latency matters. It combines multimodal understanding of text, images, audio, and video with a 1M-token context window and up to 128K tokens of output. It supports deep thinking, tool calling, JSON mode, and prompt caching.
- MiniMaxMiniMax H3H3 is a next-generation open-weights, general-purpose multimodal video model. Rather than being limited to specialized tasks such as generating, editing, or referencing, H3 understands multimodal contexts that bring together text, images, video, and audio. This enables it to interpret creative intent in a unified way and deliver more natural, coherent generation and expression.
- MiniMaxMiniMax H3 MaxMiniMax H3 Max is a next-generation general-purpose multimodal video model ranking #1 for overall quality, prompt understanding, and aesthetics in all evaluations, while generating a 5-second video in under 3 seconds.
- MiniMaxMiniMax M2MiniMax-M2 redefines efficiency for agents. It is a compact, fast, and cost-effective MoE model (230 billion total parameters with 10 billion active parameters) built for elite performance in coding and agentic tasks, all while maintaining powerful general intelligence.
- MiniMaxMiniMax M2.1MiniMax 2.1 is MiniMax's latest model, optimized specifically for robustness in coding, tool use, instruction following, and long-horizon planning.
- MiniMaxMiniMax M2.1 LightningMiniMax-M2.1-lightning is a faster version of MiniMax-M2.1, offering the same performance but with significantly higher throughput (output speed ~100 TPS, MiniMax-M2 output speed ~60 TPS).
- MiniMaxMiniMax M2.5MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. It is capable of handling the entire development process of various complex systems. It covers full-stack projects across multiple platforms including Web, Android, iOS, Windows, and Mac, encompassing server-side APIs, functional logic, and databases.
- MiniMaxMiniMax M2.5 High SpeedM2.5 highspeed: Same performance, faster and more agile (output speed approximately 100 tps)
- MiniMaxMiniMax M2.7M2.7 delivers outstanding performance in real-world software engineering, including end-to-end full project delivery, log analysis and bug troubleshooting, code security, machine learning, and more.
- MiniMaxMiniMax M2.7 High SpeedM2.7 Highspeed: Same performance, faster and more agile (output speed approximately 100 tps)
- MiniMaxMiniMax M3MiniMax-M3 is a frontier-class foundation model that unites the three capabilities defining today's frontier: a 1M-token context window, frontier coding and agentic performance, and native multimodality — the first open-weight model to deliver all three in a single system.
- MistralMinistral 14BMinistral 3 14B is the largest model in the Ministral 3 family, offering state-of-the-art capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. Optimized for local deployment, it delivers high performance across diverse hardware, including local setups.
- MistralMinistral 3BA compact, efficient model for on-device tasks like smart assistants and local analytics, offering low-latency performance.
- MistralMinistral 8BA more powerful model with faster, memory-efficient inference, ideal for complex workflows and demanding edge applications.
- MistralMistral CodestralMistral's cutting-edge language model for coding released end of July 2025, Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation.
- MistralMistral EmbedGeneral-purpose text embedding model for semantic search, similarity, clustering, and RAG workflows.
- MistralMistral Large 3Mistral Large 3 2512 is Mistral’s most capable model to date. It has a sparse mixture-of-experts architecture with 41B active parameters (675B total).
- MistralMistral Medium LatestMistral's frontier-class multimodal model optimized for agentic and coding use cases.
- MistralMistral NemoA 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese, Korean, Arabic, and Hindi. It supports function calling and is released under the Apache 2.0 license.
- MistralMistral SmallMistral Small currently runs Mistral Small 4, a multimodal model combining instruction following, reasoning, and coding with a 262,144-token context window.
- MorphMorph V3 FastMorph offers a specialized AI model that applies code changes suggested by frontier models (like Claude or GPT-4o) to your existing code files FAST - 4500+ tokens/second. It acts as the final step in the AI coding workflow. Supports 16k input tokens and 16k output tokens.
- MorphMorph V3 LargeMorph offers a specialized AI model that applies code changes suggested by frontier models (like Claude or GPT-4o) to your existing code files FAST - 2500+ tokens/second. It acts as the final step in the AI coding workflow. Supports 16k input tokens and 16k output tokens.
- MetaMuse Glimmer 30B
- MetaMuse Image 1.0Muse Image is the first image generation model from Meta Superintelligence Labs, it uses advanced reasoning to understand complex prompts, seamlessly blending multiple photos into high-quality creations you can download and share anywhere.
- MetaMuse Spark 1.1Muse Spark 1.1 is strongest at agentic performance, tool use, and computer use. It does well on long-running tasks with 1M token context window, can delegate execution to sub-agents running in parallel, and is trained to use computer interfaces on desktop, mobile, or browser.
- MetaMuse Spark 1.2A coding-optimized model purpose-built for agentic workflows. Improvements to code generation, debugging, and codebase understanding — with a 1M context window that handles your entire project in one session.
- MetaMuse Spark 1.2 ContributorA coding-optimized model with pricing designed for builders. Same model, same capabilities — up to 95% less than Standard. Your inputs and outputs are used to train and improve Meta's AI models.
- MetaMuse Spark 1.3Muse Spark 1.3 is Meta’s multimodal reasoning model for long-horizon agentic and coding workflows. With a 1M-token context window, reliable tool calling, higher first-attempt accuracy, and native understanding of video, images, and documents, it helps developers build capable coding agents and AI development workflows with fewer unnecessary turns and cleaner output.
- MetaMuse Spark 1.3 ContributorMuse Spark 1.3 Contributor is Meta’s cost-optimized API option for agentic coding workflows. It offers a 1M-token context window, dependable tool calling, and multimodal perception at $0.10 per million input tokens and $0.20 per million output tokens, with usage permitted to improve Meta’s products.
- GoogleNano Banana (Gemini 2.5 Flash Image)Nano Banana (Gemini 2.5 Flash Image) is Google's first fully hybrid reasoning model, letting developers turn thinking on or off and set thinking budgets to balance quality, cost, and latency. Upgraded for rapid creative workflows, it can generate interleaved text and images and supports conversational, multi‑turn image editing in natural language. It’s also locale‑aware, enabling culturally and linguistically appropriate image generation for audiences worldwide.
- GoogleNano Banana Pro (Gemini 3 Pro Image)Nano Banana Pro (Gemini 3 Pro Image) builds on Nano Banana's generation capabilities into a new era of studio-quality, functional design to help you create and edit high-fidelity, production-ready visuals with unparalleled precision and control. Improvements include enhanced world knowledge and reasoning, dynamic text and translation, and studio level controls.
- NVIDIANemotron 3 Nano 30B A3BNVIDIA Nemotron 3 Nano is an open reasoning model optimized for fast, cost-efficient inference. Built with a hybrid MoE and Mamba architecture and trained on NVIDIA-curated synthetic reasoning data, it delivers strong multi-step reasoning with stable latency and predictable performance for agentic and production workloads.
- NVIDIANemotron 3 UltraA 550B parameter (55B active) open reasoning model from NVIDIA, built for long-running agent workflows. It uses a hybrid Mamba-Transformer MoE architecture and supports a 1M token context window.
- NVIDIANemotron 3.5 Lightning 30BNVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 is a large language model (LLM) trained by NVIDIA. The model employs a hybrid Mixture-of-Experts architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. The Lightning 3.5 model is released alongside a number of speculative decoding methods for faster text generation. The model has 3B active parameters and 30B parameters in total.
- AmazonNova 2 LiteNova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text.
- AmazonNova LiteA very low cost multimodal model that is lightning fast for processing image, video, and text inputs.