Browse all AI Gateway models
Every model available on Vercel AI Gateway, with API access, pricing, and a playground. 363 models · Page 3 of 7.
Search and filter all models →- OpenAIGPT-3.5 TurboOpenAI's most capable and cost effective model in the GPT-3.5 family optimized for chat purposes, but also works well for traditional completions tasks.
- OpenAIGPT-4 Turbogpt-4-turbo from OpenAI has broad general knowledge and domain expertise allowing it to follow complex instructions in natural language and solve difficult problems accurately. It has a knowledge cutoff of April 2023 and a 128,000 token context window.
- OpenAIGPT-4.1GPT 4.1 is OpenAI's flagship model for complex tasks. It is well suited for problem solving across domains.
- OpenAIGPT-4.1 miniGPT 4.1 mini provides a balance between intelligence, speed, and cost that makes it an attractive model for many use cases.
- OpenAIGPT-4.1 nanoGPT-4.1 nano is the fastest, most cost-effective GPT 4.1 model.
- OpenAIGPT-4oGPT-4o from OpenAI has broad general knowledge and domain expertise allowing it to follow complex instructions in natural language and solve difficult problems accurately. It matches GPT-4 Turbo performance with a faster and cheaper API.
- OpenAIGPT-4o miniGPT-4o mini from OpenAI is their most advanced and cost-efficient small model. It is multi-modal (accepting text or image inputs and outputting text) and has higher intelligence than gpt-3.5-turbo but is just as fast.
- OpenAIGPT-4o mini TranscribeGPT-4o mini Transcribe is a speech-to-text model that uses GPT-4o mini to transcribe audio. It offers improvements to word error rate and better language recognition and accuracy compared to original Whisper models. Use it for more accurate transcripts.
- OpenAIGPT-4o TranscribeGPT-4o Transcribe is a speech-to-text model that uses GPT-4o to transcribe audio. It offers improvements to word error rate and better language recognition and accuracy compared to original Whisper models. Use it for more accurate transcripts.
- OpenAIGPT-5GPT-5 is OpenAI's flagship language model that excels at complex reasoning, broad real-world knowledge, code-intensive, and multi-step agentic tasks.
- OpenAIGPT-5 miniGPT-5 mini is a cost optimized model that excels at reasoning/chat tasks. It offers an optimal balance between speed, cost, and capability.
- OpenAIGPT-5 nanoGPT-5 nano is a high throughput model that excels at simple instruction or classification tasks.
- OpenAIGPT-5 proGPT-5 pro uses more compute to think harder and provide consistently better answers. Since GPT-5 pro is designed to tackle tough problems, some requests may take several minutes to finish.
- OpenAIGPT-5-CodexGPT-5-Codex is a version of GPT-5 optimized for agentic coding tasks in Codex or similar environments.
- OpenAIGPT-5.1-CodexGPT-5.1-Codex is a version of GPT-5.1 optimized for agentic coding tasks in Codex or similar environments.
- OpenAIGPT-6 AstraGPT-6 Astra is OpenAI's most capable model for complex reasoning, coding, computer use, research, and document creation.
- OpenAIGPT-6 LunaGPT-6 Luna is OpenAI's efficient reasoning model for focused, high-volume tasks. It accepts text and images, generates text, and supports a 1,050,000-token context window with up to 128,000 output tokens. Configurable reasoning effort, structured outputs, prompt caching, and Responses API tools support cost-sensitive applications, coding tasks, and automated workflows.
- OpenAIGPT-6 SolGPT-6 Sol is OpenAI's reasoning model for complex coding and agentic workflows. It accepts text and images, generates text, and supports a 1,050,000-token context window with up to 128,000 output tokens. Configurable reasoning effort, structured outputs, prompt caching, and Responses API tools make it suited to software development and multi-step automation.
- OpenAIGPT-Live 1Our premier model for natural, expressive voice conversations with smooth interruption handling.
- OpenAIGPT-Realtime miniGPT-Realtime mini is capable of responding to audio and text inputs in realtime over WebRTC, WebSocket, or SIP connections.
- OpenAIGPT-Realtime-1.5GPT-Realtime-1.5 is our flagship audio model for voice agents and customer support.
- OpenAIgpt-realtime-2GPT Realtime 2 is our most capable realtime voice model. It supports speech-to-speech interactions with configurable reasoning effort, stronger instruction following, and more reliable tool use for complex voice-agent workflows.
- OpenAIgpt-realtime-2.1GPT-Realtime-2.1 updates GPT-Realtime-2 with improved alphanumeric recognition, silence and noise handling, and interruption behavior. It supports speech-to-speech interactions with configurable reasoning effort, instruction following, and tool use for complex voice-agent workflows.
- OpenAIgpt-realtime-whisperGPT Realtime Whisper is a streaming speech-to-text model for applications that need low-latency transcript deltas from live audio. It is designed for realtime use cases where developers need to tune latency and accuracy. GPT Realtime Whisper is priced by audio duration rather than text tokens.
- SpaceXAIGrok 4.1 Fast Non-Reasoning
- SpaceXAIGrok 4.1 Fast Reasoning
- SpaceXAIGrok 4.20 Beta Non-ReasoningGrok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering consistently precise and truthful responses.
- SpaceXAIGrok 4.20 Beta ReasoningGrok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering consistently precise and truthful responses.
- SpaceXAIGrok 4.20 Multi Agent BetaMultiple agents collaborate in parallel to perform deep research tasks.
- SpaceXAIGrok 4.20 Multi-AgentMultiple agents collaborate in parallel to perform deep research tasks.
- SpaceXAIGrok 4.20 Non-ReasoningGrok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherence, delivering consistently precise and truthful responses.
- SpaceXAIGrok 4.20 ReasoningGrok 4.20 Beta is the newest flagship model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherence, delivering consistently precise and truthful responses.
- SpaceXAIGrok 4.3Grok 4.3 is a new model matching the scale of Grok 4.20 with an improved architecture and a December 2025 knowledge cutoff.
- SpaceXAIGrok 4.5SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
- SpaceXAIGrok 4.6Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work. It stays with complex tasks across many steps, whether researching a topic, analyzing information, working across a codebase, or turning an idea into a polished application or work artifact.
- SpaceXAIGrok 4.7Grok 4.7 is SpaceXAI’s advanced AI model for coding and professional knowledge work, built to tackle complex, multi-hour tasks with improved self-verification and long-context handling. It strengthens software engineering, document creation, and presentation workflows while maintaining Grok 4.6’s speed and pricing.
- SpaceXAIGrok Build 0.1xAI's fast coding model trained specifically for agentic coding.
- SpaceXAIGrok ImagineState-of-the-art video generation across quality, cost, and latency. Grok Imagine is x.AI's most powerful video-audio generative model yet. Bring an image to life, start from a simple text prompt, or even refine a complex cinematic sequence.
- SpaceXAIGrok Imagine ImageGenerate high-quality images from text prompts with xAI's imagine API.
- SpaceXAIGrok Imagine Image 2.0
- SpaceXAIGrok Imagine Video 1.5
- SpaceXAIGrok STTTranscribe audio to text in 25 languages with batch and streaming modes.
- SpaceXAIGrok TTSGenerate speech with 5 expressive voices, speech tags, and telephony codecs.
- SpaceXAIGrok Voice Think Fast 1.0Build real-time voice applications powered by Grok. Stream audio and text bidirectionally via WebSocket for voice assistants, phone agents, and interactive voice systems.
- SpaceXAIGrok Voice Think Fast 2.0Grok Voice Think Fast 2.0, our next-generation voice model with improved intelligence, transcription accuracy, and conversational capabilities.
- Tencent CloudHy3
- Thinking MachinesInklingInkling is a multimodal MoE model (975B total, 41B active, 256k context) reasoning over text, image, and audio inputs.
- Thinking MachinesInkling SmallInkling-Small is a lighter-weight model with 12B active parameters, trained with a similar recipe, to Inkling that achieves strong performance with even lower cost and latency.
- InterfazeInterfaze BetaInterfaze is an AI model built on a new architecture that merges specialized DNN/CNN models with LLMs for developer tasks that require deterministic output and high consistency like OCR, scraping, classification, STT and more.
- TypeSafe AIJevJev is TypeSafe AI’s System One evaluation model for fast, structured decisions in software. It evaluates shared state against typed questions and returns choices, scores, and boolean probabilities, supporting classification, routing, rubric-based assessment, and automated verification. Multiple questions can be evaluated in parallel within a single request.
- Moonshot AIKimi K2 InstructKimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the Muon optimizer, Kimi K2 achieves exceptional performance across frontier knowledge, reasoning, and coding tasks while being meticulously optimized for agentic capabilities.
- Moonshot AIKimi K2 ThinkingKimi K2 Thinking is an advanced open-source thinking model by Moonshot AI. It can execute up to 200 – 300 sequential tool calls without human interference, reasoning coherently across hundreds of steps to solve complex problems. Built as a thinking agent, it reasons step by step while using tools, achieving state-of-the-art performance on Humanity's Last Exam (HLE), BrowseComp, and other benchmarks, with major gains in reasoning, agentic search, coding, writing, and general capabilities.
- Moonshot AIKimi K2.5kimi-k2.5 is Kimi's most versatile model to date, featuring a native multimodal architecture that supports both visual and text input, thinking and non-thinking modes, and dialogue and agent tasks.
- Moonshot AIKimi K2.6Kimi K2.6 demonstrates particularly strong performance in long-horizon coding tasks and produces professional-grade design with code and vision.
- Moonshot AIKimi K2.7 CodeKimi-K2.7-Code is a coding model from Moonshot AI. It has improved coding & agent performance over K2.6, more reasoning efficiency with less overthinking, and improved instruction following for long-horizon coding.
- Moonshot AIKimi K2.7 Code High SpeedKimi K2.7 Code HighSpeed is the high-speed version of Kimi K2.7 Code, the same model as Kimi K2.7 Code, but with an output speed of approximately 180 Tokens/s and up to 260 Tokens/s in short context scenarios, delivering a more extreme coding experience.
- Moonshot AIKimi K3Kimi’s flagship model for long-horizon coding and end-to-end knowledge work, with a 1M-token context window.
- Moonshot AIKimi K3 FastFast version of Kimi’s flagship model for long-horizon coding and end-to-end knowledge work, with a 1M-token context window.
- Kling AIKling v2.5 Turbo Image-to-VideoKling 2.5 Turbo is a major update to the AI video generation model focused on significantly improving speed, video quality, temporal stability, and creative control for creators, making professional-grade AI-generated video faster, more coherent, and easier to direct from text prompts.
- Kling AIKling v2.5 Turbo Text-to-VideoKling 2.5 Turbo is a major update to the AI video generation model focused on significantly improving speed, video quality, temporal stability, and creative control for creators, making professional-grade AI-generated video faster, more coherent, and easier to direct from text prompts.