Full-modal models × Harness upgrade, unlock new Agent experiences
OpenAI · Text
GPT-5.6 Terra, a mid-tier model in the 5.6 series, with a 272K context window and 128K output. Well-rounded performance for general production tasks.
OpenAI · Text
GPT-5.6 Sol, OpenAI's flagship model in the 5.6 series, with a 272K context window and 128K output. Top-tier reasoning, multimodal understanding, and coding capabilities.
OpenAI · Text
GPT-5.6 Luna, a fast and cost-efficient model in the 5.6 series, with a 272K context window and 128K output. Ideal for high-concurrency and cost-sensitive scenarios.
OpenAI · Text
OpenAI GPT-5.5 Pro, the high-compute pro tier of GPT-5.5 for the most demanding frontier reasoning.
OpenAI · Text
OpenAI's latest flagship, GPT-5.5, with a 272K context window and 128K output. Leading across multimodal understanding, reasoning, and coding.
OpenAI · Text
OpenAI GPT-5.4 Mini, a compact variant of GPT-5.4 balancing cost and capability for everyday tasks.
OpenAI · Text
GPT-5.4, with a 272K context window and 128K output. The predecessor to GPT-5.5, offering well-rounded performance for general production tasks.
OpenAI · Text
OpenAI GPT-5.4 Pro, the high-compute pro tier of GPT-5.4 for demanding reasoning workloads.
OpenAI · Text
OpenAI GPT-5.4 Nano, the smallest variant of GPT-5.4 for low-cost, high-volume lightweight tasks.
OpenAI · Text
GPT-5.3 Instant, with a 272K context window and 128K output. An ultra-fast version that balances quality and speed, ideal for cost-sensitive scenarios.
Zhipu · Text
Zhipu's latest flagship GLM-5.3, with major advances in complex software engineering and agent tasks. Same base model as GLM-5.2, all gains from post-training: 50% coding improvement and emergent cybersecurity / vulnerability-discovery capability. 1M context, 128K output, always-on thinking (reasoning_effort: low/high/max).
OpenAI · Text
GPT-5.3 Codex, optimized specifically for programming tasks, with standout long-context code generation and multi-file coordinated editing.
OpenAI · Text
GPT-5.2, with a 272K context window and 128K output. A previous-generation flagship, a versatile choice for everyday Q&A, content creation, and lightweight coding.
Zhipu · Text
Zhipu's latest flagship, GLM-5.2, with a 1M context window and 128K output. Excels at long-horizon tasks and consistently ranks at the top of multiple authoritative benchmarks. Use the model name glm-latest to access the newest version.
OpenAI · Text
OpenAI GPT-5.2 Pro, the high-compute pro tier of GPT-5.2 for advanced reasoning and research.
OpenAI · Text
OpenAI GPT-5.2 Codex, a coding-optimized variant of GPT-5.2 with strong multi-file code generation.
OpenAI · Text
OpenAI GPT-5.1, 272K context + 128K output. A previous-generation general-purpose model for everyday production tasks.
Zhipu · Text
Zhipu GLM-5.1, with a 200K context window and 128K output. The predecessor to GLM-5.2, with well-balanced reasoning and coding capabilities.
Anthropic · Text
Claude Opus 5, Anthropic's latest flagship model with a 1M context window and 128K output. Next-gen deep reasoning, coding, and agent capabilities. Thinking is on by default and can be turned off.
Anthropic · Text
Claude Sonnet 5, with a 1M context window and 128K output. The best balance for coding and everyday agent tasks. Thinking is on by default and can be turned off.
OpenAI · Text
OpenAI GPT-5 Codex, optimized for programming with standout multi-language code generation and agentic coding.
Anthropic · Text
Claude Fable 5, with a 1M context window and 128K output. A mid-tier option that balances capability and cost, with thinking on by default.
Zhipu · Text
Zhipu GLM-5 Code, optimized specifically for programming tasks, with standout multi-language code generation and debugging capabilities.
Doubao · Image
ByteDance Doubao Seedream 5.0 Pro, high-quality text-to-image and image editing model, strong Chinese scene understanding.
OpenAI · Text
OpenAI GPT-5 Pro, the high-compute pro tier of GPT-5 for the most demanding reasoning and research tasks.
Doubao · Image
Volcano Engine image generation model
OpenAI · Text
OpenAI GPT-5 Nano, 272K context. The smallest, lowest-cost model for lightweight classification and routing.
OpenAI · Text
OpenAI GPT-5 Mini, 272K context. A compact, affordable model for high-volume everyday tasks.
OpenAI · Text
OpenAI GPT-5, 272K context + 128K output. A previous-generation flagship with strong multimodal, reasoning and coding.
Anthropic · Text
Anthropic's flagship Opus 4.8, with a 1M context window and 128K output. The top choice for deep reasoning and long-horizon agent tasks. Thinking is on by default and can be turned off.
Zhipu · Text
Zhipu GLM-4.7 Flash, with a 200K context window and 128K output. Ultra-fast responses at low cost, ideal for high-concurrency scenarios.
Anthropic · Text
Anthropic Claude Opus 4.7, 200K context, deep reasoning and coding for complex agent tasks. Thinking toggleable.
Zhipu · Text
Zhipu GLM-4.7, with a 200K context window and 128K output. A general-purpose production-grade model that balances capability and cost.
Anthropic · Text
Anthropic most capable model for coding and complex reasoning
Anthropic · Text
Most powerful model for highly complex tasks
Anthropic · Text
Anthropic Claude Opus 4.5, 200K context, a previous-generation flagship for highly complex tasks.
Anthropic · Text
Claude Haiku 4.5, 200K context, extreme speed and cost-efficiency, ideal for high-concurrency lightweight tasks.
Anthropic · Text
Anthropic Claude Sonnet 4.5, 200K context, balanced coding and everyday reasoning, a previous-generation model.
Grok · Text
xAI Grok 4.5, 256K context. The latest Grok flagship with leading reasoning, coding and real-time data access.
Doubao · Image
ByteDance Doubao Seedream 4.5, a previous-generation text-to-image and image editing model, per-image billing.
Grok · Text
xAI's latest flagship, Grok 4.3, with a 1M context window and 1M output. Integrates real-time data from the X platform, with standout reasoning and coding capabilities.
Grok · Text
Grok 4.20 multi-agent reasoning edition, with an ultra-long 2M context window and 2M output. The top choice for multi-step reasoning and long-horizon agents. Thinking is on by default.
Grok · Text
xAI Grok 4.1, 256K context. A general-purpose model with strong reasoning, coding and multimodal understanding.
OpenAI · Text
OpenAI GPT-4.1 Nano, the smallest variant of GPT-4.1 for low-cost, high-volume lightweight tasks.
Grok · Text
Grok 4.1 Fast, with a 2M context window. Ultra-fast responses at low cost, ideal for high-concurrency production tasks.
OpenAI · Text
OpenAI GPT-4.1, 1M context + 32K output. A previous-generation model strong at coding and instruction following.
OpenAI · Text
OpenAI GPT-4.1 Mini, a compact variant of GPT-4.1 balancing speed and cost for everyday tasks.
DeepSeek · Text
DeepSeek-V4-Pro flagship model, with a 1M context window and 384K output. Significantly enhanced agent capabilities and rich world knowledge. Thinking is on by default and can be turned off.
Grok · Text
Grok Code Fast, a fast version optimized specifically for programming tasks, with standout code generation and completion speed.
Grok · Text
xAI Grok 4 Fast, a faster, lower-cost variant of Grok 4 for high-concurrency production tasks.
DeepSeek · Text
DeepSeek-V4-Flash, with a 1M context window and 384K output. Delivers faster, more economical API service. Thinking is on by default and can be turned off.
Grok · Text
Grok 4, with a 256K context window. Offers deep reasoning and multimodal understanding, well suited for complex agent tasks.
OpenAI · Text
Affordable and intelligent small model
OpenAI · Vision
OpenAI multimodal flagship model with vision support
Qwen · Text
Tongyi's latest flagship, Qwen3.8, with a 992K context window and 64K output. Leading in reasoning and multilingual capabilities, well suited for complex tasks.
Qwen · Multimodal
Alibaba Qwen3.7 Plus, cost-effective multimodal agent model with upgraded vision-language capabilities. Supports GUI operation and visual-reference code generation. 1M context + 64K output. Thinking on by default.
Qwen · Multimodal
Qwen3.7 Flash, a native vision-language Flash model with 1M context and 64K output. Upgraded multimodal understanding and agent execution over 3.6 Flash.
Qwen · Text
Alibaba Qwen3.7 Max, next-generation flagship model for the agent era. Standout programming, office productivity, and long-horizon autonomous execution. 1M context + 64K output. Thinking on by default.
Qwen · Text
Alibaba Qwen3.6 Max Preview, the most capable model in the Qwen3.6 series. Standout vibe coding and coding agent capabilities. 256K context + 64K output. ⚠️ Retiring Oct 10.
Gemini · Text
Gemini 3.6 Flash, Google's latest Flash model with a 1M context window and 64K output. Balances speed and quality, well suited for general production tasks.
Qwen · Multimodal
Alibaba Qwen3.6 Plus, native vision-language Plus model. Upgraded agentic coding, frontend programming, multimodal recognition and OCR. 1M context + 64K output. Supports thinking.
Qwen · Multimodal
Alibaba Qwen3.6 Flash, native vision-language Flash model. Upgraded agentic coding, math reasoning and spatial intelligence. 1M context + 64K output. Fast and cost-effective.
Qwen · Text
Qwen3.5 Flash, 992K context + 64K output. An ultra-fast, low-cost model for high-concurrency production tasks.
Gemini · Text
Gemini 3.5 Flash, with a 1M context window and 64K output. Balances speed and quality, well suited for general production tasks.
Gemini · Text
Google Gemini 3.5 Flash Lite, 1M context. Extremely lightweight and low-cost, ideal for high-concurrency scenarios.
Qwen · Text
Qwen3.5 Plus, with a 992K context window and 64K output. A general-purpose production-grade model that balances capability and cost.
DeepSeek · Text
DeepSeek V3.2, 128K context, balanced reasoning and output efficiency, suitable for daily and production tasks.
DeepSeek · Text
DeepSeek V3.1, 128K context. A balanced model with strong reasoning and output efficiency for daily and production tasks.
Gemini · Video
Google Veo 3.1, high-quality text-to-video and video-to-video generation with audio, ideal for cinematic clips.
Gemini · Image
Gemini 3.1 Flash image generation model, with strong text-to-image and image editing capabilities. Ideal for marketing assets and design assistance.
Gemini · Text
Gemini 3.1 Pro Preview, with a 1M context window and 64K output. Google DeepMind's flagship model, leading in multimodal understanding and long-context reasoning.
Gemini · Text
Gemini 3.1 Flash Lite, with a 1M context window and 64K output. Extremely lightweight and low-cost, ideal for high-concurrency scenarios.
Grok · Text
xAI Grok 3, 128K context. A previous-generation flagship with strong reasoning and real-time X platform data.
Qwen · Text
Qwen3 235B, 256K context. The flagship open-weight MoE model with strong general reasoning and multilingual capability.
Qwen · Text
Qwen3 Coder 480B, 256K context. The flagship open-weight coding model with strong multi-language generation and long-context code understanding.
DeepSeek · Text
DeepSeek OCR, a document understanding model that extracts text and structure from images and scanned documents.
Moonshot · Text
Kimi K3, Moonshot AI's latest flagship model with a 1M token context window and 128K output. Cutting-edge reasoning, long-context processing, and coding capabilities. Thinking is on by default and can be turned off.
Qwen · Vision
Qwen3 VL 235B, 128K context. The open-weight multimodal model with strong image, document and OCR understanding.
MiniMax · Text
MiniMax M3, a next-generation flagship language model with a 1M context window and 128K output. Achieves industry-leading results in coding and agent benchmarks, well suited for agent reasoning and tool calling.
Qwen · Text
Qwen3 Coder Plus, with a 998K context window and 64K output. Optimized specifically for programming tasks, with standout multi-language code generation and long-context code understanding.
Qwen · Vision
Qwen3 VL Plus vision understanding model, with a 260K context window. Supports image and document understanding along with multimodal reasoning.
Gemini · Text
Google Gemini 3 Flash, 1M context + 64K output. A previous-generation fast model for general production tasks.
Qwen · Text
Qwen3 Coder Flash, with a 998K context window and 64K output. An ultra-fast coding model well suited for high-concurrency code completion and generation.
DeepSeek · Text
Cost-effective model with strong coding ability
Gemini · Image
Gemini 3 Pro image generation flagship, delivering high-quality text-to-image and complex instruction parsing. Ideal for professional visual creation.
Grok · Text
xAI Grok 3 Fast, a faster, lower-cost variant of Grok 3 for high-concurrency production tasks.
MiniMax · Text
MiniMax M2.7, 1M context. A previous-generation general-purpose model balancing capability and cost for production.
Moonshot · Multimodal
Moonshot AI next-generation flagship model Kimi K2.7, multimodal multilingual 256K context. Standout long-context and coding capabilities. Thinking is on by default and can be turned off.
Moonshot · Text
Moonshot AI's next-generation model, Kimi K2.6, with a 262K context window. Standout long-context and coding capabilities. Thinking is on by default and can be turned off.
MiniMax · Text
MiniMax M2.5, with a 1M context window. A general-purpose production-grade model that balances capability and cost.
Moonshot · Text
Kimi K2.5, with a 262K context window. Balances reasoning capability with output speed, well suited for general agent tasks. Thinking is on by default and can be turned off.
Gemini · Embedding
Gemini Embedding 2, with 8K input. Converts text into high-dimensional vectors for semantic search and RAG scenarios.
google · Text
Fast and versatile model with long context
google · Vision
Google most capable model with 1M context window
Gemini · Text
Google Gemini 2.5 Flash Lite, 1M context. The lowest-cost Flash variant for lightweight high-volume tasks.
OpenAI · Video
OpenAI 视频生成专业版。
OpenAI · Video
OpenAI 视频生成,文生/图生视频。
wan · Video
单图+音频驱动,生成说话/唱歌/表演数字人视频。
wan · Video
通过指令对视频进行编辑,支持局部/整体编辑、视频重塑、视频复刻等
wan · Video
参考视频中的人或物,精准保持形象和声音,支持多参考合拍
wan · Video
图片生成视频内容,稳定保持图像主体、风格和文字等细节信息
wan · Video
文字生成视频内容,丝滑动态能力,电影美学控制,精准指令遵循
MiniMax · Text
MiniMax M2.1, 1M context. A legacy general-purpose model for cost-sensitive, high-volume production tasks.
Doubao · Text
ByteDance Doubao Seed 2.1 Turbo, balanced speed and quality, 256K context, cost-effective for production tasks.
Doubao · Text
ByteDance Doubao Seed Evolving, rapid iteration model, 256K context, continuously updated with latest capabilities.
Doubao · Text
ByteDance Doubao Seed 2.1 Pro, enhanced capability flagship, 256K context, strong reasoning and agent capabilities.
wan · Video
多模态全能参考,灵活可控生成视频内容,轻松实现故事创作、创意表达、营销物料制作等
happyhorse · Video
HappyHorse-Video-Edit支持视频编辑,自然语言指令编辑视频,可参考最多5张图片局部或全局编辑视频元素,能够精准复刻视频动态过程,实现更强表现能力。
Qwen · TTS
Qwen Audio 3.0 speech synthesis Plus, high-quality natural voices.
Doubao · Text
ByteDance Doubao Seed Character, a roleplay-oriented model with strong persona consistency and character interaction, 256K context.
Doubao · Video
ByteDance Doubao video generation model Seedance 2.0 Mini, supports video generation and editing, ideal for short video and motion graphics.
Doubao · Video
ByteDance Doubao Seedance 2.0 Fast, a faster variant of Seedance 2.0 for rapid video generation at lower cost.
happyhorse · Video
Alibaba HappyHorse 1.1 R2V, reference-to-video generation supporting up to 9 reference images, with stronger subject/scene/style consistency. Per-second billing.
Doubao · Text
Doubao Seed 2.0 Lite balances generation quality with response speed, well suited for production work like unstructured information processing, content creation, and data analysis. Thinking is on by default and can be turned off.
Doubao · Text
Doubao Seed 2.0 Mini targets low-latency, high-concurrency, and cost-sensitive scenarios with blazing-fast inference. Thinking is on by default and can be turned off.
Doubao · 3D
ByteDance Doubao 3D generation model Seed3D 2.0, image-to-3D, ideal for 3D assets and game scene modeling.
Doubao · Video
ByteDance Doubao Seedance 2.0, high-quality text-to-video and video-to-video generation, ideal for short cinematic videos.
Qwen · Image
Tongyi Qwen image generation 3.0 Pro, supporting 1K/2K resolution text-to-image and image editing, with input and output billed separately. Ideal for high-quality marketing assets and design creation.
Qwen · Image
Tongyi Qwen image generation 3.0, supporting 1K/2K resolution text-to-image and image editing. A cost-effective choice.
Doubao · Video
新一代视频创作模型,单次 30 秒长叙事、全模态参考、视频编辑与延长。
Doubao · 3D
Hyper3D Gen2 by Yingmou (影眸科技), image-to-3D generation hosted on Volcengine Ark. Ideal for 3D assets and game scene modeling.
fun · music
Tongyi music generation, billed by output audio seconds.
Qwen · ASR
Qwen speech recognition Flash for audio file transcription.
Qwen · ASR
Qwen Audio 3.0 speech recognition Flash for audio file transcription.
fun · TTS
CosyVoice 3.5 Plus speech synthesis, high-quality multilingual voices.
fun · TTS
CosyVoice 3.5 Flash speech synthesis, cost-effective.
Qwen · Embedding
Latest Qwen text embedding model, suited for retrieval and semantic matching.
Qwen · Text
Tongyi Qwen 3.8 open-source flagship, 2.4T-parameter MoE (95B active), supporting non-thinking and thinking modes.
wan · Video
Wan video face swap: replaces the character in a video, standard/pro modes.
Qwen · Text
Tongyi Qwen 3.8 open-source 27B dense model, a balanced and cost-effective choice.
Qwen · Text
Qwen long-context model, designed for ultra-long context scenarios.
Qwen · Text
Qwen deep research model with automated multi-step retrieval, analysis and report generation.
Qwen · TTS
Qwen Audio 3.0 speech synthesis Flash, low latency and cost-effective.
Qwen · Vision
Qwen OCR vision model for document recognition and structured information extraction.
kling · Image
Kling V3 image generation with 1K/2K/4K resolutions.
kling · Video
Kling V3 video generation, supporting silent/audio and 720P/1080P/4K resolutions.
wan · Video
Wan image-to-motion: drives portrait video generation with a reference motion, standard/pro modes.
Doubao · 3D
Hitem3D 2.0 by Shumei (数美万物), image-to-3D generation hosted on Volcengine Ark. Ideal for 3D assets and scene modeling.
Doubao · Text
ByteDance Doubao Seed Code, a coding-optimized model with strong multi-language generation and debugging, 256K context.
Qwen · Image
Tongyi Wanxiang Qwen Image 2.0 Pro image generation model, delivering high-quality text-to-image and image editing. Ideal for marketing assets and design creation.
happyhorse · Video
Alibaba Cloud HappyHorse 1.1 T2V video generation model. Text-to-video with upgraded dynamics, texture and audio. More fluid and natural character motion and scene atmosphere.
happyhorse · Video
Alibaba HappyHorse 1.1 I2V, image-to-video generation with improved texture, motion and cross-clip consistency. Per-second billing.
Doubao · Text
ByteDance Doubao Seed Translation, a translation-specialized model for high-quality multilingual translation, 256K context.
Doubao · Text
Doubao Seed 2.0 Pro, a flagship all-around general-purpose model with enhanced multimodal understanding, long-context reasoning, structured generation, and tool execution. Thinking is on by default and can be turned off.
Doubao · Text
Doubao Seed 2.0 Code builds on Seed 2.0's agent and VLM capabilities to strengthen coding, with outstanding front-end work and multi-language support. Non-thinking by default, with deep thinking available on demand.
Doubao · Text
ByteDance Doubao Seed 1.8, lightweight cost-effective model, 256K context, suitable for high-concurrency scenarios.
Doubao · Vision
ByteDance Doubao Seed 1.6 Vision, a previous-generation multimodal vision model, 256K context, image and document understanding.
Doubao · Text
ByteDance Doubao Seed 1.6 Flash, an ultra-fast lightweight variant of 1.6, 256K context, ideal for high-concurrency scenarios.
Doubao · Text
ByteDance Doubao Seed 1.6, a previous-generation general-purpose model, 256K context, cost-effective for everyday tasks.
Doubao · Vision
ByteDance Doubao 1.5 Vision Lite, a legacy lightweight multimodal vision model, 256K context, cost-effective image understanding.
Doubao · Video
ByteDance Doubao Seedance 1.5 Pro, a previous-generation video generation model supporting audio and silent video output.
Qwen · TTS
通义千问语音合成 Instruct Flash 版。
Qwen · TTS
通义千问语音合成 Flash 版。
Doubao · Embedding
Doubao's large-scale embedding model converts text into high-dimensional vector representations for semantic search and RAG scenarios.
Qwen · Embedding
Qwen3 Reranker 8B, an open-weight reranker that reorders retrieved documents by relevance for higher-quality RAG.
Doubao · Embedding
ByteDance Doubao Embedding Vision, multimodal embedding model supporting both image and text input.
Qwen · Text
Qwen Turbo, with a 1M context window and 16K output. Extremely lightweight and low-cost, ideal for high-concurrency simple tasks.
Qwen · Embedding
Qwen3 Embedding 8B, an open-weight text embedding model producing high-quality vectors for semantic search and RAG.
Doubao · TTS
High-quality text-to-speech model.
Doubao · Embedding
ByteDance Doubao Embedding, standard text embedding model, high-quality vector representation for semantic search and RAG.
Doubao · ASR
豆包流式语音识别模型 2.0,火山引擎语音团队基于大模型语音识别能力全新升级,依托业界领先的自研语音识别技术和海量的语音行业大数据优势,语音识别大模型拥有更加灵敏的耳朵+更加聪明的大脑。
Auto · Text
Smart routing model that intelligently matches compute and model combinations across both quality and speed. Get early access to the latest models from ByteDance and its ecosystem.