OpenAI-compatible API at gptapi.net.cn/v1. Sourced from China’s top AI labs. All specs from official vendor documentation.
| Model ID | Capabilities & Specs |
|---|---|
| DeepSeek · 2 models | |
| deepseek-v4-flash | 284B MoE (13B active) · 1M context, 384K max output. Thinking mode, JSON output, tool calls, FIM. Fast & efficient. |
| deepseek-v4-pro | 1.6T MoE (49B active) · 1M context, 384K max output. Thinking mode, JSON output, tool calls, FIM, context caching. Flag. |
| Alibaba / Qwen · 9 models | |
| qwen-image-2.0-pro-2026-06-22 | Professional image generation · Infographics, photorealism · ~$0.07/image · Contact support for access. |
| qwen3.6-35b-a3b | 35B (3B active) MoE · 256K context · Cost-efficient. $0.25/$1.50 per 1M. |
| qwen3.6-flash | 1M context · Fast tier · High throughput. $0.17/$1.00 per 1M. |
| qwen3.6-plus | 1M context · Strong generalist · Balanced cost-performance. $0.28/$1.67 per 1M (tiered ≥256K). |
| qwen3.7-max | 1M context · Qwen flagship, agent-centric · Coding, office productivity, tool orchestration. $1.67/$5.00 per 1M (official, currently 50% off). |
| qwen3.7-max-2026-06-08 | 1M context · Qwen flagship latest iteration. Continuous improvements. |
| qwen3.7-plus | 1M context · Cost-effective Qwen 3.7 · Text+image input. $0.28/$1.11 per 1M (official, 20% off). |
| qwen3.7-plus-2026-05-26 | 1M context · Qwen 3.7 Plus updated. |
| qwen3.7-text-embedding | 32K input · Embedding model · 1024-dim vectors. $0.07/1M. |
| Moonshot / Kimi · 12 models | |
| kimi-k2.5 | 256K context · Open-source SoTA · ⚠️ Deprecated, shutting down Aug 31, 2026. Available until then. |
| kimi-k2.6 | 256K context · Multimodal · Thinking + non-thinking modes · Vision, tool calls, streaming. |
| kimi-k2.7-code | 256K context · Code-specialized · Thinking-only · Text+image+video input · Tool calls, structured output. |
| kimi-k2.7-code-highspeed | 256K context · ~5-6x faster than K2.7 Code · Optimized for low-latency code completion & agent IDEs. |
| kimi-k3 | 2.8T params · 1M context · Vision input · Always-on reasoning (low/high/max) · Tool calls, web search, JSON Mode, context caching. Kimi's strongest. |
| moonshot-v1-128k | 128K context. ⚠️ Deprecated, shutdown Aug 31. |
| moonshot-v1-128k-vision-preview | 128K context, vision. ⚠️ Deprecated, shutdown Aug 31. |
| moonshot-v1-32k | 32K context. ⚠️ Deprecated, shutdown Aug 31. |
| moonshot-v1-32k-vision-preview | 32K context, vision. ⚠️ Deprecated, shutdown Aug 31. |
| moonshot-v1-8k | 8K context. ⚠️ Deprecated, shutdown Aug 31. Legacy foundation model. |
| moonshot-v1-8k-vision-preview | 8K context, vision. ⚠️ Deprecated, shutdown Aug 31. |
| moonshot-v1-auto | Auto routing · ⚠️ Not an official Kimi model. Third-party alias. |
| Zhipu / GLM · 10 models | |
| glm-4.5 | 355B MoE (32B active) · 128K context · ⚠️ Deprecated in favor of 4.7. |
| glm-4.5-air | 106B (12B active) MoE · 128K context · Most affordable GLM. High-volume workloads. |
| glm-4.5v | Vision-language · ⚠️ Deprecated. Still available with pricing. |
| glm-4.6 | 355B MoE (32B active) · 200K context · Stable general-purpose · Supports thinking mode. |
| glm-4.6v | 106B · 128K context · Vision model · Image understanding, visual reasoning. |
| glm-4.7 | 128K context · Latest 4.x · Enhanced coding, creative writing, role-play. ⚠️ Deprecating 4.6. |
| glm-5 | 1M context · Strong general reasoning & agent capabilities. |
| glm-5-turbo | Speed-optimized GLM 5 · High RPM for production throughput. |
| glm-5.1 | 1M context · Tiered pricing (≤32K / >32K) · Deep reasoning, agentic workflows. |
| glm-5.2 | 1M context · Zhipu flagship · Long-horizon agents, software engineering, deep reasoning · Tool calls, structured output. |
| ByteDance / Doubao · 5 models | |
| doubao-embedding-text-240715 | Doubao Embedding · 2560-dim vectors · Semantic search & RAG. |
| doubao-lite-128k-240428 | Doubao Lite · 128K context · Cost-effective long-text processing. |
| doubao-lite-32k-240628 | Doubao Lite · 32K context · Fast, affordable Chinese model. |
| doubao-pro-128k-240628 | Doubao Pro · 128K context · Top Chinese language performance · Ultra-long context. |
| doubao-pro-32k-240615 | Doubao Pro · 32K context · Strong Chinese NLP · Function calling. |
| MiniMax · 3 models | |
| MiniMax-M2.7 | Strong general-purpose dialogue · Stable production performance · Context caching. |
| MiniMax-M2.7-highspeed | High-speed variant · Real-time, low-latency applications. |
| MiniMax-M3 | 1M context · Multimodal (text+image+video) · Long-horizon agents, coding, creative writing · Tool calls. |
| Baidu / ERNIE · 2 models | |
| ernie-4.5-turbo-128k | Baidu ERNIE 4.5 Turbo · 128K context · Fast inference for long documents. |
| ernie-5.1 | Baidu ERNIE 5.1 · Deep Chinese language understanding · Enterprise-grade · Function calling, structured output. |
| Tencent / Hunyuan · 2 models | |
| hy3 | Tencent HY3 · 295B MoE (21B active), 192 experts · 128K context · Reasoning, agentic, production-grade. |
| hy3-preview | Tencent HY3 Preview · Early access · Configurable thinking depth. |
| StepFun · 2 models | |
| step-3.5-flash | Entry model · Fastest response, minimal cost per token. |
| step-3.7-flash | Fast model · General-purpose, competitive performance at low cost. |
| Xiaomi / MiMo · 2 models | |
| mimo-v2.5 | Xiaomi MiMo V2.5 · 128K context · Natural Chinese fluency · General-purpose. |
| mimo-v2.5-pro | Xiaomi MiMo V2.5 Pro · 128K context · Strong coding & agentic capabilities · Benchmark leader. |