2026-09-02
Official August announcements: Qwen3.8, GLM-5.3, MiniMax H3, and DeepSeek Vision
August 2026 packed several Chinese-lab launches that matter if you run an OpenAI-compatible gateway toward LatAm. This is an original fact summary from official sites (not pasted WeChat), with dates and URLs, plus NexoRouter catalog status on 2 September 2026.

3 Aug — Alibaba Qwen3.8-Max
Alibaba Cloud Community published Qwen3.8-Max: 2.4T total / 95B active parameters, up to 1M context, multimodal, API via Model Studio / QwenCloud. On 27 August the same channel detailed Qwen3.8-Flash-Next as an architecture preview toward Qwen4 (125B + n-gram embeddings, 6B active; production served as qwen3.8-flash on QwenCloud at vendor rates).
On NexoRouter (2 Sep): qwen3.8-max was listed with Pending price. Qwen/Qwen3-Max had $0.35 / 1M input and $1.40 / 1M output. Use the exact card ID.
3 Aug — MiniMax H3 open source
MiniMax published Open General Intelligence: MiniMax H3 Is Now Open Source on 3 August 2026: omni-modal video system, up to 15s, 2K via regeneration, 32 kHz stereo audio, FL2VA / Ref2VA checkpoints under a community license.
On NexoRouter: MiniMax-H3 showed Pending. For chat with a published price that day, MiniMax-M3 showed $2.04 / $0.41 / $8.17 per 1M tokens.
14 Aug — Z.ai GLM-5.3
GLM-5.3 (14 Aug 2026): post-training on the GLM-5.2 base, coding-agent focus and public benchmarks cited by Z.ai; thinking required with reasoning_effort.
On NexoRouter: glm-5.3 ($1.22 / $0.30 / $4.26) and glm-5.3-flash ($0.12 / $0.03 / $0.43) are live. Do not confuse with glm-5.2 / glm-5.2-sg (different IDs).
13–21 Aug — DeepSeek V4 GA and Vision Exp
The DeepSeek changelog on 13 Aug confirms V4-Pro GA and official API peak/off-peak pricing (effective 16 Aug UTC). 21 Aug adds deepseek-v4-flash-vision-exp.
On NexoRouter: deepseek-v4-flash-0731 has a published price; several -sg / vision / pro-0813 variants are Pending or differently priced. Check /models?q=deepseek-v4.
What to do from LatAm (news/guides tone, not an internal log)
1. Open the official post (table below) and note the vendor model name.
2. Search the ID on Models. If missing or Pending, do not budget.
3. Smoke-test in Playground.
4. Integrate Chat Completions per /docs and first API call.
5. For capability vs price, use choose models.

Official sources table
| Date | Vendor | Topic | URL |
| --- | --- | --- | --- |
| 2026-08-03 | Alibaba / Qwen | Qwen3.8-Max | https://www.alibabacloud.com/blog/alibaba-unveils-qwen3-8-max-its-largest-and-most-capable-flagship-model-to-date_603420 |
| 2026-08-27 | Alibaba / Qwen | Qwen3.8-Flash-Next | https://www.alibabacloud.com/blog/603501 |
| 2026-08-03 | MiniMax | H3 open source | https://www.minimax.io/news/minimax-h3-open-source |
| 2026-08-14 | Z.ai | GLM-5.3 | https://z.ai/blog/glm-5.3 |
| 2026-08-13 / 21 | DeepSeek | V4-Pro GA + Vision Exp | https://api-docs.deepseek.com/updates/ |
| 2026-04-24 | DeepSeek | V4 Preview | https://www.deepseek.com/en/news/v4-preview/ |
| 2026-07-16 | Moonshot | Kimi K3 (flagship context) | https://www.moonshot.ai/ |
WeChat (mp.weixin.qq.com) often blocks bots; we prefer fetchable official sites and docs. If the catalog adds an ID tomorrow, that card beats any blog summary.
What not to do
Do not paste long WeChat or vendor-blog paragraphs onto the site: that creates copyright and Google-duplicate risk. Summarize the fact (date, ID, announced capability), attribute with a deep link, and send readers to /models for the price string.
Do not assume “it is on OpenRouter” or “it is on UCloud” authorizes inventing a NexoRouter SKU. If the directory does not show the ID, the blog must not pretend availability.
Do not compare TTFT or claim “cheaper than X” without a live, comparable capture. This post series only cites NexoRouter list prices and facts from official pages.
For teams that already keep a single OpenAI client in Mexico or the Southern Cone, the value is lower friction: one base_url, one key vault, and a filterable directory. When Qwen, Z.ai, DeepSeek, or MiniMax publish the next ID, the checklist above repeats in minutes—not a week of SDK rewrites.