# Open-weight models from China: the state of play in July

[Skip to content](#lm-inhoud)Network/[NL](/en/open-weight-modellen-juli-2026)EN[Hubhub.llmnet.nlCompare models on task, language, cost and license.](https://hub.llmnet.nl/en/)[Communitycommunity.llmnet.nlPrompt techniques, patterns and system prompts.](https://community.llmnet.nl/en/)[APIapi.llmnet.nlLLMs in production: rate limits, routing, structured output.](https://api.llmnet.nl/en/)[Consultancyconsultancy.llmnet.nlRolling out AI in an organization, pilot to production.](https://consultancy.llmnet.nl/en/)[Newsnieuws.llmnet.nlAI developments, explained for the Netherlands.](https://nieuws.llmnet.nl/en/)[Benchmarkbenchmark.llmnet.nlMeasure AI quality yourself, on your own tasks.](https://benchmark.llmnet.nl/en/)[Careersvacatures.llmnet.nlAI roles, salaries and career paths in the Netherlands.](https://vacatures.llmnet.nl/en/)[Learnleren.llmnet.nlAI concepts in plain language, beginner to builder.](https://leren.llmnet.nl/en/)[Guidegids.llmnet.nlRun AI privately on your own Mac, PC, NAS or home server.](https://gids.llmnet.nl/en/)[Directorydirectory.llmnet.nlMapping the AI ecosystem: tools, models, companies.](https://directory.llmnet.nl/en/)[Radarradar.llmnet.nlSignals from X, research and communities for indie developers.](https://radar.llmnet.nl/en/)[Appsapps.llmnet.nlReviews of AI apps and open-source repos, with tips for builders.](https://apps.llmnet.nl/en/)[llmnet.nl — main site](https://llmnet.nl/en/)[](https://x.com/intent/post?url=https%3A%2F%2Fradar.llmnet.nl%2Fen%2Fopen-weight-modellen-juli-2026&text=Open-weight%20models%20from%20China%3A%20the%20state%20of%20play%20in%20July)[](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fradar.llmnet.nl%2Fen%2Fopen-weight-modellen-juli-2026)[](https://www.reddit.com/submit?url=https%3A%2F%2Fradar.llmnet.nl%2Fen%2Fopen-weight-modellen-juli-2026&title=Open-weight%20models%20from%20China%3A%20the%20state%20of%20play%20in%20July)[](#)[](https://x.com/intent/post?url=https%3A%2F%2Fradar.llmnet.nl%2Fen%2Fopen-weight-modellen-juli-2026&text=Open-weight%20models%20from%20China%3A%20the%20state%20of%20play%20in%20July)[](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fradar.llmnet.nl%2Fen%2Fopen-weight-modellen-juli-2026)[](https://www.reddit.com/submit?url=https%3A%2F%2Fradar.llmnet.nl%2Fen%2Fopen-weight-modellen-juli-2026&title=Open-weight%20models%20from%20China%3A%20the%20state%20of%20play%20in%20July)[](#)By Ivo Donker — created with AI assistance (Claude & Gemini) · Last updated: July 27, 2026

 
 
 [← Back to AI-Radar](/)
 [Model Routing API](https://api.llmnet.nl/en/model-routing)
 
 
# Open-weight models from China: the state of play in July 2026

 
 Published on July 24, 2026 · By the editors of AI-Radar
 
 

 
 
 The landscape of open-weight language models is developing at a breakneck pace in the summer of 2026. Particularly from China, we are seeing a stream of releases challenging the established order. For indie developers building their own AI agents, automating workflows via tools like n8n, and maintaining their own homelab, this wave brings both opportunities and infrastructural challenges. 
 

 
 The common thread this month is clear: the absolute top models are simply becoming too large to run locally on consumer hardware or a modest home server. Models with parameters in the trillions (T) require a strategic shift. Instead of trying to host everything locally, the best practice for indie stacks is shifting toward deploying a local LLM gateway like LiteLLM, paired with external API providers that offer these open-weight giants at highly competitive rates. This allows you to maintain control over your agent logic and workflow orchestration while benefiting from the computing power of the latest frontier models.
 

 
## Key releases and signals

 
 
 [Kimi K3: Open weights release on July 27 focuses on coding](https://www.techtimes.com/articles/321499/20260724/kimi-k3-open-weights-drop-july-27-near-frontier-coding-undisclosed-hallucination-risk.htm)
 
 
 Moonshot is about to release the full weights of Kimi K3 on July 27 under its own Kimi K3 license, which is permissive for most commercial use but carries its own conditions. This model, boasting a massive 2.8 trillion (T) parameters, has claimed first place in the Frontend Code Arena, outperforming established names like GLM-5.2 and GPT-5.6 Sol. 
 

 
 Why this is relevant: For developers building coding agents, this is a significant step forward. Due to its size, self-hosting on an average home server is unrealistic. However, the model becomes immediately interesting as soon as it becomes available through API providers and can be integrated into your local gateway.
 

 
 Signal: high
 Action: watch
 kimi
 moonshot
 coding
 api
 
 

 
 
 [DeepSeek V4-Flash: The new standard for agentic pipelines](https://openrouter.ai/blog/insights/the-open-weight-models-that-matter-june-2026/)
 
 
 With V4-Flash, DeepSeek has released an MIT-licensed Mixture-of-Experts (MoE) model. With a total of 284 billion parameters, of which only 13 billion are active per token, and a context window of 1 million tokens, this model is optimized for speed and efficiency. Scoring 79.0% on SWE-bench Verified, it is one of the most capable open-weight models for software engineering.
 

 
 Why this is relevant: This is currently the most concrete choice for active agentic pipelines. Thanks to the low active parameter count, API costs for the DeepSeek V4 Pro variant (as of August 2026: $0.435 per million input tokens and $0.87 per million output tokens on a cache miss) are extremely low. This makes it the ideal engine for complex, long-running agent workflows controlled via a local gateway.
 

 
 Signal: high
 Action: build
 deepseek
 mit
 agentic
 litellm
 
 

 
 
 [GLM-5.2: The efficient and audited MIT model from Zhipu](https://www.orcarouter.ai/blog/qwen-3-8-vs-glm-5-2)
 
 
 With GLM-5.2, Zhipu has released a 753B MoE model (of which ~40B are active) under an MIT license on Hugging Face. The model excels in reasoning tasks with a score of 91.2% on GPQA Diamond and 62.1% on SWE-bench Pro. In the community, it is frequently labeled as one of the best all-round open-source LLMs at the moment.
 

 
 Why this is relevant: Just like DeepSeek V4, the MoE architecture ensures that operational costs with API providers remain low. It is an excellent alternative or backup model in your routing setup for tasks that require deep logical reasoning.
 

 
 Signal: medium-high
 Action: idea
 glm
 zhipu
 mit
 litellm
 
 

 
 
 [Qwen 3.8-Max: Alibaba showcases 2.4T giant at WAIC](https://www.yottalabs.ai/post/qwen-3-8-max-release-date-specs-how-to-access-2026)
 
 
 Alibaba previewed Qwen 3.8-Max at the World Artificial Intelligence Conference (WAIC) in Shanghai. This sparse MoE model of 2.4 trillion parameters is fully multimodal and, according to initial claims, performs just below commercial giants like Claude Fable 5. Alibaba has promised to release the weights "soon."
 

 
 Why this is relevant: Although the performance is impressive, the weights are not yet available at this time, and the licensing structure is unknown. This is a model to watch closely, but not yet to include in active workflows.
 

 
 Signal: medium
 Action: watch
 qwen
 alibaba
 multimodal
 preview
 
 

 
 
 [Qwen 3.6 (27B/35B): Running locally on Apple Silicon](https://apxml.com/posts/best-local-llms-apple-silicon-mac)
 
 
 For those who still prefer a fully local setup without dependency on external APIs, the smaller Qwen 3.6 variants (27B and 35B) offer a solution. Optimized for Apple Silicon via MLX, these models achieve 16 to 72 tokens per second, respectively.
 

 
 Why this is relevant: These are among the few recent Chinese models that actually fit locally on consumer hardware. Keep in mind, however, that for a comfortable workflow with larger contexts, you will quickly need 64GB+ of unified memory. For lighter hardware, local inference of this size is often just too heavy, meaning the API route remains preferred.
 

 
 Signal: low-medium
 Action: watch
 qwen
 mlx
 self-host
 
 

 
## How can you use this?

 
 The trend is unmistakable: the line between 'open-weight' and 'locally runnable' has definitively blurred. The most powerful open-weight models are simply too large for a standard homelab. As an indie developer, you now get the most return on your setup by focusing on smart orchestration instead of heavy local hardware.
 

 
 
- Set up a central LLM gateway: Use a tool like LiteLLM on your home server. This allows you to easily switch between local models (for privacy-sensitive basic tasks) and external APIs of DeepSeek V4 or GLM-5.2 for heavier reasoning work.
 
- Optimize your agent architecture for MoE: Because models like DeepSeek V4-Flash and GLM-5.2 use Mixture-of-Experts, they are ideally suited for tasks that require a lot of interaction and context, without quickly draining your budget.
 
- Prepare for Kimi K3: Keep an eye on the July 27 release. Once this model is available at API routers, it could be an excellent upgrade for your coding and development agents.
 
 

 
 
 Compiled from public sources, July 2026.
