# Retrospective: which July signals came true?

[Skip to content](#lm-inhoud)Network/[NL](/en/terugblik-welke-juli-signalen-zijn-uitgekomen)EN[Hubhub.llmnet.nlCompare models on task, language, cost and licence.](https://hub.llmnet.nl/en/)[Communitycommunity.llmnet.nlPrompt techniques, patterns and system prompts.](https://community.llmnet.nl/en/)[APIapi.llmnet.nlLLMs in production: rate limits, routing, structured output.](https://api.llmnet.nl/en/)[Consultancyconsultancy.llmnet.nlRolling out AI in an organisation, pilot to production.](https://consultancy.llmnet.nl/en/)[Newsnieuws.llmnet.nlAI developments, explained for the Netherlands.](https://nieuws.llmnet.nl/en/)[Benchmarkbenchmark.llmnet.nlMeasure AI quality yourself, on your own tasks.](https://benchmark.llmnet.nl/en/)[Careersvacatures.llmnet.nlAI roles, salaries and career paths in the Netherlands.](https://vacatures.llmnet.nl/en/)[Learnleren.llmnet.nlAI concepts in plain language, beginner to builder.](https://leren.llmnet.nl/en/)[Guidegids.llmnet.nlRun AI privately on your own Mac, PC, NAS or home server.](https://gids.llmnet.nl/en/)[Directorydirectory.llmnet.nlMapping the AI ecosystem: tools, models, companies.](https://directory.llmnet.nl/en/)[Radarradar.llmnet.nlSignals from X, research and communities for indie developers.](https://radar.llmnet.nl/en/)[Appsapps.llmnet.nlReviews of AI apps and open-source repos, with tips for builders.](https://apps.llmnet.nl/en/)[llmnet.nl — main site](https://llmnet.nl/en/)[](https://x.com/intent/post?url=https%3A%2F%2Fradar.llmnet.nl%2Fen%2Fterugblik-welke-juli-signalen-zijn-uitgekomen&text=Retrospective%3A%20which%20July%20signals%20came%20true%3F)[](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fradar.llmnet.nl%2Fen%2Fterugblik-welke-juli-signalen-zijn-uitgekomen)[](https://www.reddit.com/submit?url=https%3A%2F%2Fradar.llmnet.nl%2Fen%2Fterugblik-welke-juli-signalen-zijn-uitgekomen&title=Retrospective%3A%20which%20July%20signals%20came%20true%3F)[](#)

[](https://x.com/intent/post?url=https%3A%2F%2Fradar.llmnet.nl%2Fen%2Fterugblik-welke-juli-signalen-zijn-uitgekomen&text=Retrospective%3A%20which%20July%20signals%20came%20true%3F)[](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fradar.llmnet.nl%2Fen%2Fterugblik-welke-juli-signalen-zijn-uitgekomen)[](https://www.reddit.com/submit?url=https%3A%2F%2Fradar.llmnet.nl%2Fen%2Fterugblik-welke-juli-signalen-zijn-uitgekomen&title=Retrospective%3A%20which%20July%20signals%20came%20true%3F)[](#)

# Retrospective: which July signals came true?

By Ivo Donker — compiled with AI assistance · Last updated: August 8, 2026

This is the first retrospective for the AI Radar. The 22 articles published in July contained dozens of signals — statements about what is going to happen, paired with an estimated certainty and recommended actions. However, signals have never been formally evaluated: no article has looked back to assess whether the predictions were accurate. This retrospective does just that for the four biggest themes of July: open models, the MCP protocol, the price war, and observability. The format is new for this subdomain, but it is precisely how radar builds authority rather than merely summarizing: the predictions were made here, so this evaluation cannot happen anywhere else.

## How this retrospective evaluates signals

A signal is only testable when three things are defined: exactly what was predicted, when, and how to verify if it came true. That is why this article uses the same structure for each signal. First, the prediction, linked directly to the July article where it originally appeared. Next, the current status, backed by a source and a date — a signal without a source is not a factual statement, but merely an opinion. Finally, the verdict: realized, partially realized, or still pending. The criteria separating these three are intentionally strict: a release that arrives a week later than predicted counts as realized, whereas a release that was only announced but is not yet available does not. This approach also means certain signals are deliberately left unevaluated for now: a signal with a three-month horizon should not be judged after just two weeks.

## Realized: Kimi K3 is here, and its scale was no exaggeration

The July article on [open-weight models from China](https://radar.llmnet.nl/en/open-weight-modellen-juli-2026) predicted that Moonshot would release the weights for Kimi K3 under a modified MIT license, boasting 2.8 trillion parameters—the largest open model in history—and that this model would definitively push past the limits of what can be run locally. On July 26, that became reality: Moonshot published the weights, and early reviewers ranked K3 at frontier-level for agentic coding. The second part of the prediction also came true: the hardware requirements are so steep that virtually no organization runs it in-house, making cloud providers the only practical path for access. The verdict is twofold: the release materialized, and the caveat noted in the article—large does not mean self-hostable—was confirmed. For builders, this changes nothing about the optimal approach: keep proprietary agent logic local, and route heavy models via API. The underlying calculation can be found in [the guide on local vs. API-based hosting](https://gids.llmnet.nl/en/lokaal-of-via-een-api-de-rekensom-over-drie-jaar).

## Confirmed: MCP goes stateless, with mandatory authentication

The [July article on MCP and agent tooling](https://radar.llmnet.nl/en/mcp-agent-tooling-juli-2026) reported that the specification for the biggest MCP update since its launch would be finalized on July 28: a stateless architecture (dropping Mcp-Session-Id), hardened OAuth/OIDC authorization with six SEPs, and the removal of tools like tools/list-scoping that cannot operate securely without sessions. The release candidate had already been locked on May 21, 2026, and validated for ten weeks by SDK maintainers and client implementers; on July 28, 2026, the final 2026-07-28 specification was published—right on schedule. The core premise was confirmed: the protocol has formally broken away from its stateful past toward a stateless model, and the authorization requirement is now codified. What remains open is adoption across existing implementations and the migration away from the experimental Tasks API from 2025-11-25—not the status of the specification itself. The advice from July—start auditing now, because postponement means accumulating technical debt—remains fully valid.

## Confirmed: the price war became structural, faster than expected

The [article on the price war](https://radar.llmnet.nl/en/model-prijzenoorlog-juli-2026) argued that the decline in API prices was no longer a temporary marketing stunt, but a structural phase made possible by inference engineering. That assertion has been confirmed twice over in two weeks. First, Anthropic launched Claude Opus 5 on July 24, which comes close to Fable 5 intelligence for roughly half the price—a frontier model deliberately positioned on price rather than quality alone. Second, OpenAI sharpened its pricing even further (reference date August 24, 2026). Whereas in July the pricing ladder was still Sol at $5.00, Terra at $2.50, and Luna at $1.00 per million input tokens, on July 30 OpenAI slashed GPT-5.6 Luna by 80 percent to $0.20 (output to $1.20) and GPT-5.6 Terra by 20 percent to $2.00 per million input tokens (output to $12.00), while flagship Sol remained at $5.00 input and $30.00 output. This current state of affairs puts even greater pressure on the lower end of the market than the July article predicted. What has also materialized as a result is the implicit prediction from [the article on token savings](https://radar.llmnet.nl/en/token-besparing-juli-2026): model routing remains the biggest financial lever for developers. That article demonstrated that the difference between DeepSeek and an expensive frontier model per million input tokens is so substantial that routing offers the first saving you can implement without compromising quality (reference date August 24, 2026: DeepSeek V4-Flash cost a flat $0.14 per million input tokens cache-miss until August 16, but has since adopted a peak rate of $0.44 and $0.22 off-peak, while V4-Pro costs $0.66 off-peak and $1.32 during peak hours, with peak hours running from 01:00-04:00 and 06:00-10:00 UTC Monday through Friday; the $0.44 figure therefore applies only to the peak rate of V4-Flash). The [explanation of model routing on the api subdomain](https://api.llmnet.nl/en/model-routing) provides the practical implementation for this.

## Partially materialized: observability consolidation, but the wave is not over yet

The [July article on AI observability](https://radar.llmnet.nl/en/ai-observability-juli-2026) signaled a wave of consolidation: major players are dividing the playing field, self-hosted alternatives are gaining ground, and the choice of an observability stack is strategic. The article described how ClickHouse, with its January 2026 acquisition of Langfuse, solidified its position as the standard for analytics databases, while Langfuse itself reported processing over 10 billion observations per month. That consolidation has partly materialized — the acquisition stands, the MIT license and the option for self-hosting remained intact, and the market has not reverted to the proliferation of small SaaS tools. But the most significant new signal lies elsewhere: MCP is gaining built-in OpenTelemetry support, meaning that observability for agent conversations will soon be a protocol standard rather than a custom-built add-on. That was not a prediction from July, but a development confirming the direction of the July article. The verdict is therefore "partial": the consolidation itself has been realized, but the next phase — observability as a standard component of the tooling layer — has yet to prove itself. Anyone looking to track costs per user will find in [the api article on allocating costs per user](https://api.llmnet.nl/en/kosten-per-gebruiker-toerekenen) a concrete first step.

## Still open: signals with a longer horizon

Three signals from July have deliberately not yet been settled. First, Gemini 3.5 Pro: the rumor that Google is rebuilding its flagship after failed tests has been confirmed by the lack of a release on July 21, but there is no new date — a delayed signal is not yet a realized signal. Second, Anthropic's IPO: investor meetings in mid-July point to a possible IPO in October, but that date lies beyond this retrospective and an IPO can always be postponed. Third, the EU AI Act: the transparency obligation took effect on August 2, but the question of whether it will see enforcement in practice can only be answered months from now. These three remain on the radar, and the next retrospective will evaluate them again.

## Partly realized: token saving as an ongoing practice

The [article on token saving](https://radar.llmnet.nl/en/token-besparing-juli-2026) was not about a single announcement but about five community techniques, which makes evaluating it different from a product release. The technique highlighted in the article as the greatest lever—dynamic model routing—has since been doubly confirmed: the RouteLLM framework from the article demonstrated up to 85 percent cost savings while retaining 95 percent of quality, and recent weeks' price movements (Opus 5 at half price, Luna as a budget tier) make price disparities between models wider rather than narrower. The remaining four techniques—shortening context, caching, smaller models for routine tasks, and deliberately omitting unnecessary data—are not one-off events but ongoing habits; they can only be verified against your own billing, not the news. The verdict is therefore "partially realized as a signal, lasting as a practice": the claim that routing yields the largest savings is confirmed, but whether readers actually adopt the five techniques is a behavioral question rather than a news item. Anyone looking to measure the savings firsthand can use the approach that [is on the benchmark subdomain for latency measurements](https://benchmark.llmnet.nl/en/latency-percentielen-meten): measure before and after a change using the same setup, and judge based on the numbers rather than gut feeling.

## What this means for upcoming editions

The first retrospective yields three lessons that sharpen future editions of the radar. First: predictions improve when formulated in a testable way. A signal with a clear horizon and observable criteria ("available within two weeks") can be held accountable; a signal without a date cannot. Second: the quality of the evaluation stands or falls with sources—this article relied on release announcements and early August monthly summaries, and every verdict points to a dated event. Third: a retrospective is not a scorecard, but a corrective mechanism. The signals that materialized (K3, MCP, pricing) confirm July's selections; the signals still pending determine what will be worth watching in the coming weeks.

## Conclusion

Of the five evaluated July themes, three materialized and two partially did. Kimi K3 was released and indeed proved too large for local hardware; MCP shifted on schedule to stateless with mandatory authentication; and the price war proved structural, with Claude Opus 5 as the strongest evidence. Observability consolidation was partially realized, but the next phase—observability as a default in the protocol—has yet to prove itself, and token reduction was confirmed as a key lever but remains an ongoing practice rather than a completed milestone. That is a strong outcome for a first edition: the radar doesn't just make predictions, it keeps itself accountable. The next retrospective will arrive in a month, covering longer-horizon signals for Gemini, Anthropic, and the EU AI Act.

[Radar](/) — part of the llmnet.nl knowledge network on AI and LLMs.
