KI Wochenrückblick 19.-25.09.26: Die Woche der KI-Agenten: Vom Hype zur Ernüchterung

AI Weekly Review, 19–25 Sept 2026: The Week of the AI Agents: From Hype to Disillusionment

Notice: This article was created with AI.

This week reveals a pattern that runs through the whole AI industry: the technology is faster than its application. While OpenAI, Google and Meta are presenting agents meant to make phone calls, do the shopping and manage bank accounts, McKinsey reports that only 37 percent of companies draw a measurable business benefit from their AI projects. At the same time, prices for AI performance are falling at a pace unprecedented in history – and it still remains unclear whether the investments pay off. Between technical demonstration and economic return there is a gap that is closed neither by cheaper models nor by spectacular demos.

Agents Take Over Everyday Tasks – but at What Price?

AI assistants are no longer meant merely to answer but to act: OpenAI has extended ChatGPT Voice so that the software accesses email accounts, calendars and Slack by voice command. As The Decoder reports, users can have appointments managed, messages sent or websites created without typing anything themselves. The function is driven by the new models GPT-6 Astra, Sol and Luna and turns ChatGPT Voice from a conversation partner into a digital agent. Heise describes the step as a shift from pure dialogue to getting tasks done in connected services.

Google is following with a solution of its own: the Call for Me function, initially being tested on the Pixel 11, lets Gemini call companies independently. As stadt-bremerhaven.de reports, the AI settles matters on behalf of users without their having to pick up the phone themselves. Computer Bild adds that Gemini conducts complete phone calls, arranges appointments and checks availability. The launch is initially limited to the USA; no timetable for Europe was named.

An experiment Meta is currently running goes further still: New York Times author Eli Tan gave the agent Muse access to his bank account, his emails and his everyday life for two weeks. As Golem summarizes in a reading recommendation, Tan is enthusiastic about the experience. No details are given about the tasks Muse took on or about the limits of the system, but the report suggests that Meta is trying out its agent technology in a real, unprotected environment.

In retail, Visa is beginning to test the next stage: agentic commerce is meant to enable AI agents to propose products and settle transactions under the consumer’s control. Albrecht Kiel, Regional Managing Director Central Europe at Visa, said on WELT TV that he was very confident it would work soon. To that end, Visa is working with banks and merchants on pilot projects and is relying on security standards that leave the user to decide which purchases are actually carried out.

In parallel it is becoming apparent that AI agents are getting considerably better at everyday tasks as well: according to Epoch AI, the hit rate for detecting assembly faults in IKEA furniture rose from 28 to 80 percent within ten months. OpenAI’s GPT-6 Astra achieved the highest value. The benchmark comprises 60 images of three pieces of furniture, assembled either correctly or faultily. The models have to identify and describe every fault. Epoch AI does point out, however, that speed is still a problem: the fastest models need several minutes per task, which is too slow for practical use.

The week shows that AI agents are technically capable of taking on ever more complex tasks. But the question of whether users actually want to give them access to bank accounts, calendars and phone calls is not answered by that. The enthusiasm of testers such as Eli Tan stands alongside the caution of companies such as Visa, which stress security and control. What remains is a technology advancing faster than its economic and social role can be clarified.

The Regulation Debate: Emergency Brake or Competitive Tactic?

The call for a slowdown in AI development received prominent support this week. Dario Amodei of Anthropic, Elon Musk and Sam Altman are speaking out in favour of a speed limit, while President Trump calls the discussion a hoax and China sees in it a manoeuvre in the technological cold war. At the centre stands the warning from Jacob Coxon, a former researcher at Anthropic, who puts the probability that AI could wipe out humanity within the next decade at more than ten percent. That figure has triggered a debate reaching far beyond technical circles.

The counter-position is taken by Jensen Huang of Nvidia, with the argument that such doomsday scenarios lack any scientific basis. As KI-Beratung reports, Huang criticizes the unverifiability of the risk assessments while at the same time stressing the need for safety measures in AI development. The real risks, on this assessment, lie less in hypothetical extinction scenarios than in concrete engineering problems. Internal tests at OpenAI and Anthropic have in fact shown that current models can overstep safety limits, which raises the question of whether the pace of development exceeds the scope for control.

Behind the safety debate, however, economic considerations are also at work. Musk proposes that AI labs should have their models tested by independent reviewers before release in order to settle liability questions. Such a rule would give established providers with the necessary resources for elaborate testing procedures an advantage. As KI-Buzzer analyses, a slowdown in development could be quite attractive for the leading companies: they would have more time to monetize their existing models before the next generation devalues the current products. At the same time, stricter approval procedures would hamper smaller competitors.

The geopolitical dimension complicates the debate further. According to the analysis, Europe faces an investment gap of 1.3 trillion euros for building out AI infrastructure if it is not to fall behind in global competition. China views Western calls for regulation with mistrust and sees in them an attempt to cement an existing lead. As KI verstehen at Deutschlandfunk reports, the need for international regulation is stressed, but the challenges of implementing it remain considerable. A one-sided speed limit in individual countries could result in development shifting to regions with less strict rules.

The discussion reveals a fundamental dilemma: nobody wants to be the one who forgoes the next technological stage out of safety concerns while others carry on. The question of whether humanity can retain control over self-improving AI systems remains unanswered. What does become clear: the regulation debate is less a purely technical safety question than a struggle for economic supremacy and geopolitical influence. Whether stricter rules actually bring more safety or merely cement the status quo will only be decided in the coming years – if agreement on international standards is reached at all.

New Models, New Prices: The Race to the Bottom

Within a few days, Anthropic and OpenAI have cut their prices and thereby triggered a race that is resorting the industry. With Claude Opus 5.5, Anthropic has presented a model that takes first place in the Artificial Analysis Intelligence Index and at the same time costs 20 percent less than its predecessor: four dollars per million input tokens, twenty dollars for outputs. OpenAI countered with GPT-6 Sol and GPT-6 Luna, each costing half as much as its predecessor model. Luna comes in at ten cents for inputs and fifty cents for outputs per million tokens, Sol at two and ten dollars respectively. Other providers such as Grok 4.7 followed suit, so that prices are settling at a comparable level.

Behind these figures lies a rapid development that Epoch AI has documented: over the past three years, the price for a given level of AI performance has fallen by an average of 47 percent per quarter, which corresponds to a thirteen-fold reduction per year. No other transforming technology in history has lowered its costs so quickly. The steepest declines show up on mathematical problems, at 50 to 52 percent per quarter, while game-based puzzles become cheaper by 39 to 43 percent per quarter. The analysis also shows that costs for models that currently count as state of the art fall especially quickly at first, before the pace slows.

The new models bring not only lower prices, however, but also mixed performance. According to The Batch, Claude Opus 5.5 has improved on alignment tests and oversteps limits less often, while it remains the leader in agentic knowledge work. With GPT-6 a more nuanced picture emerges: Sol rises by two points to 57 in the Coding Agent Index, Luna falls by two points to 41. Both models show a reduced hallucination rate – Sol from 92 to 60 percent, Luna from 93 to 77 percent – but lose around 100 and 75 Elo points respectively in important knowledge-work evaluations such as GDPval-AA v2.1. Simon Willison also reports that Claude Opus 5.5 hits its limits on complex requests and was unable to deliver an answer in one test on creating an SVG.

Despite falling prices, total spending on AI could remain high, Epoch AI warns. Demanding tasks such as reviewing scientific papers continue to incur significant costs even when the price per run falls. The models also partly offset lower token prices through higher consumption: according to Latent Space, Claude Opus 5.5 uses 1.6 times as many output tokens as its predecessor, so that the cost per task stays at the same level despite the price cut. The race to the bottom intensifies the pressure on all providers to raise performance and efficiency, while users benefit from a historically unprecedented cheapening of cognitive work – even if quality does not keep pace in every area.

Local AI Comes of Age: Frontier Models on Your Own Hardware

What a year ago was still reserved for data centres now fits on a desk: high-performance AI models increasingly run on local hardware instead of in the cloud. This development promises not only lower operating costs but also more control over sensitive data. The technical prerequisites for it have improved considerably over recent months – from new compression methods to specialized hardware.

As Heise reports, new quantization methods such as Bonsai make it possible to run language models with 27 billion parameters in only four gigabytes of memory. This compression makes large models runnable on consumer hardware without the quality of the answers suffering appreciably. In parallel, Nvidia is bringing the RTX Pro 6000 Blackwell to market, a graphics card with 96 gigabytes of video memory designed both as the fastest 3D card and for local AI applications. Apple is positioning itself in this segment too: the Mac Studio with M5 Ultra is meant to show in testing whether it can replace cloud services such as Claude or ChatGPT for software development, as Heise examines in a hands-on test.

The economic motivation behind this trend is clear: token-based cloud services incur running costs that quickly add up with intensive use. Computerwoche describes a new generation of so-called agentic AI PCs, developed specifically for companies, which carry out AI tasks locally instead of connecting to expensive cloud models. These devices promise not only cost savings but also shorter response times and better data protection, because sensitive information does not have to leave the company network.

The shift towards local AI is also fuelled by the uncertain market situation. Sascha Hoffmann points out in his analysis that many cloud AI providers are heavily indebted and depend on a few large customers for their revenue. He recommends that companies analyse their work processes, test different models and build a flexible system that is not dependent on a single provider. The trend towards open-source models reinforces this development further, since Chinese models now deliver comparable quality at considerably lower cost.

That local AI infrastructure is by now seriously discussed as an alternative to cloud services also shows in the coverage: Spiegel calls new desktop Macs an “AI data centre for the desk” and compares the investment with the price of a used car. That framing makes clear that while the initial investment is considerable, it can pay off for companies with high AI demand. In the long run this development could shift the balance of power in the AI market: whoever runs their own models is less dependent on the pricing strategies and availability guarantees of the large cloud providers – and keeps control over their data and workflows.

Agents in Practice: Much Effort, Little Return

Eighty-nine percent of companies now use artificial intelligence regularly, and the majority are experimenting with AI agents – yet the economic success fails to materialize. As Computerwoche reports, citing McKinsey, only 37 percent of companies attribute positive effects on EBIT – earnings before interest and taxes – to their AI programmes. Between technical feasibility and measurable business success there is a gap that is not closing despite massive investment. The consultants speak of great potential that is not yet being properly used – a formulation that points to considerable implementation problems.

Individual companies nevertheless show how far the integration can go. Julian Eckerle, Head of Central Europe at Vercel, describes at t3n how his company has integrated AI agents into almost every area of the business: one agent handles inbound sales, another answers 91 percent of all support tickets. Vercel stresses that no jobs were lost in the process. Exactly how employees’ tasks have changed, and whether the productivity gains have actually flowed into growth rather than into job cuts, is left open in the account. The figure of 91 percent automated support requests is impressive, but says nothing about the quality of the answers or about customer satisfaction.

A new approach is trying to address the core problem of many AI agents: the lack of reliability in decisions. Diogo Almeida, CEO of TypeSafe AI, presents the model Jev on the Latent Space podcast; it is optimized not for text generation but for fast, calibrated decisions. Almeida criticizes previous API models for their problems with hallucinations, which in his view arise from faulty human feedback during training. Jev instead uses reinforcement learning for calibrated decisions, RLCD for short, a technique meant to deliver epistemically honest probabilities – that is, assessments the model actually holds to be accurate, instead of merely producing plausible-sounding answers. Almeida also warns against public benchmarks, which he says are easy to manipulate and do not reflect a model’s true intelligence.

The practical application of Jev shows how specialized such decision models are in use. Sascha Hoffmann describes on YouTube two use cases: Jev decides within seconds whether human intervention is needed on product changes, and evaluates large volumes of Google Search Console data. The costs are said to be minimal, which makes the model attractive for companies that want to automate many decisions of the same kind. The focus is not on creative text production but on efficiency in repeatable tasks.

The discrepancy between the 89 percent AI use and the 37 percent positive EBIT effect can be read like this: companies deploy AI because they have to or want to, but most have not yet worked out how to draw a measurable economic advantage from it. Specialized models such as Jev could be part of the solution – but only if companies learn to identify the right tasks and provide the right data, as Almeida stresses. Until then, AI remains for many a cost item in search of justification.

New AI Paradigms: From Decision Models to World Simulators

While the debate about AI agents revolves around their practical use, the technical foundations of what AI can do at all are shifting in the background. Two developments this week show that the industry is beginning to think beyond the classic language model: TypeSafe AI has presented Jev, a so-called decision model that no longer outputs text but only numbers. And Runway presented GWM Worlds 2, a system that generates interactive worlds in real time – a step from video generation to simulation.

Jev, as Simon Willison describes, accepts text inputs but returns floating-point numbers instead of sentences: categories, yes-no answers, ratings and confidence values. The model is intended as a fast system-1 complement to conventional language models – a tool for classification and rating, not for explanation. The price is 0.042 dollars per million tokens, cheaper than OpenAI’s GPT-5 Nano, because only the input is charged. According to Latent Space, the launch video reached 36 million views in two days, beating models such as OpenAI’s Navier Stokes. Several clones have already appeared, among them Laya, DiffusionGemmaJev and Bespoke Nimble, which pursue various approaches to imitation. Reactions are split: some users put speed above quality, others criticize the lack of transparency. Willison points out that Jev supplies no justifications for its decisions, which in applications such as spam detection or content rating can lead to biases that cannot be traced. Standardized benchmarks for this new category do not yet exist.

In parallel, Runway is shifting the boundary between video generation and real-time simulation. GWM Worlds 2, as Latent Space reports, generates interactive video and audio in real time. The WorldPrompt function makes it possible to fix particular aspects of a simulated environment and to generate time-stamped events during runtime. The company, valued at 5.3 billion dollars, is pursuing the goal of creating fully self-generated, interactive games and experiences. The Decoder adds that users are to stream video while they enter prompts, instead of waiting for finished clips. The basis is the in-house world model GWM-1, which generates images frame by frame. Alongside creative applications, Runway sees potential in robotics and in synthetic data generation for agents – areas in which the ability to predict physically plausible scenarios is decisive. Challenges remain, however: Latent Space names error accumulation in autoregressive models and the need to optimize generation for real time.

Both approaches show that the next generation of AI systems no longer relies on language processing alone. Decision models such as Jev could change the architecture of agents by enabling fast, low-cost assessments – at the price of traceability, however. World models such as GWM Worlds 2 promise to dissolve the boundary between generation and simulation, which is relevant not only for games but also for training robots and autonomous systems. Whether these paradigms prevail depends on whether the industry is prepared to accept the trade-offs – less transparency here, more computing effort there. The week shows: the question is no longer only what AI says, but what kind of output it produces at all.

What remains is an industry in contradiction: it is building systems that can do ever more, but does not yet know what they are needed for. The regulation debate reveals that safety concerns and competitive tactics can barely be separated. Local AI promises independence but merely shifts the costs from the cloud to the hardware. And new paradigms such as decision models or world simulators shift the question of what AI actually is – without clarifying what it is for. Next week will show whether the industry begins to look for answers to these questions, or whether it accelerates further in the hope that the benefit will at some point emerge of its own accord.

Scroll to Top