AI Published 9 min read

There is no best AI model: what actually matters in 2026

OpenAI, Claude, Grok, Gemini and increasingly capable open-weight models are taking different paths. What should businesses compare, and what do computer use, compute limits and the AGI debate really mean?

AI model trade-offs and compute demand displayed in a data-centre workspace

A few years ago, asking which AI model is best sounded like a sensible question. In 2026 it has become the wrong one. OpenAI, Anthropic, Google, xAI and a growing group of Chinese and open-weight developers are all producing extremely capable systems that are not optimised for exactly the same work. One may be stronger at long reasoning, another faster and cheaper at high volume, another interesting because it can operate software, and another because a company can run more of the stack itself.

For a small or medium-sized business, that is good news. You do not need to choose an AI brand and remain loyal to it. You need to understand your own work well enough to choose the right balance of intelligence, speed, reliability, control and cost for each job. That is increasingly how we think about models when building with AI at ARKARA: which logo sits at the top of a benchmark this month matters far less than whether the model does the actual job reliably inside the system around it.

The strongest model is rarely right for every job

The top of the market is moving unusually fast. OpenAI released GPT-6 Astra on 3 September 2026 and highlighted computer use as a central capability. OpenAI, GPT-6 Astra Anthropic shipped Claude Fable 5.1 on 1 September. Anthropic, models overview Google made Gemini 3.8 Flash generally available on 2 September. Google, Gemini API changelog xAI’s Grok 4.6 arrived in August with a 500,000-token context window. xAI, models An article declaring a permanent winner among those systems would age faster than the article itself.

Price is a more durable part of the comparison. GPT-6 Astra lists at $10 per million input tokens and $50 per million output tokens, while cheaper models in the same provider’s catalogue cost a small fraction of that. OpenAI, pricing Anthropic and xAI also price different capability tiers very differently. If a business needs to classify thousands of routine enquiries, extract standard fields or generate short internal summaries, maximum reasoning on every request can be wasteful. The reverse is also true: using a weak model for a genuinely hard task can be a false economy once retries, human correction and expensive mistakes are counted.

We would therefore start with the workload rather than the brand. How difficult is the reasoning? What happens if the answer is wrong? How much volume passes through the system? How fast must it respond? Does it need tools, long context, vision or computer use? Those questions often matter more than a few points on a benchmark.

The providers are becoming different tools

The major providers increasingly have distinct strengths, even if they change with each release. OpenAI has pushed toward reasoning, tools and work across applications. Anthropic’s Claude family is prominent in coding and agent work, and its computer-use tooling left beta in August 2026. Anthropic, release notes Google combines Gemini with a broad search, productivity and cloud ecosystem, while xAI positions Grok 4.6 for coding and agentic work. xAI, release notes

Those are not permanent labels. The useful comparison is to run the same real task across several candidates and measure quality, failure rate, latency, tool reliability and cost. Benchmarks can help form the shortlist; they should not make the purchasing decision.

Open weights change who controls the safeguards

Capable open-weight models have changed the discussion because model choice is no longer only about which hosted API to call. DeepSeek published DeepSeek-V4-Pro under the MIT licence in August 2026, while families such as Alibaba’s Qwen and Moonshot’s Kimi have continued releasing large open-weight models for coding, reasoning and agentic work. DeepSeek-V4-Pro model card The attraction is control: an organisation may be able to run a model on infrastructure it chooses, keep more data inside an environment it governs, adapt the system to a specialised domain and reduce dependence on one provider.

That freedom has a safety tradeoff. The restrictions users experience in a hosted service are not all contained inside the model itself. They also come from the surrounding harness: policies, classifiers, monitoring, access controls, tool permissions and the provider’s ability to suspend or change a service. When an operator controls both the model weights and the surrounding system, those restrictions can be modified far more deeply than a normal hosted-service user could modify them.

This is not merely theoretical. OpenAI has published research in which it deliberately fine-tuned an open-weight model to remove refusal behaviour so it could estimate worst-case frontier risk before release. OpenAI, Estimating worst-case frontier risks of open-weight LLMs The International AI Safety Report 2026 makes the balanced version of the same point: open weights support research and innovation, while making safeguards easier to remove, use harder to monitor and released weights impossible to recall. International AI Safety Report 2026

That does not make open models a bad idea. Concentrating advanced AI inside a few private companies creates another risk because they control access, policy and price. Open models distribute capability and control, including misuse risk; closed models allow stronger central enforcement while concentrating power. The right choice depends on what an organisation needs to protect and what responsibility it can take on itself.

Closed frontier AI has an access question too

AI feels unusually accessible today. Consumers and small businesses can use systems capable of coding, research and tool use for the price of ordinary software subscriptions, sometimes for free. It would be risky to assume that the highest end of AI will always be priced that way.

The frontier is expensive to build and operate, while providers are competing aggressively for users and developers. Published list prices already show a useful fact: within the same provider, the highest-capability model can cost many times more per token than a cheaper tier.

If advanced reasoning remains compute-intensive while providers eventually need sustainable margins, the market may separate. Efficient models could become very cheap for routine work, while long-running frontier agents and large compute budgets become premium resources. Competition, open models, specialised hardware and better inference could push the other way. The point is simply that today’s highly competitive prices do not prove unlimited frontier intelligence will remain equally affordable to everyone.

Computer use expands the territory of automation

Computer use matters because it reaches software that APIs do not. If a CRM exposes a reliable API, that remains the cleaner integration. But many businesses still depend on portals or older systems with poor APIs or none at all. A computer-use model can work with the graphical interface a person already uses, reading the screen, clicking and typing.

That can unlock processes that were previously stuck as manual work, but it is less predictable than an API. Interfaces move, login sessions expire and unexpected states appear. Both OpenAI and Anthropic ship computer-use capabilities with guidance around isolation, allowed actions, step limits, outcome verification and user control for hard-to-reverse actions. OpenAI, computer use Our view is that computer use will not replace APIs. It will extend automation into places where a dependable integration does not exist.

The next bottleneck may be physical

Software can improve in months. Semiconductor fabs, advanced packaging, data centres, transformers, transmission lines and new electricity generation operate on much longer timelines. That difference matters as AI demand shifts from short chat responses toward reasoning systems and agents that can consume far more compute over longer periods.

The International Energy Agency estimated that data centres consumed around 415 TWh of electricity in 2024, about 1.5% of global electricity use, and projected consumption to more than double to roughly 945 TWh by 2030, with AI a major driver. It also warned that around 20% of planned data-centre projects are at risk of delay and that wait times for critical grid components such as transformers and cables have doubled in three years. IEA, Energy and AI The scale of investment is equally striking: Stanford’s 2026 AI Index recorded $285.9 billion of private AI investment in the United States in 2025. Stanford HAI, 2026 AI Index

That creates a plausible mismatch: model capability may improve faster than enough compute and electricity can be built economically. If so, efficiency becomes as important as raw intelligence; a model reaching almost the same result with half the computation may be more valuable than one winning a benchmark by a small margin.

The counterargument is that infrastructure is expanding and hardware and algorithms can improve dramatically. We should not assume the bottleneck permanently outruns the industry. Yet cheaper computation can also increase total demand: if an agent becomes ten times more efficient and companies run one hundred times more agents, power use still rises. Future AI prices will depend on chips, energy, grids and efficiency as well as smarter models.

Are we entering the AGI era?

AGI, Artificial General Intelligence, broadly refers to a system with general rather than narrow intelligence that can learn, reason and apply knowledge across a wide range of tasks. There is no universally accepted definition and no agreed test that tells us when it has arrived, which is why researchers have proposed frameworks of levels rather than a single threshold. Morris et al., Levels of AGI

The optimistic view is that increasingly general models, tool use, coding, multimodal reasoning and computer operation show rapid movement toward that threshold. The opposing view is that broad competence and strong benchmark results are not the same as general intelligence: today’s systems still fail in surprising ways and depend heavily on the scaffolding around them.

For business, the definition may matter less than the capability. A system that can understand a goal, find information, use tools, operate software and complete hours of useful cognitive work can change a company whether or not researchers call it AGI.

Even if a laboratory builds something broadly accepted as AGI, unlimited access would not follow automatically. We would still have to ask who controls it, what safeguards surround it, how much compute it consumes and what that compute costs. If frontier intelligence remains compute-intensive, access could be shaped by capital, chips and electricity as much as by software.

That is why we do not think the future is one model defeating every other model and becoming the permanent answer. The likelier outcome is a layer of models with different strengths, prices and degrees of openness underneath products that users increasingly judge by results rather than by model name. The durable question is not which provider is winning this month. It is what work we want software to become capable of doing, how much responsibility we are prepared to give it, and what it will cost to provide that capability reliably.