;

Mapping the frontier

Models can now support context windows exceeding 1 million tokens, sufficient to embed entire codebases, knowledge bases, or conversation histories in a single request.

Despite well publicised recent calls by the figureheads of Anthropic and OpenAI, to ‘pace the frontier’, the landscape of Frontier LLM remains extremely dynamic and indicates no signs of slowdown. Leading models—including Qwen 3.8 max, DeepSeek V4-Flash-0731, Inkling-Small, and Gemini 3.7 Flash – have shifted the competitive axis from raw capability to cost, latency, and fit for specific operational constraints.

Long-horizon execution—the ability to maintain goal-directed action across extended sequences of steps remains the highest-value capability for enterprise agents. Models can now support context windows exceeding 1 million tokens, sufficient to embed entire codebases, knowledge bases, or conversation histories in a single request.

This is where the criticality of multimodality comes into play. Every major frontier model accepts text, image and document inputs natively and this has had a considerable impact on cost and price efficiency. Open-weight models (such as DeepSeek or Qwen3) now close the capability gap with proprietary systems while enabling self-hosting for organisations with data sovereignty requirements.

Let’s take a closer look at some of those below:

Qwen3.8-Max:

Vendor-stated; requires independent benchmark validation

Reported by Alibaba as a 2.4-trillion-parameter mixture-of-experts model with around 95 billion active parameters, a one-million-token context window and improvements across coding, professional work, research and long-horizon agentic tasks. Several benchmark claims are vendor-reported and should be treated as requiring independent validation.

DeepSeek V4-Flash-0731:

Vendor-stated/reported; claims should be treated cautiously

Positioned as a re-post-trained 284-billion-parameter mixture-of-experts model with roughly 13 billion active parameters and DSpark speculative decoding support. Published sources describe improved agentic and coding performance, but the strongest benchmark comparisons are vendor-stated and should be presented with caution.

Inkling-Small:

Reported signal; requires primary-source confirmation

A 276-billion-parameter open-weights mixture-of-experts model from Thinking Machines Lab, with 12 billion active parameters, native text, image and audio input support, variable thinking effort and a context window of up to one million tokens.

Gemini 3.7 Flash:

Vendor-stated; some external reporting repeats Google’s claims

Google says Gemini 3.7 Flash is its “most intelligent workhorse model yet for coding and agents”, with improvements in software engineering, web development, knowledge work, multi-step planning and tool calls. This is a useful signal that frontier-model competition is shifting from chatbot quality toward autonomous task execution.

Frontier AI safety: why the off switch still works

Our Global Head of Cybersecurity, Davey McGlade, did a deep dive on Frontier AI safety, exploring recent AI agent incidents and posing a simple but critical question:

'Should organisations be focusing less on slowing AI down and more on how to contain it when things go wrong?'

He highlights why controls like compute limits, least privilege, segmentation and automated circuit breakers deserve just as much attention as AI safety guardrails. Read the full piece here:

Read blog

Agent infrastructure

Previous page

Security and governance

Next page