Agent infrastructure is the new battleground
The Model Context Protocol (MCP) emerged in November 2024 as the first standardised interface for connecting AI agents to external systems.
It was originally unveiled by Anthropic as an open standard and was subsequently adopted by OpenAI, Google, and other major AI players. The reason for this widespread adoption is obvious: MCP is what many industry observers refer to as a universal connector or the 'USB-C for AI applications'.
MCP has become a shared integration layer, connecting AI agents to enterprise tools, data sources and external systems. Adoption is widespread, thousands of MCP servers now exist, but as the number of connected tools proliferates in this rapid expansion, loading all tool definitions upfront and passing intermediate results through the context window slows down agents and increases costs.
Increasingly, it appears as if multi-agent orchestration is moving toward more structured production workflows. Recent studies indicate that stateless agents – and LLMs are stateless by design – relying on a central LLM orchestrator have limits to their efficiency when coordinating across multiple tools and data sources. However, standardised frameworks for multi-agent coordination remain in their infancy.
Some of the hyper scalers, like Microsoft for example, have structured teams of planner, architect, QA and executor agents. This supports a move away from a single monolithic agent toward more deterministic, role-based workflows. Here are some instructive examples of agent infrastructure:

DeepSeek Harness
Vendor-stated; external coverage corroborates release
DeepSeek Harness is a strong agent-infrastructure signal: it is described as an open-source, plugin-first agent framework built around tools, skills, sessions, sandboxes, scheduling, loops and storage. These are core building blocks for agent systems rather than simple chatbot interactions.

Nvidia’s agentic stack
Reported signal; vendor-stated if sourced from NVIDIA materials
Nvidia’s Nemotron 3.5 Lightning and NeMo Switchyard point to agent orchestration as a competitive layer. NVIDIA describes the stack in terms of routing work across models to optimise cost, speed and capability; the escalation-router concept is especially agentic, but should be framed as an orchestration design pattern rather than a fully validated production outcome.

Muse Glimmer and local agents
Vendor-stated; secondary reporting consistent
Meta’s Muse Glimmer is notable because it is described as an open-weight model for local agent workflows, including tool calling, coding assistance, working with screenshots and documents, and recovery from failed tool calls. Broader examples such as scheduling, drafting messages and organising files should be presented as the kind of local-agent workflow this direction enables, rather than confirmed product guarantees.
AI Labs Spotlight:
Can Counterfactual AI Make Visual Inspections More Explainable?

Read Filippo’s full deep dive here: LLMs vs Machine Learning and Deep Learning for Structured Data | Version 1
Large language models (LLMs) have made extraordinary progress in recent years. They write code, reason through complex problems, summarise documents and handle tasks that would have previously required significant specialist effort. It is no surprise that teams are now considering whether they can be applied to the kinds of structured data problems that have traditionally been the domain of machine learning and deep learning.
It is a legitimate question and an important one. LLMs are increasingly capable and the boundary between what they can and cannot do is shifting at pace. But capability in general does not always translate to the right tool for a specific problem. The best way to answer the question is to run the experiments.
Our AI Labs team recently demonstrated how counterfactual explanations can be applied to image classification problems.
Using CF Proto, the approach identifies the image regions that influence a model's decision and generates a counterfactual example that would lead to a different classification outcome. While originally demonstrated using a manufacturing quality inspection scenario, the team noted that the same explainability approach could be applied anywhere image-based models are used. Discussion during the session highlighted that the primary value lies in understanding model decisions rather than changing real-world outcomes.