Most, if not all, large-scale manufacturers are juggling a plethora of systems in their digital backbones. Often, these will all be operating at once on the shop floor: a predictive maintenance agent watches vibration sensors for safety patterns, a vision-based quality control agent flags product defects and a scheduling agent juggles changeovers as an inventory agent keeps a tab on procurement. Each has a clear operative task, but more often than not, these systems are not talking to one another.
It’s that failure to collaborate that undermines agentic AI in large-scale manufacturing. Whether an agent functions or not is one thing, but another significant issue at hand for large manufacturing companies is that they are functioning in almost total isolation.
And, while the number of large-sized organizations scaling AI agents in at least one function has risen from 27% to 40% year over year, they are honing in on specific functions such as IT and software engineering instead of evenly distributing agents across the operation as a coordinated system. Meanwhile, although the majority of organizations plan to deploy autonomous agents, only about a third believe their infrastructure is ready to support them. What’s more, research from 2025 found that 79% of multi-agent LLM system failures stem from issues surrounding coordination and specification.
Manufacturers are adding autonomous decision-makers to the floor and production line faster than they’re building the connectivity between them. This is exacerbating into a coordination and collaboration problem that undermines how these models perform. Ultimately, this shows up in three distinct parts of the AI stack. Here’s where, and what manufacturers should do to ensure collaborative, well-connected agentic systems.
Collaboration’s Biggest Friction Points
The first and perhaps biggest barrier is data, which continues to be a significant sticking point for many large organizations looking to scale AI. Most manufacturers are running on decades of accumulated infrastructure (and data tied to that) in the form of MES and ERP systems that do not share a common schema with other streams of data such as communication and human resources. Because of this, agents struggle to communicate with each other, if at all; isolated or poorly formatted data poses an almost unbreachable barrier between multiple agents.
This pans out on the factory floor as an agent monitoring machine health and performance on one line not having any reliable way of knowing what the scheduling agent two systems over plans for the coming changeover, even though the two decisions and workflows directly affect each other.
The next friction point is protocol. It’s one element that does not get enough consideration when approaching AI collaboration. Agentic AI is very new, and until recently, there was no real standard way for agents built on different frameworks, foundations, or vendors to communicate and collaborate.
Fortunately, that is starting to change. For example, Google’s Agent2Agent (A2A) protocol, which launched last year, standardizes how independent agents (potentially built on entirely different platforms) discover each other’s capabilities and delegate tasks. Crucially, this is explicitly designed to complement Anthropic’s Model Context Protocol (MCP), which standardizes how each agent accesses tools and data sources. The common goal here is to overcome isolation and ensure collaboration, even among multi-agent settings, at scale.
Finally, many manufacturers are still contending with the human oversight question. Even a perfectly coordinated agent network will fail if the people overseeing it do not trust its decisions or have unrealistic expectations surrounding its capabilities. Analysts from MIT Sloan have raised warnings around inflated expectations and eroded confidence surrounding agentic AI, alongside a potential threat of human disengagement with AI agents. On the shop floor, these gaps and friction points can be catastrophic: the human thread connecting the agentic ecosystem is absolutely vital to safe, reliable operations.
Bridging the Trust Gap
It’s tempting to treat the human side of this as a purely cultural issue that can be solved with more training sessions. However, these efforts will fall short unless they are coupled with durable architectural fixes. Oversight design must account for when human judgment should defer to the system, not just defaulting to overriding the agent’s suggestion because the person on the loop disagrees with it.
Build in reasoning transparency, which means there is always visibility and the ability to break down why an agent made a decision or flagged a suggestion beyond simply observing the output. And, when operators disagree, ensure there is a clear and efficient escalation path that enshrines transparency and traceability surrounding decision-making processes.
Rebuilding a Sense of Wholeness
Agents, much like the teams they are built to support, need more than access to a narrow, specific task. They need what amounts to a sense of wholeness; a working model of how their piece of operation connects to everything else. A scheduling agent optimizing changeover time without any visibility into what the maintenance agent sees about aging equipment is only making a somewhat acceptable decision that might not be the right one when put against other workflow demands.
This is where the protocol layer and the sense of wholeness problem converge. Shared context is the practical means by which an agent’s individual task is anchored to the wider plant’s actual state. In manufacturing, there is an existing partial answer to this in the form of standards such as ISA-95, which define common means for how enterprise and control systems essentially communicate across a plant. The novelty now, though, is extending this idea of a shared operational vocabulary from human-readable integration into information that agents can process.
Without that shared model, every additional agent deployed adds a new isolated point of reasoning and decision-making that is not part of a coherent whole. That is exactly how agent sprawl worsens instead of resolving itself as adoption scales.
Keeping a Pulse on the Right Measurements for Success
Another trap is simply measuring agentic AI the same way any other software rollout is monitored: focusing on efficiency numbers and making sure they keep increasing. For coordinated and collaborative multi-agent systems, this is not necessarily beneficial and can actually backfire when it becomes the only metric of performance.
The first thing worth confirming is not simply how much money an agent saves but whether it is safe and secure to operate as designed. Ensure that it acts within its intended boundaries, has predictable failures, and allows for human intervention and escalation accordingly. Only once that is settled should quality take the center stage, such as whether the agent’s output is actually correct and consistent and not just instant.
This is the order of priority that manufacturers should follow. Speed, throughput, and cost savings are the lowest priorities. A productivity gain built on top of an unsafe or unreliable agent is a liability to teams, technology, production lines, relationships, and business reputation.
Test at a small scale before going larger. A single agent handling a single well-defined task can yield a clean before-and-after comparison. However, when the lens gets widened to encapsulate dozens of interacting agents across a whole plant before any have been validated individually, it’s much more difficult to determine which part of the system is driving the result. Complexity should be added to the measurement approach only at the same pace it is added to the deployment itself.
Here, collaboration and measurement start to co-exist. An agent’s output will only be as trustworthy as the context it had access to when it acted. This means that the aforementioned shared-context problem is actually a precondition for measurement. Manufacturers who have not solved for a unified operations picture across systems have no reliable way of knowing whether a positive or negative result reflects the agent’s judgment or a gap in what it is allowed access to. The metrics that matter are readable when the sequencing and shared context are locked in.
Coordination and Collaboration at the Center of Capability
Ultimately, large manufacturers’ structural advantages (their massive production volumes, historical datasets, process consistencies, and talent and technology resources) are undoubted. These advantages, however, only compound if manufacturers first resolve the collaboration problem.
Looking to the future, as AI capabilities evolve to possibly total autonomy, and both systems and production lines grow, this collaboration and coordination are an absolute necessity. Treating interoperability, protocol standards, shared context, calibrated human oversight, and a sense of wholeness as architectural prerequisites ahead of scaling and deploying is a must. It’s the dividing line between manufacturers who scale agentic AI successfully and safely, and those who are stuck in pilots.
Credit: Source link


























