Hands-on agent testing - the management burden
I have spent a fair amount of time (and $$$) recently testing and validating a range of LLM and agentic systems, including Claude Cowork and Claude Code, Gemini and Spark, Hermes Agent from Nous Research, MuleRun, OpenClaw variants such as Pokee and PokeeClaw, and tools including Kimi and Qwen.
The productivity gains are real. Research that once took hours can be compressed into minutes, writing and analysis can begin from a far stronger first draft, and complex information can be processed at a speed that was difficult to imagine only a few years ago.
Yet the more I use these systems, the more I find myself thinking about the management layer they require.
An agent still needs a clear brief, access to the right data and active supervision, while its outputs need to be tested for accuracy, logic, tone and format. Its decisions and assertions need to be questioned, its integrations need to be configured, and the credentials or secret keys that enable useful access need to be managed securely. When the system drifts, someone has to recognise it and bring it back.
At times, improving my personal productivity feels as though it is becoming a full-time job.
For leaders, this raises a broader economic question. If the time saved in execution is absorbed by briefing, monitoring, correction and governance, then automation may shift work rather than remove it. Unless organisations redesign roles, controls and decision rights around these systems, the promised benefit can disappear into an expensive new layer of oversight.
I do not see that as an argument against the technology. If anything, it offers an early view of how deeply the technology may reshape organisations once reliability, integration and control improve.
The workforce impact will not be limited to people doing the same jobs faster. Some roles will narrow, others will expand, and new management responsibilities will emerge as people learn to direct, challenge and govern digital workers alongside human ones. The organisations that gain most will probably be those that treat this as an operating-model change rather than another software rollout.
The potential is already clear, although the economics remain less settled. What we are seeing now is not the finished model of work, but the beginning of a much larger redesign of work and the workforce around it.