Can an AI Chief of Staff Build Its Own Team?
Inside our experiment with an AI Chief of Staff that designed and delegated work to specialist agents, including what broke and what we will test next.
Jordan Whiting
Founder and CEO, DataMust · 17th July 2026
Experiment
Agentic organisation design and autonomous task delegation
We gave one AI agent a broad business objective, context about DataMust and permission to build the specialist team it thought it needed. It designed a marketing function surprisingly well. Keeping that team working over a long chain of tasks was the harder part.
Lab summary
- Status: In progress
- Question: Can one AI Chief of Staff interpret a business goal, create a specialist team and coordinate the work with a human setting direction?
- Hypothesis: A well-contextualised strategic agent can design and delegate the work without every role and task being specified in advance.
- Stack: Hermes Agent, Multica and multiple LLMs
- Early observation: The system was stronger at organisation design than dependable, long-running orchestration.
- Next test: Move the organisation into a persistent cloud environment and separate strategic reasoning from faster execution work.
The question
What happens if, instead of building individual AI agents and manually assigning each one a job, you start with a single AI Chief of Staff?
Give it a business goal. Give it context about the company. Give it access to tools. Then give it the ability to create and manage its own team.
Could it work out which capabilities it needs, create specialist agents, delegate work between them and allow a human to operate primarily at the direction-setting level?
That is what I have been testing at DataMust.
How the experiment was set up
The starting point was an AI Chief of Staff with context about DataMust: our business goals, brand, strategy and existing technology environment.
Rather than prescribing exactly how it should achieve an objective, I wanted to see whether it could work out an appropriate structure and plan for itself.
The initial objective was deliberately broad:
Build a marketing function for DataMust.
The Chief of Staff could create additional agents with their own roles and responsibilities. Hermes Agent provided the agent environment. Multica acted as the project-management layer, making it possible to see tasks being created, delegated and moved between the emerging team.
The intended human role was closer to a business owner than an operator: set the direction, observe the work and intervene when the system needed to change course.
What happened
Structurally, the system understood the assignment.
The Chief of Staff created an AI Director of Marketing. That agent examined DataMust's strategy and brand context, then identified the capabilities it believed a functioning marketing department would need.
From there, the organisation expanded. Content needed to be created. Brand consistency needed an owner. Website content needed to be published. Other roles appeared to coordinate campaigns, create assets and manage parts of the workflow.
Before long, the system had designed something approaching a complete marketing department with roughly ten roles.
For a small consultancy, ten human roles would be excessive. Creating another specialist agent is comparatively inexpensive, but every new role still adds coordination, review and context-management work.
The first experiment therefore raised a question rather than answering one: should an AI organisation copy the lean structure of a small business, or can greater specialisation improve the work enough to justify the orchestration overhead?
Why specialist roles changed the output
One of the more useful observations had less to do with autonomy and more to do with how roles shape an LLM's response.
My own background is weighted towards data, technology and systems. When I use a general-purpose model directly, my questions naturally come from that perspective.
The AI Director of Marketing approached the business from a different professional frame. It used marketing structures, terminology and assumptions that I would not naturally reach for first. It surfaced considerations I had missed and exposed gaps in how I was thinking about the function.
The specialist agent did not have access to a different underlying intelligence. Its role, context and objective encouraged it to explore a different part of the model's knowledge.
That may be one of the more useful reasons to create specialist agent teams. The value is not only task division. A carefully defined role can change the questions the model asks and the knowledge it brings forward.
Where the system broke down
The biggest issue was not the quality of the planning. It was the feedback loop.
The experiment ran locally through Hermes Desktop, with Multica providing the cloud-based project-management layer. We were also testing a slower, reasoning-heavy model.
A Chief of Staff interprets the objective. It delegates to a Director. The Director creates a plan. The plan generates tasks for specialists. Their work then needs to be reviewed, returned or passed to another role.
Each step can be sensible on its own. Once enough steps are connected, latency compounds. A locally centred setup also makes long-running work difficult when the laptop remains part of the operating environment.
The system could determine what it wanted to do. Allowing it to keep executing that plan without the workflow becoming cumbersome was much harder.
Good organisational reasoning is not the same as dependable autonomous execution.
What version two will test
The next iteration will move towards a persistent cloud environment so the organisation can continue working on longer objectives without relying on a local laptop.
The model hierarchy also needs to become more deliberate.
A capable reasoning model can sit at the strategic layer, interpreting goals, challenging assumptions and deciding what should happen next. Narrow execution tasks can then move to faster models that are better suited to contained work.
Think deeply at the top. Execute quickly at the edges.
I also want the Chief of Staff to understand the company's broader technology environment, rather than only the tools already connected to it.
If a content agent needs image generation, it should identify that capability. If another agent needs publishing or scheduling software, it should be able to raise the requirement.
The useful long-term pattern may be to make those capabilities available through governed Model Context Protocol connections. That would avoid building a separate bespoke integration for every agent while keeping access visible and controlled.
The next interface will also be simpler. Telegram is one option for giving the Chief of Staff direction and receiving updates without turning the agent organisation into another dashboard that needs constant management.
What this experiment does not prove
This is one internal experiment, not a production benchmark.
It does not establish the optimal number of agents. It does not show that a ten-role organisation is better than a smaller one. It does not prove that human management or review can be removed.
It also does not yet demonstrate dependable asynchronous execution over hours or days. That is the main question for the next iteration.
The observations are useful because they narrow the problem. Team design appears achievable. Orchestration, model selection, access and human intervention still need much more testing.
Early conclusion
An AI Chief of Staff appears capable of taking a broad organisational goal and designing a coherent specialist team around it. The first attempt was over-engineered, but the underlying organisational logic was strong enough to continue exploring.
The harder problem is orchestration.
How should agents work asynchronously over long periods? Where should humans intervene? Which models belong at each level? How many specialist roles improve the outcome before coordination becomes the bottleneck?
The next test is a persistent, multi-model version of the organisation with a simpler direction-setting interface and more deliberate governance around its tools.
The first experiment does not solve autonomous organisation design. It does show that the idea warrants further testing.
About the author
Jordan WhitingFounder and CEO, DataMust
Jordan leads DataMust's client work with a practical, commercial lens. He helps teams turn Microsoft Fabric, Power BI and AI-ready data foundations into decisions people can use in production.
Exploring a practical use case?
Discuss a governed AI use case
A 30-minute fit call helps determine whether the use case is ready to implement or needs a clearer roadmap first.