A .NET Agent Orchestration Framework - Introduction
Introduction
In this new series of posts I will talk about a multi-agent orchestration framework that I've been developing in the last few months. This framework (name yet to be defined) is provider and model-agnostic and the responsibility to connect to the models of choice is outside its scope; anything works as long as it has a Microsoft.Extensions.AI wrapper. As with most of the code I write, it is being developed in .NET and this post explains some of the core concepts and technologies used to build it; I will expand on each concept (the orchestrator, the agents, the training, the knowledge management, etc) in subsequent posts. I also plan to eventually make the code available in GitHub and NuGet.
This is somewhat linked to my PhD research, but I won't get too academic here (I am not!), this is essentially a proof of concept that seems to work, and something I'm having fun with!
The posts in the series I plan to write are, as of now:
- Introduction (this one)
- Orchestrator (TBD)
- Agents (TBD)
- Agent training (TBD)
- Agent scheduling (TBD)
- Agent review and revise (TBD)
- MCP tools (TBD)
- Putting it all together (TBD)
- Extensibility and next steps (TBD)
Do keep in mind that, as the framework is evolving, there may be changes. I will try to keep all posts updated.
Why a Multi-Agent Framework?
Let's start from here: why build a multi-agent framework - isn't a single agent good enough? Well, there are some good reasons, IMO:
- Specialisation beats generalism: a single agent juggling database schema, UI conventions, .NET idioms, and business rules tends to produce shallower output in each domain than agents each prompted/tuned for one thing
- Context window hygiene: each specialised agent only needs the context relevant to its domain, not everything. Keeps prompts smaller, more focused, and less prone to the model losing track of instructions buried in a huge system prompt
- Parallelism: independent subtasks (e.g., drafting the DB schema while another agent drafts the UI component) can run concurrently instead of sequentially through one agent
- Built-in review/critique: peer review between agents catches errors a single self-reviewing agent is more likely to miss, since it's a genuinely different "pass" rather than the same model re-reading its own output
- Easier debugging/iteration: when something's wrong, you can inspect which agent's output was faulty and improve just that prompt/tool config, rather than re-tuning one monolithic prompt
- Standardized agent/tool abstractions: IChatClient, function-calling middleware, etc. give us a consistent way to define what an agent is without reinventing message-passing, tool-call parsing, and streaming handling for each specialist
- Orchestration primitives you'd otherwise rebuild: task distribution, result aggregation/synthesis, retry/timeout handling, and conversation-state threading between agents are all things a framework typically handles - and can get subtly wrong if hand-rolled (e.g., losing partial results on a failed agent call)
- Observability/tracing hooks: multi-agent systems are hard to debug blind - which agent said what, in what order, with what tool calls. Frameworks generally wire up tracing (OpenTelemetry-style) so you can see the whole exchange, which matters a lot once you add peer review (now you have agent→agent traffic, not just agent→user)
- Composability: if today it's DB/UI/.NET/business-logic and tomorrow you add a testing or security-review agent, a framework gives you a slot to plug it into rather than restructuring bespoke glue code
Let's see how this is addressed in what I built.
Basic Concepts
In this framework we have essentially agents and an orchestrator. There are, of course, lots of other smaller components, but these are the most important ones.
The orchestrator receives a prompt, dynamically decides how to distribute the work to be done by the different specialised agents it knows, and then collects and summarises the responses. It also takes care of logging, collecting metrics, etc. My implementation broadly follows the Orchestrator-Worker (Hierarchical/Magentic) pattern and the evaluator/critique pattern if you are curious. The pipeline is: plan → execute → (review → revise)* → synthesize.
Agents that can use whatever model they want (different agents can use different models, possibly picking the one that is the best fit for their specialisation) but are trained, specialised, for a specific technology or field of studies. These fields of studies are, for now, a closed set, and include:
- .NET development
- Database development
- DevOps
- Documentation writing
- Testing
- Planning
- Synthesis
- ...
Besides the specialisation, an agent has a distinct name, and some tags that more finely describe it (e.g., "SQLServer", "PostgreSQL, etc). It can make use of some services, of which I'll talk in a moment. Agents wrap an instance of IChatClient (possibly a different one for each agent), and are instantiated with different parameters. The actual training is performed by a different class. The agent's prompt is built from its original prompt and their training when the orchestrator executes it. After processing, the work of each agent is reviewed by another agent, and its feedback is used for revising it. Agents can load state from past invocations and, at the end of a successful processing, can save the new state too. They can also share state amongst all the agents and have access to a shared knowledge base, which can be used to enhance its prompt.
The following picture summarises this:
Core Assumptions
All of the services implement an interface that describes its contract, and there is always a minimum, default, implementation. It should be simple to provide new implementations. All relevant aspects can be configured, core services replaced or extended. Always Dependency Injection (DI)-friendly.
All of the steps - planning, execution, review, revise, synthesis - are executed by agents.
All of the calls are asynchronous and can be cancelled, thus terminating the entire orchestration process.
All of the agent work is implemented by LLM models, as provided by an IChatClient implementation, which must be built outside the scope of this framework. The usage of tokens is controlled and logged.
All work is logged using the standard .NET constructs, including the time it took, tokens consumed, and activity traces are also produced.
All work occurs in a session and is identified by a session id. All agent operations have a task id and exist in a session.
Class Diagram
These are the essential interfaces and their implementing classes, plus the relationships between them. Some implementing classes are not shown.
Here's a brief summary of these types, there are many more, but I don't find them relevant enough to be listed here:
- IAgent/SpecializedAgent: the core agent interface and its only provided implementation
- IOrchestrator/Orchestrator: the agent orchestrator that coordinates the work of the different agents; it makes use of many different services
- IAgentKnowledgeStore: a store for keeping the history of each agent's operations after they are done as well as loading it initially to build the agent prompt
- IAgentScheduler/ParallelAgentScheduler/SequentialAgentScheduler: used by the Orchestrator to schedule the agents; there are two implementations: one for parallel and the other for sequential execution
- IPlanner: used by the Orchestrator to plan the agent assignments
- ISynthesizer: also used by the Orchestrator to synthesise the final agent's responses
- IAgentSelector: the algorithms to use for selecting an agent for execution or for reviewing
- IAgentTrainingProfile: the contract for the agent training; used by the SpecializedAgent to train the agent before actually doing any work
- IKnowledgeStore: a general-purpose knowledge base that is indexed by a specialisation and stores string content and their corresponding vector embeddings
- ISharedState: a key-value store bound to a session, which should be thread-safe
- ISharedStateFactory: builds the shared state. Usually used by the ISharedStateManager implementation to build the shared state for a session when it does not already exist
- ISharedStateManager: returns the shared state for a given session
- IMcpToolsProvider: registers additional MCP tools for a SpecializedAgent or Orchestrator
Technologies Used
Besides .NET 10, the framework uses the Microsoft.Extensions.AI framework to provide a low-level, provider-agnostic abstraction. I probably could have used the Microsoft Agent Framework, I have played with it before, and at some point I will certainly have a go in porting my framework to it; they are both great choices, part of Microsoft's AI ecosystem. For now, it was a somewhat conscious choice which I can revisit later. I also make use of MCP for file operations and a sharing of state between agents. Finally, I also use the ONNX Runtime for generating vector embeddings but it doesn't have to be used if we provide an alternative implementation.
Conclusion
This is what I wanted to talk about now. Stay tuned for the next posts in this series. As always, feel free to reach out and ask any questions or make any comments! This is a very important topic to me, and sure would appreciate comments on this, as we go along!
Comments
Post a Comment