RaidGuild Cohort
Public

From Roleplay to Operating System: RaidGuild’s First Multiplayer Agent Experiments

Author

Dekan

Date Published

A guild game master coordinates distinct agents around a luminous council table and shared decision paths.

About two years ago, I started running governance and coordination simulations with large language models inside RaidGuild. We were testing OpenAI models alongside Llama 3 because open-source models mattered to us. The setup was simple: Python scripts, the OpenAI API, Venice AI, and a lot of prompt revisions.

The work began as a search for practical uses. We wanted to see where agents could contribute, where they became unreliable, and which constraints made the difference. Those experiments helped inform more member experiments, product builds, and client work across RaidGuild as models and tooling improved.

I revisited that history during the RaidGuild roundtable on July 30, 2026. The details still matter because many problems now described as “agent orchestration” were already visible in those early simulations.

Building worlds that could produce a decision

We ran many variations on the same pattern. They were governance and coordination simulations built around fictional worlds, made-up personas, and a scenario with enough tension to force a discussion. Hogwarts was one setting among several.

A typical simulation might include a pragmatist, a collectivist, and an individualist. Each persona had a distinct outlook and a reason to disagree with the others. The Game Master established the world, introduced the immediate problem, and kept a longer-term goal in view.

One scenario opened with a dark portal appearing in the Hogwarts Grand Hall. The characters had to decide how to respond while balancing safety, collective responsibility, individual freedom, and the wider intent of the group. The fantasy setting gave us distance from real governance disputes while preserving the mechanics of conflict and decision-making.

The tension was essential. Without it, the agents had very little to deliberate. They produced pleasant conversation, then drifted toward agreement before they had examined the tradeoffs.

The agents wanted harmony

The most persistent failure was excessive agreement. Long stretches of dialogue sounded like: “Yes, you are right.” “That is absolutely correct.” “Yes, that is the next right thing.”

A persona described as an individualist would abandon its position after one polite reply. The collectivist would accept the pragmatist’s proposal without testing its consequences. Everyone tried to be helpful, and helpfulness erased the convictions that made the simulation useful.

Prompt honing became an exercise in giving the personas enough backbone to sustain disagreement. The Game Master had to inspect their responses, reject weak character work, and ask for revisions with stronger conviction. Even then, personality prompts alone could not carry the process to a decision.

A workflow finally got us to a vote

The simulations began producing useful outcomes once we gave them a defined sequence. The Game Master introduced the scenario and intent. Each persona received its own turn to deliberate. The group recorded a soft vote, negotiated around the disagreements it exposed, and then moved to a final vote.

Every turn also needed an exit condition. Without a clear stop, agents could continue responding to one another indefinitely, lose the original context, and spin into hallucinated details. The exit was part of the design, not cleanup after the fact.

This workflow gave the Game Master something concrete to validate. It could check whether each persona had spoken from its assigned position, whether the soft vote reflected the discussion, and whether unresolved tension required another negotiation round. A final vote became possible because the process defined how to reach it.

What those simulations taught us

Most of this work was early prompt and context engineering. We were learning how strongly a role needed to be defined, how context changed behavior, how loops formed, and where orchestration had to intervene. Model capability mattered, but process design often determined whether the simulation produced a coherent decision.

The experiments also clarified the limits of simulated governance. An agent could play a delegate, facilitator, or voter inside the scenario. That performance did not grant authority outside it. Human intent, review, and ownership still had to remain visible.

As new models arrived, we repeated the same practical test: could this technology help people coordinate, and could we contain its failure modes well enough to trust the result? Better models improved the conversations. They did not remove the need for roles, stages, validation, and exits.

How the lessons spread through RaidGuild

These experiments gave other members patterns they could challenge and extend. The work moved into new experiments, product development, and client delivery. Each setting exposed a different part of the problem: preserving context, making agent activity inspectable, assigning authority, and handing work back to people.

Agent proposals move through deliberation, negotiation, voting, and visible exit paths.

The current working frame reflects that accumulated experience. Buzz is where members and agents collaborate. Prism preserves context and runs auditable workflows. Portal presents reviewed outcomes to members. Their shape will continue to change, but the underlying concerns came directly from watching early agent groups agree too quickly, lose context, and fail to finish a decision process.

The strongest lesson was practical: coordination emerged from deliberate structure. We had to create tension, protect each persona’s turn, validate its conviction, record intermediate decisions, and stop the loop at the right moment.

Next in the series

The next article will follow the shift from these Game Master simulations into persistent context, memory, and auditable workflows. That is where short-lived roleplay began turning into infrastructure that members could inspect, reuse, and adapt. A link will be added here when the next installment is ready.

This series is a field report from the people doing the experiments. We will keep the failures in the story because they explain why the systems developed the way they did.

Illustrations for this article were created for the series from original prompts using generative tools.

Current program

open

Join Cohort 15

Collaborative, agent-assisted content creation

25 days until Cohort 15 starts

A small collaborative cohort for turning complex Web3, AI, and coordination ideas into useful public artifacts with agent-assisted research, production, distribution, and review.

Join the cohort

Comments

No comments yet.

Log in to leave a comment. Log in