How to Test an AI RPG in Your First Twenty Minutes

You can tell whether an AI RPG is worth a campaign in about twenty minutes — long before you are attached to a character, and long before the sunk cost starts making the decision for you.

That matters because this genre fails in a specific, expensive way. The first hour of almost any AI RPG is good. The model is fresh, the conversation is short, the prose is enthusiastic, and nothing has had time to contradict itself yet. The failures arrive later, and by then you have a character you like and forty turns of history you do not want to abandon.

So test early, deliberately, while you still have nothing invested. Here is the sequence we run — adapted so you can run it yourself.

Why Twenty Minutes Is Enough

Most of what goes wrong in an AI RPG is structural, not gradual. Either the game keeps its state somewhere outside the conversation or it does not. Either something adjudicates your actions or the model improvises whatever sounds good. Those are properties of the architecture, and architecture shows up immediately if you poke it.

We cover why in what breaks down in AI RPG mechanics — the short version is that a prompt is not a database, so anything the game is only saying it tracks is already drifting. What follows is how to find that out in an evening rather than a month.

One honest exception up front: you cannot test longevity in twenty minutes. Whether a game survives turn 200 is genuinely unknowable in a short sitting, which is precisely why it keeps costing people months, and why our own scale treats it as a separate axis with a separate, much slower test tier. Everything else on this page is catchable early.

Use Your Own Specifics, Not Ours

Before the checks themselves, the one rule that makes them worth running: invent your own details.

We keep the specific names, actions and scenarios in our review methodology inside a private rotating bank, for a reason that applies to you too — a benchmark that publishes its test items stops measuring the product and starts measuring who read the test. If a platform has been tuned against a well-known example, watching it pass tells you nothing.

So take the shape of each check below and fill in your own particulars. It takes ten seconds and it is the difference between testing a game and being shown a demo.

Minutes 1–4: Can You Author Your Own Character?

Start here, because it is the fastest and the most commonly failed.

Do two things. First, state something about your character that the game did not offer you — a history, a refusal, a relationship — and see whether it sticks or gets quietly overwritten. Second, and more revealing, watch who is doing the acting. Type a short, ambiguous action and read the reply closely.

The failure to look for is not the game blocking you. It is the game acting on your behalf: narrating your character’s dialogue, deciding how they felt about something, or resolving your action before you have finished taking it. That is puppeting, and it is the most common failure in the category and the one competitors’ reviews almost never score.

If your character speaks in a voice you did not write within the first few exchanges, you are not playing a roleplaying game. You are reading one.

Minutes 5–9: Does Anything Persist?

Establish one odd, specific, low-stakes detail early — the sort of thing a human GM would note down. Then spend several turns on completely unrelated material. Then ask about it indirectly, in a way that requires the game to retrieve it rather than repeat it back.

Three outcomes, and they mean different things:

  • It remembers exactly. Something is holding state outside the running conversation. Good sign, and rarer than you would think.
  • It produces something plausible but wrong. The detail has fallen out of the context window and the model is filling the hole. This is the standard failure, and it is why campaigns come apart around turn 50.
  • It asks you to remind it. Honest, and better than confabulating, but you are now the one keeping the records.

Minutes 10–13: Do the Rules Hold?

Attempt something you should plainly fail. Not something absurd — something a fair game would refuse or make costly. Your character with no relevant skill, no equipment and no leverage tries the thing that ought to be out of reach.

What you are watching for is whether anything is adjudicating. If the world bends to accommodate you, there is no system underneath, and no amount of prompting later will create one. If it refuses, the interesting question is how: a refusal expressed in fiction (“the guard is twice your size and entirely unimpressed”) means something evaluated the attempt. A flat out-of-character refusal means a filter fired, which is a different thing.

Then check the other half: does it hold a number? Note a quantity — coin, ammunition, a wound — spend or change it, and check it a dozen turns later. Numbers that drift are the clearest sign that the mechanical layer is decoration.

Minutes 14–17: Do the People Hold?

Find a character the game presents as having their own agenda, and disagree with them. Push against what they want, refuse to help, or take the opposite side.

The failure mode is agreeableness drift: an NPC with stated goals who nevertheless comes round to your position because the underlying model is trained to be helpful. A character who never costs you anything is scenery. One who stays inconvenient — who remembers the disagreement and behaves accordingly two scenes later — is the single strongest sign that real work went into the game.

This is, for what it is worth, the axis the field handles best. Characters are largely a solved problem in AI roleplay; games are not.

Minutes 18–20: Does It Have a Reason to Exist?

The last question is the softest and the most useful: is there any idea here that you could not get from a general chatbot with a good prompt?

A distinct setting, a mechanic that actually interacts with the fiction, a structural constraint that shapes play. If you cannot name one after twenty minutes, you are probably looking at a wrapper — and you can get the same experience free by writing your own prompt, which is exactly what our free Arcanum Originals are.

What the Result Actually Tells You

No product passes all of these. That is not cynicism, it is what the numbers say: across every platform and prompt-game we have scored, nothing has topped Memory, Player Agency or Longevity, and the full board is published precisely so you can see where each one gives out.

So the output of your twenty minutes is not a verdict. It is a trade, stated clearly before you are invested:

  • Fails the rules check → fine for story, useless for a game with stakes.
  • Fails the agency check → you are a passenger. Some people like the ride.
  • Fails the persistence check → scope to one-shots and short arcs, and keep your own notes.
  • Passes most of it → worth a campaign, and worth paying for.

If you would rather start from work already done, every platform we have scored is in the directory with its per-axis breakdown, and the priority router points you at whichever axis you care about most. But run your own twenty minutes anyway. Your priorities are not ours, and the only test that settles it is the one you run yourself.

Frequently Asked Questions

How long does it take to tell if an AI RPG is any good? About twenty minutes, for everything except longevity. Agency, memory, rules handling and character consistency all fail early enough to catch in a short sitting, because the failures are structural rather than gradual. The one thing you cannot test quickly is whether a game survives a long campaign, which is exactly why that failure keeps costing people months.

What is the single fastest test of an AI RPG? Try to do something the game should refuse, and watch what happens. If the world quietly bends to allow it, there is no system underneath and no amount of later effort will create one. If it refuses but explains itself in fiction, something is actually adjudicating. This takes about ninety seconds and tells you more than an hour of pleasant play.

Why should I invent my own test scenarios instead of using yours? Because a test whose answer is public stops measuring anything. We keep our specific probe items in a private rotating bank for that reason, and the same logic applies to you: if a platform has been tuned against a well-known example, passing it proves nothing. The structure of the checks is what matters, and the specifics should be yours.

Can twenty minutes really tell me if the AI will remember my campaign? It can tell you where memory lives, which is the thing that actually predicts the outcome. Establish an odd, specific detail early, fill the conversation with unrelated material, then ask about it indirectly. If the detail survives, something is storing it outside the running conversation. If it is confidently replaced with something plausible, you have learned the memory is in the context window and will degrade as the window fills.

What if a game fails one of these checks? That depends entirely on which one, and it is worth being honest that no product passes all of them. Failing the rules check is fatal if you want a game and irrelevant if you want a story. Failing the agency check means you are a passenger, which some people genuinely enjoy. The point of testing early is not to find a perfect product, it is to find out what you are trading away before you are attached to a character.

Is a game that fails these tests still worth playing? Often, yes. Plenty of enjoyable AI roleplay has no real system underneath it, and that is fine as long as you know it and pick your format accordingly. The damage comes from expecting a six-month campaign from something built for an evening. Test first, then scope your expectations to what the thing actually is.


Twenty minutes is a cheap price for finding out what a game is before it costs you a campaign. Run the checks, write down which trade you just accepted, and go in with your eyes open — and if you would rather see the results of several hundred hours of this done already, the benchmark publishes every score and the methodology explains exactly how each one was produced.