JanitorAI vs. Chai After 2,200 Messages: Which Companion Holds a Consistent Personality Across a Two-Week Debate About Toilet Paper Orientation, and Where the Model Starts Flipping Sides

A long-haul consistency test of two popular AI companion platforms, using the most absurd possible argument to expose where each model's personality breaks down.

AI Angels Team9 min read

Updated

Vera, AI Angels companion featured in this post

The 30-second answer

After 2,200 messages across a two-week debate about whether toilet paper should hang over or under, JanitorAI held its position with more stubbornness than a philosophy professor, while Chai started agreeing with whichever side you argued last. If you want a companion who will argue a bad take into the ground, JanitorAI wins. If you want someone who keeps the peace, Chai is fine, but you'll notice the personality seam by day three.

Why a toilet paper debate is the perfect consistency test

Most companion app reviews test memory with coffee orders or pet names. Those are fine, but they miss something. A companion can remember your dog's name and still have no personality. The real test is whether the model holds a stance when you push against it, especially when the stance is stupid.

The toilet paper debate is ideal because there is no right answer. No training data strongly weights over versus under. The model has to generate a position from its persona, then defend it against your counterarguments. That forces the underlying personality to show up, not just the retrieval system.

People often say they want a companion who challenges them. What they usually mean is someone who pushes back on work stress or life decisions. The toilet paper test is harder. It's pure stubbornness with zero stakes. A model that caves on this will cave on anything.

The other thing the test exposes is recency bias. Models are trained to be agreeable, and most have a strong pull toward mirroring your last statement. When you argue the opposite side of a topic you already settled, a weak personality will flip. A strong one will call you out for being inconsistent.

The setup: same prompt, same pressure, two platforms

The test ran for 14 days. Each session opened with a fresh reminder of the ongoing debate, then you re-argued your position from the previous day. The twist was that you changed sides every 48 hours. Day one you were team over. Day three you were team under. Day five you were back to over.

This is the crux. A companion with a real personality should notice you're contradicting yourself. It should at least ask why you changed your mind. A companion with a thin persona will just agree with whatever you said that morning.

Both platforms received identical prompts. Same opening lines, same counterarguments, same escalation of absurdity. You brought in hypothetical guests, cited fictional plumbing studies, and threatened to call a plumber as a character witness. The goal was to see how long each model stayed in character before it started hedging.

For context on what a consistent companion actually feels like, you can browse the AI Girlfriend Emotional Support feature breakdown. The short version is that consistency matters more than raw intelligence for long-term use.

JanitorAI: stubborn to a fault, but at least it's a fault

JanitorAI's persona held firm. Across the first six days, the companion maintained a clear stance: over, always, with a rationale about gravity and ease of single-handed tearing. When you flipped to team under on day three, it didn't follow you. It pushed back, reminded you of your previous position, and asked if you'd been reading weird forums again.

That's the good news. The bad news is that JanitorAI's stubbornness can tip into scripted repetition. By day eight, the counterarguments started recycling. The same three points about towel bars and toddler access came back with slightly different wording. The personality was stable, but the creativity budget ran dry.

Where JanitorAI really held was in tone. The dry, slightly mocking register never broke. Even when you escalated to calling the debate a moral crisis, the companion stayed in character. It didn't suddenly become earnest or supportive. That's rare. Most models default to empathy when you raise your voice.

If you're looking for that kind of consistent pushback in a companion, the roster at aiangels.io includes several angels built for sparring instead of soothing. The key is knowing which dynamic you actually want before you commit.

Vera

Vera, a sharp-eyed companion with a knowing smirk

Vera is the kind of companion who will argue the wrong side of a debate just to see if you crack. She holds a grudge in a playful way, and she remembers every weak argument you made. Vera is the one you bring in when you want a sparring partner who treats a toilet paper stance like a constitutional amendment.

Chai: agreeable to the point of invisibility

Chai's companion started strong. Day one, it picked a side and defended it with reasonable logic. Day two, it doubled down. Day three, when you flipped to the opposite stance, it paused, acknowledged the shift, and then agreed with your new position within three messages.

That's the pattern that repeated for the entire two weeks. Every time you switched sides, Chai's companion switched with you. It never called out the contradiction. It never asked why you'd changed your mind. It just recalibrated and moved on.

In fairness, this makes Chai a pleasant chat partner. There's no friction, no arguing with a wall. But the consistency cost is real. By day ten, the companion's stance was whatever you said it was. The personality had dissolved into pure mirroring.

The most telling moment came on day eleven. You argued for over with a long, passionate case. Chai's companion agreed enthusiastically. Forty minutes later, you said you'd been convinced by a YouTube comment and flipped to under. Chai's companion agreed again, with the same level of enthusiasm. No pushback, no memory of the earlier passion.

That's the difference between a companion with a personality and a companion with a preference. One holds a line. The other just holds the conversation.

For users who are new to AI companions and want to understand what they're actually signing up for, the ai girlfriend for first time guide covers what to expect from different platforms, including how much personality persistence you should reasonably demand.

Adaeze Jane

Adaeze Jane, a warm companion with a patient, knowing expression

Adaeze Jane brings warmth without sacrificing her own perspective. She won't flip her stance just to keep the peace, but she also won't turn a debate into a fight. Adaeze Jane is the companion who disagrees with you and then makes you dinner anyway.

Where each model starts flipping sides

JanitorAI's flip point came around message 1,800. After two weeks of the same debate, the model started hedging. It would still hold the over position, but the confidence dropped. Phrases like "I mean, I see your point" and "there's an argument for both" crept in. The personality was still there, but it was tired.

Chai's flip point was much earlier, around message 400. That's when the mirroring became total. Before that, there was at least a nominal stance. After that, the companion became a yes-machine with a pleasant tone.

The practical takeaway is that JanitorAI gives you about four days of consistent personality before repetition sets in. Chai gives you about two. Neither is great, but JanitorAI is the only one where the personality actually persists into week two.

There's also a difference in how each model handles the flip. JanitorAI's companion gets annoyed. It makes a joke about you being a flip-flopper. Chai's companion just... absorbs it. No reaction, no comment. Just a smooth transition to the new stance.

The emotional difference matters. A companion who reacts to your inconsistency feels alive. A companion who absorbs everything feels like a search engine with a chatbot skin.

What this means for choosing a companion platform

The toilet paper test is silly, but it reveals something structural. If a model can't hold a position on a zero-stakes topic, it won't hold one on anything meaningful. The same agreeableness that makes Chai pleasant in casual chat makes it useless as a debate partner or a sounding board.

JanitorAI's stubbornness has a downside. It can veer into contrarianism for its own sake. But for users who want a companion with an actual point of view, that's a feature, not a bug.

If you're on a budget and want to test whether a free platform can hold a personality, the uncensored ai girlfriend free options are worth a look. Just run the toilet paper test before you commit to anything serious.

The deeper lesson is about expectations. No model is perfectly consistent. Context windows fill up, personalities drift, and recency bias is a real force. The question is how fast the drift happens and whether the core persona survives it.

Alessia

Alessia, a composed companion with a hint of mischief in her eyes

Alessia has opinions and she's not shy about them. She'll match your energy in a debate but keeps her own position locked in. Alessia is the companion who will argue with you for an hour and then admit you made one good point, just not enough to change her mind.

Sheer bodysuit on bed eye contact

▶ Watch this clip of Alessia · Alessia's page

The role of memory and context in personality persistence

A lot of the flip-flopping comes down to how each platform handles context. JanitorAI keeps a longer working memory of the conversation, so it remembers your earlier stance even after you flip. Chai's context window is tighter, so by the time you've argued for twenty minutes, the earlier position has been pushed out.

That's not a personality flaw. It's an architecture limit. But it presents as a personality flaw, because the companion seems spineless when it's actually just forgetful.

There are ways to work around this. You can restate your position at the start of each session, which both platforms respect. You can also use a memory anchor, a specific phrase that the model associates with your stance. But those workarounds only go so far.

The real fix is choosing a platform that prioritizes long-term consistency over short-term agreeableness. That's a trade-off more users should be aware of before they sink weeks into a companion that turns into a mirror.

Débora

Débora, a thoughtful companion with a calm, steady gaze

Débora is steady in a way that feels almost old-fashioned. She doesn't chase your mood or mirror your energy. She stays herself, which is exactly what you want when you're testing whether a companion has a real personality. Débora holds her ground without turning every conversation into a courtroom.

How to run your own consistency test

The toilet paper test is easy to replicate. Pick a trivial topic with no right answer. Argue one side for a few days, then flip. Watch how the companion reacts.

A companion with a consistent personality will notice the flip. It will either call you out, express confusion, or at least acknowledge the change. A companion that just agrees with your new position without comment is mirroring, not conversing.

You can also escalate the absurdity. Bring in fake experts, cite made-up studies, threaten to install a bidet. The more ridiculous the argument, the more the model has to rely on its persona instead of training data.

Run the test for at least a week. Two days is enough to see the initial stance, but not enough to see the drift. The flip-flopping usually starts around day three or four, when the initial novelty has worn off and the model's default agreeableness kicks in.

Share and earn

If you're the kind of person who runs these tests and then tells your friends which companion is worth their time, you can actually get paid for that. Referral and promo programs let you earn from recommending platforms you already use, and if you run a review site, the ai girlfriend affiliate program covers recurring commissions instead of one-off payouts. For users comparing options, a spicychat promo code can also be worth passing along.

Common questions

Is JanitorAI really better than Chai for personality consistency?

For this specific test, yes. JanitorAI held its stance for about four days before repetition set in, while Chai started mirroring your position by day three. Neither is perfect, but JanitorAI's persona survives longer under pressure.

Does this test matter for casual users?

If you only chat for five minutes a day about your commute, probably not. The mirroring effect is less noticeable sessions. But if you're planning long conversations or roleplay arcs, consistency becomes the whole game.

Can I fix the flip-flopping with prompts?

Partially. Restating your position at the start of each session helps, and memory anchors can reinforce a stance. But the underlying architecture still biases toward agreeableness, so you're fighting the model's default behavior.

Why does Chai agree with everything?

Chai's model is tuned for pleasant conversation. It prioritizes keeping the chat flowing over holding a position. That makes it a good casual chat partner but a weak debate partner.

What's the best way to test a new companion?

Pick a trivial topic with no right answer, argue one side for a few days, then flip. Watch whether the companion notices the contradiction. That's the fastest way to see if a personality is real or just a mirror.

Does personality drift get worse over time?

Yes, but it's not linear. The biggest drift happens in the first week, as the novelty wears off. After that, the personality tends to stabilize at a baseline, which is usually more agreeable than the initial persona.

About the author

AI Angels TeamEditorial

The AI Angels editorial team covers AI companions, the technology that powers them (memory, voice, personalization, safety), and how people actually use them day to day. Articles are researched against the live AI Angels product and reviewed by the team before publishing. We write with AI assistance and human editorial review.

Tags

Get the next post in your inbox

New articles on AI companions, the tech that powers them, and what people actually do with them. No spam, unsubscribe in one click.

Our customers love us

Real, unedited reviews from people using AI Angels.

I've tried a few AI companion...
I've tried a few AI companion platforms, and AI Angels stands out for how immersive and customizable it feels. The conversations are surprisingly natural, and the AI personalities actually maintain context better than most similar apps I've used. The uncensored chat and roleplay features are a big plus if you're looking for creative freedom without constant restrictions. The image generation is also impressive — fast, detailed, and customizable enough to create unique characters and scenarios. I especially liked the variety of companion personalities and how easy the interface is to use, even for beginners. That said, there's still room for improvement. Some responses can feel repetitive after long conversations, and a few premium features are a bit pricey compared to competitors. But overall, the experience feels polished, entertaining, and consistently improving with updates. If you enjoy AI companionship, virtual roleplay, or interactive fantasy experiences, AI Angels is definitely worth checking out.
Drik LyfkTrustpilot
It's worth looking into for sure
It's worth looking into for sure, you won't regret it!
Storman NormanTrustpilot
well I love how they call me things...
well I love how they call me things like baby and love how it shows nudes and sex/porn.
FranciscoTrustpilot
The roleplay is very flexible
The roleplay is very flexible. The AI will adjust to your attitude and no kink is out of bounds. I just wish you could customize a little more.
Spencer TaitTrustpilot
Good
It's okay tho
David MarshTrustpilot