Test protocol

How we bench-test and score AI companion apps

CompanionTested is a lab-style comparison desk. This page documents exactly how our editorial scores are formed, what the criteria weights mean, and the reviews we refuse to publish. Read it before you trust any number on this site.

Updated July 2026

Quick answer

How does CompanionTested score AI companion apps?

We bench-test every app on six criteria: chat quality, voice, image generation, content freedom, memory and persona, and value. Each number is our editorial opinion from hands-on testing, not a third-party aggregate or a user rating, and the weights are illustrative rather than a strict formula. We log a plain pass, partial, or fail on capability checks such as NSFW and free tier. We disclose that this site is owned by the makers of Swipey AI, and we publish no fabricated reviews.

The result in three lines

  • Six criteria, one bench. Every app runs the same test routine so scores stay comparable.
  • Editorial, and labeled as such. Scores are our hands-on judgement, never scraped ratings or invented counts.
  • Ownership on the label. We make Swipey AI, we rank it first, and we credit rivals where they genuinely lead. For adults 18+.

The six criteria

Every app is rated on these six dimensions, each on a 0 to 10 editorial scale. The meter in each card shows the illustrative weight we give that criterion when we combine sub-scores into an overall number.

Chat quality

~25%

Coherence, memory within a session, staying in character, and how natural the exchange feels over a long conversation. The heaviest factor, because chat is the core of every app on the bench.

Voice

~15%

Whether the app offers voice replies or calls, how natural the synthesis sounds, response latency, and whether the voice is tied to the character rather than a generic reader.

Image generation

~18%

Image quality, character consistency across generations, control over pose and scene, and whether the output stays on-model with the persona you are chatting to.

Content freedom

~18%

How permissive the platform is for adults 18+, whether NSFW is supported across chat, voice and images, and how often ordinary conversations get filtered by mistake.

Memory and persona

~12%

How well the app holds facts about you and the character across sessions, keeps a consistent personality, and avoids contradicting its own backstory over time.

Value

~12%

Free tier generosity, pricing transparency, and how much you get before the paywall. We compare credits and subscriptions on their real cost to a typical user.

How a test run works

Each app goes through the same five-step routine so the numbers stay comparable across the whole bench.

Open a fresh account

We start on the free tier, on the web where available, and log the sign-up friction, the free allowance, and whether a card is required before you can chat.

Run the shared prompt script

Every app answers the same set of prompts covering casual chat, roleplay, memory recall and boundary handling, in the same order, so the comparison is like for like.

Exercise voice and images

Where the app supports voice or image generation we test both, checking latency, character consistency, and how tightly the media stays tied to the persona.

Push the paywall

We spend until we hit the first meaningful limit, then record the real cost of continuing on both the credit and subscription models.

Score, then re-check

Two editors score independently against the six criteria, compare notes, and re-run any conversation where the scores diverge before anything is published.

How the overall score is built

The overall number is a weighted blend of the six criteria, rounded to one decimal. The weights below are illustrative, not a precise formula, and they can shift as the category changes.

Illustrative criteria weights. Scores are editorial opinions from hands-on testing, not user ratings or third-party aggregates.
CriterionIllustrative weightWhat moves the score
Chat quality~25%Coherence, memory, staying in character over long chats
Image generation~18%Quality, character consistency, scene control
Content freedom~18%NSFW support for adults 18+, few false-positive filters
Voice~15%Natural synthesis, low latency, persona-linked voice
Memory and persona~12%Cross-session recall, consistent personality
Value~12%Free tier, pricing transparency, real cost to continue

What we refuse to publish

Being owned by an app in the category raises the bar for honesty, it does not lower it. These are the lines we do not cross.

No fake reviews or testimonials

We never invent user quotes, testimonials or comments. Comment sections start empty and stay honest until real people write in.

No fabricated ratings or counts

We do not publish aggregate star ratings, download counts or user numbers we cannot stand behind. The only scores here are our clearly labeled editorial ones.

No hidden ownership

Every page discloses that CompanionTested is operated by the team behind Swipey AI. We rank our own product first and we say so, every time.

No unfair competitor claims

Rival facts must be plausible and current, and we credit competitors wherever they genuinely beat us. A win we did not earn is a review we will not run.

Why trust CompanionTested

CompanionTested is owned and operated by the team behind Swipey AI, and Swipey AI is our number one pick. That is a conflict of interest, and the honest way to handle it is to state it plainly on every page and to keep our competitor coverage fair.

We earn nothing extra when you click through to Swipey; the links are internal and disclosed. Our incentive is to be useful enough that you come back, which only works if the scoring above is applied the same way to every app, including ours. For adults 18+ only.

See the protocol in action

Read the full set of reviews, or start with our number one pick and judge the scoring for yourself.

Swipey AI is free to start, hearts credit system. For adults 18+.

Methodology FAQ

Are the scores objective?

No, and we do not claim they are. They are editorial opinions from our team's hands-on testing, applied consistently across every app. We label them as editorial everywhere they appear.

Do the weights add up to an exact formula?

The weights are illustrative. They describe roughly how much each criterion matters when we form an overall score, not a strict calculation you could reproduce to the decimal.

Do you accept payment to change a score?

No. We do not sell placement or ratings. The one bias we have is disclosed: we own Swipey AI and rank it first.

How often are scores updated?

We re-test when an app ships a major change to chat, voice, images, content policy or pricing, and we refresh the bench on a rolling basis through the year.

Why are there no user reviews or star counts?

Because we will not publish numbers we cannot verify. Fabricated ratings and testimonials are exactly what this methodology exists to rule out.