makepocketpussy.com

How we review sleeves

The testing process behind every review on this site, including what we measure, who tests, how we score, and what we refuse to do.

Every review on this site follows the same process. We publish it here so you can judge whether our verdicts are worth anything, and so you can tell us when we have got something wrong.

We buy most of what we test

The majority of the sleeves we review are bought at retail with our own money. Where a brand sends us a free sample, the review says so in the testing panel at the top of the page, under “Unit source”. We do not accept payment for a review, we do not accept edits from a brand, and we do not let a brand see a review before it publishes.

We are an affiliate of Fleshlight and several other retailers. That means if you buy through a link on this site, we earn a commission at no extra cost to you. It does not change what we write. The clearest evidence of that is the number of products on this site we tell you not to buy.

What we measure

Before anyone uses a sleeve, it gets measured and photographed. We record:

  • Canal length and entry diameter, measured by hand rather than copied from the box
  • Material and case type
  • Price paid, and the date we paid it
  • Internal texture, mapped zone by zone from entry to end

That last one is the part most reviews skip. A sleeve is not one sensation, it is three or four in sequence, and which zone dominates depends on your length. Our texture map shows you the sequence so you can work out which part you will actually be in contact with.

How long we test for

A sleeve gets a minimum of six sessions across at least three weeks before we write anything. There is a reason for the gap: SuperSkin and comparable materials change. A texture that felt aggressive when new often softens, and a sleeve that cleaned easily on day one can start holding water in a chamber by week three. A same-day review cannot tell you any of that.

Where a verdict depends on anatomy, a second tester repeats the cycle. Tightness in particular is not an absolute property of a sleeve, it is a relationship between the sleeve and the person using it, and one tester’s “very tight” is another’s “barely there”.

How we score

How the score is built, and what it deliberately leaves out

Most sites in this category rate a sleeve across a dozen or so parameters and then average all of them into a single number. We think that is wrong, and it is worth explaining why, because it changes what our scores mean.

Consider realism. Some of the most highly regarded textures ever made score close to zero on realism, and their owners love them for the reason they score badly. Averaging realism into an overall score marks a product down for succeeding at something else. The same applies to tightness: high is not better, it depends entirely on your own girth. And to intensity: overwhelming is exactly right for a short session and exactly wrong for a long one.

So we sort every parameter into one of three classes and only score two of them.

  • Character parameters describe what a sleeve is. They never touch the score. They exist so you can match a sleeve to yourself, and we show them as a position on a spectrum with both ends labelled rather than as a mark out of five.
  • Performance parameters are the ones where better is genuinely better for everybody.
  • Cost of ownership parameters are running costs — lube, noise, cleaning, drying. We rate the burden, so a low number is the good outcome, and we invert them before scoring.

The headline score is 65% performance and 35% cost. Nothing from the character class is in it. If we have not owned a sleeve long enough to judge its running costs, we publish a performance-only score and say so on the review.

The consequence worth stating plainly: a sleeve can score badly here and still be exactly right for you, and our character profile is how you would find that out. A score is a summary, not a substitute for reading.

Character — described, never scored

Criterion Weight
Parameter What 1 and 5 mean
Intensity 1 = Gentle throughout
5 = Overwhelming, hard to last in
Tightness 1 = Loose, easy entry
5 = Very tight, resistant entry
Texture density 1 = Mostly smooth walls
5 = Structure packed along the whole canal
Realism 1 = Obviously a toy
5 = Closest to partnered sex
Variation 1 = One sensation end to end
5 = Distinct zones that feel different
Smoothness 1 = Grabby, high friction
5 = Glides

Performance — higher is better

Parameter What 1 and 5 mean Weight
Stimulation quality 1 = Muted, hard to feel
5 = Every element registers clearly
30%
Finish quality 1 = Flat, or has to be abandoned
5 = Strong, and usable through the finish
25%
Suction control 1 = All or nothing
5 = Cap adjusts it across a usable range
20%
Durability 1 = Softens or tears within months
5 = Holds its texture after long use
25%

Cost of ownership — lower is better

Parameter What 1 and 5 mean Weight
Lube consumption 1 = One application lasts a session
5 = Needs topping up repeatedly
25%
Noise 1 = Quiet with the cap closed
5 = Audible through a wall
20%
Cleaning effort 1 = Rinse through and done
5 = Must be everted and worked at
30%
Drying time 1 = Dry in about an hour
5 = Still damp a day later
25%

A parameter we have not properly assessed is left blank rather than guessed, and we do not publish an overall score until at least half the weighted parameters are filled in. If a review has no score, that is why.

Measuring the canal

Textured sleeves are not uniform. A canal is a sequence of zones — a tight entry ring, a wide chamber, a constriction, a ribbed run to the end — and which of those zones you actually engage depends on how long you are.

This is the single most decision-relevant fact about a textured sleeve, and almost nobody publishes it. We measure every zone: where it starts, where it ends, its narrowest diameter, and what is inside it. Those measurements drive the scale diagram, the fit calculator, and the fit-and-reach section on every review.

The fit calculator runs entirely in your browser. The measurements you enter are not sent to us, not stored, and not logged. There is no analytics on that tool by design.

Two honest caveats. Silicone stretches, so a canal narrower than you is not a barrier — it is where the grip comes from, and our girth figures indicate firmness rather than fit. Depth is the harder limit: a zone behind you is a zone you paid for and will not feel, though it still generates suction.

Why we do not publish performer biographies

Several sites in this category attach a personal dossier to each sleeve — date of birth, star sign, nationality, ethnicity, height, weight, body measurements. We do not, and it is a deliberate policy rather than an omission.

Three reasons. First, none of it helps you choose a sleeve. A star sign cannot tell you whether a texture will suit you, and we would only be publishing it to catch search traffic, which is precisely the kind of padding that makes a page worse.

Second, these are real people. A performer licensed a mould of their body to a manufacturer. They did not agree to have a third-party affiliate site compile their date of birth alongside their physical measurements, and that combination is exactly the profile used by people who harass them. Much of the biographical data circulating on adult wikis is also simply wrong, and republishing an incorrect date of birth about a named individual is not a small error.

Third, our readers keep saying they do not want it. Scroll the user reviews on any large sleeve database and you will find experienced reviewers explicitly asking others to stop writing about the performer and write about the product.

What we publish instead is product provenance: whose body the orifice was cast from, as credited by the manufacturer; whether the texture was designed for them or licensed in; when it first shipped; and whether it has been revised since. That last point matters more than any of the biographical material — if a texture was reworked in 2016, every review written before then is describing something you can no longer buy. We track those revisions and we tell you which version we had.

We do not scale scores to make a product look good. There is no rule that says the average has to land near eight.

Where AI fits, and where it does not

We use AI as an editing tool. It takes the testing notes and measurements our team has recorded and helps us structure them into consistent sections, so that every review answers the same questions in the same order and you can compare two pages easily.

It does not decide verdicts, invent measurements, or write anything from scratch. The tool we built for this is configured to return “insufficient data” rather than fill a gap, and no section appears on the site until a human editor has read it and approved it. Anything with a factual error in it is our fault, not the tool’s, and we would like to hear about it.

What we will not do

  • Review a product nobody on the team has physically handled
  • Publish a score for something we have not tested through a full cycle
  • Describe a sensation we did not experience
  • Leave a review up unchanged when the product has been reformulated
  • Write about performers’ bodies rather than the products molded from them

Tell us when we are wrong

If your experience of a sleeve contradicts ours, that is useful information and we want it. Our corrections policy explains what happens next.