Skip to content
boop.

AI can generate infinite options. Humans still decide what feels right.

Making things used to be the bottleneck. You had one designer, one week, and one idea you could afford to execute properly. The scarce resource was production, so that’s what tools optimised for.

That constraint is gone. A founder with no design background can produce nine credible homepages in an afternoon. The hard part moved somewhere less comfortable: you now have nine options, a strong personal preference, and no idea whether a stranger would agree with you. Taste didn’t get automated. It got exposed.

The obvious move is to ask a model which one is better. You get a fluent, well-structured answer in three seconds, and it’s worth almost nothing — not because the model is stupid, but because it isn’t doing the thing you need. It’s predicting what a design critique sounds like. Nobody looked at your work and chose.

A preference is a different kind of fact.

When 55 of 82 people pick B, that number isn’t an opinion about B. It’s a record of 82 small decisions made by people with their own histories, their own screens and their own reasons for being unimpressed. It can be wrong about the wider world, and Boop is careful to say so — but it cannot be fluent nonsense. Something actually happened.

That’s the whole product. Two versions, one question, and a count. The interesting part is what sits underneath: the reason tags that tell you it won on comprehension rather than beauty, and the eleven people who bothered to write a sentence explaining why your favourite version confused them.

Where we do use a model.

Exactly one job: reading written feedback and grouping it into themes, so ninety comments become three things you can act on. It never casts a vote, never writes one, and never changes a count. Every number on a results page came from a person who clicked. Where a summary is model-written, it says so on the page.

What we’re honest about.

  • The sample is not representative. Boop’s voters are internet builders who wandered into a feed. That’s the right room for “does this read as trustworthy” and the wrong room for anything requiring a specific profession or demographic. Audience filters are planned and clearly marked as not built.
  • Small samples are small. Confidence describes how consistent the observed preference is at this sample size, using a 95% Wilson score interval. It is not a claim about the wider population — a taste test tells you what these humans picked, not what everyone would pick.
  • Anonymous voting is a speed bump. We hash an anonymous session identifier server-side and enforce one vote per test. Someone determined to skew a result with fresh browsers can. We don’t fingerprint, and we don’t pretend otherwise — the response count is always on the page so you can judge for yourself.
  • Sample tests are fictional. Boop ships with a seeded corpus so the product isn’t an empty room on your first visit. Every one of those tests is marked Demo.

Why it’s a loop, not a subscription.

A preference-testing platform with no voters is a database. The credit system exists to make giving feedback worth something: 1 credit for every 3 boops you give, and 1 credit for every human response you request. Buying a pack is a shortcut for people in a hurry. It is not the entrance.

The design of this site is the argument, too. It’s quiet and mostly white and gets out of the way, because the thing you uploaded is what people are supposed to be looking at.