How a ShipScore is computed

Every verdict on this site comes out of the same pipeline, on the same harsh scale, with the evidence printed underneath. This page is how it works and where it has been wrong — two frameworks, one referee, and the receipts.

Part I — Why this reads differently
Survivorship

Almost every source of startup numbers is a collection of self-reported success stories.

Founders who made it write them up; founders who didn't, don't. So the picture you get is filtered twice — first by who survived, then by what they chose to say about themselves. It is a genuinely useful library and a systematically flattering one.

We track the numbers instead of collecting the stories. Ten thousand products, revenue read independently and refreshed every morning — the ones compounding, the ones flat, and the ones quietly losing the money while their ratings stay high. Nobody writes a case study about the third kind, which is exactly why it is the useful kind.

That is also why our own record is unflattering: 294 verdicts and not one has said "build". A source that only shows you winners can't produce a number like that.

Two frameworks, one verdict

Most validators answer one question: is this a good business? That question already has a good answer — MJ DeMarco's CENTS. But a business can pass CENTS today and be worthless in three years, because the thing that kills a good idea is rarely the idea. It's that somebody else's advantage was more durable than yours.

So we run a second reading over the first: Warren Buffett's moat — a durable competitive advantage, and crucially the direction it is moving. DeMarco asks whether the business is worth building. Buffett asks whether the advantage lasts. Nobody else scores an app idea against both, because the second one needs something almost nobody has: the revenue history of the competitors, over time.

A ShipScore is those two answers held against each other — on live money rather than opinion.

01 — CENTS, adapted (not borrowed)

Five commandments — Control, Entry, Need, Time, Scale — scored by one referee against one evidence standard, then composited into the ShipScore.

C — ControlDo you own the customer, or does a gatekeeper own them for you?
E — EntryHow hard is this to copy once it works?
N — NeedIs anyone already paying for this problem to go away?
T — TimeDoes it earn while you sleep, or only while you work?
S — ScaleCan it reach a lot of people without costing a lot more?

What we changed. DeMarco wrote CENTS as a self-assessment — a founder honestly grading their own plan. Self-assessment is exactly where optimism lives. We rebuilt it as an adversarial instrument:

Framework by MJ DeMarco (The Millionaire Fastlane · UNSCRIPTED). The adaptation, the evidence standard and the calibration are ours.

02 — The moat layer

Buffett's test is whether a business has a durable competitive advantage — and he cares far more about whether the moat is widening or narrowing than about this year's earnings. That test is almost never applied to small software, because the numbers aren't published.

We can apply it, because we track them. For every rival with a real revenue history, the engine reads the direction of the moneywidening, stable, narrowing, eroding — and your verdict is scored against it.

A rating tells you how loved a product is. The moat trend tells you whether the money is staying.

Those come apart constantly: a beloved rival whose revenue is quietly draining is an opening; a mediocre-looking rival whose revenue compounds is a wall. Ratings alone point you at the wrong one — which is why "check the reviews" is such expensive advice.

Part II — How a verdict is made
The short version

Your idea is read into a product description and a set of competitor signatures. Those are matched against our corpus — 14,000+ tracked products, refreshed daily by our own collectors, carrying revenue histories, funding, exits and complaint volume. The matched evidence, not your description, is what gets scored. A frontier model does the judging under the rubric above; a different one writes the re-spins, so the thing proposing an alternative is never the thing that graded the original.

The output is one of three stamps — STOP · SKIP, DEAD SLOW · CAUTIOUS, FULL AHEAD · BUILD — plus the kill switches, the pace test, the wedge, and the cheapest next step. The telegraph needle on every report lands on the verdict; the word after the dot is its name; the number beside it is the composite.

What we don't publish: the rubric itself, the thresholds, the matching rules, the model line-up and the sources our collectors run against. That is deliberate. A scale you can reverse-engineer is a scale you can write for, and a score written for is worth nothing to the next person who reads one. What we publish instead is the evidence under every verdict — so you can check the reasoning even where you can't see the ruler.

Part III — The evidence standard
Not all evidence is equal, and we grade it

Every figure in a report carries its provenance, because "revenue" from a founder's tweet and revenue from independent tracking are not the same claim.

Owner-verifiedRead once from the company's own Stripe, with their permission, and computed by us. The strongest thing we hold — and the only tier where the number is not somebody's report of itself.
Independently verifiedRevenue tracked by a third party, not self-reported. Strong, but still a reading taken from outside.
DisclosedOpen dashboards and founder statements, shown as such and dated.
ReportedPress and public announcements, cross-checked before they count.
Platform signalsRatings, reviews and upvotes. Useful, gameable, weighted accordingly.
Market outcomesWhat comparable businesses actually changed hands for — asking prices and completed sales.
Demand noisePublic complaint volume. A signal that a problem is felt, never proof anyone will pay.
The one definition, published

Stripe's own documentation confirms that MRR is configurable per merchant — whether recurring discounts come off, whether a subscriber counts from signup or from first payment. Two identical businesses can therefore publish different "MRR" and both be telling the truth. So we do not read anyone's reported figure. We compute it ourselves, the same way for everyone, and we print the rule:

Sum of active subscriptions' recurring amounts, normalised to a month (annual ÷ 12, weekly × 52/12). Excludes one-time charges. Excludes trials — a trialing subscription has not paid. Recurring discounts subtracted. Computed by Will It Ship, identically for every verified account.

Version v1-2026-08-15 — stamped on every stored reading, so a figure can always be traced to the rule in force when it was taken. Change the wording and the version changes with it.

Every owner-verified figure carries the date it was read. We keep no customer records, no emails and no charge-level data from a connected account — only the one number, and in Phase 1 the access is dropped the moment it is computed.

In the evidence directory the tiers print as the labels on each row: stripe-verified is independently verified (a third party reads the company's Stripe); open-dashboard, founder-reported and self-reported (X) are disclosed; press-reported and estimate are reported. Only the claim flow on an app's own page produces owner-verified — that is our read of the owner's Stripe, taken once, with permission, and dropped.

A tier is never silently upgraded. If the only figure we hold is founder-reported, the report says so — and an idea is never scored as unwanted merely because our corpus is thin there; curated category leaders exist precisely so "no evidence here" can't masquerade as "no demand anywhere".

Part IV — Where we've been wrong
The misses stay in the record

Every verdict joins the public archive unedited, including the engine's own. Four corrections worth naming, because a validator that has never been wrong in public is one that hasn't been checked:

Part V — The pledges
Privacy
Money-back
Honesty
Part VI — The questions
01 — Is the validator really free?

Yes. ShipScore verdicts and the daily idea at noon are free, forever. Re-spins and the full evidence pack are Pro.

02 — Are the verdicts public?

Yes. Every report joins the public record, unedited — including the engine's own misses.

03 — What is ShipScore?

One verdict from two frameworks: DeMarco's CENTS, adapted into an adversarial referee that only moves on tracked evidence, plus a Buffett-style reading of whether each rival's advantage is widening or narrowing. The weakest commandment is always named.

04 — What is the Evidence Engine?

The pipeline behind ShipScore: your idea becomes a set of competitor signatures, those are matched against a corpus of 14,000+ tracked products our collectors refresh daily, and the matched evidence — not your description — is what gets scored. The model line-up, the matching rules and the sources behind the corpus stay in-house.

05 — Does a skip verdict mean "don't build"?

It's a lane reading, not a command. A skip says the 2026 lane is closed on the evidence shown. The engine scores the market, not you — and it publishes its own record so you can judge the judge.

06 — Who is it for?

Indie builders and founders at idea stage who want the truth with receipts — and anyone tired of cheerleader validators.

See it run tomorrow morning

That's the method. The cheapest way to judge it is to watch it work: one gap a day, scored on this scale, with the evidence printed underneath — and the misses published alongside.

Free, one email a morning, unsubscribe in a click.

ShipScore is scored on a rubric inspired by the CENTS framework by MJ DeMarco (The Millionaire Fastlane · UNSCRIPTED). Will It Ship is not affiliated with or endorsed by MJ DeMarco.
Privacy · Terms · support@willitship.app