Methodology

How the number is built.

Anyone can hand a manuscript to a language model and print whatever comes back. That isn't an assessment. It's one opinion, with nothing to tell you how much to trust it, and no way of knowing whether it would say the same thing tomorrow. This page describes what we do instead, and what we deliberately don't.

A panel, not a model

Every manuscript is read in full — not sampled, not summarized — by multiple independent AI readers drawn from different companies and different model families. They do not see each other's scores. They score against a shared, explicit standard, and they frequently disagree.

The disagreement is the point. One model gives you one opinion and no way to judge it. Several different models give you a range — and a range is something you can measure, correct, and learn from. We use models built by different companies on purpose: models from the same family tend to make the same mistakes, so when they agree it can look like confirmation when it's really just an echo.

A separate supervisor model sits above the panel. It classifies the work, builds the standard the readers score against, chairs a structured deliberation in which panelists must argue from the text, and issues the final judgment. It doesn't just count votes. It weighs how good each argument is, and when the panel disagrees with its own reading of your book, it has to settle the question and explain its reasoning.

The value isn't in calling a model. It's in everything that sits between the promise and the models to make the answer consistent.

You don't tell us what kind of book you wrote

Most services make you pick your genre from a dropdown before they'll look at anything. We don't, for a simple reason: writers are often wrong about their own books, and being measured against the wrong yardstick is the fastest way to get feedback that's no use to you.

So the first thing that happens is a quick pass that reads your manuscript and works out what it actually is — the category, the sub-category, and who it's for. That decides which standards we hold it to, because different kinds of book fail in completely different ways. A mystery is judged on whether the clues were there for a reader trying to solve it. A thriller works the other way round — you're often meant to know more than the hero, and that's the source of the tension, not a mistake. Judge one by the other's rules and you'll tell a perfectly good book it's broken.

We then tell you what we decided, and what our second choice was. If we've read your book as the wrong kind of thing, you'll see it straight away rather than wondering why the notes feel off. We'll also tell you the single category we'd recommend you use when you query agents, since most submission forms only let you pick one.

A framework built from how the industry actually judges

The scoring framework was not invented from intuition. It was assembled from the standards agents, acquiring editors, developmental editors, contest judges and award bodies actually apply.

Where those sources agreed, we built the agreement in. Where they flatly contradicted each other — and they often do — we left the rule out rather than pick a side. A rule half the industry disagrees with isn't a standard. It's somebody's taste.

Universal core plus a genre layer

Every book is judged on the same core set of things, so a score means the same thing whatever shelf your book belongs on — that's what lets an agent compare a thriller to a literary novel and have the numbers mean something. On top of that, each genre gets its own additional test, because what counts as a flaw genuinely differs from genre to genre.

Some of those divergences are direct inversions. A mystery is judged on whether clues are placed fairly for a reader competing with the detective; a thriller works the opposite way, generating tension from a reader who knows more than the protagonist — so judging a thriller by mystery conventions penalises it for a rule it was never playing by. Science fiction and fantasy penalise a speculative system that is under-explained; horror penalises a threat that is over-explained, because specification kills dread. A framework that applied one standard to both would be wrong in opposite directions at once.

We keep the genre-specific part deliberately small. Load in too many genre rules and you stop measuring how good a book is and start measuring how well it follows convention — which would punish exactly the books that break a rule on purpose and are right to.

Gates are separate from scores

Some things aren't about quality at all — they're about what kind of book you've written. A romance that doesn't end happily isn't a bad romance; it's not a romance. The useful thing to tell you is which shelf your book actually belongs on. So we handle those separately and tell you plainly, instead of quietly docking your score for it.

How we check our own scoring

This is the part that separates an assessment from a guess.

The number of dimensions was tested, not assumed

There's a known trap in scoring anything this way: give someone a list of categories to score and they'll often just repeat their overall impression down the whole list. You end up with a page of numbers carrying one number's worth of information. So we tested it — running the same books through competing versions of our framework and measuring whether each category was really telling us something new.

The answer surprised us, and it changed the framework. We kept the version where each category measured something the others didn't, merged the ones that turned out to be measuring the same thing, and rewrote the ones that were bleeding into each other.

Per-reader bias is measured and corrected

Readers have habits. One is consistently a soft touch; another is consistently tough. If we ignored that, your score would depend partly on which readers happened to pick up your book — which is luck, not judgment.

So we measure each reader's habit against the group, across a set of books we know well, and correct for it before its opinion counts. A reader who always scores high isn't a bad reader — it's a reader whose thumb we know is on the scale, and we can simply take it off. Ignoring that reader instead would throw away a genuinely useful opinion.

Separately, we look at how steady each reader is once that correction is made. A reader whose opinions bounce around counts for a little less. But we never give a reader more say just because it tends to agree with the others — that would reward going along with the crowd, which defeats the purpose of having a panel. And no reader is ever silenced completely.

The same manuscript gets the same score

The strongest version of this isn't a statistic, it's a rule. A score belongs to one specific version of a book. Send us the same file again and we hand back the score it already earned — we don't run it again. So the thing you might reasonably worry about, that you get a number you don't like, resubmit, and get a different one, can't happen. There is no second number to go looking for.

For a genuinely new draft, we do run it fresh, so the question becomes how steady the scoring is. We test that rather than assume it. We took one manuscript and put it through the full system six separate times. The quality score came back 77, 80, 77, 77, 77, 79 — identical on four of the six, and every run inside a three-point spread. Every run used the same panel of readers. The sample report on our home page is one of those six.

That is the moderation layer doing its job. Underneath it the readers genuinely disagree, and on one of those six runs they disagreed violently — the spread between the highest and lowest reader was over twenty-five points. The final number still landed on 77. That is the whole design: individual AI opinions wobble, and the job of the layer above them is to stop that wobble reaching you.

This is also why the quality score is the only one that decides whether a book clears. It is the measure we hold to the tightest standard and the one we can stand behind run to run — we would rather gate on one dependable number than on three.

One honest limit on all of this: it is six runs on one book. It is real evidence and it is more than we had yesterday, but it is not yet a general promise, and we will keep measuring across more manuscripts and update this page with whatever we find — including if it gets worse.

Reproducibility & discrimination

Same file, resubmittedsame score, by rule
Quality score, 6 independent runs77, 80, 77, 77, 77, 79
— identical runs4 of 6
— widest miss3 points
Models per evaluation4 or more
Minimum for a sealed score4

If fewer finish, we mark the score unofficial and it can't qualify for the marketplace at any number. Losing a reader doesn't just make the result fuzzier — it can move it in one direction, which makes it a different measurement rather than a rougher one.

We work out what your book is trying to do — before we judge it

This is the part we think matters most, and we only built it because our own system got something badly wrong.

One of the novels we scored ends by revealing that the main character was never real. That reveal is the point of the book — it's what the whole thing is about. Our readers flagged it as a mistake. They said it undercut everything the reader had invested in him.

They were right that it undercuts your investment. That's what it's for. But they counted it as a flaw, because nothing in our system had told them the author did it on purpose.

That's the trap, and it's an easy one to fall into. Hand an AI a book and a list of rules, and it will check one against the other and call every difference a mistake. It can't tell the difference between a writer who slipped and a writer who aimed. And the books most likely to get punished for it are the ambitious ones — the ones breaking a rule deliberately, which are exactly the books worth finding.

So we establish intent first

Before any score is given, we work out three things about your manuscript and write them down:

Then we measure the gap between that and what is on the page. The score is the distance between your book and its own best self.

What this changes about the criticism you get

Our readers are told plainly: judge the execution, never the intent. If the book withholds something on purpose, the question is whether the withholding pays off — not whether you should have withheld it. Where we think a deliberate choice doesn't land, we have to say what would make it land, instead of telling you to take it out.

You will still get hard notes. A choice can be deliberate and still not work, and we will say so and say why. But there is a real difference between "this doesn't work yet, and here's what it needs" and "you shouldn't have done this" — and only one of those is any use to someone with a draft in front of them.

You get judged on the book you set out to write, not the one a rulebook expected.

An absolute bar, not a curve

The bar is a fixed standard, not a ranking against whoever else submitted that week. If it were a ranking, the same book would pass in a quiet month and fail in a busy one — and there'd be nothing you could do about which month you landed in. A fixed bar means your score means the same thing this year as it will in five years.

We check that bar against books whose fate we already know — and the important test isn't the acknowledged masterpieces. It's the books that sold well without being literary landmarks, because those are the ones a screening system is most likely to get wrong. If our gate would have turned away a novel that readers actually loved, the gate is broken, not strict. That test caught a genuine mistake in our own scoring, and fixing it is why the bar sits where it does.

We will publish what share of submissions clear, and we'll start the moment that share means anything. We're at the very beginning of the beta, and a percentage drawn from a handful of books would be noise dressed up as a statistic — which is the opposite of the point. The point is that clearing the bar is only worth something to an agent if it's genuinely hard, and only worth something to you if we're straight about how hard.

Three scores, and only one is a gate

Quality

How well the book is written, measured against the other manuscripts agents actually receive. This is the only score that decides whether your book reaches the marketplace. A book nobody expects to be a bestseller still gets in if the writing earns it.

Commercial

How sellable it looks — the hook, where it sits on the shelf, who buys books like it. Agents see this. It never blocks you.

IP value

What could be built beyond the book itself — film and TV, sequels, spin-offs. This is deliberately a different question from whether it will sell. A brilliant standalone novel that makes one great film can outsell a franchise and still score lower here, and that's not a knock on the book. For non-fiction we ask a different set of questions entirely.

Keeping these three apart matters. Roll them into one number and a commercial hook could buy its way past weak writing — which is exactly what already goes wrong in the current system.

They're also reported differently, on purpose. Quality is a number out of 100, because it decides admission and it's the one we hold to that standard — it's built from a set of anchored dimensions scored by several calibrated readers and then arbitrated. Commercial and IP value are reported as a band: low, average, high, or extremely high. Those two are single holistic judgments rather than panel-arbitrated ones, and a band is the resolution they honestly carry. It also reflects what the number is for — no agent reads a book differently at 47 than at 51.

What we deliberately do not do

Things we won't do, on purpose

  • We don't care how many followers you have. Not your mailing list, not your social media, not how a previous book sold. None of that is in your manuscript, so any score we gave it would be made up. And a system that rewards writers who already have an audience is exactly the problem we're trying to fix. If nobody has heard of you and your book is good, you can clear our bar. That's the entire point of this.
  • We don't check whether your facts are right. We can't. Real fact-checking means phoning sources and pulling records — you can't do it by reading the pages. Telling you we'd done it would be a lie. What we can do, for non-fiction, is tell you whether you've backed up what you claim and whether you've pushed a claim further than your evidence goes. That's a different question, and it's one we can actually answer.
  • We won't tell you you're the fourth-best book we've read. Here's a thing that turns out to be true of human agents too: people are far better at agreeing a book isn't ready than at agreeing which of two good books is better. So we'll tell you plainly whether you're above the line or below it. We won't rank the books that made it, because we'd be inventing the order.
  • We judge by what publishers are buying now. Not by what makes great literature, and not by what sold thirty years ago. If you're sending your book to agents this year, this year's standards are what you're actually up against. Those standards aren't the only way to judge a book and they'll change — we'd rather say that out loud than pretend ours is the only measure.
  • Your book is never used to teach an AI to write. No model is trained on it. We use commercial AI providers under their business/API terms and available privacy controls; your work is encrypted while we hold it, and no agent sees a word unless you decide they can. One honest exception, because it isn't nothing: books we still hold get re-scored when we change the system, so we can check the scoring hasn't drifted. That's the machine grading itself against work it has already graded. Nothing about your writing goes into a model.

Still in beta, and we will say so

We keep track of which version of our system produced every score. Scores from different versions aren't directly comparable, and we won't pretend they are — when we change how we judge, we change what the number means.

During the beta a person reads every report before it goes out. We've tuned the system against published books whose reputations we already know, which proves it can tell strong writing from weak — but published books aren't the same as the manuscripts that land in an agent's inbox, and that's the range that matters most. Early submissions are how we close that gap. That's the honest reason the first hundred are free.