Anyone can hand a manuscript to a language model and print whatever comes back. That isn't an assessment. It's one opinion, with nothing to tell you how much to trust it, and no way of knowing whether it would say the same thing tomorrow. This page describes what we do instead, and what we deliberately don't.
Every manuscript is read in full — not sampled, not summarised — by multiple independent AI readers drawn from different companies and different model families. They do not see each other's scores. They score against a shared, explicit standard, and they frequently disagree.
The disagreement is the point. One model gives you one opinion and no way to judge it. Several different models give you a range — and a range is something you can measure, correct, and learn from. We use models built by different companies on purpose: models from the same family tend to make the same mistakes, so when they agree it can look like confirmation when it's really just an echo.
A separate supervisor model sits above the panel. It classifies the work, builds the standard the readers score against, chairs a structured deliberation in which panelists must argue from the text, and issues the final judgment. It doesn't just count votes. It weighs how good each argument is, and when the panel disagrees with its own reading of your book, it has to settle the question and explain its reasoning.
The value isn't in calling a model. It's in everything that sits between the promise and the models to make the answer consistent.
The scoring framework was not invented from intuition. It was assembled from the standards agents, acquiring editors, developmental editors, contest judges and award bodies actually apply — published editorial assessment criteria, acquisition memo structures, contest scoring sheets with disclosed weightings, and the craft frameworks that working editors reference by name.
Where those sources agreed, we built the agreement in. Where they flatly contradicted each other — and they often do — we left the rule out rather than pick a side. A rule half the industry disagrees with isn't a standard. It's somebody's taste.
Every book is judged on the same core set of things, so a score means the same thing whatever shelf your book belongs on — that's what lets an agent compare a thriller to a literary novel and have the numbers mean something. On top of that, each genre gets its own additional test, because what counts as a flaw genuinely differs from genre to genre.
Some of those divergences are direct inversions. A mystery is judged on whether clues are placed fairly for a reader competing with the detective; a thriller works the opposite way, generating tension from a reader who knows more than the protagonist — so judging a thriller by mystery conventions penalises it for a rule it was never playing by. Science fiction and fantasy penalise a speculative system that is under-explained; horror penalises a threat that is over-explained, because specification kills dread. A framework that applied one standard to both would be wrong in opposite directions at once.
We keep the genre-specific part deliberately small. Load in too many genre rules and you stop measuring how good a book is and start measuring how well it follows convention — which would punish exactly the books that break a rule on purpose and are right to.
Some things aren't about quality at all — they're about what kind of book you've written. A romance that doesn't end happily isn't a bad romance; it's not a romance. The useful thing to tell you is which shelf your book actually belongs on. So we handle those separately and tell you plainly, instead of quietly docking your score for it.
This is the part that separates an assessment from a guess.
There's a known trap in scoring anything this way: give someone ten categories to score and they'll often just repeat their overall impression ten times. You end up with ten numbers carrying one number's worth of information. So we tested it — running the same books through competing versions of our framework and measuring whether each category was really telling us something new.
The answer surprised us, and it was the opposite of what we expected. Fewer, broader categories were worse — a big vague category invites a single gut reaction, while narrow specific ones force a real judgment each time. So we went with what the evidence showed, merged the categories that turned out to be measuring the same thing, and rewrote the ones that were bleeding into each other.
Readers have habits. One is consistently a soft touch; another is consistently tough. If we ignored that, your score would depend partly on which readers happened to pick up your book — which is luck, not judgment.
So we measure each reader's habit against the group, across a set of books we know well, and correct for it before its opinion counts. A reader who always scores high isn't a bad reader — it's a reader whose thumb we know is on the scale, and we can simply take it off. Ignoring that reader instead would throw away a genuinely useful opinion.
Separately, we look at how steady each reader is once that correction is made. A reader whose opinions bounce around counts for a little less. But we never give a reader more say just because it tends to agree with the others — that would reward going along with the crowd, which defeats the purpose of having a panel. And no reader is ever silenced completely.
We test this rather than just claim it. Run the same book through twice and the score lands within about a point. In one test the same manuscript got the exact same score from two different sets of readers — which matters more, because it means the result held up even when we changed who was doing the reading.
And a score belongs to a specific version of a book. Send us the same file again and you'll get back the score it already earned, not a fresh roll of the dice. You can't shop for a better number.
| Same manuscript, repeat run — mean difference | ~1 pt |
| Same manuscript, different panel composition | identical |
| Readers per evaluation | 4 |
| Minimum readers for a sealed score | 3 |
If fewer than three readers finish, we mark the score unofficial and it can't qualify for the marketplace at any number. Losing a reader doesn't just make the result fuzzier — it can move it in one direction, which makes it a different measurement rather than a rougher one.
The bar is a fixed standard, not a ranking against whoever else submitted that week. If it were a ranking, the same book would pass in a quiet month and fail in a busy one — and there'd be nothing you could do about which month you landed in. A fixed bar means your score means the same thing this year as it will in five years.
We check that bar against books whose fate we already know — and the important test isn't the acknowledged masterpieces. It's the books that sold well without being literary landmarks, because those are the ones a screening system is most likely to get wrong. If our gate would have turned away a novel that readers actually loved, the gate is broken, not strict. That test caught a genuine mistake in our own scoring, and fixing it is why the bar sits where it does.
We publish what share of submissions clear. That number is the whole point. Clearing the bar is only worth something to an agent if it's genuinely hard, and only worth something to you if we're honest about how hard.
How well the book is written, measured against the other manuscripts agents actually receive. This is the only score that decides whether your book reaches the marketplace. A book nobody expects to be a bestseller still gets in if the writing earns it.
How sellable it looks — the hook, where it sits on the shelf, who buys books like it. Agents see this. It never blocks you.
What could be built beyond the book itself — film and TV, sequels, spin-offs. This is deliberately a different question from whether it will sell. A brilliant standalone novel that makes one great film can outsell a franchise and still score lower here, and that's not a knock on the book. For non-fiction we ask a different set of questions entirely.
Keeping these three apart matters. Roll them into one number and a commercial hook could buy its way past weak writing — which is exactly what already goes wrong in the current system.
We keep track of which version of our system produced every score. Scores from different versions aren't directly comparable, and we won't pretend they are — when we change how we judge, we change what the number means.
During the beta a person reads every report before it goes out. We've tuned the system against published books whose reputations we already know, which proves it can tell strong writing from weak — but published books aren't the same as the manuscripts that land in an agent's inbox, and that's the range that matters most. Early submissions are how we close that gap. That's the honest reason the first hundred are free.