SEO 14 min read Updated August 2026

Content Scoring: How to Evaluate Quality Before Publishing

Most weak content isn't published because someone decided it was good. It's published because nobody stopped to ask whether it was. Content scoring is the habit of asking — running each piece through a consistent check before it goes out, so "is this good enough?" stops being a gut call made under deadline and becomes a decision you can make the same way every time. It's the difference between publishing on purpose and publishing by default.

Everything drafted, and what a scoring gate lets through

  1. Every draft that reaches "finished" the only question asked so far is "is it done?"
  2. The ones aimed at a real reader relevance
  3. The ones saying something only you could say depth
  4. The ones worth your name on them publish
Fig 01 · The gate, and what it is forWithout the gate, all four bands publish, because "finished" and "good" were never asked as separate questions. The bands that don't make it are not waste — they are the standard doing its job.

Why score content at all

The case for scoring is simple: content quality is a decision, and undecided decisions default to "publish." When there's no standard, the piece that's merely finished gets treated the same as the piece that's genuinely good, because the only question anyone asked was "is it done?" A score forces the better question — "is it worth it?" — at the one moment it still matters, before the thing is live under your name.

Everything a business published last year, sorted honestly

  • Genuinely good someone decided it was, and was right
  • Fine — nobody checked finished, competent, and asked nothing of anyone
  • Should not have gone out and quietly taxed everything around it

The shape of the problem, not a measurement

Fig 02 · The middle band is the whole argumentNobody publishes the third band on purpose. The second band is the one that does the damage, because every piece in it passed the only test anyone applied: it was finished.

This matters more now than it used to, because the cost of thin content has gone up. Search rewards depth and consistency; a pile of forgettable posts doesn't just fail to help, it dilutes the authority of the good pages around it. AI answer engines simply skip content that says nothing specific. So the weak piece isn't neutral — it's a small tax on everything else you've published. Scoring is how you stop paying it, one piece at a time.

There's a quieter benefit, too. A shared standard turns "good" from one person's taste into something a team can agree on and repeat. Without it, quality lives in whoever happens to review a piece that day, and it drifts every time that person changes. With it, the same bar applies whether the founder, a writer, or a new hire is holding the draft. Scoring doesn't just catch weak pieces — it makes "what we consider good" explicit enough that other people can hit it without you in the room, which is the only way quality survives handing the work off.

Scoring is not an audit

It's worth drawing a clean line between two things that get confused, because they solve different problems. Scoring happens before you publish — it's a gate a new piece has to pass. An audit happens after, across everything you've already put out — it's an inventory of a library you built before you had a standard. One keeps the quality of what's coming; the other fixes the quality of what's there.

When the check happens, against how much it looks at

Before publishing · one piece

Scoring

A gate. Stops the next weak piece going out, and costs a few minutes per draft.

After publishing · one piece

A post-mortem

Useful once, occasionally. Too slow to be a standard and too late to change anything.

Before publishing · the whole plan

Editorial strategy

Decides what gets written at all. Upstream of scoring, and a different job.

After publishing · everything

An audit

Drains what already pooled. Necessary if you built a library before you had a standard.

Left column: it happens before publishing

Top row: it looks at one piece

Fig 03 · Two of these four are the pair you needScore to stop the bleeding, audit to drain what has already pooled. Doing only one of them is the common mistake, and which one you skipped decides which problem you have.

You need both, and they reinforce each other. If you only audit, you're forever cleaning up messes you keep making, because nothing stops the next weak piece from going out. If you only score, your back catalogue stays full of the thin content you published before the gate existed. Score to stop the bleeding; audit to drain what's already pooled. We cover the second half of that in how to conduct a content audit; this piece is about the gate.

The five dimensions of quality

A useful score measures a few things that actually predict whether a piece will earn anything, not a long checklist nobody completes. Five dimensions cover most of it. Relevance: is this written for a specific reader with a specific question, or for "everyone," which is another word for no one? Depth: does it say something only you could say, from real experience, or does it restate what's already everywhere?

"Five tips for better email marketing" — a finished draft, scored honestly

  • Relevance. Aimed at everyone with an inbox, which is another word for no one.
  • Depth. The same five tips as every other page on the subject.
  • Clarity. Well-formatted, easy to follow, point near the top.
  • Search fit. A real query, answered by a thousand near-identical pages already.
  • Business value. Nobody finishes it and thinks of you as the obvious partner.
Fig 04 · Competent, finished, and pointlessThree clear fails and the draft still looks perfectly publishable — which is precisely the piece a score exists to catch. Nothing here is fixed by another editing pass.

Clarity: can a reader — and a machine — follow it and lift the point, or is the answer buried under a wall of text? Search fit: does it resolve a genuine question people actually ask, structured so it can be found and cited, without keyword stuffing? Business value: does it build trust and open a path toward a conversation, or is it a pageview that leads nowhere? A piece doesn't need a perfect ten across all five. But a piece that's weak on most of them isn't ready, however finished it looks.

To see the five working together, score a hypothetical. Say you've drafted "Five tips for better email marketing." Relevance: weak — it's aimed at everyone with an inbox, no specific reader. Depth: weak — the five tips are the same five on every other blog. Clarity: fine, it's well-formatted. Search fit: crowded — a thousand near-identical pages already answer this. Business value: near zero — nobody finishes it and thinks of you as the obvious partner. Three clear fails, two passes, and the piece still looks perfectly publishable. That's exactly the content a score exists to catch: competent, finished, and pointless.

The scorecard

Here's the whole thing as a table you can hold a draft against. For each dimension, decide honestly which column the piece is closer to.

Dimension A weak piece A strong piece
RelevanceWritten for "everyone," no clear reader or questionAnswers one real question for one real reader
DepthRestates what's already everywhereSays something only you could, from experience
ClarityBuries the point under a wall of textAnswer-first, clean structure, easy to quote
Search fitTargets no real query, or stuffs keywordsResolves a genuine intent, structured to be found
Business valueA pageview that leads nowhereBuilds trust and opens a path to a conversation

Two things about using it in practice. The first is that the columns are deliberately written as descriptions rather than as numbers, because a number invites averaging and averaging is how a piece with two catastrophic weaknesses and three strengths comes out looking acceptable. You are not adding anything up. You are reading five sentences and deciding which one each dimension is closer to.

The second is that the honest answer is often "somewhere in between," and that's fine as long as you say which side of the middle. A dimension you genuinely cannot place is usually a signal that the piece hasn't decided something — most often who it's for, which is relevance wearing a disguise. In that case the useful output of the score isn't a verdict at all; it's the note that sends the draft back with one specific question attached.

The two dimensions people fudge

In practice, three of the five are easy to judge honestly. Clarity, search fit, and relevance are fairly visible — you can see whether the point is buried, whether there's a real query behind the piece, whether it's aimed at someone specific. The two that get fudged, every time, are depth and business value, because they're the ones that ask uncomfortable questions.

How easy a dimension is to judge, against how much it predicts

Clarity

Search fit

Relevance

Depth · value

Across: how much it predicts whether the piece earns anything

Up: how visible it is on a quick read

Fig 05 · The two that matter are the two that get fudgedDepth and business value sit top-right: highly predictive and genuinely uncomfortable to answer. Score them strictly or the whole exercise launders weak content through a rigorous-looking process.

Depth is uncomfortable because the honest test is: "would anyone who already knows this subject learn anything here?" A lot of content fails that quietly — it's competent, well-formatted, and adds nothing a dozen other pages don't already say. Business value is uncomfortable because the honest test is "does this actually move someone toward working with us, or does it just exist?" Content that scores well on the easy three and poorly on these two is the exact trap most businesses fall into: polished, plentiful, and inert. Score those two strictly, and you'll cut the pieces that were always going to underperform, no matter how clean the formatting.

Setting the publish threshold

A score is only useful if it forces a decision, so the point of scoring is the threshold, not the number. Decide in advance what a passing piece looks like — for most teams, something like "strong on at least four dimensions, and never weak on depth or business value." Anything that clears it, publish. Anything that doesn't lands in one of two buckets: revise, if a clear fix would get it there, or cut, if the piece was never going to earn its place no matter how much you polish it.

One week of drafts, scored, against a threshold agreed in advance

MonTueWedWedThuFriFri

The dashed rule is the threshold, agreed before anyone read a draft

Left to right: the week's drafts

Bottom to top: how the piece scored

Fig 06 · The rule is the decisionThree publish, four don't. The value isn't the heights — it's that the line was drawn before anybody had a draft they were fond of. A threshold set afterwards is not a threshold.

That second bucket is the one that makes scoring worth doing. The willingness to cut a finished piece — to spend the effort of writing it and still not publish it — is what separates a real standard from a rubber stamp. It feels wasteful in the moment. It isn't. Publishing a weak piece costs more than not publishing it, because it dilutes everything around it and earns nothing itself. The draft that dies at the threshold did its job: it told you where your standard is.

Scoring one real draft

Scoring is quick enough that describing it takes longer than doing it, so here is one draft going through the gate and the record it leaves behind. The piece is 1,900 words on client handovers, written from a real position, and the scoring took about four minutes.

The thing worth noticing isn't the verdict. It's that the score got written down next to the draft rather than said out loud in a review call — because a scoring habit that leaves no trace can't tell you anything six months later about which weakness keeps recurring.

What one scored draft leaves on disk

  • drafts/handovers/ one piece, one folder
  • draft.md 1,900 words, finished Tuesday
  • score.md five lines, four minutes, one verdict
  • score-notes.md why depth passed — the position it argues
  • cut/ drafts that did not clear the threshold
  • email-tips.md cut on depth · kept, not deleted
  • email-tips.score.md the reason, in one line

Nine months of these is a dataset about your own standard — and it is the only way to see a pattern in what you keep cutting.

Fig 07 · Cut, not deletedThe cut folder is the useful half. A draft that failed on depth is a record of where your standard actually sits, and three of them failing for the same reason is a finding about the system upstream.

The verdict on this one was publish: strong on relevance, depth and business value, fine on clarity, unremarkable on search fit. Four minutes, one line of reasoning per dimension, and a note explaining the depth call — because depth is the dimension a future reader of that file will most want justified.

You can automate the record-keeping without automating the judgment, and that split is the whole trick. An n8n job can open the score file when a draft is marked finished, prompt for the five lines, and file the result; Claude Code can pre-fill the mechanical observations — whether the piece answers a real query, whether the point is near the top, whether it repeats something already published. What no tool should do is issue the verdict, because a model asked whether a draft is good will tell you yes.

Making scoring a habit, not a bottleneck

The risk with any quality gate is that it turns into bureaucracy — a form people fill in without thinking, or a delay that makes publishing feel like paperwork. Keep it light. Scoring should take a few minutes, not a meeting. Five dimensions, an honest read, a decision. If it becomes a ritual where the number is the point and the judgment is skipped, it's actively harmful, because it launders weak content through a process that looks rigorous.

What happens to a scoring habit over its first year

  • Week one

    "This is obviously useful"

    The first cut piece is a small shock and settles the standard faster than any amount of discussion.

  • Month two

    "Most things pass"

    Correct, and not a sign the gate is redundant — the drafts changed because the gate exists.

  • Month four

    "Do we still need to do this?"

    The dangerous moment. Scoring becomes a form somebody fills in afterwards to justify a decision already made.

  • Month six onward

    "It takes three minutes"

    If it survived month four, it is now just how a draft gets finished, and nobody debates it.

Fig 08 · Month four decides itA gate that becomes paperwork is worse than no gate, because it launders weak content through a process that looks rigorous. The test is whether anything still gets cut.

The way to keep it real is to keep it about the decision, not the record. You're not producing a grade for a spreadsheet; you're answering one question — publish, revise, or cut — as honestly as you can, the same way every time. Done right, scoring doesn't slow good content down at all; it clears quickly. What it slows down is weak content, which is exactly what you want slowed.

One structural safeguard is worth building in: whoever wrote the piece should not be the only person who scores it. Not because writers can't be honest, but because four minutes after finishing something is the worst possible moment to ask whether it says anything new. A second reader with the same five dimensions in front of them takes about the same four minutes and catches the two dimensions that get fudged. On a team of one, the substitute is time — score it the next morning, not the moment you stop typing.

What a score cannot tell you

Being honest about the limits keeps the habit from turning into superstition. A score is a judgment made before anyone has read the piece except you, which means there are things it structurally cannot see — and knowing which things stops you trusting the number further than it deserves.

Four questions, in order of how confidently a pre-publish score can answer them

  1. 01Is it clear?

    Answerable on the page. A score is genuinely reliable here.

  2. 02Is it ours?

    Answerable if you know the business. Reliable, and the one people fudge.

  3. 03Will anyone find it?

    A reasonable guess. The query is real; whether you rank for it is not yours to decide.

  4. 04Will it change anyone's mind?

    Unknowable in advance. This is the one that actually matters, and no rubric reaches it.

Fig 09 · Confidence falls as importance risesThe score is most certain about the question that matters least. That isn't an argument against scoring — it's the reason a score is a floor and never a prediction.

Two failure modes follow from that. The first is treating a high score as a promise: the piece cleared the bar, so it must perform, and when it doesn't the standard gets blamed. It shouldn't be — clearing the gate means the piece deserved publishing, not that it will work. The second is scoring in a way that optimises for the rubric rather than the reader, which is how you end up with content that is measurably strong on all five dimensions and somehow still inert.

The correction for both is the same: the score decides publish, revise or cut, and nothing else. What actually happened to a piece is a separate question, answered months later by looking at what people did with it — and that answer belongs in the audit, not in the gate.

Why a system raises every score

Here's the thing scoring reveals over time: the same weaknesses keep showing up. Pieces are thin because the knowledge behind them was thin. They're irrelevant because there was no clear reader in mind. They're hard to quote because there was no structure to write into. Scoring catches these one at a time — but a good system prevents most of them before a draft exists.

The same score distribution, at three stages of building the system

Before any system

Everything from excellent to unpublishable, with no pattern to it.

Knowledge organised

Depth stops being the recurring failure, because the drafts have something to draw on.

Voice and structure too

The floor rises. Most drafts now arrive at the gate already past it.

Three shapes, not three measurements — left to right in each plot is the score, and height is how many pieces landed there

Fig 10 · The gate stops being where quality is decidedScoring never gets better at its job. What changes is what arrives at it — which is why a run of cuts for the same reason is a message about the system, not about the drafts.

This is also the answer to the objection that scoring is just gatekeeping your own work. It would be, if the score were the end of the process. It isn't — it's an instrument pointed upstream. Every cut carries a reason, and reasons accumulate into a pattern: three pieces cut on depth in a quarter is not three bad drafts, it's a knowledge layer that hasn't been fed. Four cut on relevance is a topic-selection problem, not a writing problem. The gate is where you find that out cheaply.

That's the real relationship between a score and a Content OS. When your knowledge is organized, your topics are chosen deliberately, and your structure is a given, most drafts arrive at the scoring gate already strong — because the system did upstream what the score only measures at the end. Scoring stays valuable as the honest final check. But if you find yourself cutting piece after piece for the same reason, the score isn't the problem to fix. The system that keeps producing that weakness is. A good score is a symptom of good work, and good work is what a system is for.

Who does the scoring

There's a question underneath all of this that gets skipped surprisingly often, and it decides whether any of the rest works: who is holding the scorecard? A standard is only as reliable as the person applying it, and the obvious answer — the writer scores their own draft — has a known and well-documented problem.

You cannot see the gap between what you meant and what you wrote. The draft reads clearly to you because you have the missing context in your head, supplying it silently every time your eye passes over a thin sentence. That is not carelessness; it is how writing works for everybody, and it is exactly why self-scored specificity and usefulness run consistently high. If one person is the only reader before publication, expect the scores to be roughly a point generous on the two dimensions that matter most, and calibrate the threshold accordingly rather than pretending otherwise.

The cheap fix is a second pair of eyes that is not required to be an expert. A colleague who knows nothing about the subject is better at judging clarity than one who knows it well, because they cannot fill in the gaps for you — if they can't say what the piece argues after one read, the score on clarity is wrong regardless of what you gave it. Five minutes of somebody else's attention catches more than an hour of your own re-reading, and it doesn't need seniority to do it.

For a solo founder with nobody to ask, the substitute is time and a change of medium. Score it the next morning rather than the night you finish, and read it somewhere other than the document — on a phone, or out loud. Both work by breaking the familiarity that hides the gap. A model can help here too, and its useful question is not is this good — it will say yes — but what does this piece claim, and what would someone need to already know for it to make sense? The answer tells you what you left in your head.

Frequently asked questions

What is content scoring?

Content scoring is evaluating a piece of content against a consistent set of quality criteria before you publish it, so you catch weak content before it goes out rather than after. It turns "is this good enough?" from a gut call into a repeatable check across a few dimensions — relevance, depth, clarity, search fit, and business value — that everyone applies the same way.

How is content scoring different from a content audit?

Scoring judges one piece before it's published; an audit reviews your whole existing library after the fact. Scoring is a gate you pass new work through; an audit is an inventory of what's already out there. You use scoring to keep quality high going forward, and an audit to find and fix what was published before you had a standard.

Should I use a numeric score?

A rough number helps if it forces a decision — publish, revise, or cut — but the number matters less than the dimensions behind it. The goal isn't a precise grade; it's a consistent, honest check that stops weak content going out under your name. If a score ever becomes a box to tick rather than a real judgment, it's doing harm, not good.

Doesn't scoring slow down publishing?

A little, and that's the point. The businesses that win with content publish less but better. A short check that catches a thin piece before it's published saves the far larger cost of diluting your authority with content that was never going to earn attention, a ranking, or a citation. Scoring trades a few minutes now for not spreading your credibility thin later.

What are the five dimensions of content quality?

Relevance, depth, clarity, search fit and business value. Relevance asks whether the piece is written for one real reader with one real question. Depth asks whether it says anything only you could say. Clarity asks whether a reader or a machine can lift the point without digging. Search fit asks whether it resolves a genuine query and is structured to be found. Business value asks whether it opens a path toward a conversation. A piece does not need to be strong on all five, but a piece weak on most of them is not ready however finished it looks.

Who should score a draft?

Ideally not only the person who wrote it. Four minutes after finishing something is the worst moment to judge whether it says anything new, so a second reader working from the same five dimensions catches the two that get fudged — depth and business value. On a team of one the substitute is time: score it the next morning rather than the moment you stop typing.

Can AI score content for you?

It can do the mechanical observations and should not issue the verdict. A tool like Claude Code is reliable on whether a piece answers a real query, whether the point sits near the top, and whether it repeats something you already published — and an n8n job can open and file the score record so the habit leaves a trace. What it cannot do is tell you whether a draft is good, because a model asked that question will say yes. Automate the record-keeping, never the judgment.

Where to start

Don't build an elaborate system. Take the next piece you're about to publish and run it through the five dimensions honestly — relevance, depth, clarity, search fit, business value — and be strict on the two people fudge. Decide: publish, revise, or cut. Do that for a few pieces and two things happen. Your published work gets noticeably stronger, and you start to see the patterns in what you almost published — the recurring weaknesses that point back to something upstream worth fixing. Scoring is a small habit with a large effect: it makes "good enough" a decision you make on purpose, instead of one that gets made for you by a deadline.

Keep reading

Publishing more and hearing less?

If more content isn't turning into more results, the problem usually isn't volume — it's that nothing's stopping the weak pieces before they go out. We build the standard and the system behind it, so what you publish is worth publishing. Tell us what you're putting out, and we'll show you where the quality is leaking.

Start a Conversation