The Same Resume, Two Different Scores

Here is an experiment worth running with any AI-powered resume checker. Scan your resume against a job description, note the score, then scan the exact same resume against the exact same job description a few minutes later. With a surprising number of tools, the number moves — sometimes by a little, sometimes by a lot.

For a job seeker, that should be unsettling. If a score can swing from a 72 to an 88 without a single word changing, then which number was real? Did you just barely fail a threshold on a bad roll of the dice? A resume score is only useful if you can trust it to mean something — and a number that changes when nothing else did means nothing at all.

This isn't a fringe edge case. It's a structural property of how a lot of AI scoring is built. And understanding why it happens is the key to understanding what a trustworthy resume score should look like.

An Open Experiment the Whole Field Should Thank

In 2026, HackerRank did something genuinely valuable for the industry: they open-sourced an AI resume-grading system and, to their credit, published honest data about how it behaved. That transparency is rare and worth applauding — most companies bury results like these.

What the data showed became a widely-discussed lesson. When the same resume was scored many times over, the results spread across a wide band rather than landing on a stable number. A resume that should have had one quality could score anywhere across a range of tens of points from run to run. And critically, the usual engineering fix — turning the model's "temperature" (its randomness setting) down to zero — did not make the problem go away.

The team themselves described the effect in memorable terms: closer to a luck filter than a quality assessment. That is an unusually candid thing to say about your own system, and it moved the conversation forward for everyone building in this space. We treat it as a gift: a public, well-documented demonstration of a trap that is otherwise easy to fall into quietly.

Why It Happens: When the Judge Is a Language Model

The root cause is architectural. If you hand an entire resume and a rubric to a large language model and ask it to return a final score, you have made the model the judge. And language models, by their nature, are probabilistic. They sample from a distribution of possible outputs. Even with randomness dialed down, the sheer complexity of scoring a whole document against a whole rubric leaves enough room for the output to wander. Small, invisible differences in how the model attends to the text produce different verdicts.

There are two other traps that tend to travel with this one:

  • Unanchored rubrics. If a scoring dimension is defined in a sentence or two with no concrete examples of what a 3 versus an 8 looks like, the model has no stable reference. It fills the gap differently each time, and a category like "experience" ends up producing nearly the same score regardless of how senior the candidate actually is.
  • Weighting built for one kind of career. A rubric tuned for software engineers — heavy on open-source contributions and public code — will quietly score a nurse, a structural engineer, a teacher, or a financial analyst near zero. Not because those people are unqualified, but because the yardstick was never built for them.

Put together, you get a score that is unstable, hard to explain, and biased toward a narrow definition of merit. For a tool that people make real career decisions with, that is a serious problem.

The Three Principles We Built Around

When we designed Hyrenora's scoring, these failure modes were the design brief. We wanted a score a person could actually reason about. That led to three non-negotiable principles.

1. Deterministic: the same input always gives the same score

This is the big one. In Hyrenora, an AI model is allowed to extract and understand — to read a job description and pull out the skills it's really asking for, including the ones implied but not stated. But the model is never the final judge. The actual scoring is done by deterministic logic: given the same resume and the same job description, you get the same score, every time. We even test this directly — the scorer is run repeatedly against a fixture set and required to produce identical output, as a gate on every change we ship. No luck filter.

2. Explainable: every score shows its evidence and a fix

A number on its own — "68" — tells you nothing you can act on. So every score in Hyrenora comes with its reasoning. Which job-description keywords you matched and which you're missing. Which of your bullets lack quantification. Whether your contact details are in a form an applicant-tracking system can actually read. The point of a score isn't to rank you; it's to show you exactly what to change and why. If we can't explain a number, we don't show it.

3. Role-agnostic: driven by the job, not by a template of "good"

Hyrenora scores your resume against the specific job description you're targeting — not against a fixed idea of what a strong candidate looks like. A registered nurse applying to an ICU role is measured against what that role asks for. A structural engineer is measured against structural-engineering requirements. There is no hidden assumption that everyone should have a GitHub profile. The job defines the target, so the tool works for the whole labour market, not one slice of it.

What That Looks Like When You Use It

Instead of one opaque number, Hyrenora gives you three honest, separate scores, because they answer three genuinely different questions:

  • JD Match — how well your resume matches this job. This is where semantic matching earns its keep: it recognises that "reinforced concrete" on your resume covers a job description asking for "RC", or that "finite element analysis" answers a call for "FEA" — matches a plain keyword search would miss — while still being fully repeatable.
  • Resume Quality — how well the resume is written: action verbs, quantified impact, structure. Writing advice, honestly labelled as writing advice.
  • ATS Parse-ability — whether a real applicant-tracking system can cleanly read your data at all. This is the actual technical risk, kept separate from how the resume reads to a human.

Splitting them matters. A resume can be beautifully written, parse perfectly, and still be a weak match for a particular job — and you deserve to see that as three clear signals, not blended into a single figure that hides which lever to pull.

The Takeaway

The industry learned something important from an openly-shared experiment: an AI that acts as the final judge of a resume produces scores you can't fully trust, can't easily explain, and can't safely apply across every kind of career. Credit to the teams willing to surface that in public — it made the whole field smarter.

Our conclusion was to design in the other direction: let AI do what it's brilliant at — reading and understanding language — and keep the scoring itself deterministic, explainable, and driven by the real job. A resume score should be something you can reason about, act on, and get the same answer from twice. That's a lower-drama promise than a magic number. We think it's a far more useful one.