A1 Jobs
Field Notes · Verified Prep

Benchmark · 10 real CVs · June 2026

The CV looks great. About 40% of it isn’t true.

We took the workflow most job-seekers already use — paste your CV into a general-purpose AI chatbot and ask it to tailor — and graded every single bullet against the candidate’s real history. The chatbot writes the best-looking résumé in the room. It also fabricates the most.

By The A1 Jobs team Method: same rubric, per-bullet evidence grader Claude Opus 4.8 chatbot workflow vs A1’s pipeline

Updated 8 July 2026 — product description refreshed; benchmark run June 2026 on the engine noted.

40.6% of the bullets Claude Opus 4.8 wrote were not grounded in the candidate’s real CV.
+24 authenticity points: A1’s grounded single-shot (80.6) over Opus 4.8 (56.4), pooled.
9/10 CVs where A1’s grounded tailoring was more honest than the chatbot.

Tailoring your CV to a job is now a 30-second habit — and for most job-seekers it happens in a general-purpose AI chatbot, not in a CV tool: paste your résumé into ChatGPT or Claude, paste the job description, ask it to “tailor my CV to maximise fit for this job.” The chatbot rewrites your summary, reshuffles your experience toward what the posting asks for, and salts in the right keywords for the applicant-tracking system. The output is genuinely impressive — clean, confident, on-message. To be clear about what we tested: that paste-into-a-chatbot workflow is the baseline in this study — it is what people do today outside A1. A1 itself is not a chatbot and never has been.

So we ran the experiment properly. We took 10 real, publicly-published CVs — de-identified by role, spanning new-grad to 20-plus-year leaders — and paired each with a real job description from public company job boards (Stripe, Databricks, GitLab, Figma, via Greenhouse). Then we tailored every CV two ways — once through that chatbot workflow, once through A1’s pipeline — and scored both on identical instruments. One question drove everything: when an AI makes your CV look better, where do the new claims come from?

01 / Method

What we measured, and how

Every tailored CV went through the same two graders, so the comparison is apples-to-apples.

A 6-dimension quality rubric (“how good does it look”)

A 0–100 score across skill-keyword coverage, experience relevance, achievement quantification, section completeness, ATS compatibility, and a holistic authenticity impression. (In the product UI that sixth dimension is now labelled Faithfulness, to keep it distinct from the per-bullet authenticity score below.) This is the surface read — the thing a recruiter or an ATS reacts to in the first six seconds.

A per-bullet, evidence-grounded authenticity grader

The honest part. Every tailored bullet is checked against the candidate’s real CV and labelled one of three ways:

  • grounded — directly supported by the real CV.
  • reframed — the same fact, given new emphasis for the role. Still fair.
  • unsupported — a claim, metric, tool, or scope that simply isn’t in the real CV.

For this June 2026 run the authenticity score weighted grounded bullets fully and reframed bullets at 0.6: grounded × 1.0 + reframed × 0.6, divided by graded bullets, × 100. We also report the raw share of bullets that are unsupported. (The live product score has since been retuned to be attestation-aware: grounded × 1.0, reframed × 0.85, and an unsupported bullet earns 0.6 only when the candidate explicitly confirms they can defend it — the grade itself stays unsupported, so provenance is never rewritten.)

The two tailorers — a chatbot workflow vs A1’s pipeline

On one side, the baseline: a general-purpose AI chatbot — Claude Opus 4.8, Anthropic’s most capable model — role-playing exactly what a job-seeker does today outside A1: paste CV, paste JD, ask it to tailor. That is the everyday ChatGPT/Claude workflow, not an A1 feature.

On the other, A1’s tailoring pipeline, in the three versioned engine variants of this June 2026 run: structured v1 (the platform default when the benchmark ran), structured-grounded v2 (v1 plus hard anti-fabrication grounding at generation time — promoted to the platform default on 29 June 2026 after beating v1 on authenticity on all 10 CVs), and grounded single-shot (a benchmark-only experiment). A1 is not a chatbot — there is no free-text prompt box. It is a structured pipeline: parse the CV into evidence, generate a rewrite constrained by that evidence, then grade every bullet against it — so each bullet carries a verdict on whether it’s real, reframed, or needs review. The comparison, in one line: A1’s pipeline vs the paste-into-a-chatbot workflow.

02 / The core finding

Polished and dishonest are not the same axis

Plot every tailorer on two axes — how good it looks (quality) against how honest it is (authenticity) — and the story separates cleanly. The frontier chatbot sits top-right on looks and bottom on honesty. A1’s grounded variants give up a few points of polish to climb the honesty axis.

Figure 1 · Pooled means, all 10 CVs

Best-looking ≠ most honest

Each point is one tailorer. Right = looks better; up = more honest. Claude Opus 4.8 wins on looks (78.8) and loses on honesty (56.4); A1’s grounded single-shot tops the honesty axis at 80.6 for a few points of surface polish. Engine names are the June-2026 versions; grounded v2 has been A1’s default engine since 29 June 2026.
The frontier chatbot writes the best-looking CV in the room — and the most fabricated one.

That trade isn’t a rounding error. Opus 4.8’s 40.6% unsupported rate means that, on average, two of every five tailored bullets introduced a claim, number, tool, or scope the candidate never actually wrote down. A reframe is fair game. An invented metric is a landmine in an interview.

03 / Per-CV

The gap is widest where it’s most dangerous

Pool-level averages hide the texture. Broken out CV by CV — Opus 4.8 against A1’s grounded single-shot — A1 is the more honest tailorer on 9 of 10. And the honesty gap stretches widest exactly where you’d least want a confident invention.

Figure 2 · Authenticity per CV (0–100)

A1 grounded vs Claude Opus 4.8, all 10 CVs

Sorted by A1’s lead. The honesty gap is largest on thin / junior CVs — design director (+45.5), new-grad SWE (+44.5), research SWE (+43.6) — the candidates most likely to trust a confident rewrite. A1’s single loss is on an unusually dense senior CV (data engineer, −7.0).

Read that pattern again. The candidates with the least material to work from — the new grad, the career-changing design director, the researcher with a short industry track record — are precisely the ones the chatbot “helps” most enthusiastically, by inventing the experience they don’t have. They’re also the candidates least able to spot it, and most likely to get caught defending it in an interview.

04 / Pooled

Honesty, side by side, all four

Across all four tailorers, the relationship holds: as authenticity rises, the unsupported share falls. The frontier chatbot anchors the dishonest end; A1’s grounded single-shot anchors the honest one.

Figure 3 · Four tailorers, pooled across 10 CVs

Authenticity vs unsupported bullets

Higher authenticity, lower fabrication is the goal. A1’s grounded single-shot reaches 80.6 authenticity versus the chatbot’s 56.4; the chatbot leaves 40.6% of bullets unsupported, nearly double A1’s grounded variants.
05 / Cost & speed

The honest output is also the cheap, fast one

You might expect honesty to cost more. It’s the opposite. A1 runs its grounded tailoring on a small, fast model — roughly 50× cheaper per token than Opus 4.8. In this benchmark, the grounded single-shot variant produced a graded, grounded CV in a single pass of about five seconds; the production pipeline spends a few more steps on the same small model (parse → grounded generation → per-bullet grading) — still a fraction of the cost of one Opus-class call. (Pricing is list price and approximate; these numbers move over time.)

Claude Opus 4.8

Input / 1M tokens$5.00
Output / 1M tokens$25.00
For scale — Haiku 4.5$1.00 / $5.00
Per-bullet provenance— none —

Best-looking output, no grounded grade.

A1 — grounded tailoring

Input / 1M tokens~$0.10
Output / 1M tokens~$0.40
Relative cost~50× cheaper
Single-shot benchmark pass~5s

Cheaper, faster — and it ships the grade.

A1 is not trading honesty for cost. The grounded output is the cheap, fast one — and it’s the only one that hands back a per-bullet verdict a single chatbot pass structurally can’t produce.

06 / Caveats

What this benchmark is — and isn’t

Read the data honestly, too

  • The grader is strict. It also flags generic skills-list lines and boilerplate as unsupported for both sides, so the absolute unsupported-% is conservative. The robust signal is the gap between tools, not the raw number.
  • n = 10 real CVs. This is an illustrative study, not a giant benchmark. Treat it as directional.
  • Different judge, on purpose. The grader runs on a different model family (Google Gemini) than the Opus baseline it judges. If anything, that strengthens the finding — an independent model flags the same fabrication.
  • This is the June 2026 run. Engine names are the versions tested then. Since the run, structured-grounded v2 has been A1’s default engine (promoted 29 June 2026 after beating v1 on authenticity on all 10 CVs), and the live authenticity score has been retuned to credit reframed bullets at 0.85 and candidate-confirmed bullets at 0.6 (see Method).
What A1 ships

A grade you can trust, not just prose you can’t.

The general-purpose chatbot most candidates already paste their CVs into makes them look great and quietly fabricates around 40% of their bullets. A polished CV you can’t defend in an interview isn’t an asset — it’s a liability with good formatting.

A1’s moat is Verified Prep: tamper-evident, candidate-consented proof that a tailored CV is grounded in real evidence. Every bullet carries a verdict — grounded, reframed, or flagged for review — so candidates ship CVs they can stand behind, and recruiters get a trustworthy signal instead of a confident guess. That’s the thing a single LLM pass can’t give you.

See how the badge and the signed API work on the Verified Prep page, try the pipeline from the A1 Jobs home page, or see pricing.

Verified Prep · grounded in real evidence