Tailoring your CV to a job is now a 30-second habit — and for most job-seekers it happens in a general-purpose AI chatbot, not in a CV tool: paste your résumé into ChatGPT or Claude, paste the job description, ask it to “tailor my CV to maximise fit for this job.” The chatbot rewrites your summary, reshuffles your experience toward what the posting asks for, and salts in the right keywords for the applicant-tracking system. The output is genuinely impressive — clean, confident, on-message. To be clear about what we tested: that paste-into-a-chatbot workflow is the baseline in this study — it is what people do today outside A1. A1 itself is not a chatbot and never has been.
So we ran the experiment properly. We took 10 real, publicly-published CVs — de-identified by role, spanning new-grad to 20-plus-year leaders — and paired each with a real job description from public company job boards (Stripe, Databricks, GitLab, Figma, via Greenhouse). Then we tailored every CV two ways — once through that chatbot workflow, once through A1’s pipeline — and scored both on identical instruments. One question drove everything: when an AI makes your CV look better, where do the new claims come from?
What we measured, and how
Every tailored CV went through the same two graders, so the comparison is apples-to-apples.
A 6-dimension quality rubric (“how good does it look”)
A 0–100 score across skill-keyword coverage, experience relevance, achievement quantification, section completeness, ATS compatibility, and a holistic authenticity impression. (In the product UI that sixth dimension is now labelled Faithfulness, to keep it distinct from the per-bullet authenticity score below.) This is the surface read — the thing a recruiter or an ATS reacts to in the first six seconds.
A per-bullet, evidence-grounded authenticity grader
The honest part. Every tailored bullet is checked against the candidate’s real CV and labelled one of three ways:
- grounded — directly supported by the real CV.
- reframed — the same fact, given new emphasis for the role. Still fair.
- unsupported — a claim, metric, tool, or scope that simply isn’t in the real CV.
For this June 2026 run the authenticity score weighted grounded bullets fully and reframed bullets at 0.6: grounded × 1.0 + reframed × 0.6, divided by graded bullets, × 100. We also report the raw share of bullets that are unsupported. (The live product score has since been retuned to be attestation-aware: grounded × 1.0, reframed × 0.85, and an unsupported bullet earns 0.6 only when the candidate explicitly confirms they can defend it — the grade itself stays unsupported, so provenance is never rewritten.)
The two tailorers — a chatbot workflow vs A1’s pipeline
On one side, the baseline: a general-purpose AI chatbot — Claude Opus 4.8, Anthropic’s most capable model — role-playing exactly what a job-seeker does today outside A1: paste CV, paste JD, ask it to tailor. That is the everyday ChatGPT/Claude workflow, not an A1 feature.
On the other, A1’s tailoring pipeline, in the three versioned engine variants of this June 2026 run: structured v1 (the platform default when the benchmark ran), structured-grounded v2 (v1 plus hard anti-fabrication grounding at generation time — promoted to the platform default on 29 June 2026 after beating v1 on authenticity on all 10 CVs), and grounded single-shot (a benchmark-only experiment). A1 is not a chatbot — there is no free-text prompt box. It is a structured pipeline: parse the CV into evidence, generate a rewrite constrained by that evidence, then grade every bullet against it — so each bullet carries a verdict on whether it’s real, reframed, or needs review. The comparison, in one line: A1’s pipeline vs the paste-into-a-chatbot workflow.
Polished and dishonest are not the same axis
Plot every tailorer on two axes — how good it looks (quality) against how honest it is (authenticity) — and the story separates cleanly. The frontier chatbot sits top-right on looks and bottom on honesty. A1’s grounded variants give up a few points of polish to climb the honesty axis.
Figure 1 · Pooled means, all 10 CVs
Best-looking ≠ most honest
The frontier chatbot writes the best-looking CV in the room — and the most fabricated one.
That trade isn’t a rounding error. Opus 4.8’s 40.6% unsupported rate means that, on average, two of every five tailored bullets introduced a claim, number, tool, or scope the candidate never actually wrote down. A reframe is fair game. An invented metric is a landmine in an interview.
The gap is widest where it’s most dangerous
Pool-level averages hide the texture. Broken out CV by CV — Opus 4.8 against A1’s grounded single-shot — A1 is the more honest tailorer on 9 of 10. And the honesty gap stretches widest exactly where you’d least want a confident invention.
Figure 2 · Authenticity per CV (0–100)
A1 grounded vs Claude Opus 4.8, all 10 CVs
Read that pattern again. The candidates with the least material to work from — the new grad, the career-changing design director, the researcher with a short industry track record — are precisely the ones the chatbot “helps” most enthusiastically, by inventing the experience they don’t have. They’re also the candidates least able to spot it, and most likely to get caught defending it in an interview.
Honesty, side by side, all four
Across all four tailorers, the relationship holds: as authenticity rises, the unsupported share falls. The frontier chatbot anchors the dishonest end; A1’s grounded single-shot anchors the honest one.
Figure 3 · Four tailorers, pooled across 10 CVs
Authenticity vs unsupported bullets
The honest output is also the cheap, fast one
You might expect honesty to cost more. It’s the opposite. A1 runs its grounded tailoring on a small, fast model — roughly 50× cheaper per token than Opus 4.8. In this benchmark, the grounded single-shot variant produced a graded, grounded CV in a single pass of about five seconds; the production pipeline spends a few more steps on the same small model (parse → grounded generation → per-bullet grading) — still a fraction of the cost of one Opus-class call. (Pricing is list price and approximate; these numbers move over time.)
Claude Opus 4.8
Best-looking output, no grounded grade.
A1 — grounded tailoring
Cheaper, faster — and it ships the grade.
A1 is not trading honesty for cost. The grounded output is the cheap, fast one — and it’s the only one that hands back a per-bullet verdict a single chatbot pass structurally can’t produce.
What this benchmark is — and isn’t
Read the data honestly, too
- The grader is strict. It also flags generic skills-list lines and boilerplate as unsupported for both sides, so the absolute unsupported-% is conservative. The robust signal is the gap between tools, not the raw number.
- n = 10 real CVs. This is an illustrative study, not a giant benchmark. Treat it as directional.
- Different judge, on purpose. The grader runs on a different model family (Google Gemini) than the Opus baseline it judges. If anything, that strengthens the finding — an independent model flags the same fabrication.
- This is the June 2026 run. Engine names are the versions tested then. Since the run, structured-grounded v2 has been A1’s default engine (promoted 29 June 2026 after beating v1 on authenticity on all 10 CVs), and the live authenticity score has been retuned to credit reframed bullets at 0.85 and candidate-confirmed bullets at 0.6 (see Method).
A grade you can trust, not just prose you can’t.
The general-purpose chatbot most candidates already paste their CVs into makes them look great and quietly fabricates around 40% of their bullets. A polished CV you can’t defend in an interview isn’t an asset — it’s a liability with good formatting.
A1’s moat is Verified Prep: tamper-evident, candidate-consented proof that a tailored CV is grounded in real evidence. Every bullet carries a verdict — grounded, reframed, or flagged for review — so candidates ship CVs they can stand behind, and recruiters get a trustworthy signal instead of a confident guess. That’s the thing a single LLM pass can’t give you.
See how the badge and the signed API work on the Verified Prep page, try the pipeline from the A1 Jobs home page, or see pricing.
Verified Prep · grounded in real evidence