A verified, tested personalisation skill

Prepare once, test blindly, adjudicate independently, and iterate only when the evidence requires it.

One main path, with bounded repair loops

1 · Freeze the benchmark
Lock 200 leads: 100 accepted unchanged and 100 pre-correction cases.
Success criteria
  • Every case has adequate website evidence and a stored hash.
  • The candidate renders inside “Láttam, hogy [p1] foglalkoztok.”
Constraints
  • No duplicates, leaked corrections, or Personalisation 2 scoring.
  • The set cannot change after testing starts.
Benchmark valid?200 cases · evidence · hashes · no leakage No
Repair the benchmark
Success: replace only invalid cases and rehash.
Constraint: do not inspect model outputs yet.
Yes 2 · Run both existing skills once
Run the current skill and recovered Claude skill on identical evidence.
Success criteria
  • Every case receives one output or a recorded provider failure.
  • Outputs are stored unchanged with model settings.
Constraints
  • No content retries, prompt edits, examples, or campaign uploads.
  • Retry infrastructure failures only, with the same inputs.
Run valid?all outputs or recorded provider failures No
Rerun failed cases
Success: provider failures are resolved or retained.
Constraint: inputs and prompt stay frozen.
Yes 3 · Judge independently
A fresh judge reviews every non-SEND output plus a random SEND sample.
Success criteria
  • SEND is truthful, representative, grammatical, and natural.
  • Alternative good wording passes; disagreements cite a failed property.
Constraints
  • The judge cannot see the first verdict, correction, skill, or score.
  • Style preference alone cannot cause REWRITE.
inconsistent
Clarify the rubric
Success: conflicts are reconciled and rejudged.
Constraint: do not change generated outputs.
Adjudicated pass?≥190/200 SEND · each group ≥90 · zero WRONG Yes · no iteration No 4 · Improve the stronger skill
Fix observed error classes on separate development cases, then run one terminal locked test.
Success criteria
  • Terminal result reaches ≥190/200 SEND, each group ≥90, zero WRONG.
  • The winning prompt and files are frozen.
Constraints
  • Maximum five rounds or stop after three rounds without improvement.
  • Locked cases cannot guide edits; no repeated terminal tests.
Terminal test passes?same locked thresholds · one terminal attempt No
Stop with evidence
Success: preserve the best version, score, errors, and next hypothesis.
Constraint: do not promote an unverified skill.
Yes 5 · Release the exact tested skill
Install the frozen winner as the canonical personalisation skill.
Success criteria
  • Installed files match tested hashes; a clean invocation succeeds.
  • Benchmark, settings, verdicts, and adjudication are retained.
Constraints
  • No untested rewrite during installation.
  • No lead upload, campaign modification, or email sending.
Exact tested release?hash match · clean invocation · receipt saved No
Repair the installation
Success: restore exact files and rerun clean check.
Constraint: tested content cannot change.
Yes Done · verified skill released
Success: adjudicated thresholds passed and installed hashes match.
Constraint: campaign activation remains separate.

SEND means true, representative, grammatical, and natural. REWRITE means the fact is usable but the wording is weak. WRONG means false or materially misleading. NO_EVIDENCE means the source cannot support a safe claim.