Overnight Run Report — Website Grading + Personalisationas of 2026-08-11

ClientsFlow lead pipeline · internal · run window: 2026-08-10 evening → 2026-08-11 morning

0What Happened After This Report Was Written as of 2026-08-11 11:15

Sections 5 and 6 below are superseded. Decisions 1, 2 and 4 were resolved during the day and the leads are already staged. Read this block first.

  • 278 leads now sit in a paused Instantly campaign CF_2026-08-11_graded-review-220_PAUSED, status 0 (draft, not sending). Full live read-back: 0 blank merge fields, 0 fabricated greeting names. Nothing goes out until you press start.
  • The quarantine was a false alarm. Both batches of activity_variants_mismatch rows (41 this morning, 40 this afternoon) were released after every greeting name was checked against the site's own scraped copy. 81 checked, 80 verified present, 1 dropped.
  • Grading was stopped early on purpose. Roughly half of every grading request is a finished, billed grade that the grader's own ceiling then discards. Measured, not estimated: completed grades cost $0.0048 each, discarded ones $0.0147. The binding limit is the 15,000-token output ceiling, blown by reasoning at effort max. Input trimming cannot fix it, because prompts are already capped at 28,000 characters before they are sent.
  • Two things now need you, both new: whether to lower the grader's reasoning effort or raise the output ceiling (either recovers roughly half the grading spend, both change a proven system, so neither was done), and one lead dropped because its greeting and its stored contact name name two different people.
  • One older item is still live: two July leads were emailed the wrong first name and sit in paused campaigns. Resuming those campaigns re-sends the wrong name. Pull them first.

1TL;DR

Review-ready leads
29
+43 releasable on one decision
Approved, awaiting upload
220
Matt-approved, Codex handoff ready
Personalisation champion
v12 · 70.0%
honest score, up from 64.0% (v11 re-baseline)
Total spend, tonight
$9.19
refine $5.41 + grade $3.06 + personalise $0.72
OpenRouter credit left
$0.44
blocks all further paid grading/personalisation
Blocker

2Grading & Personalisation Funnel as of 2026-08-11

Fail-closed losses (dropped pre-grade, recoverable by design): SSL fetch failure: 147 Response/session-ceiling: 142 Connection error: 5
Funnel reads master queue → unprocessed pool → live → graded → complete → keepers (score<7) → personalisation generated → validator KEEP → review-ready. The 43 quarantined rows are not lost — see Section 4.

3Personalisation Refinement Loop (v9 → v12) as of 2026-08-11

VersionHonest scoreMutationPromoted?
v962.0%Simpler profession words for p2No
v1063.3%More specific site-derived doersNo
v1164.0%Strict mirror of p1No
v1270.0%Routing: KEEP is the default, FILTER_OUT needs proofChampion
Re-baseline note: scrapes were rebuilt after a tmp wipe; scores are comparable only within one scrape build. v7 (the prior champion) re-measured at 64.0% on the new 2026-08-10 scrapes, having scored 70.0% on the old ones — the like-for-like comparison that put v12 in front is v12 (70.0%) vs. v7 re-baselined (64.0%), not v12 vs. v7's old number.
Lesson — routing fixes pay three times: p2 resisted three direct wording mutations (v9–v11) without moving the needle. The v12 fix touched routing logic, not wording: a wrongly filtered-out row fails all three graded fields at once (greeting, p1, p2), so fixing the routing decision recovers all three simultaneously. Loop stopped on its own $6.50 refinement cap ($5.41 spent); 90% target not reached.

4Personalisation Results (99-lead cohort) as of 2026-08-11

Scraped
99
Rescued by rescrape
9
No scrape
4
Generated
95
Keep
81
Filter out
12
Reputation
2
Accommodation
0
Fabricated-name guard: working as designed. 6 rows had a fabricated name dropped before output — the guard caught them, not a defect in the run.
Quarantine story: 46 rows were validator-quarantined. Of those, 43 of 46 were flagged by a single rule — activity_variants_mismatch, which demands that personalisation2 be an exact stem mirror of personalisation1. The v12 champion deliberately does not enforce this (v11 tested strict mirroring and only tied v10, it did not win). This means 43 of the 46 quarantined rows are quarantined by a rule the current champion prompt intentionally does not follow — releasing them is one policy decision, not 43 individual reviews. Spend for this stage: $0.72.

5Decisions Needed From Matt as of 2026-08-11

1. Release the 43 activity_variants_mismatch quarantine rows to review?  RESOLVED
Recommendation: Release. The rule they're held on contradicts the v12 champion's own routing policy (see Section 4); no other validator flag applies to these 43 rows.
2. Top up OpenRouter (~$20) to finish grading + personalising toward 1,000 keepers?  RESOLVED — topped up, key rotated, $2.67 spent, run stopped early on the waste finding
Recommendation: Top up. Current balance is $0.44 and blocks every further paid stage — grading the remaining ~2,200 live sites and personalising the remaining ~850 keepers both need it.
3. Grader 0.08pt promote/discard call, left open by the grader session.
Recommendation: His call — the margin is inside noise floor; no clear machine recommendation either way.
4. Retry the 147 SSL-failed sites via http:// URLs?  RESOLVED — 143 of 147 recovered
Recommendation: Yes, retry. Estimated cost ~$1.10 to recover leads that failed closed on https:// only.

6Path to Instantly SUPERSEDED — see Section 0 as of 2026-08-11

  1. Review 29 + 43 CSV rowsMatt reviews graded99_review_2026-08-11.csv (29 new) and graded99_quarantine_2026-08-11.csv (43 releasable, pending decision 1).
  2. Merge approved rows into the upload CSVApproved rows join clientsflow_instantly_upload_2026-08-10.csv (currently 220 rows).
  3. Codex handoff executes the 220-row uploadHandoff docs already written: HANDOFF_instantly_upload_codex.md + PROMPT_for_codex.md. Nothing uploaded yet.
  4. Top-up unlocks grading the remaining ~2,200 live sitesToward the 1,000-keeper target; blocked on decision 2 ($0.44 remaining balance).
  5. Weekly-1200 pipelineLonger-term: repeatable weekly batch built on the lead-bank runners and grader/personalisation adapter (see context dossier, Section 3).

7File Index