Files
justvitamin/content_population_exports/apify_gapfill_runbook.md
T
Omair Saleh 056c47581f feat: editorial review dashboard + elite-grade pilot batch (5 SKUs)
Ships the second dashboard surface — a Pattern Library + Preview Theatre — that
presents the 4-section PDP pilot batch back to Umar, compliance, and the board
in an editorial format. Adds the full data layer that drives it: 5 source-backed
per-SKU drafts at QA 100/100, 15 competitor PDP semantic extracts, PubMed
evidence packs, EFSA claims library extension, JV brand voice guide, hand-curated
product FAQs, and the Matrixify-ready CSV exports for Lewis.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-02 18:50:09 +08:00

1.5 KiB

Apify selective gap-fill packet

Generated: 2026-05-20T16:41:07.121Z

Top-3 scraping budget raised for hundreds of items. Run capped first-batch targets, inspect the dataset, then decide whether to scale toward 1000 items per product/source.

Spend guardrails

  • DataForSEO Stage 1 first; Apify Stage 2 only for named gaps.
  • One actor lane, first 3 brands, max 100 requests per brand, max depth 1, then stop for dataset inspection.
  • Do not run all 7 competitor domains in one go.
  • Save raw datasets under data/sources/apify/raw/ and inspect before any normalization/scaling.

First batch only if needed

Hold until first-batch review

Files

  • content_population_exports/apify_gapfill_targets.csv - all Stage 2 targets with run/hold status.
  • content_population_exports/apify_competitor_pdp_input_template.json - no-spend actor input template for the first batch only.
  • data/sources/apify/raw/_apify-output-template.json - expected local raw output shape after a run.

Stop / scale rule

Scale only if the first dataset contains useful PDP/review/claim evidence with source URLs and the cost is acceptable. Otherwise keep Apify off and stay with DataForSEO/manual source drops.