# JV Content Population — Plan v2 (2 June 2026) > Supersedes `CONTENT_POPULATION_PLAN.md`. The earlier plan still applies for the dashboard audit, Matrixify schema, and competitor research lanes; this v2 narrows scope to what Umar called out in the 21 May + 1 June Looms and rebases the dashboard around the 8 metaobjects that actually drive the new PDP. --- ## 1. The narrow brief (what Umar actually asked for) Per the 21 May Loom (`55449ee0013148e7936344b83caf9ba8`) and the 1 June Loom (`ce4501e9de834cfdae628c60fe9e41d3`), the immediate scope across the JV catalogue is **4 PDP sections** built from **8 metaobject types**: | PDP section | Metaobject types | Reference design (from 21 May Loom) | |---|---|---| | **Scientific Studies** | `scientific_study` | Wild Nutrition "Food-Grown Magnesium" — headline + paragraph + 2-3 stat tiles + image. **Currently missing from JV template — Umar flagged to Lewis 1 June.** | | **Results / Timeline** | `timeline_phase` + `timeline_block` | Spacegoods 90-Day; JV's own Figma — "Results / What you can expect" with linked scientific study + Month 1/2/3/4+ phases | | **Comparison Tables** | `comparison_column` + `comparison_row` + `comparison_table` | Grüns "Us vs. Them" (3-col, ✓/✗); Club EarlyBird "Brain Drops" (4-col, values + notes) | | **FAQs** | `faq_pair` + `faq_block` | JV Figma — 3 product-specific + 4 standard pairs; collapsible rows; "Still unsure? Get in touch" CTA | Plus the broader Matrixify metaobjects from the original template (still needed but lower urgency): dietary_tag, health_goals, key_ingredients, benefits, clinically_shown_to, promo_card. **Umar's stated priority order (21 May email):** 1. Meta Objects + Content Population 2. NPD Research (Collagen) 3. Social Media Content / Activation 4. Influencer Marketing Scraping --- ## 2. The 8 new metaobject schemas ```text scientific_study - name_internal (handle) - headline (e.g. "Scientifically supported sleep") - body_copy (paragraph; cites study source) - link_url (NIH/PubMed/EFSA citation) - link_label (e.g. "Independent scientific human study") - stat_1_value, stat_1_label - stat_2_value, stat_2_label - stat_3_value, stat_3_label (optional) - image (file_reference) timeline_phase - phase_label (e.g. "Month 1", "Day 30") - body_copy (single paragraph) timeline_block - name_internal - headline (e.g. "Results") - subtitle (e.g. "What you can expect") - intro_copy (1-2 lines that sit above the phases) - study_link_label (e.g. "independent scientific human study") - study_link_url - phases (list.metaobject_reference → timeline_phase) comparison_column - name (e.g. "Grüns", "Generic Multivitamin") - is_us (boolean) - accent_colour (hex) - image (optional thumbnail) comparison_row - feature_label (e.g. "Cost Per Serving", "Taste") - values (json: index → { mark: ✓/✗/star/text, value, note }) comparison_table - name_internal - section_title (e.g. "Us vs. Them", "Not All Brain Fuel Is Created Equal") - subtitle - columns (list → comparison_column) - rows (list → comparison_row) faq_pair - scope ("standard" | "product") - question - answer (multi_line) - source_citation_id (optional → conversion-blocker rank or study id) faq_block - section_title ("Frequently asked questions") - cta_label ("Still unsure? Get in touch") - cta_url - pairs (list → faq_pair) ``` These extend (do not replace) the original Matrixify template's `clinically_shown_to`, `key_ingredients`, `benefits`, etc. The original `faq.heading_one/two/three` trio is **deprecated** — replaced by `faq_block` referencing a list of `faq_pair`s. Lewis confirmation needed that Shopify metaobject `list.metaobject_reference` is supported in his template. --- ## 3. Data source hierarchy (ground truth → directional) Ranked by trustworthiness for citation purposes: | Tier | Source | Status today | Use for | |---|---|---|---| | 1 | **JV website DB** (`tblEcomProduct`, `tblEcomSKUs`, `tblEcomProductCategories`, `tblEcomStockLevels`, `images_by_guid.json`) | Already in `JV Migration to shopify/full_db/` | Product names, ingredients, strengths, prices, variants, categories, stock — ground truth, every comparison/PDP fact lives here | | 2 | **NIH / PubMed** (E-utilities API, free, no auth) | Not built | Peer-reviewed evidence for `scientific_study` stat tiles. Source: https://www.ncbi.nlm.nih.gov/books/NBK25500/ | | 3 | **EFSA Health Claims Register** (UK/EU-approved nutrient/health claims) | Not built | The **legally allowable wording** for any nutrient claim. Critical compliance gate. Public CSV. | | 4 | **Cochrane Library** + **Examine.com** | Not built | Meta-analyses + ingredient evidence summaries. Cochrane open, Examine free abstracts. | | 5 | **Feefo** (JV first-party reviews) | Pulled for 3 SKUs (`_raw_reviews.json`) | Customer voice. Umar's compliance ask: every generated benefit/FAQ traces back here. Folder shared 10 Mar. | | 6 | **Trustpilot — JV service reviews** | Manual files exist | Service signal (delivery, customer service). Re-mapping flagged April. | | 7 | **Loox** (post-migration) | App deal secured, install pending | Product reviews going forward (replacing Trustpilot for product-level) | | 8 | **Amazon — JV products** | Not built | Verified-purchase volume + photo content for `pdp.results` / lifestyle | | 9 | **Amazon — competitor products** (Elavate, Wild Nutrition, Ancient & Brave, Absolute Collagen, Sunna) | Not built | Competitor signal — Umar explicitly asked: *"top competitor — scrape their negative TrustPilot reviews"* (Elavate, 21 May) | | 10 | **Trustpilot — competitors** | Not built | Same as #9 | | 11 | **Reddit** (r/Supplements, r/Nutrition, UK fitness subs) | Not built | Directional sentiment + language patterns for `pdp.who_its_for`, FAQs | | 12 | **Competitor websites** (Wild Nutrition, Dirtea, Heights, Spacegoods, Grüns, Club EarlyBird, MoonBrew, Lumity, 8Hours, Heights, Nothing Fishy, Primal Queen, LIT, Ritual, Elavate, Absolute Collagen, Sunna) | Not built | Design pattern reference + claims they use (benchmark, not copy) | **Compliance rule:** every `scientific_study.body_copy`, every `comparison_row` JV-favouring claim, every health-related FAQ answer **must cite tier 1-4 sources only**. Customer reviews (tier 5-9) are evidence for *consumer experience* claims ("customers report fewer joint aches") but never for *clinical* claims ("reduces inflammation"). --- ## 4. Per-section data routing (where each field comes from) ### `scientific_study` - `headline`, `body_copy` → drafted by AI from PubMed abstract(s) + EFSA-approved phrasing → human-edited - `link_url` → PubMed DOI or NIH publication URL (preferred); EFSA register entry as fallback - `stat_X_value` / `stat_X_label` → extracted from the cited study's abstract/findings; reviewed by human - `image` → Shopify Files (lifestyle photography produced separately, briefed via `photo-brief.json`) ### `timeline_phase` + `timeline_block` - `phase_label` → fixed schedule per category (collagen: "Month 1/2/3/4+"; energy: "Day 1/30/60/90"; magnesium: "Week 1/2/4/8"). Decide cadence per supplement class. - `body_copy` → AI synthesises from: - PubMed onset-of-effect data - Feefo reviews mentioning timeframes ("after 3 weeks I noticed…") - Amazon reviews same - `intro_copy` → category-level template - `study_link_*` → tier-2 source ### `comparison_table` (`comparison_row` / `comparison_column`) - `columns` → JV product + 1-3 competitor reference products - `comparison_row.feature_label` → product attributes from `tblEcomProduct` (strength, form, dietary tags, etc.) + sentiment-derived features ("Taste", "Easy to swallow") - `comparison_row.values` → JV column from DB ground truth; competitor columns from competitor PDP scrape (`competitors///pdp.json`) - All values reviewed before publish (compliance + fairness) ### `faq_pair` - `scope: "standard"` (4 base questions per the 21 May Loom): - "How long until I see results?" — answered from `timeline_block` + reviews - "Can I take this with other supplements?" — answered from product DB (form, ingredients) + standard guidance - "Is this safe long term?" — answered from EFSA + standard caveats - Product-specific catch-all (e.g. "Why include BioPerine?") — from `tblEcomProduct.description` + key_ingredients data - `scope: "product"` (top 3 questions on the page) — auto-suggested from `conversion-blockers.json` (already exists per SKU) --- ## 5. Dashboard rebuild (v2, narrow scope) The current dashboard has 7 views built around *intelligence*. The v2 dashboard is built around *production*. See PROPOSED_DASHBOARD_v2.md (to be written) for the full UI design — high level: - **Pipeline** — per-SKU status across all 8 metaobjects (draft → ai-generated → reviewed → approved → exported) - **Sources** — ingest health: DB row counts, NIH coverage per ingredient, Feefo coverage per SKU, Amazon coverage, Trustpilot, Reddit, competitor PDPs - **Editors** — one workspace per metaobject type with source-citation panel, regenerate, approve - **Per-SKU view** — all 8 sections for one SKU, single approval gate per SKU - **Compliance Trail** — every claim → its tier 1-4 citation chain - **Export** — build + validate + ship CSVs to Lewis Existing views to **keep** as read-only reference: ReviewInsights, StrengthsWeaknesses, ConversionBlockers, ImageAudit, Improvements. They feed the editors as source-citation data. Existing views to **retire**: CatalogOverview (replaced by Pipeline), ConversionDriver (logic absorbed into editors). --- ## 6. Scraping / extraction lane (concrete scripts to add) All under `jv-dashboard/scripts/` following the existing Bun + TypeScript pattern of `tag-reviews.ts`: | Script | Source | Output | |---|---|---| | `ingest-db.ts` | `JV Migration to shopify/full_db/*.xlsx` | `data/pipeline/catalog.json` (extended with full ingredient panel, strength, form) | | `ingest-feefo.ts` | Feefo CSV/folder | `data/intelligence//_raw_reviews.json` (already pattern) | | `pubmed-evidence.ts` | NIH E-utilities API | `data/evidence/.json` | | `efsa-claims.ts` | EFSA Health Claims Register CSV | `data/evidence/efsa-claims.json` | | `examine-evidence.ts` | Examine.com (scraped abstracts) | `data/evidence/-examine.json` | | `scrape-amazon.ts` | Amazon UK (JV + competitor SKUs) | `data/reviews//amazon.json`, `data/competitors///amazon.json` | | `scrape-trustpilot.ts` | Trustpilot brand pages | `data/reviews//trustpilot.json`, competitor coverage | | `scrape-reddit.ts` | Reddit API per ingredient/brand | `data/research//reddit.json` | | `scrape-competitor-pdps.ts` | 16 competitor sites listed by Umar | `data/competitors///pdp.json` | | `generate-scientific-study.ts` | PubMed evidence + EFSA + Feefo | `data/metaobjects/scientific_study/.json` | | `generate-timeline.ts` | PubMed onset + Feefo time-mentions | `data/metaobjects/timeline_block/.json` | | `generate-comparison.ts` | DB + competitor PDP scrape | `data/metaobjects/comparison_table/.json` | | `generate-faq.ts` | conversion-blockers + standard set | `data/metaobjects/faq_block/.json` | | `build-metaobjects-csv.ts` | All `data/metaobjects/**` | Matrixify-format CSVs per definition | | `build-products-csv-v5.ts` | Extend `build_products_csv_v4.py` | Adds the 19+8 metafield columns | | `validate-assets-manifest.ts` | Cross-check filename references vs Shopify Files API | Validation report | **Tooling notes:** - Scrapers use Playwright (already in stack via `tools/loom-capture/`) - Amazon/Trustpilot harder — consider Apify actor budget (see Stage 1 DataForSEO section in `CONTENT_POPULATION_PLAN.md`) - LLM generation: Claude Opus for synthesis with citations; Gemini Flash for bulk tagging (existing `tag-reviews.ts` pattern) - Every generated artefact carries `sourceIds: [...]` mapping back to source files for the dashboard's Compliance Trail --- ## 7. Honouring Umar's email-thread asks | Umar's ask | When | How v2 addresses it | |---|---|---| | Prioritise research (product, competitor, Reddit, Amazon reviews) — Feefo as starting point | 13 Mar | Source hierarchy §3; scraping lane §6 | | Customer language from authentic reviews | 1 May | `faq_pair.source_citation_id` + Compliance Trail view | | RAG over Feefo + JV DB | 2 May | Editors do retrieval-grounded generation: every regenerate pulls top-N Feefo/DB chunks for the SKU as context | | Predictive AI for shipping (net postage cost/income, retention) | 2 May | **Defer to post-migration**; track as separate skill | | Collagen NPD priority | 6 May | Phase 1 still active — `flavor-intelligence.json`, `competitor-comparison.json`, `audience-profile.json` for collagen | | 4-section PDP scope (Scientific Studies, Results, Comparison, FAQs) | 21 May + 1 June | This entire plan | | Standardise timelines + comparison tables across the range | 22 May | `timeline_block` per category cadence (collagen / energy / mineral / vitamin classes share phase schedules); `comparison_table` reused across SKUs in same category | | Shareable front-end URL with blocked-out sections | 27 May | Wait for Lewis — track as dependency | | Elavate negative Trustpilot scrape | 21 May | `scrape-trustpilot.ts` includes Elavate brand page; sentiment-filtered output | --- ## 8. What changes vs the original plan - ✅ Scope narrowed from "all 19 product metafields + 6 metaobjects" to "8 new metaobjects driving 4 PDP sections" as immediate target - ✅ Data hierarchy made explicit, with NIH + EFSA as compliance-safe sources for clinical claims - ✅ Dashboard rebuild planned to focus on *production*, not just intelligence - ✅ All 8 new metaobject schemas defined - ✅ Per-section data routing made explicit - ⏸ Original metaobjects (dietary_tag, health_goals, key_ingredients, benefits, clinically_shown_to, promo_card) deferred — still required, after the 4-section batch - ⏸ Collagen NPD brief still active, runs parallel - ⏸ Predictive shipping AI deferred --- ## 9. Loom captures (decoded reference) Stored under `tools/loom-capture/`: - `1-june-pdp-walkthrough/` — 45 frames @ 3s, contact-sheet.jpg, manifest.json — 1 June Loom (2:12) - `21-may-uiux-walkthrough/` — 54 frames @ 3s, contact-sheet.jpg, manifest.json — 21 May Loom (2:42) - Both `transcript.txt` files are empty (`{"phrases":[],"schemaVersion":"1.1.3"}`) — Umar didn't enable Loom auto-transcription on either video - See `tools/loom-capture/README.md` for re-capture instructions