Case Measured

Keeping a 64,000-Product Online Catalog Accurate

We rebuilt how a ~64,000-product online catalog is updated. Supplier price lists, datasheets and stock files now go through one pipeline — parse, match, diagnose, human approval, bulk write with rollback, read-back. In production, price updates verified 386 of 386 after writing and stock updates 556 of 556, with zero products outside scope touched.

The problem

Suppliers refresh price lists and product data in inconsistent formats, with contradictory status columns and widely broken image links. A full-site audit found:

812

live products with broken images

3,948

live products with no price

7,793

in-stock products hidden from sale

24

brand pages that opened empty

Fixing thousands of listings by hand was not feasible. Bulk-editing without a way back was not safe.

The rebuilt flow

  1. Supplier price list, datasheet and inventory file
  2. Read-only snapshot of the live catalog
  3. Parse and clean the supplier files
  4. Decide what should be on sale
  5. Match to existing listings
  6. Price-change diagnostic
  7. Human approval
  8. Bulk write with rollback
  9. Read back and reconcile
  10. Exception list and supplier data request

Results Measured

  • About 64,000 products across 56 brands audited in full
  • One brand rollout: a 5,887-SKU supplier price list matched against 7,073 listings; 5,204 products brought live
  • Price update verified after writing: 386 of 386 correct, 0 products outside scope touched
  • Stock update verified after writing: 556 of 556 correct
  • 3,378 supplier image URLs tested with real requests; 2,829 of 2,838 uploads verified (99.68%)
  • Storefront brand pages consolidated from 55 to 12
  • Every production write script requires explicit human release — 34 of 34
  • The same process has since been reused across further brands

What the project covers

The update pipeline above is built from thirteen capabilities. Each one has run on the live catalog.

CapabilityWhat it doesVerified at
Read-only catalog snapshotPulls the live catalog into a local baseline without touching production5 related tables in one backup
Supplier file parsingReads messy supplier spreadsheets: repeated page headers, dates stored as numbers, #N/A, conflicting status columns5,887-SKU price list
Sellable-scope decisionDecides what should be on sale from the price list, stock list and current site, and checks every model is accounted for898 models, all accounted for
Exact model matchingMatches supplier SKUs to listings without “cleaning away” real differences between variants5,887 SKUs vs 7,073 listings
Price-change diagnosticTells a small adjustment from a full repricing before anything is written3,286 comparable products
Reversible bulk writesEvery change ships with an execute script, a rollback script and a per-item change record; preview by default2,000+ rows in one run
Read-back verificationRe-reads the catalog to prove that what was written equals what was calculated386 / 386 prices, 556 / 556 stock
Image pipelineTests every link with a real request, sorts failures into seven types, normalizes images to one format, uploads and verifies each file3,378 URLs tested; 2,829 / 2,838 uploads
Category inferenceAssigns categories to new products from existing patterns; anything uncertain goes to a personFull mapping for one brand
Full-site auditBuckets problems across the whole catalog: broken images, missing prices, hidden in-stock items, empty brand pages~64,000 products
Supplier data requestsTurns “why this can’t go live” into a ready-to-send English request sheet for the supplier327 and 102 items for two brands
Storefront cleanupMerges duplicate brands, retires empty brand pages, changes listing status in bulkBrand pages 55 → 12
Platform changesDeploys site modifications with a rollback point, step-by-step checks and a smoke testOne version upgrade

Designed next, not built yet: detecting what changed between two supplier price lists, and running the pipeline on a schedule with automatic checks and rollback.

Where people stay in control

Six decisions are never automated:

  1. Which price tier to use
  2. Which product categories to sell
  3. Category assignment when automatic inference fails
  4. Releasing each write
  5. Resolving model-number conflicts
  6. Anything irreversible

Three rules behind it

  • Measure, don’t assume. Image availability is decided by a real request, not by the file name.
  • Every write ships with a rollback, and runs in preview first.
  • Disable, don’t delete. Deleting breaks links and order history, and cannot be undone.

Key principle The system calculates every option and its consequence. The business owner makes the call.

Which task takes too much time?

Tell us how your team does it today and which software and files are involved. We can work through where a tool or a different process would help.

Discuss Your Project