Skip to content
StoreProfiles

Dataset · DOC-0006

The registry, described as data

Everything the registry knows is public, free and re-usable under CC BY 4.0. This page is the dataset card: what the fields mean, how often they change, what the data deliberately does not cover, what it holds about people, and how to cite it. The endpoint reference lives on /api-docs.

1,139store records171,089products10,101daily snapshots26–85score range, median 58

What is in a record

FieldTypeMeaning
slugstringStable key of the record and the URL segment. Never reused, never changed once assigned.
domainstringThe store's own domain, lowercased, no www. The join key against any other dataset.
nicheenumOne of the registry's verticals, assigned from the store's own product_type values rather than from a third-party category.
trust_scoreinteger 0–100The published score under the stated rules version. Null while a first audit is pending.
score_breakdownobjectOne {got, max} pair per component. The components sum to trust_score — that identity is asserted by the deploy gate on every release.
products / distinct_productsintegerRows recorded and distinct titles among them. They differ when a feed repeats a product across variants, and the gap is itself a signal.
catalog_totalintegerReal catalog size from walking every page of the feed, not the size of our stored sample. Null before a store has been re-ingested.
min_price / currencydecimal / stringCheapest product priced at 2 or above, in the store's modal currency. Mixed-currency catalogs never blend.
history[]arrayOne daily snapshot per store: products, average price, trust score, rules_version. Compare scores only within one rules_version.

Scoring weights, thresholds and the worked arithmetic are on /how-we-score. The maxima are 5 + 25 + 15 + 12 + 8 + 15 + 20 = 100.

How often it changes

Every store is re-ingested, re-scored and re-audited once every 24 hours, and one snapshot per store is written per day. Coverage runs from 2026-07-18 to 2026-07-29. External facts — policy pages, a sampled review rating, a timed server response, domain registration — are collected on their own slower schedules, so a brand-new record can carry a score before its verified tier has anything in it.

What it does not cover

It is not a review dataset
There are no customer reviews here, and none are planned. Every value is computed from what a store publishes, so the dataset says nothing about what buyers experienced after checkout.
Coverage is Shopify stores we have ingested
Not a census. Stores enter through discovery or an owner submission, and every candidate is checked for a live public product feed before it is admitted, so the set skews toward stores whose feed is open.
Scores are comparable only within a rules version
The formula is versioned and the version travels with every snapshot. A change in weights is our step, not the store's, and mixing versions in one comparison produces movement that never happened. Current: rules v0.8.
Absence of a fact is never a penalty
A missing policy page or an uncollected review rating scores zero, never a deduction. A low score means little was proven, not that something bad was found.

People, and what we hold about them

The dataset is about businesses, not people.Records are built from a store’s public product feed, its published policy pages, its homepage and public domain-registration data. We do not collect names, staff details, customer data or buyer reviews, and there is no user-generated content anywhere on the registry — nothing here was submitted by a visitor.

Two addresses are the exception and both are given to us deliberately: an owner may leave an email when claiming a profile, and a reader may leave one when reporting a wrong fact. Neither is published, sold or shared, and neither appears in the CSV, the API or this dataset.

A fact is wrong

Report it from the profile. Reports are read against the source we cited, and a fact we cannot stand behind is removed rather than argued.

How facts are sourced
You own the store

Claim the profile to reply publicly under the audit. Claiming never edits the audit or the score — it adds your answer beside them.

Claim a profile
You want the record gone

Use the report form on the profile and pick “Something else”. Requests land in the same queue as wrong facts, which is read on the daily cycle; there is no email inbox to write to, and pretending otherwise would be worse than saying so.

Find the profile

How to cite it

CC BY 4.0 asks for attribution, so here is the wording, filled in with today’s figures. The date is the date of retrieval, not of publication: the registry is recomputed every 24 hours and a number without its date means nothing.

Plain
StoreProfiles (2026). Independent registry of e-commerce stores, rules v0.8: 1,139 records. Retrieved 2026-07-29 from https://storeprofiles.com. Licensed CC BY 4.0.
BibTeX
@misc{storeprofiles_2026,
  title        = {StoreProfiles: an independent registry of e-commerce stores},
  author       = {{StoreProfiles}},
  year         = {2026},
  note         = {Rules v0.8; 1,139 records; retrieved 2026-07-29},
  howpublished = {\url{https://storeprofiles.com}},
  license      = {CC BY 4.0}
}

Citing a single store instead? Every profile carries its own line, with that store’s score and the rules version it was computed under. Machine-readable licence terms for AI crawlers are at /.well-known/rsl.xml (RSL 1.0).

Niches covered: Beauty & Skincare · Fashion & Apparel · Supplements & Nutrition · Home & Kitchen · Food & Drink.