Dataset · DOC-0006
The registry, described as data
The registry’s own data — scores, facts, audit text, research figures — is public, free and re-usable under CC BY 4.0; product names, images and descriptions belong to the stores (terms, section 3). This page is the dataset card: what the fields mean, how often they change, what the data deliberately does not cover, what it holds about people, and how to cite it. The endpoint reference lives on /api-docs.
What is in a record
| Field | Served in | Type | Meaning |
|---|---|---|---|
| slug | CSV · API | string | Key of the record and the URL segment; the store endpoint repeats it as id. It can change — a rename, or a duplicate merged into another record. The old profile address then answers with a permanent redirect, but /api/store/{old slug} answers 404, so keep the domain beside it. |
| domain | CSV · API | string | The store's own domain, lowercased, no www. The join key against any other dataset. |
| niche | CSV · API | enum | One of the registry's verticals. Assigned from the store's own product_type values when they point clearly to one vertical; otherwise it stays as filed by the third-party store catalog where discovery found the store. |
| status / claimed | CSV · API | string / boolean | status (CSV only) is ingested or claimed; claimed is true once the owner has proved control of the domain. |
| trust_score | CSV · API | integer 0–100 | The published score, computed under the current rules (v0.10). Null while a first audit is pending. |
| score_breakdown | API store | object | One {got, max} pair per component. The components sum to trust_score — that identity is asserted by the deploy gate on every release. |
| products / distinct_products | CSV | integer | Rows we store per store — up to 250 from each crawl of its feed, plus rows the feed stopped returning within the last 14 days — and distinct titles among them. They differ when a feed repeats a product across variants, and the gap is itself a signal. |
| product_count | API | integer | The store's real catalog size, from walking every page of its feed, not the size of our stored sample; the stored row count for a store not yet walked. Inside history[] the same name means stored rows — see below. |
| catalog_truncated | API | boolean | True when the crawl hit the feed's page cap, so product_count is a lower bound. Pages print such figures with a “+” or the word “at least”, never as exact. |
| products_sampled | API | integer | How many products of that catalog we actually store — up to 250 from each crawl; rows the feed has not returned for 14 days are dropped. Every price figure is computed over this sample, not over the whole catalog. |
| min_price / currency | CSV · API | decimal / string | CSV: the cheapest product priced at 2 or above, in the store's modal currency (USD when meta.json returned none), so mixed-currency catalogs never blend. API catalog object: the lowest price above zero in the sample, with the currency its rows carry. Currency is what the store's own meta.json returned to our crawler; Shopify Markets can serve a different one by region. |
| avg_price | API | decimal | The average of sampled prices above zero. |
| history[] | API store | array | Daily snapshots, oldest first: day, product_count (rows in our sample), catalog_total (the real catalog that day — null before 5 August 2026, when the column started), avg_price, min_price, trust_score, rules_version. Compare scores only within one rules_version; a snapshot without a score has no version. |
CSV is /registry.csv; API is /api/registry and /api/store/{slug} (“API store” — the second only). Profiles also print facts neither of them carries: return and delivery windows, free-shipping thresholds, the low-end price and the domain’s registration date, each with its source. A policy fact parsed under withdrawn parsing rules is withheld rather than published.
Scoring weights, thresholds and the worked arithmetic are on /how-we-score. The maxima are 5 + 25 + 15 + 12 + 8 + 15 + 20 = 100.
How often it changes
Every store is re-ingested, re-scored and re-audited once every 24 hours, and one snapshot per store is written per day. Coverage runs from 2026-07-18 to 2026-09-15. External facts — policy pages, a sampled review rating, a timed server response, domain registration — are collected on their own slower schedules, so a brand-new record can carry a score before its verified tier has anything in it.
What it does not cover
- It is not a review dataset
- There are no customer reviews here, and none are planned. Every value is computed from what a store publishes, so the dataset says nothing about what buyers experienced after checkout.
- Coverage is Shopify stores we have ingested
- Not a census. Stores enter through discovery or an owner submission, and every candidate is checked for a live public product feed before it is admitted, so the set skews toward stores whose feed is open.
- Scores are comparable only within a rules version
- The formula is versioned and the version travels with every scored snapshot. A change in weights is our step, not the store's, and mixing versions in one comparison produces movement that never happened. Current: rules v0.10.
- Absence of a fact is never a penalty
- A missing policy page or an uncollected review rating scores zero, never a deduction. A low score means little was proven, not that something bad was found.
People, and what we hold about them
The dataset is about businesses, not people. Records are built from a store’s public product feed, its published policy pages, its homepage and public domain-registration data. It holds no customer data, staff details or buyer reviews. Where a store is run by one person under their own name, the store name can identify them. What a verified owner adds — the store’s name and description, social links, partner terms and a public reply — is published on the profile.
Outside the dataset the registry keeps the email addresses people give it: to claim a store, join a store’s partner waitlist, subscribe to Registry Weekly or report a fact, plus the business contact addresses stores publish on their own sites. Waitlist details go to the store’s verified owner, who runs the queue. None of these addresses is sold or appears in the CSV, the API or this dataset; the privacy notice lists each one, why it is kept and for how long.
Report it from the profile. Reports are read against the source we cited, and a fact we cannot stand behind is removed rather than argued.
How facts are sourced →Claim the profile to reply publicly under the audit. Claiming adds a fixed +15 to the score and your answer beside the audit; it never edits the audit text or any other component.
Claim a profile →The verified owner can have a profile removed: claim it, then email [email protected] from the address you claimed with, giving the profile address.
How removal works →How to cite it
CC BY 4.0 asks for attribution, so here is the wording, filled in with today’s figures. The date is the date of retrieval, not of publication: the registry is recomputed every 24 hours and a number without its date means nothing.
StoreProfiles (2026). Independent registry of e-commerce stores, rules v0.10: 3,081 records. Retrieved 2026-09-15 from https://storeprofiles.com. Licensed CC BY 4.0.
@misc{storeprofiles_2026,
title = {StoreProfiles: an independent registry of e-commerce stores},
author = {{StoreProfiles}},
year = {2026},
note = {Rules v0.10; 3,081 records; retrieved 2026-09-15},
howpublished = {\url{https://storeprofiles.com}},
license = {CC BY 4.0}
}Citing a single store instead? Every profile carries its own line, with that store’s score and the rules version it was computed under. Machine-readable licence terms for AI crawlers are at /.well-known/rsl.xml (RSL 1.0).
Niches covered: Beauty & Skincare · Fashion & Apparel · Supplements & Nutrition · Home & Kitchen · Food & Drink.