Mobile Proxy Ecommerce Data
Ecommerce data answers commercial questions rather than technical ones: where your assortment has gaps, who is winning the shelf, and whether your advertised price is holding. This guide is about those questions, not about how to crawl.
- Assortment beats price for strategy — what a competitor stocks and stopped stocking moves slowly and means more.
- Share of shelf is a per-market number — search ranking on a marketplace is assembled per visitor and per country.
- Advertised price is a channel problem — a breach is usually a distribution failure wearing a pricing costume.
- The mechanics live next door — crawl cadence and proxy tier are covered in the ecommerce scraping guide.
See the catalogue and ranking a local shopper is shown.
Stable collection turns a catalogue into a trend.
Ecommerce Data and What It Answers
There is a companion guide on this site about how to crawl ecommerce sites — crawl cadence, geo-priced catalogues, which proxy tier a job needs, and how not to burn a pool. This page is deliberately not that. It is about the commercial questions the crawling exists to answer, which is a separate discipline and the one that usually gets skipped.
Those questions are narrower than people expect. Where does my range have gaps against the competition. Who is winning visibility in the categories I care about. Is my advertised price holding in every channel. Is a competitor testing something in one market before rolling it out. None of them are answered by a price feed, and all of them are answered by data you would collect anyway.
Price gets the attention because it is a single number that changes often. Assortment moves slowly and means considerably more: it is the difference between knowing this week’s tactics and knowing next year’s strategy.
Assortment and Catalogue Coverage
Assortment analysis asks what a competitor lists that you do not, and what they have stopped listing. The second half is the valuable one and it only exists if you were collecting before the product disappeared. A delisting is a decision somebody made — about margin, about supply, about a category they are exiting — and it is announced nowhere.
| What you compute | What it tells you | Cadence |
|---|---|---|
| Assortment gap | Products competitors carry in a category and you do not | Weekly |
| Delisting events | Quiet exits from a category, and supply problems before they are visible | Weekly |
| Share of shelf | Proportion of visible listings in a category that are yours | Weekly, per market |
| New product introductions | Launches, and pilots run in one market first | Weekly |
| Availability duration | How long items are actually out of stock, not just whether they are | Daily |
Share of shelf deserves a note, because it is where single-vantage collection fails most quietly. Marketplace ranking is assembled per visitor and per country, so the same category query run from two markets returns different products in a different order. Measured from one place, share of shelf describes one market’s shelf while carrying a name that implies the category.
Advertised Price and Channel Compliance
For a brand rather than a retailer, the recurring ecommerce problem is not competitor pricing at all. It is whether your own products are being advertised at the price you agreed, by the sellers you authorised, in the markets you intended. A breach found on a marketplace is almost always a distribution failure wearing a pricing costume — stock reached a channel it should not have, and the price is simply the symptom that became visible first.
That is why the seller identity matters as much as the number. Recording who is selling, where they ship from, and whether they appear on your authorised list turns a pricing report into something a channel manager can act on. Collecting the price alone produces a dashboard nobody can do anything with.
Building a Usable Ecommerce Dataset
Two structural decisions determine whether an ecommerce dataset is still useful in a year, and both are cheap to get right at the start and impossible to fix later.
- Key on the product, not the URL — Retailers restructure their sites regularly. A URL-keyed series silently restarts every time somebody reorganises a category.
- Record the market on every row — Not as a property of the run. The moment two runs are merged, an unlabelled observation becomes unusable.
- Separate the catalogue job from the price job — Different cadences, different volumes. Running both at price frequency is how a programme becomes expensive and conspicuous.
- Store availability beside price — A price on out-of-stock inventory is not a price, and reporting it as one costs credibility quickly.
- Keep the raw page — Parsers break silently when markup changes. The capture is what lets you re-extract a field you did not originally take.
For the collection mechanics this page assumes, see the ecommerce scraping guide, and for the pricing half specifically, price monitoring.
Collect Each Storefront Locally
Live PXM2 locations — pick the markets whose catalogues and rankings you track and collect from inside each:
France
India
Singapore
Frequently Asked Questions
How is this different from the ecommerce scraping guide?
That guide covers the mechanics: geo-priced catalogues, crawl cadence, which proxy tier a job needs, and how not to burn a pool. This page covers the commercial questions those mechanics serve — assortment coverage, share of shelf, advertised-price compliance and what to do with the answers. If you are deciding how to crawl, read that one. If you are deciding what is worth crawling, read this one.
What is worth collecting beyond price?
Assortment and its changes, first of all. Which products a competitor lists, which they quietly delisted, which they introduced and where. Then availability, because a low price on out-of-stock inventory is not a price. Then position — where products appear in category and search listings — and the seller behind the listing on a marketplace. Between them those describe a competitor’s strategy, while price alone describes only this week’s tactics.
What is share of shelf and why does it need local collection?
It is the proportion of visible listings in a category or search result that belong to a given brand or seller. It matters because visibility, not catalogue size, determines what gets bought. It needs local collection because marketplace ranking is assembled per visitor and per country: the same query run from two markets returns different products in a different order, so a single vantage point measures one market’s shelf and calls it the category.
Should ecommerce data be collected signed in?
Usually not, and the decision should be deliberate rather than accidental. Signed-in sessions surface member pricing and personalised ranking, which is a different measurement with different comparability properties — and it means holding an account relationship you now have to reason about. Most competitive questions are answered perfectly well by an anonymous shopper, which is also the simplest thing to keep consistent over years.
How often should catalogues be collected?
Far less often than prices. Assortment changes on a scale of weeks, so a weekly catalogue sweep alongside a faster price job covers both properly at a fraction of the request volume. Running the whole catalogue at price cadence is the usual way an ecommerce programme becomes expensive and starts attracting attention it did not need.
Related Mobile Proxy Guides
This page is the commercial half; the scraping guide next door is the mechanical half.
Business use cases
Core mobile proxy guides
See Every Storefront as a Local Shopper
Dedicated 4G/5G modems with unlimited bandwidth and unlimited rotations — carrier IPs that return the catalogue and ranking each market actually gets.
Get a Mobile Proxy