Market research used to mean surveys, focus groups, and syndicated reports that arrived quarterly and were already out of date. Web data changed the economics of it. You can now observe what an entire market is doing β what competitors charge, what customers complain about, which products are gaining traction, how demand varies by region β continuously, at a fraction of the cost, and with a sample size no survey could reach.
What you cannot do is skip the methodology. Web-derived market research fails in a specific and recognisable way: it produces large volumes of plausible-looking data that support confident conclusions which happen to be wrong, usually because the collection was geographically or demographically skewed in a way nobody accounted for.
This guide covers what web data can and can't tell you about a market, how to design collection that produces defensible conclusions, why geographic accuracy matters more here than almost anywhere else, and how to turn observational data into something a decision can rest on.
What Web Data Actually Tells You
Being precise about this prevents most of the analytical mistakes that follow.
What it observes well:
- Competitive positioning. What competitors offer, how they price it, how they describe it, and how that changes over time.
- Assortment and availability. What products exist in a market, who carries them, and what's in stock.
- Expressed sentiment. What customers say publicly about products and companies, in reviews, forums, and social discussion.
- Demand signals. Search interest, review velocity, ranking movements, and discussion volume as proxies for attention.
- Market structure. Who the players are, how concentrated the market is, and how that shifts.
- Regional variation. How all of the above differs by geography, which is frequently the most valuable output.
- Change over time. The one thing traditional research does worst and continuous collection does best.
What it does not observe:
- Actual sales volumes. Review counts and rankings are proxies with unknown and inconsistent conversion rates. They are not sales data.
- Private or B2B transactions. Anything negotiated rather than listed is invisible.
- Motivation. You can see what people said, not why they chose what they chose.
- The silent majority. Reviewers are a self-selected minority with distinctive characteristics.
- Intent. Browsing and searching are not buying.
- Anything offline. For many markets this is most of the market.
The discipline that matters: be explicit about which of these you're measuring. A great deal of bad market research consists of treating a proxy as if it were the thing itself, and the failure mode is a chart that looks authoritative and means something other than what its label claims.
Designing Collection That Produces Valid Conclusions
The research design questions, which should be settled before any code is written.
Define the market boundary explicitly. Which products, which categories, which channels, which geographies, which price bands. Vague boundaries produce datasets that can be made to say almost anything.
Define your sampling frame. What is the universe you're trying to describe, and how does your collection relate to it? Scraping the top hundred results from one marketplace is not a sample of a market β it's a sample of one marketplace's ranking algorithm, which is a different thing entirely.
Decide what constitutes a unit of observation. A product? A listing? A seller? A review? These produce different datasets and different conclusions, and mixing them silently is a common error.
Establish your comparison basis. Over time, across competitors, across regions, or against a baseline. The comparison determines what you need to hold constant.
Plan for what you'll hold constant. If you're comparing regions, everything except region must be identical β same products, same collection time, same methodology. Any other variable that differs contaminates the comparison.
Decide your collection frequency before you start. Changing it mid-series creates discontinuities that look like market movements.
Write down what would falsify your hypothesis. Research designed to confirm a belief usually succeeds, and the success means nothing.
> Tip: The most common failure in web market research is a dataset that answers a slightly different question from the one being asked. Write the question down first, in one sentence, and check the collection design against it.
Why Geography Dominates This Use Case
More than in any other application of web data, geographic accuracy determines whether your conclusions are valid.
Markets are localised in ways that aren't visible from one location:
- Pricing varies by region, sometimes substantially, and often deliberately.
- Assortment varies. Products available in one market simply don't exist in another.
- Search and recommendation results are localised, so competitive visibility differs by location.
- Advertising differs by region, so competitive messaging you observe is location-specific.
- Currency, tax, and shipping presentation all change what a customer actually sees.
- Availability and stock status are frequently regional.
- Language and localisation affect both content and positioning.
The methodological consequence: collecting all your data from one location produces a picture of that location, presented as if it were a picture of the market. If you're researching a multi-country market from a single vantage point, your findings are wrong in ways you can't detect from inside the dataset.
What good geographic practice looks like:
- Collect from each market you're describing, using addresses genuinely located there
- Match granularity to the question. Country level for national comparisons; city level where local intent, delivery zones, or regional pricing matter
- Tag every record with its collection location. Without this you cannot separate regional variation from noise
- Collect the same items across all markets so comparisons are valid
- Collect simultaneously where possible, since time and geography confound each other
- Verify your geographic targeting works rather than assuming it does. Check that the pages you receive actually reflect the intended location β currency, language, and availability are usually the tells
This is why residential proxies with city-level targeting matter for market research specifically. Datacenter addresses geolocate to a small number of hosting hubs, which is fine when geography is irrelevant and useless when it's the variable you're studying.
Choosing Your Proxy Setup for Research
Market research has an unusual proxy requirement profile, and matching it properly matters more than in most use cases.
Where geography is the variable, residential with granular targeting is the requirement. If you're comparing markets, your collection addresses have to be genuinely located in those markets. Datacenter addresses geolocate to a small number of hosting hubs, which means a "German" datacenter address may return the same page as a Dutch one. For any finding that turns on regional difference, this isn't a cost optimisation β it's a methodological requirement.
Match granularity to the question. Country-level targeting is adequate for national market comparisons. City-level matters where delivery zones, local availability, regional pricing, or local search results are part of what you're measuring.
Where geography is irrelevant, use datacenter. Industry publications, regulatory filings, job listings, trade press, and most B2B sites have no meaningful bot detection and no regional variation worth capturing. Routing these through residential addresses is straightforward waste, and at continuous-collection volumes the difference is substantial.
Segment your sources by requirement. Maintain a per-source configuration recording which tier it needs, determined by testing rather than assumption. Run five hundred requests through datacenter addresses against each significant source and record the status distribution. The results usually split sources into a large permissive majority and a small demanding minority.
For search-derived data, consider a managed API. Localised search results are among the hardest things to collect reliably, and rank-tracking infrastructure is a maintenance burden that has nothing to do with your research question. Buying the data is frequently better value than building the collector.
Use sticky sessions where the source requires them. Marketplaces with server-side pagination, faceted navigation, or session-scoped currency settings need the address held for the duration of the sequence, or your results become internally inconsistent.
Verify targeting rather than trusting it. Check that the pages you receive actually reflect the intended market. Currency, language, availability messaging, and shipping options are the usual tells. A geo-targeting configuration that silently fails produces a dataset that looks fine and describes the wrong place.
Sources Worth Collecting From
Marketplaces and retailers. Assortment, pricing, availability, seller counts, and review data. The richest single source for consumer markets, and the one most likely to require residential addresses.
Competitor websites. Product pages, pricing pages, feature comparisons, positioning language, case studies, and careers pages. Careers pages in particular are an underused signal β hiring patterns reveal strategic direction earlier than announcements do.
Review platforms. Sentiment, feature complaints, comparison mentions, and velocity. Rich and heavily biased, which is manageable if you account for it.
Forums and communities. Where genuine unprompted discussion happens. Lower volume, higher signal, harder to collect systematically.
Search results. What surfaces for category queries reveals competitive visibility and how the market is framed. Localised, so geography applies.
Industry publications and trade press. Announcements, market commentary, and regulatory developments.
Job listings. Hiring volume, roles, locations, and required skills. A reliable leading indicator of investment and direction.
Regulatory and public filings. Where available, the most reliable data in the entire mix, and frequently overlooked because it's less convenient than scraping a marketplace.
App stores. Rankings, review velocity, update cadence, and feature descriptions for mobile-first markets.
Aggregators and comparison sites. Convenient and second-hand. Useful for discovery, less reliable as primary data since you inherit their coverage decisions.
Common Research Questions and How to Answer Them
Concrete methodologies for the questions people actually ask, since the abstract advice only goes so far.
"Who are the main players in this market, and how concentrated is it?"
Collect listings or products across the category from several independent sources β marketplaces, comparison sites, search results in each target market. Resolve entities to companies. Count presence and visibility rather than assuming any one source's ranking reflects reality. Concentration measures computed from a single marketplace describe that marketplace, so use at least three sources and report where they disagree.
"How do competitors price against us?"
Match products across competitors using identifiers where they exist and careful attribute matching where they don't. Collect from the same markets, at the same time, on the same schedule. Record list price, promotional price, shipping, and any conditions separately β a headline price without shipping and conditions is not comparable. Track changes rather than snapshots, since the timing of changes reveals strategy.
"Is demand for this category growing?"
No single web signal measures demand. Triangulate: review velocity across the category, new product introductions, seller counts, search interest, discussion volume in relevant communities, and hiring in the sector. Where several independent proxies move together, the signal is worth something. Where one moves alone, it's probably an artefact of that platform.
"What do customers complain about?"
Collect reviews across products and competitors, extract complaint themes, and normalise by volume so a high-selling product doesn't dominate purely through review count. Segment by rating band, since three-star reviews are typically more informative than one- or five-star ones. Remember that reviewers are a self-selected minority.
"Should we enter this market?"
The multi-market comparison. Collect the same category from each candidate market, from addresses genuinely located there. Compare assortment breadth, price levels, competitor count and concentration, review volume, and whether the incumbents appear to be local or international. The geographic discipline matters most here, because the entire question is about differences between markets.
"What is this competitor about to do?"
Leading indicators rather than announcements: job listings by role and location, careers page changes, new product pages appearing before launch, changes in positioning language, patent and regulatory filings, and expansion of shipping or service regions. This is slow, cumulative work and it produces the earliest signals available.
"How is our category being framed?"
Collect search results, category pages, and comparison content for the terms customers actually use. What surfaces reveals how the market is defined by the intermediaries customers encounter, which is often different from how participants define it internally.
Building a Continuous Research Programme
One-off research answers a question. Continuous collection answers questions you haven't asked yet, and that's where the compounding value sits.
Start with a defined tracked set. Products, competitors, categories, and markets, explicitly listed. Ad hoc collection produces data you can't compare over time.
Hold methodology constant. Same sources, same parameters, same schedule, same processing. Changes create discontinuities that read as market movements. When you must change something, start a new series rather than continuing the old one.
Version everything. Collection parameters, extraction logic, and taxonomy. When a figure looks odd six months later, you want to know whether the market changed or your pipeline did.
Store raw responses. Market research questions evolve. Being able to reprocess history against a new question is worth a great deal, and re-collecting the past is impossible.
Separate collection from analysis. Collect broadly and consistently; analyse narrowly and specifically. Coupling them means every new question requires new collection.
Build change detection, not just snapshots. The alerting layer β a competitor changed price, a new entrant appeared, assortment shifted β is usually more actionable than the periodic report.
Review the tracked set periodically. Markets change. A tracked set defined two years ago may be missing the players that matter now.
Instrument data quality. Per-source extraction rates, match rates in entity resolution, coverage completeness. Silent degradation in a research pipeline produces confidently wrong trend lines, which is worse than an obvious failure.
Qualitative Signals Worth Watching
Not everything valuable in market research is countable, and the qualitative layer is frequently where the earliest signals appear.
Positioning language changes. How a competitor describes itself on its homepage, in what order it lists benefits, and which audience it addresses. Rewrites are deliberate and usually precede a strategic shift.
Pricing page structure. Tier names, what moves between tiers, what becomes an add-on, and what quietly disappears. This reveals monetisation strategy more clearly than the numbers do.
Feature comparison pages. Who a competitor chooses to compare themselves against tells you who they think their competition is, which is often not who you assumed.
Job listing composition. Roles, seniority, locations, and required skills. A sudden cluster of hires in one function or geography is a plan being executed.
Documentation and changelog activity. For technical products, the pace and nature of shipped changes is a better development signal than any announcement.
Customer complaint themes over time. Not the volume, which tracks sales, but which issues appear and disappear. A complaint theme vanishing usually means it got fixed, which tells you where effort went.
Language and market localisation. New languages appearing on a competitor's site is expansion in progress.
What competitors stop doing. Removed pages, discontinued products, and retired features. Absence is harder to track than presence and often more informative, which is why it requires deliberate change detection rather than periodic snapshots.
Bias: The Central Analytical Problem
Every web-derived dataset is biased. The question is whether you know how.
Selection bias. Who is represented in your source? Review platforms over-represent extremes. Marketplaces over-represent sellers who chose that channel. Forums over-represent enthusiasts. None of these are the market.
Geographic bias. Discussed above, and the most consequential in this use case.
Platform bias. Each source has its own demographics, conventions, and incentives. Conclusions drawn from one platform describe that platform.
Survivorship bias. Delisted products, failed companies, and removed content are absent. Any analysis of "what works" that only looks at what's currently present has a hole in it.
Temporal bias. Collection timing interacts with promotional cycles, seasonality, and news events. A price snapshot taken during a sale period is not a price.
Ranking bias. Collecting the top results means collecting what an algorithm surfaced, which reflects that algorithm's objectives rather than the market's structure.
Language bias. Collecting only in one language systematically excludes parts of a market.
Volume bias. The most prolific sources dominate a dataset unless you weight deliberately. Three domains providing 60% of your records means your findings describe those three domains.
How to handle it:
- Document the bias rather than pretending to eliminate it. You can't remove selection bias from public review data. You can state it.
- Triangulate across sources with different biases. Where independent sources agree, confidence rises. Where they disagree, that disagreement is itself information.
- Weight deliberately rather than letting collection volume decide representation.
- Sanity check against external data β published statistics, industry reports, or anything you have with a known methodology.
- State limitations in the output. Research that acknowledges its constraints is more useful and more credible than research that doesn't.
Entity Resolution: Where the Work Actually Is
The unglamorous problem that consumes most of the engineering effort in any serious market research pipeline, and the one that quietly determines whether your findings mean anything.
The same product appears across sources as different strings. One marketplace lists it with the manufacturer's full name and a pack size; another abbreviates; a third bundles it; a fourth has a typo. Matching these correctly is the difference between comparing like with like and producing a price comparison that's silently comparing a six-pack against a single unit.
Use identifiers wherever they exist. Manufacturer part numbers, barcodes, ISBNs, model numbers, and standard product identifiers. These are reliable when present and absent more often than you'd like.
Normalise aggressively before matching. Case, punctuation, whitespace, common abbreviations, unit representations, and pack size notation. A large share of non-matches are formatting differences.
Extract structured attributes. Brand, model, size, colour, capacity, and variant. Matching on structured attributes is far more reliable than matching on full title strings.
Use fuzzy matching with a confidence score, not a binary decision. Record the score with the match so downstream analysis can filter on confidence.
Manually review the ambiguous middle. High-confidence matches and clear non-matches need no attention. The band between them is where errors live, and human review of a sample there is worth a great deal.
Handle variants deliberately. Is a different colour the same product? A different pack size? A regional edition? These are analytical decisions, not technical ones, and they should be documented rather than emerging from whatever the matching code happened to do.
Maintain a persistent entity registry. Once you've resolved a product, record the mapping so you don't redo the work and so your time series stays consistent. A product that gets matched differently between collection runs creates a fake price change.
Measure your match rate. What proportion of items resolved, at what confidence. A falling match rate is an early warning that a source changed its formatting.
Expect to iterate. Entity resolution is never finished. Budget ongoing effort rather than treating it as a build-phase task.
Turning Collection Into Analysis
Normalise before comparing. Currency, units, pack sizes, product identifiers, and category taxonomies all differ between sources. Comparing unnormalised data across sources produces nonsense that looks like insight.
Entity resolution is the hard part. The same product appears under different names, SKUs, and descriptions across sources. Matching them is where most of the analytical engineering effort actually goes, and doing it badly quietly corrupts everything downstream.
Distinguish levels and changes. A competitor's price is less interesting than the fact that they changed it. Structure your data to support change detection, not just snapshots.
Track your own metric definitions. "Market share" derived from review counts is a specific, defensible thing if you define it and a misleading thing if you don't.
Segment before aggregating. Aggregate figures hide the variation that usually contains the insight. Segment by region, category, price band, and seller type first, then aggregate deliberately.
Preserve uncertainty. Where your data is a proxy, say so in the output. A number presented without its caveat will be used without its caveat.
Build for reproducibility. Store raw responses, version your processing, and record collection parameters. When someone questions a finding six months later, you want to be able to trace it.
Validate against reality. Periodically check a handful of data points manually against the live source. Automated pipelines drift, and drift in market research produces confidently wrong conclusions rather than obvious errors.
Presenting Findings Credibly
Research that nobody trusts is research that changes nothing, and web-derived findings face more scepticism than survey data β often justifiably.
Lead with the question, not the data. What was asked, what was found, what should change. The methodology belongs in the appendix, not the opening slide.
Label proxies as proxies, every time. "Review-count-derived visibility index" is more credible than "market share," even though it's less satisfying to write.
State the sampling frame explicitly. Which sources, which markets, which time period, which product universe. A reader who understands what was and wasn't covered can calibrate; one who doesn't will either over-trust or dismiss it.
Show the variation, not just the average. Ranges, distributions, and segment breakdowns. Averages in market data conceal more than they reveal, and presenting only the average invites the first sharp question you can't answer.
Attach confidence honestly. Some findings are robust across multiple independent sources. Others rest on one platform's data. Distinguishing these in the presentation is what separates useful research from a deck.
Anticipate the "that doesn't match what I see" objection. Someone will check a data point against their own screen and get a different answer, because their view is personalised, located differently, or captured at another moment. Being able to explain exactly what you measured turns that from a credibility problem into a methodology conversation.
Keep a reproducible trail. Raw responses, collection parameters, and processing versions. When a finding is challenged six months later, tracing it in an hour rather than a week is the difference between the research standing and quietly being dropped.
Report what would change your conclusion. Research that names its own weak points is consistently trusted more than research that presents itself as complete.
Common Mistakes
Answering a slightly different question. The most common failure. A dataset that describes one marketplace's top hundred results gets presented as a description of a market. Write the question down first and check the design against it.
Collecting from one location. Produces a picture of that location presented as a picture of the market, and the error is invisible from inside the dataset.
Treating proxies as measurements. Review counts are not sales. Search interest is not demand. Rankings are not market share. Each may correlate; none is the thing itself, and labels matter.
Skipping entity resolution. Comparing prices across sources without properly matching products produces confident nonsense.
Changing methodology mid-series. Creates discontinuities that read as market movements and are actually your own pipeline.
Aggregating before segmenting. Aggregate figures hide the variation that usually contains the insight.
Ignoring survivorship. Analysing what's currently listed and drawing conclusions about what works, while delisted failures are invisible.
Single-source dependency. Every platform has its own biases. Conclusions from one source describe that source.
Not storing raw responses. Research questions evolve, and reprocessing history against a new question is only possible if you kept the source data.
Presenting findings without limitations. Research that states its constraints is more useful and more credible. A number presented without its caveat will be used without its caveat.
Collecting more often than reality changes. Generates noise, cost, and false signals without adding information.
Letting collection volume determine representation. If three domains provide most of your records, your findings describe those three domains unless you weighted deliberately.
A Realistic Build Sequence
The order that avoids the most expensive rework on a research programme.
Phase one β write the question down. One sentence. What decision will this research inform, and what would change as a result? Research without a decision attached to it tends to expand until it produces a large dataset and no conclusion.
Phase two β define the market boundary and sampling frame. Which products, categories, channels, geographies, and price bands. Be specific enough that someone else could apply the same boundary and get the same universe.
Phase three β source inventory. List candidate sources, assess what each covers, note its known biases, and rank by usefulness for the specific question. Include at least one source with a different bias profile from the others, for triangulation.
Phase four β per-source proxy requirement testing. Datacenter first, five hundred requests, status distribution recorded. Split sources into permissive and demanding, and note which ones require geographic targeting.
Phase five β small-scale collection and manual validation. A few hundred records. Read them. Check them against the live sources by hand. Almost every methodological problem is visible at this stage and expensive to fix later.
Phase six β entity resolution design. Before scaling, work out how items will be matched across sources and what confidence threshold you'll accept. Retrofitting this to a large dataset is painful.
Phase seven β storage, provenance, and reproducibility. Raw responses, collection parameters, processing versions. Build this before volume collection, since the whole point is being able to trace a finding later.
Phase eight β scaled continuous collection with quality instrumentation. Extraction rates, match rates, coverage completeness, and per-source health. Silent degradation in a research pipeline produces wrong trend lines rather than obvious errors.
Phase nine β analysis, segmentation, and triangulation. Now, with a dataset you can defend.
Phases one and five are the ones that get skipped under pressure, and they're the two that determine whether the output is a decision or a dashboard.
Legal and Ethical Considerations
- Public data collection and authentication circumvention are legally distinct. Competitive research on publicly accessible pages sits on very different ground from accessing anything behind a login you weren't granted.
- Terms of service are contracts, and many explicitly prohibit automated collection. That's commercial and reputational risk even where no criminal statute applies.
- Data protection law applies to personal data in your dataset. Reviews contain names and sometimes more; that's personal data regardless of it being publicly posted.
- Competition law considerations exist around price monitoring and information exchange, particularly in concentrated markets. Systematic competitor price collection is normal commercial practice; using it in ways that facilitate coordination is not. This is worth a conversation with counsel.
- Rate-limit as a matter of conduct. Degrading a competitor's site is a problem regardless of legality, and it's an unforced error.
- Be accurate about what your data represents when presenting internally or to clients. Overstating what a proxy measures is a professional problem before it's ever a legal one.
- Respect published crawl directives where they apply to you.
Not legal advice, and jurisdictions differ. Commercial research programmes at scale warrant a lawyer's review, particularly around competition law in concentrated markets.
Frequently Asked Questions
Can I measure market share from web data?
Not directly. You can build proxies from review counts, listing counts, or visibility, and they may correlate with share. Label them as proxies, state the assumption, and triangulate against any real data you have.
Do I need residential proxies for market research?
Where geography matters to your findings, usually yes β datacenter addresses geolocate to hosting hubs rather than to markets. Where you're collecting from sources without bot detection and geography is irrelevant, datacenter is fine and much cheaper. Test per source.
How often should I collect?
Match the rate of change in what you're measuring. Prices move daily in some categories and quarterly in others. Collecting more often than the underlying reality changes generates noise and cost without insight.
How do I handle products that don't match across sources?
Entity resolution, and expect it to consume more effort than the collection itself. Use identifiers where available, fuzzy matching with manual review where not, and record match confidence so you can filter on it.
Is review data reliable?
As a record of what reviewers said, yes. As a representation of customer opinion, it's heavily skewed toward extremes and toward whoever the platform's users are. Useful with that caveat, misleading without it.
How do I know if my sample is representative?
Usually you don't, and claiming otherwise is where credibility gets lost. Define your sampling frame explicitly, document what it does and doesn't cover, and triangulate against independent sources.
Should I collect competitor pricing continuously or periodically?
Continuously if you'll act on it, periodically if you're building a picture. Continuous collection also gives you change detection, which is usually more valuable than the levels themselves.
How do I handle sources that block me?
Work out why first. Rate limiting masquerading as blocking is common and free to fix by halving concurrency. Header quality and TLS fingerprint matter more than most people expect. Escalate the proxy tier only after those are clean and you've measured that the escalation actually helps.
Should I use one continuous pipeline or separate collections per question?
One continuous collection over a defined tracked set, with analysis layered on top. Separate per-question collections mean every new question requires new collection, and you lose the historical comparison that makes continuous research valuable.
How do I stop the tracked set from going stale?
Review it on a schedule β quarterly is reasonable for most markets. New entrants and category shifts are exactly the things a fixed tracked set will miss, and they're often the most important findings available.
Can web research replace surveys?
It answers different questions. Web data observes behaviour and public expression at scale; surveys probe motivation and reach people who don't post. The strongest research programmes use both and treat agreement between them as a confidence signal.
Getting the Collection Layer Right
Market research depends on geographic accuracy more than most web data work, which makes proxy selection a methodological decision rather than just an infrastructure one. ProxyScrape's residential proxies offer country, state, and city-level targeting, which is what lets you collect the same product set from each market you're describing rather than inferring regional variation from a single vantage point. For sources without bot detection where geography isn't the variable β industry publications, regulatory filings, job listings β their datacenter plans handle the volume far more cheaply, and for search-derived competitive visibility data their SERP API removes the maintenance burden of scraping localised results yourself. Their market research documentation covers the setup in more detail.
β Compare proxy options for multi-market data collection
Web data has made market research faster, cheaper, and continuous, and it has also made it much easier to be confidently wrong at scale. The teams that get real value from it are the ones that wrote the research question down before writing the scraper, collected from each market they intended to describe rather than from wherever their servers happened to be, documented their biases instead of pretending to have eliminated them, and labelled their proxies as proxies. The collection is the easy part. Knowing what your data actually represents is the work.