Streamforge

Methodology

How Streamforge data works

Creator data is bought on trust and used for decisions that cost real money. These pages explain how each part of it is produced — what is measured directly, what is estimated from public evidence, what is modelled, and where each of those runs out.

Why a data vendor should show its working

Almost every number in influencer marketing is an estimate wearing the clothes of a measurement. An audience demographic breakdown is inferred from the people who happen to comment. A sponsorship count depends on how often creators disclose. A media value is a reach figure multiplied by a rate somebody picked. All of these can be produced carefully or carelessly, and rendered in an interface they look identical.

That is a problem for buyers, because the difference only surfaces later — in a rate negotiated against a decayed metric, a campaign targeted at an audience that was never really there, or a board slide built on a multiplier nobody can defend. The only durable fix is for the method to be legible before the decision, not after it.

So each page here states what a field is built from, which decisions it is strong enough to carry, and which it is not. Where a claim could not be evidenced against production data, it was rejected rather than softened, and the rejections are published below alongside the replacements.

The methods

Five explanations, one per data type

Claims we do not make

Four things this data cannot support

These are standard claims in this category. Each was tested against production measurements and rejected. They are published here because a buyer comparing vendors is better served by knowing where the limits are than by a longer feature list.

Questions and answers

Common questions

Why publish limitations rather than just capabilities?

Because the failure mode for creator data is not a wrong number, it is a confident number used for a decision it cannot support. A buyer who knows which fields are measured, which are estimated, and which are modelled can use all three correctly. A buyer who is told everything is equally solid will eventually price a contract against an estimate.

How should these pages be used during a platform evaluation?

As a list of questions to put to every vendor you are considering, including this one. Ask where a demographic figure comes from, what a refresh interval covers, whether a sponsorship count includes affiliate links, and whether a media-value figure has a stated multiplier. The answers separate vendors far more reliably than a feature grid does.

Are these methods audited?

The public claims on this site are checked against an internal record that tracks each claim, its evidence, its owner, and its review date. Claims that cannot be evidenced are rejected rather than softened, which is what the section above records. Audience estimates are separately compared against first-party platform data.

Where does this data come from?

Public platform sources across YouTube, Twitch, TikTok, Instagram, and X, normalised into a shared model, with game and genre context attached from IGDB. The foundations page covers ingestion, normalisation, cross-platform identity linking, and the limits of each.

Applying this to a campaign

These pages describe how the data is built. The field guide covers what to do with it — planning a campaign, shortlisting and vetting creators, negotiating rates, briefing production, and measuring what the work actually returned.

Read the field guide