Methodology
How Streamforge data works
Creator data is bought on trust and used for decisions that cost real money. These pages explain how each part of it is produced — what is measured directly, what is estimated from public evidence, what is modelled, and where each of those runs out.
Why a data vendor should show its working
Almost every number in influencer marketing is an estimate wearing the clothes of a measurement. An audience demographic breakdown is inferred from the people who happen to comment. A sponsorship count depends on how often creators disclose. A media value is a reach figure multiplied by a rate somebody picked. All of these can be produced carefully or carelessly, and rendered in an interface they look identical.
That is a problem for buyers, because the difference only surfaces later — in a rate negotiated against a decayed metric, a campaign targeted at an audience that was never really there, or a board slide built on a multiplier nobody can defend. The only durable fix is for the method to be legible before the decision, not after it.
So each page here states what a field is built from, which decisions it is strong enough to carry, and which it is not. Where a claim could not be evidenced against production data, it was rejected rather than softened, and the rejections are published below alongside the replacements.
The methods
Five explanations, one per data type
- Data foundations
How Streamforge builds and maintains creator data
What the creator dataset contains, how platform-specific records become comparable profiles, how game context is attached, and where interpretation still matters.
- What does 48M creators mean?
- Which platforms are covered?
- How is the data refreshed?
- Audience analysis methodology
How Streamforge Analyzes Creator Audiences
A practical explanation of the evidence, modeling, validation, and limitations behind Streamforge audience demographics, psychographics, authenticity, and community-health insights.
- Are these first-party platform analytics?
- Does every creator have the same coverage?
- How are missing signals handled?
- Content value methodology
How Streamforge Estimates the Value of Creator Content
The model uses different approaches for on-demand content and live streams, then applies audience geography, industry, and format context to produce a labeled estimate—not a promise of ROI.
- Is estimated media value the same as what a creator should charge?
- Is it the same as campaign ROI?
- Do viewership stability and sponsorship saturation change the formula?
- Data freshness methodology
How Streamforge Refreshes Creator Data
Creator data does not change on one universal clock. Streamforge refreshes fields according to source availability, platform behavior, creator activity, data type, and processing needs.
- Is all Streamforge data real time?
- Why can two fields have different timestamps?
- What happens when a source has not provided enough new evidence?
- Sponsorship detection methodology
How Streamforge Identifies Sponsored Creator Content
Streamforge combines platform-native signals, disclosure language, brand context, promotional evidence, and relationship classification to identify likely sponsorships without treating every brand mention as paid.
- Can the model detect sponsorships in multiple languages?
- Does a promo code always mean sponsorship?
- Can Streamforge see private sponsorship contracts?
Claims we do not make
Four things this data cannot support
These are standard claims in this category. Each was tested against production measurements and rejected. They are published here because a buyer comparing vendors is better served by knowing where the limits are than by a longer feature list.
Profiles are updated daily
Refresh timing varies by source, platform, creator, and field. Different parts of the same profile carry different update times, and a universal interval is not something any dataset this size can honestly promise.
Read the method →Audience demographics are measured
They are estimated from public evidence about the people visible around a creator, which is a biased sample of the audience. Where the evidence is too thin to support a field, the correct output is that there is not enough data.
Read the method →Every brand mention is a sponsorship
Affiliate activity, a creator promoting their own business, and an organic mention are each classified separately. Counting them together is the most common way competitive intelligence overstates a rival brand’s spend.
Read the method →Estimated media value is campaign revenue
A value estimate multiplies reach by a rate somebody chose. It is a planning input and an internal comparison, and it does not belong beside attributed revenue in a report.
Read the method →
Questions and answers
Common questions
Why publish limitations rather than just capabilities?
Because the failure mode for creator data is not a wrong number, it is a confident number used for a decision it cannot support. A buyer who knows which fields are measured, which are estimated, and which are modelled can use all three correctly. A buyer who is told everything is equally solid will eventually price a contract against an estimate.
How should these pages be used during a platform evaluation?
As a list of questions to put to every vendor you are considering, including this one. Ask where a demographic figure comes from, what a refresh interval covers, whether a sponsorship count includes affiliate links, and whether a media-value figure has a stated multiplier. The answers separate vendors far more reliably than a feature grid does.
Are these methods audited?
The public claims on this site are checked against an internal record that tracks each claim, its evidence, its owner, and its review date. Claims that cannot be evidenced are rejected rather than softened, which is what the section above records. Audience estimates are separately compared against first-party platform data.
Where does this data come from?
Public platform sources across YouTube, Twitch, TikTok, Instagram, and X, normalised into a shared model, with game and genre context attached from IGDB. The foundations page covers ingestion, normalisation, cross-platform identity linking, and the limits of each.
Applying this to a campaign
These pages describe how the data is built. The field guide covers what to do with it — planning a campaign, shortlisting and vetting creators, negotiating rates, briefing production, and measuring what the work actually returned.
Read the field guide