Skip to content

Step 2: Assess Data Quality

1Having completed Step 1, the lead agency will have a Checklist covering the full range of data sources. Not all of these, however, will be equally reliable or complete. Step 2 goes through each source in turn to assess its availability and quality, so that later steps can distinguish strong data from data that requires further development.

  1. 1Document the characteristics of each data source.
  2. 2Score data availability, using the RAG rating (Green / Amber / Red).
  3. 3Score data reliability (High / Medium / Low).
  4. 4Verify data sources against each other and against their definitions.

1Every source in the Checklist ends this step with a documented characterisation and a paired availability and reliability score. This scored Checklist is used in Step 4 to identify and prioritise gaps.

1Before scoring a source, it helps to record how it was collected and what it actually measures — for example, whether a figure comes from a full census, a sample survey, an administrative record, or a modelled estimate, and what time period and geographic coverage it applies to. Capturing this for each source, rather than only the overall score, makes it much easier to explain later why a score was given, and to prioritise improvement efforts sensibly.

1For each data point in the Checklist, assign a RAG rating for data availability, as defined in Table 2.11. Sources rated Red or Amber should be considered for prioritised improvement in Step 4.

2Table 2.11. Scoring data availability.

RatingMeaningDescriptors
GreenAvailable and fit for purposeComprehensive data available at both national and sub-national level, collected periodically.
AmberPartially available and/or partially fit for purposeData available at a limited spatial (for example, only some regions) or temporal resolution (for example, a single snapshot rather than a regular series), restricted to specific sectors, or difficult to access consistently.
RedNot available and/or not fit for purposeNo data available, or data that does not allow plastic-relevant flows to be measured.

Access is not the same as public availability

A source that is not publicly published should not automatically be scored lower than one that is. Confidential administrative or industry data accessed through a formal data-sharing agreement — for example, a memorandum of understanding with an industry association, with the data anonymised for statistical use — can be just as available and reliable as a published dataset and is often the only way to obtain good-quality, regularly collected data on commercially sensitive categories.

Score restricted-access data on the actual comprehensiveness and regularity of the data itself. Note the access arrangement (public, restricted, or not yet accessible) separately in the Checklist, rather than folding it into the availability score.

1Reliability refers to the completeness and accuracy of the data — how far it can be counted on to be consistent and free from errors across time and sources. Score reliability as High, Medium or Low, as defined in Table 2.12.

2Table 2.12. Scoring data reliability.

RatingMeaningDescriptors
HighRepresentative at the national level and compiled using consistent, well-documented methods.For example: official government statistical surveys, customs records, or an established industry reporting system with a clear, repeatable methodology. Well-run industry or administrative data can score High even where it is not, strictly speaking, a scientific study — the test is methodological consistency, not the type of institution collecting it.
MediumPartially representative of the national picture or compiled using a less preferred method.For example: a component such as waste composition estimated from data in other regions or countries, or figures derived through proxy calculations such as apparent consumption (production plus imports minus exports).
LowLimited representativeness or methodological rigour.For example: data derived from a proxy statistic such as monetary value, or a single, unverified estimate.

3This step assesses each source’s reliability on its own terms. Comparing sources against each other, and resolving discrepancies between them, is covered separately below and carried further in Step 4.

1Data points from different sources can vary significantly, usually because of differences in methodology or in how a term is defined. For example, “municipal solid waste generated” could be interpreted as everything collected plus leakage along the supply chain, or only the material that physically arrives at a waste recovery or disposal facility — two different data points that are easy to mistake for the same measurement.

2To verify data values and sources:

  • 3check which definition is being used for the data category;
  • 4cross-validate between different sources covering the same category;
  • 5account for double-counting risk where the methodology involves summing data from more than one source; and
  • 6check that assumptions and calculations behind a figure are recorded and evidence-based.

No data is not the same as weak data

A country may genuinely have no data at all for a category. This is different from having data that is simply weak, and the Checklist should reflect that difference. In such cases, record “no data available” rather than leaving the entry blank or entering a zero, since a blank entry or an unexplained zero could be mistaken for an actual measurement of nil activity, whereas a recorded absence gives a clear signal for Step 4’s gap analysis.

7Specific quality issues and quality-assurance approaches for each plastics data category are set out in Table 2.13.

8Table 2.13. Quality issues and quality-assurance approaches, by data category.

Data categoryKey quality issuesQuality assurance approaches
Production dataTime lags in reporting. Potential missing data from small producers. Data may cover total output without distinguishing what is later exported.Cross-check between national and international databases. Compare with industry capacity data. Validate against trade flows.
Consumption dataApparent consumption does not capture stockpiling. Informal market activity is often not captured.Cross-validate apparent consumption calculations against market surveys. Compare with industry reports. Check consistency with waste generation data.
Trade dataSmall shipments may fall below reporting thresholds. Plastic waste trade is often poorly recorded. Partner countries may report the same shipment differently. Low-value shipments (for example, waste or cheap plastic products) are more prone to estimation error in official statistics.Compare mirror statistics between trading partners. Cross-validate between national and international databases. Check for reporting thresholds and exemptions.
Waste dataInformal-sector activity is not typically captured. Mixed waste streams are difficult to characterise. Measurement approaches are inconsistent between jurisdictions.Compare with consumption data. Cross-check between different stages of waste management. Validate against EPR collection data, where available.

9On the production note above: a high reported production volume is not necessarily a domestic consumption or waste problem if most of that volume is exported. Cross-check production figures against trade data before drawing conclusions about domestic plastic flows.