Step 2: Assess Data Quality
2.3 Step 2: Assess Data Quality
Section titled “2.3 Step 2: Assess Data Quality”2.3.1 Purpose
Section titled “2.3.1 Purpose”1Having completed Step 1, the lead agency will have a Checklist covering the full range of data sources. Not all of these, however, will be equally reliable or complete. Step 2 goes through each source in turn to assess its availability and quality, so that later steps can distinguish strong data from data that requires further development.
2.3.2 Key Actions
Section titled “2.3.2 Key Actions”- 1Document the characteristics of each data source.
- 2Score data availability, using the RAG rating (Green / Amber / Red).
- 3Score data reliability (High / Medium / Low).
- 4Verify data sources against each other and against their definitions.
2.3.3 Outputs
Section titled “2.3.3 Outputs”1Every source in the Checklist ends this step with a documented characterisation and a paired availability and reliability score. This scored Checklist is used in Step 4 to identify and prioritise gaps.
2.3.4 Detail and supporting guidance
Section titled “2.3.4 Detail and supporting guidance”Documenting data characteristics
Section titled “Documenting data characteristics”1Before scoring a source, it helps to record how it was collected and what it actually measures — for example, whether a figure comes from a full census, a sample survey, an administrative record, or a modelled estimate, and what time period and geographic coverage it applies to. Capturing this for each source, rather than only the overall score, makes it much easier to explain later why a score was given, and to prioritise improvement efforts sensibly.
Scoring data availability
Section titled “Scoring data availability”1For each data point in the Checklist, assign a RAG rating for data availability, as defined in Table 2.11. Sources rated Red or Amber should be considered for prioritised improvement in Step 4.
2Table 2.11. Scoring data availability.
| Rating | Meaning | Descriptors |
|---|---|---|
| Green | Available and fit for purpose | Comprehensive data available at both national and sub-national level, collected periodically. |
| Amber | Partially available and/or partially fit for purpose | Data available at a limited spatial (for example, only some regions) or temporal resolution (for example, a single snapshot rather than a regular series), restricted to specific sectors, or difficult to access consistently. |
| Red | Not available and/or not fit for purpose | No data available, or data that does not allow plastic-relevant flows to be measured. |
Access is not the same as public availability
A source that is not publicly published should not automatically be scored lower than one that is. Confidential administrative or industry data accessed through a formal data-sharing agreement — for example, a memorandum of understanding with an industry association, with the data anonymised for statistical use — can be just as available and reliable as a published dataset and is often the only way to obtain good-quality, regularly collected data on commercially sensitive categories.
Score restricted-access data on the actual comprehensiveness and regularity of the data itself. Note the access arrangement (public, restricted, or not yet accessible) separately in the Checklist, rather than folding it into the availability score.
Scoring data reliability
Section titled “Scoring data reliability”1Reliability refers to the completeness and accuracy of the data — how far it can be counted on to be consistent and free from errors across time and sources. Score reliability as High, Medium or Low, as defined in Table 2.12.
2Table 2.12. Scoring data reliability.
| Rating | Meaning | Descriptors |
|---|---|---|
| High | Representative at the national level and compiled using consistent, well-documented methods. | For example: official government statistical surveys, customs records, or an established industry reporting system with a clear, repeatable methodology. Well-run industry or administrative data can score High even where it is not, strictly speaking, a scientific study — the test is methodological consistency, not the type of institution collecting it. |
| Medium | Partially representative of the national picture or compiled using a less preferred method. | For example: a component such as waste composition estimated from data in other regions or countries, or figures derived through proxy calculations such as apparent consumption (production plus imports minus exports). |
| Low | Limited representativeness or methodological rigour. | For example: data derived from a proxy statistic such as monetary value, or a single, unverified estimate. |
3This step assesses each source’s reliability on its own terms. Comparing sources against each other, and resolving discrepancies between them, is covered separately below and carried further in Step 4.
Verifying data sources
Section titled “Verifying data sources”1Data points from different sources can vary significantly, usually because of differences in methodology or in how a term is defined. For example, “municipal solid waste generated” could be interpreted as everything collected plus leakage along the supply chain, or only the material that physically arrives at a waste recovery or disposal facility — two different data points that are easy to mistake for the same measurement.
2To verify data values and sources:
- 3check which definition is being used for the data category;
- 4cross-validate between different sources covering the same category;
- 5account for double-counting risk where the methodology involves summing data from more than one source; and
- 6check that assumptions and calculations behind a figure are recorded and evidence-based.
No data is not the same as weak data
A country may genuinely have no data at all for a category. This is different from having data that is simply weak, and the Checklist should reflect that difference. In such cases, record “no data available” rather than leaving the entry blank or entering a zero, since a blank entry or an unexplained zero could be mistaken for an actual measurement of nil activity, whereas a recorded absence gives a clear signal for Step 4’s gap analysis.
7Specific quality issues and quality-assurance approaches for each plastics data category are set out in Table 2.13.
8Table 2.13. Quality issues and quality-assurance approaches, by data category.
| Data category | Key quality issues | Quality assurance approaches |
|---|---|---|
| Production data | Time lags in reporting. Potential missing data from small producers. Data may cover total output without distinguishing what is later exported. | Cross-check between national and international databases. Compare with industry capacity data. Validate against trade flows. |
| Consumption data | Apparent consumption does not capture stockpiling. Informal market activity is often not captured. | Cross-validate apparent consumption calculations against market surveys. Compare with industry reports. Check consistency with waste generation data. |
| Trade data | Small shipments may fall below reporting thresholds. Plastic waste trade is often poorly recorded. Partner countries may report the same shipment differently. Low-value shipments (for example, waste or cheap plastic products) are more prone to estimation error in official statistics. | Compare mirror statistics between trading partners. Cross-validate between national and international databases. Check for reporting thresholds and exemptions. |
| Waste data | Informal-sector activity is not typically captured. Mixed waste streams are difficult to characterise. Measurement approaches are inconsistent between jurisdictions. | Compare with consumption data. Cross-check between different stages of waste management. Validate against EPR collection data, where available. |
9On the production note above: a high reported production volume is not necessarily a domestic consumption or waste problem if most of that volume is exported. Cross-check production figures against trade data before drawing conclusions about domestic plastic flows.