Data sources
Where our water-system information comes from, how it gets onto your report, and how fresh it is.
Primary sources
- EPA public water system records (SDWIS). The federal registry of community water systems gives us each utility's official identity — the PWSID you see on every report — plus its primary source type and state.
- EPA Community Water System Service Area Boundaries. Where states have provided authoritative service-area geometry, we use it to match street addresses to the system that actually serves them. State-provided boundaries are the strongest evidence we accept for a "confirmed" match.
- Utility Consumer Confidence Reports (CCRs) and annual Water Quality Reports. The values on your report — hardness ranges, disinfectant type, reported notices — are taken from the utility's own most recent published report, and each report page links directly to that document so you can verify us in one click.
- US Census geocoding and ZCTA geography. Address lookups are geocoded server-side with the Census Bureau's public geocoder; ZIP-level coverage is derived from the intersection of Census ZIP Code Tabulation Areas with service-area boundaries.
- State drinking-water program publications for advisories and system-level actions.
- NSF/ANSI and WQA certified-product listings for treatment claims. A product claim only counts when the specific model is certified for the specific contaminant under the relevant standard (for example NSF/ANSI 44 for softeners, 53 or 58 for health-related reduction claims).
Provenance on every data point
Each value we show carries four things: the source organization, the reporting period, the date we retrieved it, and a provenance label — Utility-provided, State-provided, EPA-reported, EPA-modeled, Third-party reference, User-entered, Laboratory result, or HWC calculation. The label appears as a badge next to the value, and the report's narrative text is generated from the same record, so the words can never claim more than the badge does.
How confidence is labeled
- Confirmed — your address falls inside a state-provided service boundary, or you explicitly confirmed the utility against your bill.
- Likely — a ZIP-level match, which can never be certain because ZIP codes and service areas overlap. Modeled boundaries are always labeled "likely," never presented as fact.
- We won't guess — if your area doesn't match verified data, the answer is a clear no-match with honest alternatives, not a plausible substitute.
Freshness
- Utility values carry the reporting period of the document they came from (annual reports, typically published each spring) and are updated when the next report is released.
- An automated weekly check verifies that every source link still resolves; anything broken or moved is flagged for review before it can mislead anyone.
- Source documents are archived at retrieval time, so a later change to a utility's website can't silently orphan a value we published.
What we don't do with data
- No substituted averages: if a current, source-backed hardness value for your system is missing, the report says "measure first" instead of borrowing a regional number.
- No blending: utility data, your reported symptoms and laboratory results stay visibly separate on every report — three kinds of evidence, three different weights.
- No safety scores derived from any of it.
Spotted an error?
Every value traces to a named document, so corrections are fast to verify. Use the contact form with the topic "data correction," include the report link, and we'll re-check the original source and fix or annotate the value.