Spatial Extreme Value Modeling for PFAS Source Apportionment

Connor McNeill

North Carolina State University

Sean O’Connor

North Carolina State University

Ana Rappold

US Environmental Protection Agency

Brian Reich

North Carolina State University

August 4, 2026

Background

What is PFAS?

  • Widely used chemicals found in many products such as plastics, personal care items, sunscreen, etc. [1]
  • Known as “forever chemicals” which break down slowly.
  • Build up in blood stream of humans when ingested.
  • Exposure over time has potential health risks such as prostate and kidney cancer, developmental delay in children, and reduced immune system response. [1]

Motivation

  • PFAS is known to spread to humans through several means, including though contaminated drinking water. [2]

  • Our goal is to determine which industries are contributing the most PFAS to California’s drinking water.

  • We also want to see how this varies across the different ecoregions of the state.

GAMA Data

Figure 1: This figure shows a map of the 8018 GAMA Program wells in California.

Industry Sites

Figure 2: This figure shows a map of the 20077 industry facilities that are the potential PFAS sources.

River Connectivity

Figure 3: Map of HUC polygons and well locations.

Methodology

Source Apportionment Model

  • Y_i is observed concentration of a type of PFAS at well \boldsymbol s_i
  • The model relating the sources to the concentrations is

Y_i = \left(C_0(\boldsymbol s_i) + \sum_{k=1}^K C_k(\boldsymbol s_i) \right) \exp\left\{\varepsilon_i - \frac{\sigma^2}{2} \right\}

  • C_0 is the baseline amount of PFAS (i.e. not from these industry sites)
  • C_k is the contribution of industry sites from industry k.
  • Note that \varepsilon_i \overset{iid}{\sim} N(0, \sigma^2) and E[\exp\{\varepsilon_i - \sigma^2/2\}] = 1

Source Contributions

We model the source contribution [3] from each industry sector k, summed across the m_k industry sites as

C_k(\boldsymbol s_i) = \sum_{j=1}^{m_k} f(\boldsymbol s_i, \boldsymbol t_{jk}) V_{jk} where our transfer function f is defined as

f(\boldsymbol s_i, \boldsymbol t_{jk}) = \exp \left\{- \left(\frac{d(\boldsymbol s_i,\boldsymbol t_{jk})}{\phi} \right)^2\right\}

Extended Generalized Pareto Distribution

  • Given the extreme values, it makes sense to model the V_{jk}’s with a heavy-tailed distribution.
  • We’re not interested in extreme values alone. Small contributions can accumulate across many industry sites.
  • V_{jk} is a latent variable, so we cannot set a threshold.
  • We use the Extended Generalized Pareto Distribution (EGPD) [4] to model the entire distribution without requiring a fixed threshold.
  • \kappa controls the left tail, \xi controls the right tail, and \tau controls the scale of the distribution.

Extended Generalized Pareto Distribution

Figure 4: This figure shows how the EGPD changes with differing \kappa values.

Bayesian Hierarchical Model

  • We use a spike-and-slab setup [5] to help reflect the prior belief that many sources emit little PFAS, i.e., below the detection limit.
  • The “slab” distribution (b = 1) models the V_{jk} values that we expect to be above the detection limit. EGPD(\kappa, \xi, \tau_k)
    • The shape parameters being the same across industries while scale parameter varies by industry.
  • The “spike” distribution (b = 0) models the V_{jk} values that are close to the detection limit. EGPD(1, 0, \tau_{DL})
    • \tau_{DL} is set so that the 95th percentile is the detection limit.

Results

Posterior Means of V_{jk}’s

Figure 5: These maps show the log-posterior means of the industry site contributions.

Proportion of PFOS Contributed

Figure 6: This forest plot shows the posterior mean and corresponding 95% credible intervals for the proportion of PFOS contributed by that industry.

PFOS Contributed by Ecoregion

Figure 7: This plot looks at the proportion of PFOS contributed across 8 different ecoregions.

Conclusion

  • Our method is able to attribute PFAS to potential industry sources via source apportionment modeling.
  • This is accomplished through our Bayesian hierarchical model which incorporates the EGPD into a mixture model prior.
  • We see that chemical plants and waste management sites are the primary sources, although it varies across different regions.

Questions? Contact Me:

ctmcnei2@ncsu.edu
connor-mcneill.com

References

[1]
US EPA, O. (2021). Our Current Understanding of the Human Health and Environmental Risks of PFAS.Available at https://www.epa.gov/pfas/our-current-understanding-human-health-and-environmental-risks-pfas.
[2]
Anon. (2023). Tap water study detects PFAS “forever chemicals” across the US | U.S. Geological Survey.Available at https://www.usgs.gov/news/national-news-release/tap-water-study-detects-pfas-forever-chemicals-across-us.
[3]
Kowalczyk, G. S., Choquette, C. E. and Gordon, G. E. (1978). Chemical element balances and identification of air pollution sources in Washington, D.C. Atmospheric Environment (1967) 12 1143–53.
[4]
Naveau, P., Huser, R., Ribereau, P. and Hannart, A. (2016). Modeling jointly low, moderate, and heavy rainfall intensities without a threshold selection. Water Resources Research 52 2753–69.
[5]
George, E. I. and McCulloch, R. E. (1993). Variable Selection Via Gibbs Sampling. Journal of the American Statistical Association 88 881–9.
[6]
de Carvalho, M., Huser, R., Naveau, P. and Reich, B. J. (2026). Handbook of statistics of extremes. Chapman & Hall/CRC, Boca Raton, FL.