This section has five parts. First, we describe the empirical model. Second, we present the construction of the green hydrogen offtaker database. Third, we outline the calculation of the network metrics used to measure spillover potential. Fourth, we describe the simulation procedure for the baseline green hydrogen demand projection and policy evaluation. Finally, we discuss limitations.
Empirical model
We begin by clarifying terminology, following Rogers’ theory of innovation53. Adoption, or adoption probabilities, refer to installation-level decisions to switch from a fossil energy source to green hydrogen or one of its derivatives. Diffusion refers to the aggregate system share of green hydrogen or its derivatives. The two are directly linked: diffusion reflects the aggregate outcome of adoption decisions made by thousands of individual installations or offtakers.
For the empirical identification of the model, we use natural gas power plants as historical analogue (see Supplementary Note 3 for details on the choice of analogue). The sample includes all natural gas power plants deployed in the EU between 1961 and February 2026 as recorded in the S&P Capital IQ Power Plant database. We opt for commercial rather than open-source data for this analysis, as, contrary to, for example, the Joint Research Center (JRC) Open Power Plants database, S&P Capital IQ reports reliable and largely complete values for power plant commissioning dates, which are required for a diffusion analysis. This dataset has been widely used to address research questions requiring power plant-level data62,63.
For specifying our empirical model, we model adoption as a function of cost competitiveness, spatial spillovers and distance to waterways and natural gas pipeline corridors,
$$P({A}_{i,t})=\phi \left({\beta }_{0}+{\beta }_{1}{C}_{i,t}+{\beta }_{2}{S}_{i,t}+{\beta }_{3}{D}_{i}^{\,\mathrm{water}}+{\beta }_{4}{D}_{i}^{\mathrm{pipe}}+{\beta }_{5}{S}_{i,t}{C}_{i,t}\right),$$
(1)
where:
-
P(Ai,t): probability of adoption for installation i in year t
-
Ai,t: adoption at installation i in year t
-
ϕ(⋅): logistic link function
-
Ci,t: cost competitiveness for installation i in year t
-
Si,t: spatial spillover (index) for installation i in year t
-
\({D}_{i}^{\,\text{water}\,}\): distance from installation i to the nearest port (inland or maritime)
-
\({D}_{i}^{\,\text{pipe}\,}\): distance from installation i to the nearest natural gas pipeline transmission corridor
-
Si,tCi,t: interaction term capturing joint effects of spatial spillovers and cost competitiveness
For all locations, adoption as the dependent variable is defined as binary and equal to 1 if a natural gas power plant has been deployed at the location and 0 otherwise. Adoption is irreversible and only considered for the set of observed plants sites, meaning that in February 2026 the adoption status for all locations is equal to 1.
The selection of independent variables is based on the theory of diffusion of energy technologies. For energy technologies, diffusion typically occurs as a new technology displaces an incumbent one. A major driver of adoption is therefore the relative cost–risk trade-off between the new and incumbent options13,64 (Supplementary Note 2). We include cost competitiveness as the primary technology-specific driver of adoption, defining it as the difference between the cost of the new technology and that of the incumbent. Details on the estimation of the cost-competitiveness measure are provided below. The rationale for considering spatial spillovers is explained extensively in the main article. In the context of regional policy interventions, spatial spillover pressure has been conceptualized as the intervention region’s size relative to the overall market44. We transfer this metric to an installation level, by defining the spatial spillover index as the level of local diffusion relative to the global average. We calculate this as the weighted adoption status of neighbouring installations within a 120-km radius, which is then demeaned cross-sectionally. This eliminates effects from an aggregate increase in average adoption65. The radius is chosen on the basis of hyperparameter optimization on a validation set maximizing the area under the receiver operating characteristic (ROC) curve (true versus false positive rate)66. Results are robust to respecifying the spatial spillover index with alternative spatial weight definitions or different cutoffs. We include distance to waterways and distance to natural gas transmission pipelines as physical infrastructure access variables, as natural gas or hydrogen transmission could occur by way of pipeline (including repurposed natural gas pipelines), or by way of water, in pressurized tanks or chemical carrier form45. Note that these are static measures, whereas the dynamic build-up of shared infrastructure is conceptually captured through the spatial spillover measure.
Our model relies on a direct link between adoption and cost competitiveness. This is consistent with the approach taken in many energy system simulations and techno-economic models that too rely on information on the cost of green hydrogen and competing technologies30,31,67. Cost competitiveness considerations also underlie—albeit only implicitly—the rate of diffusion assumptions that underpin the logistic models informing feasibility space studies35,37. Our model builds on and extends these concepts by making cost competition explicit—operationalizing cost competitiveness as the cost spread between green hydrogen and an incumbent technology—and taking an installation-level perspective that allows us to represent relational ties and spillovers between different offtakers.
Quantifying cost competitiveness requires comparing LCOEs across competing technologies. For natural gas power generation, this involves competition with coal, nuclear and renewable sources. The LCOE captures factors such as capital costs, fuel costs (including transport) or the capacity factor, many of which vary across space68. Measuring cost competitiveness as the aggregate spread between the LCOE of natural gas and one competing technology is therefore only of limited validity. Moreover, historical datasets on the LCOE of natural gas and competing technologies are extremely scarce. Existing time series combine multiple sources with data quality issues and methodological inconsistencies43. We address this by estimating cost competitiveness through a proxy based on a first-pass logit model that includes only spatial spillovers and physical infrastructure access. The residuals are then interpreted as a latent measure of cost competitiveness that is unobserved in the historical data,
$$\begin{array}{rcl}{\phi }^{-1}({A}_{i,t})&=&{\beta }_{0}+{\beta }_{1}{C}_{i,t}+{\beta }_{2}{S}_{i,t}+{\beta }_{3}{D}_{i}^{\,\text{water}}+{\beta }_{4}{D}_{i}^{\text{pipe}\,}+{\beta }_{5}{C}_{i,t}{S}_{i,t}+{\varepsilon }_{i,t},\\ {C}_{i,t}&=&{A}_{i,t}-\phi \,\left({\hat{\beta }}_{0}+{\hat{\beta }}_{1}{\hat{C}}_{i,t}+{\hat{\beta }}_{2}{\hat{S}}_{i,t}+{\hat{\beta }}_{3}{\hat{C}}_{i,t}{S}_{i,t}+{\hat{\beta }}_{4}{D}_{i}^{\,\text{water}\,}+{\hat{\beta }}_{5}{D}_{i}^{\,\text{pipe}\,}\right).\end{array}$$
(2)
Using proxies is consistent with established approaches in economic studies and is commonly used when data limitations constrain direct measurement of a variable69,70. Residuals can effectively capture a latent concept when that concept dominates the error term independently of the included variables and when the resulting proxy aligns with theoretical expectations in validation tests71. Because we control for both spatial spillovers and physical infrastructure access through distance to waterways and natural gas transmission pipelines, it is reasonable to assume that cost competitiveness drives the error term. Similar strategies have been used to measure market inefficiencies and non-competitive behaviour in power markets72 and in other contexts where variables are difficult to measure directly, for example, to identify unusual armament patterns in political science73.
To provide further evidence for the robustness of the proxy, we use an external measure of cost competitiveness, defined as the spread between the LCOE of natural gas and that of coal. For both LCOE time series, we use data collected by Way et al.43, who aggregate several historical sources to time series ranging from 1982 to 2020 for natural gas and from 1948 to 2020 for coal. The data quality of these series is generally limited, with incomplete coverage and partially inconsistent assumptions regarding generation technologies and input costs43. Furthermore, the data reflect aggregate US LCOEs and are not directly applicable to a European model. Despite its limitations, the LCOE measure is highly related to our proxy (correlation coefficient of 0.7 and R2 of 0.47).
In the second-pass regression we reintroduce the residuals to quantify the unexplained variation of the first-pass model, interpreted as a latent measure of cost competitiveness, in a form suitable for forward-looking simulation. As the residuals are a function of Ai,t, this produces perfect classification by construction. This is appropriate here, as the regression is used only to assign a coefficient to the residuals rather than generate new insights. To avoid overfitting, we introduce a random noise term with mean zero and an s.d. equal to one-fourth of the residuals’ s.d. The noise level is selected via hyperparameter optimization following the same procedure used for the distance-threshold choice66. As the residual-derived cost-competitiveness term is unitless, we rescale it to match the simulation environment by assigning the same upper bound, lower bound and zero anchors as the competitiveness measures used in the simulation65,66.
As Ai,t is cumulative and non-reversible, it is subject to autocorrelation. We therefore implement the logit model using GEE and R’s geepack package. GEE is a semiparametric generalized linear model for longitudinal data that provides consistent and unbiased coefficient estimates under weaker time-dependence assumptions than standard generalized linear models. Note that while we specify a first-order autoregressive (AR(1)) correlation structure, the estimates remain consistent and unbiased even if the real correlation structure deviates from AR(1)74.
We estimate adoption for a set of observed plant sites, with proximity to waterways and pipeline corridors included as additional controls. Our analysis thus explains the timing of adoption of natural gas power plants at locations that are relevant adoption units, rather than the estimating the unconditional probability that a plant is ever sited at an arbitrary location. This is consistent with the scope of our analysis: including locations where no plant is ever built would not identify additional adoption events and shift the analysis towards a siting problem over a large set of potential locations. Identification in our analysis relies on variation in adoption timing across locations over time. Bias from not including additional locations beyond relevant adoption units arises only in cases where adoption timing, siting and the independent variables are jointly affected by unobserved events that are both local and time-varying. While we cannot fully rule out such confounding, we are confident that it does not materially affect our results.
Results are discussed in detail above (Main). We evaluate the discriminatory power of the model using the area under the ROC curve and mean standard error66. For the training set, the area under the ROC curve amounts to 0.998, and the mean standard error to 0.0155, indicating very high discriminatory power. This result follows directly from the specification of the model, as the second pass regression includes the transformed residuals of the first-pass regression as an additional regressor. The more relevant test is therefore to evaluate model performance using alternative measures of cost competitiveness. To that purpose, we replace it with the external LCOE measure discussed above. Discriminatory power declines in this specification (area under the ROC curve of 0.79, mean standard error of 0.19) but performance remains sufficiently robust to support its validity. Marginal effects calculated on the basis of this model are similarly consistent, with the relative importance of spatial spillovers increasing at the expense of cost competitiveness (Supplementary Fig. 8). Note that we also use this LCOE time series for hyperparameter (distance threshold and random noise term) optimization on the validation set. Further robustness checks show no concerns regarding multicollinearity (Supplementary Table 4) and spatial autocorrelation (Supplementary Fig. 9). Serial autocorrelation (Supplementary Fig. 10) patterns are consistent with the specification of Ai,t as cumulative and non-reversible.
Offtaker database
For our analyses on the prospective EU hydrogen economy, we build an original database of green hydrogen offtakers across the EU, using the most recent spatially resolved open-source data available to us as of 31 December 2024. For this purpose, we define an offtaker as any installation or consumer currently using fossil energy sources that could be substituted with green hydrogen or one of its derivatives (such as e-methane, e-methanol or e-kerosene). We include offtakers from industrial, transport (aviation, shipping and heavy-duty), power and heating sectors. A breakdown by sector and country of the full integrated database is presented in Extended Data Tables 3 and 4. Industry-level emission factors and hydrogen-specific energy consumption factors are presented in Extended Data Tables 5 and 6. A detailed documentation of the database build-up is provided in the Supplementary Methods.
Network metrics
We leverage network theory to assess the structure of Europe’s prospective green hydrogen economy. In doing so, we represent offtakers as a spatial network, where each potential offtaker is a node, and edges between nodes represent influence through geographic proximity and relational ties. To characterize this network, we apply two centrality measures to reflect three different aspects of influence: degree centrality (measuring local influence and representing hydrogen valleys) and betweenness centrality (measuring mediating influence, representing hydrogen corridors)48. Details on the choice and construction of network metrics are provided in the Supplementary Methods. To make sure network metrics are not disproportionately reflected by a single sector, we recalculate them excluding heating installations, which make up approximately half of the dataset, and find no material changes (Extended Data Fig. 3).
Simulation
We build on the empirical model to generate forward-looking projections of green hydrogen demand. These projections serve as a baseline (excluding policy interventions beyond carbon pricing) against which to evaluate policy interventions. Each simulation year proceeds as follows: first, cost differentials between green hydrogen (or a derivative product) and the corresponding fossil incumbent are retrieved from forecast values and matched to each installation by sector. Second, the adoption status from the previous year is used to compute spatial spillover indices based on a 120-km distance threshold using R’s FNN and Matrix packages; as in the empirical model, these indices are demeaned cross-sectionally to remove aggregate trends in diffusion.
Third, for each installation i in sector s, the probability of adoption is computed as the product of the empirical logit prediction and a sector-level adjustment factor that aligns long-run adoption with the sector’s saturation level β6s. Formally,
$${P}_{i,t}=\phi \,\left({\beta }_{0}+{\beta }_{1}{C}_{i,t}+{\beta }_{2}{S}_{i,t}+{\beta }_{3}{D}_{i}^{\,\mathrm{water}}+{\beta }_{4}{D}_{i}^{\mathrm{pipe}}+{\beta }_{5}{C}_{i,t}{S}_{i,t}\right)\cdot {F}_{s,t},$$
(3)
where Ci,t denotes cost competitiveness, Si,t spatial spillovers and Fs,t is defined as
$${F}_{s,t}=\max \,\left(0,\ \frac{{\beta }_{6s}-{\overline{A}}_{s,t}}{1-{\overline{A}}_{s,t}}\right),$$
(4)
with \({\overline{A}}_{s,t}\) denoting the mean adoption rate in sector s at time t.
On the basis of the simulated adoption probability, adoption is then realized as
$${A}_{i,t}=\left\{\begin{array}{l}1,\qquad{\mathrm{if}}\,{P}_{i,t}\ge {U}_{i,t},\quad {U}_{i,t}\sim{\mathcal{U}}(0,1),\, s\notin {\mathcal{C}},\\ 0,\qquad{\mathrm{if}}\,{P}_{i,t} < {U}_{i,t},\quad s\notin {\mathcal{C}},\\ {P}_{i,t},\quad s\in {\mathcal{C}},\end{array}\right.$$
(5)
where \({\mathcal{C}}\) denotes a set of ‘continuous-adoption’ sectors (aviation, shipping and heavy-duty transport, where green hydrogen may be adopted partially for only a subset of planes or ships) and Ui,t is a random draw from a uniform distribution. All results (probability, adoption, cost gap and spatial influence) are stored annually for post-simulation analysis. To implement the simulation, forward-looking estimates of cost competitiveness are required. As in the empirical model, cost competitiveness is defined as the cost difference between green hydrogen (or a derivative) and the fossil incumbent in each sector and year (Supplementary Table 5). We rely on data from Odenweller and Ueckerdt6 and extend these projections to 2100 (Supplementary Fig. 11). To assess whether spatially informed policies yield superior policy outcomes relative to spatially neutral policies, we simulate demand-side policy interventions in two settings. First, we assess a one-sided CfD in a controlled setting designed to isolate the effect of spatial targeting while holding sector, technology cost conditions and installation size approximately constant. Second, we evaluate a more realistic auction design that mirrors the EHB auction mechanism. We show both approaches because they serve different purposes: the controlled setting isolates the effect of spatial targeting under comparable conditions, while the EHB-style auction assesses whether spatial targeting can also improve outcomes under a policy allocation rule closer to existing subsidy practice. Further details on the simulation procedure (including the operationalization of cost competitiveness) and the specification of the policy interventions are reported in the Supplementary Methods.
Limitations
While we conduct extensive robustness checks to ensure the validity of our work, some limitations necessarily persist. Using historical analogues is a well-established technique for empirically grounded analysis when robust historical data on the target technology—in this case green hydrogen—are unavailable and has been widely used in feasibility studies for green hydrogen35 and other technologies such as CCUS37. As with any analogue, some differences in characteristics and diffusion patterns are likely to persist. We therefore make our choice of natural gas power plants transparent and reproducible (Supplementary Note 3) and provide a robustness check by re-estimating our model using cogeneration and wind plants (Supplementary Note 4), which yields consistent results.
The empirical analysis also requires a measure of cost competitiveness. For natural gas power plants, this is the LCOE difference relative to competing technologies (coal, nuclear, hydro, wind or solar). Historical LCOE time series are of uneven quality and do not fully capture the heterogeneity of this competition43. To address this, we use a residual-based latent measure of cost competitiveness. While such approaches are common when direct measurement is difficult—including in energy research72—they are only valid when the latent signal dominates the residuals. Although we are confident this is the case here, we cannot fully exclude the persistency of some residual noise.
Our original database of 14,108 green hydrogen offtakers is a key contribution to both research and practice. The database covers approximately 70% of EU emissions and therefore most major use cases for green hydrogen. Some simplifications are unavoidable. Certain applications are excluded where alternative decarbonization options have clear economic or technological advantages (such as passenger vehicles) or where their contribution to emissions and energy use is minor (such as cargo-only flights). Given the scale of the study, we rely on (sub)industry-level emission factors and hydrogen-specific energy consumption factors, which may not fully reflect heterogeneity in activities and production processes across installations.
For the forward-looking analysis, we rely on sector-level technology pairings, which again cannot capture all plant-level differences. Production pathways vary across energy-intensive sectors. In iron and steel, for instance, upstream material can be produced via natural-gas-based direct reduction or via coke-fuelled blast furnaces75. In non-ferrous metals such as copper, zinc and lead, coal remains widely used for roasting and smelting and it also fuels rotary kilns in cement manufacturing76. Since our focus is on how geography and spatial spillovers shape diffusion rather than on detailed plant-level cost heterogeneity, we abstract from these process-specific differences.
Finally, we incorporate competition from alternative decarbonization technologies through saturation rates that cap sector-level adoption rather than embedding this competition directly into the cost-competitiveness variable. While this is the more robust approach (see the ‘Simulation’ section), it still abstracts from process- and installation-level heterogeneity in cost competitiveness. As making plant-level predictions is not the purpose of this paper, this simplification is appropriate in the present case. An analysis that extends cost competitiveness to a process or installation level in given sectors of interest could be a promising direction for future research. We provide detailed considerations of our choice of this approach and the derivation of saturation rates in Supplementary Note 6.