[{"authors":["Mengran Li"],"categories":[],"content":" Standard causal discovery breaks down precisely when it matters most — in the tails of a distribution, where rare but high-impact events occur. This project turns that obstacle into a tool. The very features that make extremes hard to model also make them directional. The core idea When a variable $X$ causally drives $Y$, predicting extreme values of $Y$ from extreme values of $X$ is systematically easier than predicting in the reverse direction. The forward-backward gap in tail prediction risk — we call it tail-induced asymmetry — is non-zero for causal pairs and vanishes for non-causal associations.\nForward Predict Y from extreme X X Y cause Risk → 0 Backward Predict X from extreme Y X Y spurious Risk \u0026gt; 0 Figure 1. Tail-induced asymmetry. When X causes Y, the forward tail prediction risk converges to zero while the backward risk stays bounded away from zero. The gap uniquely identifies causal direction from the joint tail distribution alone. This insight rests on the theory of multivariate regular variation. Under a recursive causal DAG, the angular measure characterising tail dependence encodes directional information consistent with the causal ordering — and that information can be estimated from data.\nMethod: S3ME We propose S3ME — Sparse Structure diScovery in Multivariate Extremes — a two-stage framework designed around the division of labour between which variables connect and how they are oriented.\nInput Heavy-tailed observations Stage 1 Skeleton recovery Proxy-adjusted penalised selection Stage 2 Edge orientation Forward/backward tail prediction risk DAG causal graph Figure 2. S3ME pipeline. Stage 1 recovers which variables connect, using an unpenalised proxy to absorb latent common shocks. Stage 2 orients each edge by comparing the tail prediction risk of the max-linear envelope model in both directions, penalised by EBIC. The two-stage design is deliberate. Skeleton recovery and edge orientation require different tools — penalised regression for one, tail risk minimisation for the other — and separating them lets each stage use the right machinery without compromise.\nResults Simulations S3ME maintains strong edge-recovery performance well beyond the sample size, across a range of dimensions.\n0.0 0.25 0.50 0.75 1.00 F1 score 0.84 p = 20 0.79 p = 50 0.77 p = 100 0.71 p = 200 n = 1,000 50 reps Figure 3. F1 scores for S3ME across increasing dimensionality. Performance holds up as the number of variables grows well beyond the sample size, a regime where classical causal discovery methods struggle without strong parametric assumptions. The method is robust to moderate misspecification of the tail model and to moderate levels of latent confounding.\nReal data Hydrology Danube River Network Applied to daily flow maxima at 31 gauging stations across the Danube basin. The recovered causal graph correctly reflects upstream-to-downstream flow direction at the majority of station pairs, using no geographic or hydrological prior information.\nFinance S\u0026amp;P 500 Tail Risk Applied to weekly minimum returns for 103 stocks over a 20-year period. The method identifies directional tail risk propagation across sectors, recovering known contagion pathways during historical market stress periods.\nBroader significance This work re-frames tail behaviour from nuisance to signal. Wherever extreme events propagate through a system with underlying causal structure, tail asymmetry provides a tool to recover that structure — even in regimes where classical causal discovery methods have no leverage.\nRead the paper Full theoretical development, proofs, simulation design, and the Danube and S\u0026amp;P 500 case studies are available in the arXiv preprint.\narXiv:2604.21620 PDF ","permalink":"https://mengranli.netlify.app/posts/causal_discovery/","series":[],"tags":[],"title":"Causal discovery in multivariate extremes"},{"authors":[],"categories":[],"content":"Project Overview Extreme quantile treatment effects (eQTEs) measure the causal impact of a treatment on the tails of an outcome distribution. Unlike average treatment effects, eQTEs capture how interventions affect rare, high-impact outcomes—critical for understanding climate extremes, financial crashes, or public health crises where the tail behavior matters most.\nStandard QTE methods struggle in extreme regimes due to data sparsity: by definition, extreme quantiles lie beyond observed data. Existing eQTE approaches rely on restrictive tail assumptions or apply interior-quantile theory that may not extend to the far tail. This creates a fundamental tension between the causal question (what is the treatment effect at extreme levels?) and the statistical challenge (how to estimate beyond observed data?).\nThis project develops the Tail-Calibrated Inverse Estimating Equation (TIEE) framework, which bridges this gap by combining information across quantile levels while anchoring the tail using extreme value theory.\nGoals Unified Framework: Develop an estimating equation approach that integrates interior quantile information with tail extrapolation in a principled manner.\nRobust Inference: Establish asymptotic properties and valid uncertainty quantification for extreme causal effects.\nPractical Impact: Enable causal attribution for rare events in environmental science, economics, and public health.\nMethodology: The TIEE Framework The TIEE framework addresses the core challenge of extreme causal inference through three key innovations:\nComponent Challenge TIEE Solution Tail Anchoring Extreme quantiles lack direct data Use Extreme Value Theory (EVT) models for tail extrapolation Information Borrowing Interior and tail estimates are disconnected Unified estimating equation across all quantile levels Causal Identification Treatment effects at extremes require careful handling Inverse probability weighting adapted for tail regimes Core Innovation: The TIEE estimator combines:\nInverse estimating equations for causal identification Tail calibration using generalized Pareto distributions Cross-quantile information to stabilize extreme estimates This allows valid inference at quantile levels where traditional methods fail, while properly propagating uncertainty from both the causal and extreme value components.\nKey Findings Simulation Study Our simulations evaluated TIEE under different tail behaviours (light, exponential, heavy) and model misspecifications:\nPerformance: TIEE maintains valid coverage and low bias even at extreme quantiles (e.g., 0.99, 0.995) where standard QTE methods break down. Robustness: The framework shows resilience to moderate misspecification of the tail model, thanks to information borrowing across quantile levels. Efficiency: Leveraging interior quantile data improves precision compared to pure EVT extrapolation. Real Data Application: Austrian Alps Precipitation We applied TIEE to study the causal effect of anthropogenic warming on extreme precipitation in the Austrian Alps:\nCausal Question: How has climate change altered the probability of extremely high precipitation events? Methodological Contribution: TIEE enables observational causal attribution for rare events under a counterfactual framework. Scientific Insight: The framework reveals how treatment effects vary across the outcome distribution—not just at the mean, but at the extremes where climate impacts are most consequential. Impact \u0026amp; Outputs 📄 arXiv Preprint\nTail-Calibrated Estimation of Extreme Quantile Treatment Effects\narXiv:2603.23309 | PDF\n🎯 Broader Significance\nThis framework establishes a new foundation for causal inference on rare, high-impact outcomes, with applications across:\nEnvironmental risk: Climate extreme attribution Economics: Tail risk in financial markets Public health: Rare disease treatment effects Status: arXiv preprint; journal resubmission in preparation\nPaper: Tail-Calibrated Estimation of Extreme Quantile Treatment Effects\n","permalink":"https://mengranli.netlify.app/posts/eqte/","series":[],"tags":[],"title":"Extreme quantile treatment effect (EQTE)"},{"authors":["Mengran Li"],"categories":[],"content":"Project Overview Extreme Event Attribution (EEA) is a vital field that assesses the extent to which human-induced climate change influences the probability and intensity of specific extreme weather events. As these events are becoming more frequent and intense, the need for robust statistical tools to underpin these assessments is more critical than ever.\nAttribution studies often rely on complex statistical models to characterize the behavior of rare events residing in the tail of the climate data distribution. The challenge is that climate extremes are inherently multivariate and high-dimensional. Therefore, the model choice is critical, as it must accurately capture the joint tail behaviour—the way these extreme variables occur together. Errors in modeling this tail dependence can lead to significant biases in attributing climate change\u0026rsquo;s influence.\nThis project employs a counterfactual causal inference framework, comparing the factual world (with human influence) against a counterfactual world (without human influence). Within this framework, we compare three distinct multivariate models designed to handle high-dimensional extreme data. Our objective is to estimate Attribution Ratios (ARs), which quantify the probability that human influence was necessary for the event\u0026rsquo;s occurrence.\nGoals Sensitivity Analysis: We aim to show precisely how the choice of tail model in EEA affects the resulting causal attribution conclusion.\nPractical Guidance: We will offer concrete, practical guidance for selecting and validating appropriate tail models in this context, thereby enhancing the scientific rigor and reliability of future attribution studies.\nMethodology: Comparing Three Multivariate Tail Models We employ a counterfactual causal inference framework, comparing three advanced multivariate models that represent different theoretical assumptions about extremal dependence.\nModel Core Concept \u0026amp; Assumption Key Feature Multivariate GPD (mGPD) A peaks-over-threshold model built on an exponential random vector. Assumes a rigid dependence structure (limited flexibility in how variables co-occur). Exponential Factor Copula (eFCM) Uses a shared exponential factor ($V$) added to a Gaussian process to induce spatial dependence. Provides flexible dependence via mixing, well-suited for regional extremes that decay with distance. Huser-Wadsworth (HW) Model A multiplicative model that bridges two theoretical regimes (asymptotic independence and dependence). A unified, data-driven model where parameter $\\delta$ lets the data decide the strength of the tail link. Key Research Focus: We measure the models\u0026rsquo; performance by estimating conditional dependence $\\mathbb{P}(U_2 \u0026gt; u \\mid U_1 \u0026gt; u)$ as $u \\to 1$ and observing how significant model discrepancies (misspecification) propagate to the final AR estimates.\nCore Findings from Simulation \u0026amp; Application A. Simulation Study: Impact of Model Misspecification Critical Result: Simulation Study 1 confirmed that misspecification of the extremal dependence structure leads to substantial bias in probability estimates ($\\hat{p}$). Accurate marginal distributions alone are insufficient. Performance: Study 2 showed that the eFCM consistently outperforms the mGPD in recovering the true tail probability, especially as the true underlying dependence strength increases. This validates the need for flexibility. B. Real Data Application: Multi-Region Attribution We applied the models to two distinct climate datasets: European winter precipitation maxima (CNRM model) and U.S. daily precipitation (ACCESS-CM2 model).\nModel Fit: The eFCM provided the best empirical fit for tail dependence in both regions, while the mGPD consistently showed overly strong, rigid dependence. Impact on ARs: The model choice strongly influences the spatial pattern of attribution: Rigid Tails (mGPD): Produced overly smooth maps and consistently overstated Attribution Ratios (ARs) across broad regions. Flexible Tails (eFCM/HW): Revealed realistic, regionally varying AR patterns, aligning better with physical and regional climate insights. Application Conclusion: Rigid tail assumptions severely bias attribution results. Flexible models are essential to reveal the true regional contrasts in the influence of climate change.\nConclusions and Guidance Key Conclusions Tail Dependence Matters: Attribution results depend critically on how joint extremes are modeled; misspecifying the tail structure can bias or reverse conclusions. Model Flexibility Improves Credibility: eFCM/HW fit tails well; mGPD overstates ARs. Choosing the right extremal dependence model is essential for credible EEA. Practical Guidance for Future EEA Studies Tip Rationale Prioritise Dependence Structure. Simulation results show dependence error is the dominant source of bias in probability estimation. Do Not Trust Margins Alone. Good marginal fit does not guarantee a realistic tail dependence structure. Prefer Robust Models. Choose models with built-in flexibility (like eFCM or HW) that yield more stable AR estimates and narrower uncertainty intervals. Choose Interpretability. Select models that produce scientifically defensible and usable AR patterns. Impact \u0026amp; Recognition This project has led to significant outputs and recognition within the statistical and climate communities.\n📄 Journal Article A revised manuscript, \u0026ldquo;On the importance of tail assumptions in climate extreme event attribution\u0026rdquo;, has been submitted to Climatic Change. arXiv:2507.14019 🧰 Software Release Developed and released the eFCM R package on CRAN. 🏆 Major Award: Received the 3rd Place Award in the IASC Data Analysis Competition, resulting in an invited talk and award ceremony at the ISI World Statistics Congress 2025. 🏅 Presentations: Talks (WSC 2025 - Travel grant awarded, STOR-i Extremes Workshop, RSC 2023 - Travel grant recipient) and poster (RSS 2025). ","permalink":"https://mengranli.netlify.app/posts/eea/","series":[],"tags":[],"title":"Extreme Event Attribution (EEA)"},{"authors":["Mengran Li"],"categories":[],"content":"This project showcases my work as part of the \u0026ldquo;Wee Extremes\u0026rdquo; team during the EVA (2023) Conference Data Challenge. It is summarized in the peer-reviewed article, “A wee exploration of techniques for risk assessments of extreme events” (Extremes, 2024).\nMotivation \u0026amp; Objectives Estimating risk measures for extreme events—such as high quantiles or joint exceedance probabilities—in environmental data is challenging due to:\nNon‑stationarity (covariate effects) Sparse data in the tail Multivariate dependence structure Extrapolation to unobserved extreme levels The EVA challenge provided a structured testbed (four tasks: C1–C4) involving both univariate and multivariate extremes in a simulated “Utopia” environment. The data were designed to mimic real environmental settings while preserving controlled truth for evaluation.\nOur goal was to build a composite methodology combining multiple statistical tools to robustly estimate extreme quantiles, tail probabilities, and joint behavior under uncertainty.\nMethodological Approach Key components of our approach include:\nNon-Stationary Modeling\nExtended traditional Extreme Value Analysis (EVA) by incorporating covariate effects, primarily through GAM based Generalized Pareto Distribution (GPD) parameterizations to model the conditional tail.\nTail Extrapolation\nRigorously compared different methods (e.g., GAM-GPD vs. Quantile Extrapolation) in Task C1 to select the optimal approach for predicting unobserved extreme levels.\nDependence Structure\nUsed Copula frameworks for multivariate tasks (C3, C4) to effectively model joint exceedance probabilities beyond the limitations of marginal modeling.\nUncertainty Quantification\nWe used Block Bootstrap methods to assess and validate model uncertainty. Crucially, the Block Bootstrap allowed us to perform model selection validation by generating pseudodata that preserves dependence structure, enabling us to test model performance against unobserved extreme data (far-tail extrapolation).\nStability \u0026amp; Robustness\nEmployed Model Averaging across bootstrap samples and Dimensionality Reduction techniques to mitigate overfitting and stabilize extreme quantile estimates.\nKey Findings \u0026amp; Insights 💡 Our results highlighted the critical impact of tail assumptions on risk estimates:\nValidation is Key for Extrapolation\nOur use of the Block Bootstrap was essential for validating model choices, especially when extrapolating to extreme return levels beyond the observed data range.\nTail Behavior is Critical\nMethods assuming lighter tail behavior than reality tended to severely underestimate risk, particularly for extreme return levels.\nDependence is Decisive\nErrors in tail dependence modeling can severely bias joint exceedance probability estimates (Tasks C3, C4).\nStability Matters\nModel averaging and bootstrap techniques were crucial to stabilize extreme quantile estimates and provide reliable uncertainty bounds.\nOutcome \u0026amp; Publication The work produced a peer-reviewed article:\n“A wee exploration of techniques for risk assessments of extreme events” (Extremes, 2024).\nWe contributed methodology that blends EVA, copula theory, dimensionality reduction, and bootstrap averaging.\n","permalink":"https://mengranli.netlify.app/posts/eva2023/","series":[],"tags":[],"title":"Wee Extremes: EVA 2023 Data Challenge"}]