Kernel measures of (conditional) dependence

a foundational statistical problem

Danica J. Sutherland  (she/her)  –  University of British Columbia & Amii

Based on work with Zheng He (UBC), Roman Pogodin (UCL → McGill/Mila → Google),
Antonin Schrab (UCL → Cambridge), Yazhe Li (UCL/DeepMind → Microsoft),
Namrata Deka (UBC → CMU), Arthur Gretton (UCL/DeepMind),
Nathaniel Xu (UBC), Feng Liu (Melbourne), Aaron Wei (UBC).

He et al. (2026), merging He et al. (2025) and Pogodin et al. (2024), in prep.
Pointwise testing results (Sutherland 2026): on arXiv very soon.
Mentioned extensions (Nate & Feng; Aaron): very early work.

Why care about conditional independence?

Does X tell us anything about Y once we know Z?

XYZiffPXYZ=PXZPYZ,Z-almost surely

  • Science: are cancer cell shapes associated with medication dosage, controlling for disease stage?
  • Causal discovery: keep an edge XY in a DAG only if XYZ for every candidate Z; optimal CI test gives optimal structure learning (Gao et al. 2026)
  • Fairness: do car-insurance premiums depend on neighbourhood demographics, given driver risk? (Hardt et al. 2016; Angwin et al. 2017)
  • Robustness / domain shift: do my pedestrian location predictions depend on whether I’m in Manhattan or Manitoba, given where the pedestrians actually are? (Jiang and Veitch 2022)

Comparing distributions

A problem in its own right: “two-sample testing” or Mu Zhu’s talk yesterday.

For this talk, a building block.

Maximum Mean Discrepancy

Assume a kernel k on X with RKHS H and feature map φ(x)=k(x,).

Mean embedding of a distribution P (Berlinet and Thomas-Agnan 2004): μP:=Pφ=EXPφ(X)=k(x,)dP(x)H,μP,fH=Pffor fH.

Can use this to define a distance on distributions (Gretton et al. 2012): MMD(P,Q):=μPμQH=supf:fH=1μPμQ,f=supf:fH=1PfQf.

The power function from kernel interpolation has P=δx, Q=iui(x)δxi.

For a characteristic kernel (Sriperumbudur et al. 2011) (e.g. Gaussian, …):
PμP is injective, so MMD(P,Q)=0 iff P=Q.
Characteristic kernels exist on any Polish space (Ziegel et al. 2024).

MMD in GP form

MMD2(P,Q)=EX,XPY,YQ[k(X,X)+k(Y,Y)2k(X,Y)]=EfGP(0,k)[(PfQf)2]

Worst case over the RKHS ball (kernel side) = average case under the GP prior
(Ritter 2000; Kanagawa et al. 2018, sec. 6.1)

Testing with MMD

  • Null hypothesis: P=Q.
  • Reject the null hypothesis if MMD(P^,Q^) is big.
  • Test has level α if probability of rejecting is α when null hypothesis holds.
  • Want a test with high power: probability of rejecting when PQ.
  • Can decide on “big” with a permutation test:
    • If P=Q, samples in X and Y are exchangeable.
    • Bunch of times: randomly re-assign to P~, Q~ and compute MMD(P~,Q~).
    • Reject if MMD(P^,Q^)> the 1α quantile of (distribution above plus real assignment).
    • Gives a level α test (Hemerik and Goeman 2018).
  • You should learn a (deep) kernel! (Sutherland et al. 2017; Liu et al. 2020; Wei et al. 2026)

Testing independence

Getting closer to the problem we care about!

Hilbert-Schmidt Independence Criterion

XYiffPXY=PXPYMI(X;Y)=KL(PPXPY)

Take kernels kX on X, kY on Y; kXkY is a kernel on X×Y.
Its RKHS is HXHYHS(HY,HX). Then (Gretton et al. 2005, 2012; Kanagawa et al. 2018)

HSIC(X,Y):=MMDkXkY2(PXY, PXPY)=PXY(φXφY)PXφXPYφYCXY  has  f,CXYg=Cov(f(X),g(Y))HS2=EfGP(0,kX)gGP(0,kY)[Cov(f(X),g(Y))2]

Testing with HSIC

  • HSIC^n=1n2KX,HKYHF for centring matrix H=I1n11: O(n2) time.
  • Can use same permutation idea: shuffle the Ys. Test has level α.
  • If XY, then HSIC^n=Op(1/n).
  • If HSIC>0, then HSIC^nHSIC>0 so power 1.
  • In practice, you should learn a (deep) kernel! (Xu et al. 2026)

Testing conditional independence

The problem we care about… and where things get tough.

Kernel Conditional Independence criterion

XYZiffPXYZ=ΠCI(PXYZ)for ΠCI(P)(dx,dy,dz):=PZ(dz)PXZ=z(dx)PYZ=z(dy)

ΠCI is the unique CI law matching (X,Z) and (Y,Z) marginals (Dawid and Lauritzen 1993; Massa and Lauritzen 2010). It minimizes KL(PXYZQ) over CI laws Q, and CMI(X;YZ)=KL(PΠCIP).

We can define (this version from He et al. (2026); original quantity from Zhang et al. (2011)) KCI(X,YZ):=MMDkXkYkZ2(PXYZ, ΠCI(PXYZ))=EfGP(0,kX)gGP(0,kY)wGP(0,kZ)[(EZPZ[w(Z)CovPXYZ(f(X),g(Y)Z)])2]

KCI characterizes CI

Theorem (He et al. 2026) For bounded continuous kernels on Polish spaces, the following are equivalent:

  • KCI=0 if and only if XYZ.
  • kX,kY are characteristic, and kZ is integrally strictly positive definite.
  • For bounded kernels, KCIcCMI.
  • If we take kX,kY linear and kZ1, we get (square of) the population target of GCM, ECov(X,YZ) (Shah and Peters 2020).
  • If we use kZ=ωω, get weighted GCM (Scheidegger et al. 2022).

Conditional mean embeddings (CMEs)

Have ΠCI(P)(φXφYφZ)=EZ[μXZ(Z)μYZ(Z)φZ(Z)],
using the conditional mean embedding (Song et al. 2009; Grünewälder et al. 2012; Park and Muandet 2020) μXZ(z):=E[φX(X)Z=z]HX,μXZ(z),f=E[f(X)Z=z].

So, subtracting and conditioning on Z, KCI=PXYZ[(φX(X)μXZ(Z))(φY(Y)μYZ(Z))φZ(Z)]2.

Using (KXc)ij=φX(xi)μXZ(zi),φX(xj)μXZ(zj), natural estimator KCIn=1n(n1)ij(KZ)ij(KXc)ij(KYc)ij.

Testing with KCI

  • Unfortunately we can’t do exact permutation tests: only have one sample per Z.
  • Under H0, KCIn=Op(1/n); test threshold has scale 1/n.
    • Can estimate threshold with wild bootstrap (or use concentration inequalities).
  • If KCI>0, then KCInKCI>0: the test is consistent.
  • So, all good… except that we don’t know the CMEs!

Estimating CMEs

μXZ(z)=E[φX(X)Z=z]

One option: conditional generative model to sample from XZ=z; then use Monte Carlo. Called (approximate) “Model-X” (Candès et al. 2018).

Another: regression from zZ to φX(X)HX (Grünewälder et al. 2012).

μ^XZ(z)=i=1mβi(z)φX(xi),β(z)=(KZ+mλI)1kZ(z).

  • We only need μ^(z),μ^(z)=β(z)KXβ(z): can evaluate everything.
  • Rates (Li et al. 2022, 2024): In the best case (with characteristic kX) convergence is m1/4; worst case can be arbitrarily slow.

KCI with estimated CMEs

Just plug μ^XZ and μ^YZ into KCIn: call this KCI^n. Use wild bootstrap.

Need the right kZ for good power – trying to learn it makes the problem way worse.

Is this task even possible?

No free lunch (Shah and Peters 2020)

  • X,Y,ZUnif[0,1] independent
  • D := 100th decimal digit of Z
  • X := X but first digit D; Y likewise
  • XY, but XYZ

Z 0.471829316
X 0.829157140
X 0.329157140
Y 0.104830204
Y 0.304830204
Z 0.471829716

  • Z := Z with the 100th digit re-randomized; XYZ
  • (X,Y,Z) and (X,Y,Z) are indistinguishable to any “reasonable” test that treats Z as a scalar with 10100 samples.
  • So, if we reject on (X,Y,Z) with probability α, same is true on (X,Y,Z)
  • Can construct an example like this near any alternative, for any measurable test

No free lunch (Shah and Peters 2020)

Theorem (Shah and Peters 2020, Corollary 3) For M(0,], let E be the laws on (M,M)s with a Lebesgue density. Let PE be those with XYZ, and Q=EP the rest.

If (ψn) has uniformly asymptotic level α over P, then lim supnQnψnα for every QQ.

Okay, but, who cares about the 100th decimal place?

  • This is obviously a pathological construction.
  • But what is it, really? It’s a conditional law PXZ that regression can’t learn from finite data. Hard to estimate μXZ means hard to test.
  • For KCI this is exactly the failure mode: the only unknowns are μXZ,μYZ. With them, validity is one line of Hoeffding; without them, nothing is guaranteed.
  • If we assume PXZ known (Model-X), everything is fine. If it’s wrong, Type I error inflates by the conditional total variation error (Berrett et al. 2020).
  • So, let’s think a little more about what “possible” means.

It’s actually not so bad

Three kinds of level

Tests are ψn:Vn[0,1] = rejection probability. Over a null class P, (ψn) has

  • finite-sample level α: supnsupPPPnψnα
  • uniformly asymptotic level α: lim supnsupPPPnψnα
  • pointwise asymptotic level α: supPPlim supnPnψnα

Each implies the next. Shah and Peters (2020) rule out useful tests with the first two.

Practical tests almost always only claim the third.

A not-so-scary problem

Testing H0:EW=0 among laws with finite variance.

  • Bahadur and Savage (1956) show that any test with uniformly asymptotic level α has lim supnQnψnα at every alternative.
  • The t-test has pointwise asymptotic level α and is consistent against every alternative.
  • Maybe CI testing is like this?
    • Györfi and Walk (2012) claimed yes, but proof was wrong (Neykov et al. 2021).
    • Zhang et al. (2011, Prop. 5) would imply yes, but proof sketch was wrong.
    • Stated as open by Györfi et al. (2023) and Boeken et al. (2026).

CI testing is scarier…

Theorem (Sutherland 2026) Fix M(0,] and LL0(M,s). Let P be the CI laws with density L on (M,M)s, and Q the conditionally dependent ones. If (ψn) has pointwise asymptotic level α<1 over P, infQQ lim supnQnψnα. The offending null can have a continuous density and a uniformly continuous conditional law; the alternatives can be taken C.

Corollary. If Qnψn1 for every alternative, then some null P has lim supnPnψn=1.

…but not too scary

Proposition (Sutherland 2026) Take any D0 with D(P)>0 iff XYZ, and estimators D^nD(P) in probability for every P. A partition estimate will do, or KCI with universally consistent CMEs (Tamás and Csáji 2024).

Then ψn=α+(1α)min{1,D^n/δ} has pointwise asymptotic level α, limiting power α+(1α)min{1,D(Q)/δ}>α at every alternative, and Qnψn1 whenever D(Q)δ.

  • Uniform-level CI tests have power against no alternative (Shah and Peters 2020)
  • Pointwise-level CI tests can have power 1 against some alternatives
  • …but not against all alternatives, unlike the t-test

But in practice, it’s still pretty hard.

“Eventually does the right thing” is a really weak guarantee!

Estimated CMEs break the null calibration

Fit μ^XZ=μXZ+ΔX and μ^YZ=μYZ+ΔY on independent data, plug in. Under H0, the population value of the plug-in statistic is KCI^=E[kZ(Z,Z) ΔX(Z),ΔX(Z)HX ΔY(Z),ΔY(Z)HY]0

  • μ^XZ is smooth in kZ; hopefully μXZ is too or everything is hopeless.
  • So ΔX, ΔY are smooth functions in kZand so KCI^ might be big.
  • Our test calibration is based on KCIn=Op(1/n).
  • But if KCI^>0, then KCI^n=KCI^+Op(1/n).
  • For fixed regressions, bigger n makes everything worse.
  • Can’t tell regression error from real dependence!

How bad? Pretty bad!

Synthetic rat hippocampal data. Nsplit ratio for training CMEs, rest for testing.

Reducing bias with SplitKCI

Fit two independent CMEs per variable and use (Pogodin et al. 2024) (KXc)ij=φX(xi)μ^XZ(1)(zi),φX(xj)μ^XZ(2)(zj).

Changes the bias under the null hypothesis to the (smaller) E[kZ(Z,Z) ΔX(1)(Z),ΔX(2)(Z)HX ΔY(1)(Z),ΔY(2)(Z)HY].

Can prove rates work out if the regression is much more precise than the test.

Heuristic to determine split: smallest one that says XZZ, YZZ.

Manages the regression error. Doesn’t fix it.

Ongoing work: “leave-two-out” extension that allows using the same data for training CMEs and the KCI estimate. Highlights that just throwing away data for KCI is weird…

Fundamentally, what we’re missing is the uncertainty in a Hilbert-valued regression.

Everything is about Δ=μ^μ

  • Shah and Peters (2020): an adversary puts PXZ where regression can’t see it.
  • Pointwise: even for one fixed law, regression error at small scales can be made to decay arbitrarily slowly.
  • Practice: our null calibration can’t account for regression errors. SplitKCI helps, doesn’t solve.
  • The only kind of validity that is possible – finite-sample or uniform level under an explicit smoothness class, or pointwise level with an explicit separation δ – is also exactly what uncertainty quantification for the regression would give us.

We need epistemic uncertainty quantification for Hilbert-valued kernel regression.

Hilbert-valued GPs to the rescue?

  • We’re starting to try modelling μXZGP: inputs ziZ, outputs φX(xi)HX
  • Calibrate KCI^n under the Δ posterior, instead of assuming Δ=0
  • A priori constraints on e.g. 0<[μXZ(z)](x)1 for all x,z (Holger’s talk yesterday)
  • Choosing infinite-dimensional priors/likelihoods is hard! Reasonable epistemic uncertainty in E[φX(X)Z=z] needs reasonable aleatoric covariance ΣXZ
    • λI (a) isn’t even a covariance (not trace-class) and (b) is super wrong
    • Definitely varies with z in realistic settings
    • Estimating it is another regression, at least as hard as the first
  • Maybe there’s a worst-case kernel approximation result that gives a finite-sample-valid CI test, under some explicit smoothness class?

Summary

  • MMD(P,projection of P onto H0); unconditional independence is easy
  • Conditional independence (KCI) needs a Hilbert-valued regression and is hard
    • Can’t have universally uniform level (Shah and Peters 2020)
    • Can have pointwise level and consistency against any separated class (new)
    • Ignoring rates, only slightly harder than testing a mean
  • In practice, the problem is the regression error
    • SplitKCI helps manage, definitely does not solve
    • Would probably fix it: honest epistemic uncertainty for Hilbert-valued kernel regression

Consider helping us out! 😇 · djsutherland.ml

References

Angwin, Julia, Jeff Larson, Lauren Kirchner, and Surya Mattu. 2017. “Minority Neighborhoods Pay Higher Car Insurance Premiums Than White Areas with the Same Risk.” ProPublica, April. https://www.propublica.org/article/minority-neighborhoods-higher-car-insurance-premiums-white-areas-same-risk.
Bahadur, R. R., and Leonard J. Savage. 1956. “The Nonexistence of Certain Statistical Procedures in Nonparametric Problems.” Annals of Mathematical Statistics 27 (4): 1115–22. https://doi.org/10.1214/aoms/1177728077.
Berlinet, Alain, and Christine Thomas-Agnan. 2004. Reproducing Kernel Hilbert Spaces in Probability and Statistics. Springer. https://link.springer.com/book/10.1007/978-1-4419-9096-9.
Berrett, Thomas B., Yi Wang, Rina Foygel Barber, and Richard J. Samworth. 2020. “The Conditional Permutation Test for Independence While Controlling for Confounders.” Journal of the Royal Statistical Society Series B 82 (1): 175–97. https://doi.org/10.1111/rssb.12340.
Boeken, Philip, Eduardo Skapinakis, Konstantin Genin, and Joris M. Mooij. 2026. Topological Criteria for Hypothesis Testing with Finite-Precision Measurements. arXiv. https://arxiv.org/abs/2601.13946.
Candès, Emmanuel, Yingying Fan, Lucas Janson, and Jinchi Lv. 2018. “Panning for Gold: Model-X Knockoffs for High Dimensional Controlled Variable Selection.” Journal of the Royal Statistical Society Series B 80 (3): 551–77. https://doi.org/10.1111/rssb.12265.
Dawid, A. Philip, and Steffen L. Lauritzen. 1993. “Hyper Markov Laws in the Statistical Analysis of Decomposable Graphical Models.” Annals of Statistics 21 (3): 1272–317. https://doi.org/10.1214/aos/1176349260.
Gao, Ming, Yuhao Wang, and Bryon Aragam. 2026. “Optimal Structure Learning and Conditional Independence Testing.” International Conference on Machine Learning (ICML). https://openreview.net/forum?id=uurjpxrLc5.
Gretton, Arthur, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Schölkopf, and Alexander Smola. 2012. “A Kernel Two-Sample Test.” Journal of Machine Learning Research 13 (25): 723–73. https://jmlr.org/papers/v13/gretton12a.html.
Gretton, Arthur, Olivier Bousquet, Alex Smola, and Bernhard Schölkopf. 2005. “Measuring Statistical Dependence with Hilbert-Schmidt Norms.” Algorithmic Learning Theory (ALT), 63–77. https://doi.org/10.1007/11564089_7.
Grünewälder, Steffen, Guy Lever, Luca Baldassarre, Sam Patterson, Arthur Gretton, and Massimiliano Pontil. 2012. “Conditional Mean Embeddings as Regressors.” International Conference on Machine Learning (ICML). https://arxiv.org/abs/1205.4656.
Györfi, László, Tamás Linder, and Harro Walk. 2023. “Lossless Transformations and Excess Risk Bounds in Statistical Inference.” Entropy 25 (10): 1394. https://doi.org/10.3390/e25101394.
Györfi, László, and Harro Walk. 2012. “Strongly Consistent Nonparametric Tests of Conditional Independence.” Statistics & Probability Letters 82 (6): 1145–50. https://doi.org/10.1016/j.spl.2012.02.023.
Hardt, Moritz, Eric Price, and Nathan Srebro. 2016. “Equality of Opportunity in Supervised Learning.” Advances in Neural Information Processing Systems. https://doi.org/10.48550/arXiv.1610.02413.
He, Zheng, Roman Pogodin, Yazhe Li, Namrata Deka, Arthur Gretton, and Danica J. Sutherland. 2025. “On the Hardness of Conditional Independence Testing in Practice.” Advances in Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/2512.14000.
He, Zheng, Roman Pogodin, Antonin Schrab, et al. 2026. “On Measuring and Testing Conditional Independence.” Unpublished manuscript.
Hemerik, Jesse, and Jelle Goeman. 2018. “Exact Testing with Random Permutations.” TEST 27 (4): 811–25. https://doi.org/10.1007/s11749-017-0571-1.
Jiang, Yibo, and Victor Veitch. 2022. “Invariant and Transportable Representations for Anti-Causal Domain Shifts.” ICML 2022: Workshop on Spurious Correlations, Invariance and Stability. https://arxiv.org/abs/2207.01603.
Kanagawa, Motonobu, Philipp Hennig, Dino Sejdinovic, and Bharath K. Sriperumbudur. 2018. Gaussian Processes and Kernel Methods: A Review on Connections and Equivalences. arXiv. https://arxiv.org/abs/1807.02582.
Li, Zhu, Dimitri Meunier, Mattes Mollenhauer, and Arthur Gretton. 2022. “Optimal Rates for Regularized Conditional Mean Embedding Learning.” Advances in Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/2208.01711.
Li, Zhu, Dimitri Meunier, Mattes Mollenhauer, and Arthur Gretton. 2024. “Towards Optimal Sobolev Norm Rates for the Vector-Valued Regularized Least-Squares Algorithm.” Journal of Machine Learning Research 25 (181): 1–51. https://jmlr.org/papers/v25/23-1663.html.
Liu, Feng, Wenkai Xu, Jie Lu, Guangquan Zhang, Arthur Gretton, and Danica J. Sutherland. 2020. “Learning Deep Kernels for Non-Parametric Two-Sample Tests.” International Conference on Machine Learning (ICML). http://proceedings.mlr.press/v119/liu20m.html.
Massa, M. Sofia, and Steffen L. Lauritzen. 2010. “Combining Statistical Models.” In Algebraic Methods in Statistics and Probability II, vol. 516. Contemporary Mathematics. American Mathematical Society. https://doi.org/10.1090/conm/516/10179.
Neykov, Matey, Sivaraman Balakrishnan, and Larry Wasserman. 2021. “Minimax Optimal Conditional Independence Testing.” Annals of Statistics 49 (4): 2151–77. https://doi.org/10.1214/20-AOS2030.
Park, Junhyung, and Krikamol Muandet. 2020. “A Measure-Theoretic Approach to Kernel Conditional Mean Embeddings.” Advances in Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/2002.03689.
Pogodin, Roman, Antonin Schrab, Yazhe Li, Danica J. Sutherland, and Arthur Gretton. 2024. Practical Kernel Tests of Conditional Independence. arXiv. https://arxiv.org/abs/2402.13196.
Ritter, Klaus. 2000. Average-Case Analysis of Numerical Problems. Vol. 1733. Lecture Notes in Mathematics. Springer. https://doi.org/10.1007/BFb0103934.
Scheidegger, Cyrill, Julia Hörrmann, and Peter Bühlmann. 2022. “The Weighted Generalised Covariance Measure.” Journal of Machine Learning Research 23 (273): 1–68. https://jmlr.org/papers/v23/21-1328.html.
Shah, Rajen D., and Jonas Peters. 2020. “The Hardness of Conditional Independence Testing and the Generalised Covariance Measure.” Annals of Statistics 48 (3): 1514–38. https://doi.org/10.1214/19-AOS1857.
Song, Le, Jonathan Huang, Alex Smola, and Kenji Fukumizu. 2009. “Hilbert Space Embeddings of Conditional Distributions with Applications to Dynamical Systems.” International Conference on Machine Learning (ICML), 961–68. https://doi.org/10.1145/1553374.1553497.
Sriperumbudur, Bharath K., Kenji Fukumizu, and Gert R. G. Lanckriet. 2011. “Universality, Characteristic Kernels and RKHS Embedding of Measures.” Journal of Machine Learning Research 12 (70): 2389–410. https://jmlr.org/papers/v12/sriperumbudur11a.html.
Sutherland, Danica J. 2026. “Conditional Independence Is Not Pointwise Testable.” Unpublished manuscript.
Sutherland, Danica J., Hsiao-Yu Tung, Heiko Strathmann, et al. 2017. “Generative Models and Model Criticism via Optimized Maximum Mean Discrepancy.” International Conference on Learning Representations (ICLR). http://openreview.net/forum?id=HJWHIKqgl.
Szabó, Zoltán, and Bharath K. Sriperumbudur. 2018. “Characteristic and Universal Tensor Product Kernels.” Journal of Machine Learning Research 18 (233): 1–29. https://jmlr.org/papers/v18/17-492.html.
Tamás, Ambrus, and Balázs Csanád Csáji. 2024. “Recursive Estimation of Conditional Kernel Mean Embeddings.” Journal of Machine Learning Research 25 (264): 1–35. https://jmlr.org/papers/v25/23-0168.html.
Wei, Aaron, Milad Jalali Asadabadi, and Danica J. Sutherland. 2026. “Maximum Mean Discrepancy with Unequal Sample Sizes via Generalized u-Statistics.” Transactions on Machine Learning Research. https://openreview.net/forum?id=KjXW75GHHF.
Xu, Nathaniel, Feng Liu, and Danica J. Sutherland. 2026. “Learning Representations for Independence Testing.” Transactions on Machine Learning Research. https://openreview.net/forum?id=pDvKoXRsnW.
Zhang, Kun, Jonas Peters, Dominik Janzing, and Bernhard Schölkopf. 2011. “Kernel-Based Conditional Independence Test and Application in Causal Discovery.” Uncertainty in Artificial Intelligence (UAI), 804–13. https://arxiv.org/abs/1202.3775.
Ziegel, Johanna, David Ginsbourger, and Lutz Dümbgen. 2024. “Characteristic Kernels on Hilbert Spaces, Banach Spaces, and on Sets of Measures.” Bernoulli 30 (2): 1441–57. https://doi.org/10.3150/23-BEJ1639.