TSLiNGAM: DirectLiNGAM Under Heavy Tails
收藏资源简介:
One of the established approaches to causal discovery consists of combining directed acyclic graphs (DAGs) with structural causal models (SCMs) to describe the functional dependencies of effects on their causes. Possible identifiability of SCMs given data depends on assumptions made on the noise variables and the functional classes in the SCM. For instance, in the LiNGAM model, the functional class is restricted to linear functions and the disturbances have to be non-Gaussian. In this work, we propose TSLiNGAM, a new method for identifying the DAG of a causal model based on observational data. TSLiNGAM builds on DirectLiNGAM, a popular algorithm which uses simple OLS regression for identifying causal directions between variables. TSLiNGAM leverages the non-Gaussianity assumption of the error terms in the LiNGAM model to obtain more efficient and robust estimation of the causal structure. TSLiNGAM is justified theoretically and is studied empirically in an extensive simulation study. It performs significantly better on heavy-tailed and skewed data and demonstrates a high small-sample efficiency. In addition, TSLiNGAM also shows better robustness properties as it is more resilient to contamination. Supplementary materials for this article are available online.
因果发现领域的成熟方法之一,是将有向无环图(DAGs)与结构因果模型(SCMs)相结合,用以描述结果变量对其原因变量的函数依赖关系。在给定观测数据的情况下,结构因果模型(SCMs)的可识别性,取决于针对模型中噪声变量与函数类所设定的假设条件。例如在线性非高斯无环模型(LiNGAM)中,函数类被限定为线性函数,且扰动项需满足非高斯分布。本研究提出TSLiNGAM方法,这是一种基于观测数据识别因果模型有向无环图(DAGs)的全新方案。TSLiNGAM以DirectLiNGAM为基础——后者是一种常用算法,通过简单的普通最小二乘(ordinary least squares, OLS)回归来识别变量间的因果方向。TSLiNGAM借助LiNGAM模型中误差项的非高斯性假设,实现了对因果结构更高效且稳健的估计。TSLiNGAM具备严谨的理论依据,并通过大规模仿真实验完成了实证验证。该方法在重尾分布与偏态分布数据上的表现显著更优,且展现出优异的小样本效率。此外,TSLiNGAM还具备更出色的稳健性,对数据污染具有更强的抵抗能力。本文的补充材料可在线获取。




