Statistics and Econometrics Seminars
Joint organization by ORSTAT, Faculty of Economics and Business and the Statistics Research Group, Faculty of Science; Leuven Statistics Research Center.
Practical information:
Programma
17 February 2022 Eugen Pircalabelu (UCLouvain) "Unbalanced distributed estimation and inference for (covariate-adjusted) Gaussian graphical models”
A distributed estimation and statistical inference framework is introduced for the
sparse precision matrix in the (covariate-adjusted) Gaussian graphical models under the un-
balanced splitting setting. This type of splitting arises when the datasets from different sources
cannot be aggregated on one single machine or when the available machines are of different
powers. A de-biased estimator of the precision matrix on every single machine is proposed,
and theoretical guarantees are provided. Moreover, a new de-biased estimator that is pooled
across the machines using a composite likelihood approach is proposed. It is shown to enjoy
consistency and asymptotic normality, and we provide statistical inference strategies based on
it. The performance of this estimator is investigated via simulation studies and real data exam-
ples. It is shown that the performance of this estimator is close to the non-distributed estimator,
which uses the entire dataset.
- 12:00–1:00 pm
- on-campus (KU Leuven Faculty of Economics and Business HOGM 01.85) or online via Zoom (https://us02web.zoom.us/j/88418552909?pwd=MERhSDAvSVdET0xKWnpFOUxHOVJMUT09)
24 February 2022 Elia Lapenta (ENSAE Paris) “Nonparametric instrumental variable estimation without smoothing on the instruments”
We propose a new estimation method for nonparametric models identified by Instru- mental Variables (IVs). The estimation is based on a class of Generically Comprehensively Revealing functions. Compared to methods available in the literature that smooth on the IVs, our estimation does not smooth on the instruments and thus requires the selection of less tun- ing parameters. Furthermore, it does not suffer from a curse of dimensionality on the IVs. We show that our procedure is equivalent to a classical estimation method that smooths on the Ivs but keeps the bandwidth fixed as the sample size increases.
We then extend our methodology to estimate a partly linear model with endogenous variables. We obtain convergence rates for the estimator of the nonparametric part of the model and the asymptotic normality of the estimator of the parametric components. To deal with the ill- posedness of the inverse problem, we use a Landweber-Friedman regularization. This has the advantage of being a simple iterative method that avoids the inversion of a large matrix whose dimension increases with the sample size. We finally study the implementation of our proce- dure and propose a data-driven selection of the regularization and the smoothing parameters. (Joint work with Jean-Pierre Florens.)
- 12:00–1:00 pm
- on-campus (KU Leuven room HOGM 00.85) or online
3 March 2022 Robin Fuchs (Institut de Mathématiques de Marseille) “Mixed deep Gaussian mixture model: a clustering model for mixed datasets”
Clustering mixed data presents numerous challenges inherent to the very heteroge- neous nature of the variables. A clustering algorithm should be able, despite of this heterogene- ity, to extract discriminant pieces of information from the variables in order to design groups. In this work we introduce a multilayer architecture model-based clustering method called Mixed Deep Gaussian Mixture Model (MDGMM) that can be viewed as an automatic way to merge the clustering performed separately on continuous and non-continuous data. This architecture is flexible and can be adapted to mixed as well as to continuous or non-continuous data. In this sense we generalize Generalized Linear Latent Variable Models and Deep Gaussian Mixture Models. We also design a new initialisation strategy and a data driven method that selects the best specification of the model and the optimal number of clusters for a given dataset “on the fly”. Besides, our model provides continuous low-dimensional representations of the data which can be a useful tool to visualize mixed datasets. Finally, we validate the performance of our approach comparing its results with state-of-the-art mixed data clustering models over several commonly used datasets.
- 12:00–1:00 pm
- online
10 March 2022 Cavit Pakel (Bilkent University) “Bounds on average effects in discrete choice panel data models”
Average effects in discrete choice panel data models with individual-specific fixed effects are generally only partially identified in short panels. While consistent estimation of the identified set is possible, it generally requires very large sample sizes, especially when the number of support points of the observed covariates is large, such as when the covariates are continuous. In this paper, we propose estimating outer bounds on the identified set of average effects. Our bounds are easy to construct, converge at the parametric rate, and are computationally simple to obtain even in moderately large samples, independent of whether the covariates are discrete or continuous. We also provide asymptotically valid confidence intervals on the identified set. Simulation studies confirm that our approach works well and is informative in finite samples. We also consider an application to labor force participation. (Joint work with Martin Weidner)
- 12:00–1:00 pm
- online
17 March 2022 Eric Ghysels (University of North Carolina at Chapel Hill) “Ambiguity with machine learning: an application to portfolio choice”
To characterize ambiguity we use machine learning to impose guidance and dis- cipline on the formulation of expectations in a data-rich environment. In addition, we use the bootstrap to generate plausible synthetic samples of data not seen in historical real data to create statistics of interest pertaining to uncertainty. While our approach is generic we focus on robust portfolio allocation problems as an application and study the impact of risk versus uncer- tainty in a dynamic mean-variance setting. We show that a mean-variance optimizing investor achieves economically meaningful wealth gains (33%) across our sample from 1996-2019 by internalizing our uncertainty measure during portfolio formation. (Joint work with Yan Qian and Steve Raymond)
- 12:00–1:00 pm (possibly start at 11:45 am, TBD)
- on campus
24 March 2022 Snigdah Panigrahi (University Michigan) “Approximate selective inference via maximum likelihood”
Several strategies have been developed recently to ensure valid inferences after model selection; some of these are easy to compute, while others fare better in terms of inferen- tial power. In this talk, we will address post-selection inference through approximate maximum likelihood estimation. Our goal is to: (i) efficiently utilize hold-out information from selection with the aid of randomization, (ii) bypass expensive MCMC sampling from exact conditional distri- butions that are hard to evaluate in closed forms. At the core of our new method is the solution to a convex optimization problem which assumes a separable form across multiple learning queries during selection. We illustrate the potential of our method across wide-ranging values of signal-to-noise ratio in simulated experiments.
- 12:00–1:00 pm
- online
31 March 2022 Limin Peng (Emory University) “Heterogeneous individual risk modeling of recurrent events”
Progression of chronic disease is often manifested by repeated occurrences of disease-related events over time. Delineating the heterogeneity in the risk of such recurrent events can provide valuable scientific insight for guiding customized disease management. In this paper, we propose a new sensible measure of individual risk of recurrent events and present a dynamic modeling framework thereof, which accounts for both observed covariates and unobservable frailty. The proposed modeling requires no distributional specification of the unobservable frailty, while permitting the exploration of dynamic effects of the observed covari- ates. We develop estimation and inference procedures for the proposed model through a novel adaptation of the principle of conditional score. The asymptotic properties of the proposed es- timator, including the uniform consistency and weak convergence, are established. Extensive simulation studies demonstrate satisfactory finite-sample performance of the proposed method. We illustrate the practical utility of the new method via an application to a diabetes clinical trial that explores the risk patterns of hypoglycemia in Type 2 diabetes patients.
- 5:00–6:00 pm
- online
21 April 2022 Angelo Guevara (Universidad de Chile) “Endogeneity in discrete choice models”
Endogeneity is the most severe failure that a discrete choice model can face in its
objective of making causal inference or forecasting. It occurs when the explanatory variables
are not independent of the error term and results in inconsistent estimators of the model pa-
rameters. This seminar summarizes the causes, impact, and state of the art methods to detect
and to address endogeneity in discrete choice models, which are related, but differ in relevant
aspects, from those of linear model. The emphasis is put on providing the theoretical funda-
mentals and the intuition behind the problems that may arise, as well as practical details as
on how to apply the methods to address them, stressing the main considerations that must be
taken in this endeavor. The main concepts are formally stated, and references are given to
specific recent articles for further details that may be required for advanced applications or a
deeper understanding. The seminar begins providing details on the definition, impact, causes
and examples of endogeneity in discrete choice models, followed by a revision of possible
approaches to detect this problem. Then, a deep review of the fundamentals, intuition and
practicalities of the various methods that can be used to address endogeneity is presented,
followed by a section providing insights and methods related to the obtention and validation of
instrumental variables, which are key in this effort. The seminar continues then revising the
problem of forecasting with discrete choice models that have been corrected for endogeneity,
finalizing with a summary of the main conclusions and takeaways.
- 5:00–6:00 pm
- online
28 April 2022 Cécile Adam (KU Leuven) “Local linear functional mean regression under censoring”
Among the main interests in regression analysis is to explore the influence that a functional has on a real-valued variable of interest, the response. There is some literature on flexible mean regression, in which the targeted quantity is the conditional mean of the response given the functional. Another interest in statistics is to study the regression under the condition of censoring in which the values of the response are only partially known. After an introduc- tion to mean regression function estimation in which the response variable is subject to right random censorship, different statistical methodology are presented. The finite-sample perfor- mance of the estimators is investigated via a simulation study. (Joint work with I. Gijbels and G. Claeskens.)
- 12:00–1:00 pm
- on campus
5 May 2022 Wiktor Budzinski (University of Warsaw) “Hybrid choice models vs. endogeneity of indicator variables: a Monte Carlo investigation”
We investigate the problem of endogeneity in the context of hybrid choice (integrated choice and latent variable) models. We first provide a thorough analysis of potential causes of endogeneity and propose a working taxonomy. We demonstrate that although it is widely believed that the hybrid choice framework is devoid of the endogeneity problem, there is no theoretical reason to expect that this is the case. We then demonstrate empirically that the problem exists in the hybrid choice framework too. By conducting a Monte Carlo experiment, we display the extent of the bias resulting from measurement and endogeneity biases. Finally, we propose two novel solutions to address the problem: by explicitly accounting for correlation between structural and discrete choice component error terms (or with random parameters in a utility function), or by introducing additional latent variables. Using simulated data, we demonstrate that these approaches work as expected, that is, they result in unbiased estimates of all model parameters.
- 12:00–1:00 pm
- online
12 May 2022 Wendun Wang (Erasmus Universiteit Rotterdam) “Recovering spillover structures with structural breaks in panel data models”
This paper aims at capturing time-varying spillover effects in a panel data setting. We consider panel models where the outcome of a unit not only depends on own characteristics and also the characteristics of other units (spillover effects). The effect of own characteristics can be unit-specific or homogeneous (common effects). We allow the spillover structure, i.e., which units interact with which, to be latent, and the structure and effect of spillovers may both vary over time. We model time-varying spillovers via structural breaks with unknown break points. To estimate the break points, spillover and common effects, we solve a penalized least squares optimization and employ double machine learning procedures to improve the conver- gence and inference. We establish the super consistency of the estimated break point and provide the convergence rate of estimated spillover and common effects. We illustrate the theory via simulated and empirical data.
- 12:00–1:00 pm
- on-campus
19 May 2022 Thierry Magnac (Toulouse School of Economics) “Linear models with interval-censored explanatory variables”
This paper studies the problem of inference in a set identified problem defined by linear moment restrictions with interval-censored variables. It introduces a novel and tractable inference procedure based on characterizing the identified set of the coefficient of interest in terms of the solution to convex optimization problems with equality constraints. Monte Carlo experiments evaluate the numerical performance of the novel inference procedure and compare it with existing ones.
- 12:00–1:00 pm
- on-campus
Teachers / speakers
Eugen Pircalabelu
I am a Lecturer (Chargé de cours) at UCLouvain working at the Institute of Statistics, Biostatistics and Actuarial Sciences within LIDAM at the Faculty of Science.
Prior to moving at UCLouvain, I held a Visiting professor position at Ghent University affiliated with the Department of Applied Mathematics, Computer Science and Statistics within the Faculty of Sciences and a Postdoctoral position at KU Leuven affiliated with the ORSTAT department within the Faculty of Economics and Business.
My research focuses on: Models for high-dimensional data, Probabilistic graphical models, Social network models, Copula models and Information criteria.
Elia Lapenta
Elia Lapenta received his PhD in economics from the Toulouse School of Economics in 2020 and joined ENSAE-CREST as an Assistant Professor in September 2020. His research interests focus on econometrics, with emphasis on hypothesis testing, nonparametric instrumental variables estimation, and empirical games with incomplete information.
Cavit Pakel
I am a Visiting Assistant Professor at University of Oxford Department of Economics, and also affiliated with St Antony’s College (currently on leave from Bilkent University).
I have obtained my DPhil in 2012 from the University of Oxford. In the 2016-17 academic year I was a Visiting Assistant Professor at Princeton University, Department of Economics.
My research interests lie in econometric theory, with a focus on nonlinear panel models with unobserved heterogeneity, identification of average effects, and incidental parameter bias.
Eric Ghysels
Eric Ghysels is Professor Emeritus of Economics at the University of North Carolina at Chapel Hill, Professor of Finance at Bernstein's Kenan-Flagler Business School and has been a visiting professor or researcher at several leading American, European and Asian universities.
His main areas of research are econometrics and time series finance.
Snigdah Panigrahi
I am an Assistant Professor of Statistics and also hold a courtesy appointment at the Department of Biostatistics. My work involves developing tools for replicable learning from complex data.
My current research, supported by both the National Science Foundation and the National Institutes of Health, focuses on quantifying uncertainties in outputs from machine learning methods. My work was recognized by the CAREER Award for early-career faculty by the National Science Foundation in 2023.
I have been serving as a member of the editorial board for the Journal of Computational and Graphical Statistics and as an elected member of the International Statistical Institute since 2021.
Limin Peng
I joined the Biostatistics faculty at Emory in 2005 as Rollins Assistant Professor. My methodological research has been mainly focused in survival analysis, dynamic regression, and nonparametric and semiparametric inference with my recent work tackling their overlalps with causal inference, latent class analysis, and high-dimensional statistical learning. I have also worked on developing new statistical methods driven by scientific investigations in Cystic Fibrosis, Mental Health, Environmental Health and Neurology. My methodological research programs have been supported by both NSF and NIH grants.
My collaborative research spans over areas including Diabetes, Cystic Fibrosis, Neorology, and other chronic diseases.
Angelo Guevara
Angelo Guevara is an experienced Systems Architect with 30+ years in Systems Engineering. His primary focus is on using architecture models to convey solutions, enabling a deep understanding of complex concepts and expediting digital transformation for organizations. In addition, he applies emerging technologies to automate and streamline processes, enhancing efficiency and saving costs.
Cécile Adam
Credit Risk Modeller at BNP Paribas Fortis
Wiktor Budzinski
I am a researcher focusing on modelling of consumers’ preferences, with applications in environmental, cultural and health economics. I graduated with an MA degree in Econometrics and computer science, as well as a BA degree in Mathematics. Both of them were granted by the University of Warsaw. My PhD thesis is focused on extending discrete choice modelling framework to better capture consumers’ preference heterogeneity. This involves working with spatial econometric tools in order to analyze how preferences change in space and developing new models that would allow for incorporating spatial dimension into discrete choice models. Furthermore, my work is concerned with the effect of attitudes, perceptions and other psychological factors on consumer’s choices. I worked extensively with the so-called hybrid choice models, which were developed to tackle such issues. I also have experience with modeling demand using revealed preference data. I teach Microeconometrics and Choice modelling courses at the University of Warsaw.
Wendun Wang
Wendun Wang is an associate professor at Econometric Institute, Erasmus University Rotterdam
Thierry Magnac
Thierry Magnac has been a Professor of Economics at the University of Toulouse since 2005. His research agenda is set on empirical issues in the economics of education, labor and housing, and on methodological issues in microeconometrics. He is a fellow of the Econometric Society, and of the International Association for Applied Econometrics. He received an Advanced Grant from the European Research Council in 2012 regarding the structural econometric modelling of the dynamics of household behavior. He is one of the coeditors of the Journal of Applied Econometrics and he is a Member of the Board of Directors of the International Association for Applied Econometrics.
Robin Fuchs
institut de mathématiques de marseille
Related courses
SAIAR Summer Studio: From latent to physical space
Zomerschool - Kortrijk - Howest Hogeschool, SAIARlab