Publications
100+ papers and book chapters. See also Google Scholar and Semantic Scholar.
Hyperparameter Optimization in Machine Learning
Foundations and Trends in Machine Learning, 18(6):1054-1201, 2025.
Explaining Probabilistic Models with Distributional Values
International Conference on Machine Learning (ICML), 2024.
Fortuna: A Library for Uncertainty Quantification in Deep Learning
Journal of Machine Learning Research (JMLR), Open Source Software Track, 238:1-7, 2024.
On the Choice of Learning Rate for Local SGD
Transactions on Machine Learning Research (TMLR), 2024.
Structural Pruning of Pre-trained Language Models via Neural Architecture Search
Transactions on Machine Learning Research (TMLR), 2024.
A Negative Result on Gradient Matching for Selective Backprop
NeurIPS workshop on Failure Modes in the Age of Foundation Models, 2023.
Explaining Multiclass Classifiers with Categorical Values: A Case Study in Radiography
International Workshop on Trustworthy Machine Learning for Healthcare (TML4H) at ICLR, 2023.
Geographical Erasure in Language Generation
Findings of the Association for Computational Linguistics: EMNLP 2023.
Optimizing Hyperparameters with Conformal Quantile Regression
International Conference on Machine Learning (ICML), 2023.
PASHA: Efficient HPO and NAS with Progressive Resource Allocation
International Conference on Learning Representations (ICLR), 2023.
Renate: A Library for Real-world Continual Learning
arXiv preprint, 2023.
Automatic Termination for Hyperparameter Optimization
Conference on Automated Machine Learning (Main Track), 2022. (best paper award)
Continual Learning with Transformers for Image Classification
CVPR workshop on Continual Learning in Computer Vision, 2022.
Differentially private gradient boosting on linear learners for tabular data analysis
NeurIPS workshop on Trustworthy and Socially Responsible Machine Learning, 2022.
Gradient-Matching Coresets for Rehearsal-Based Continual Learning
arXiv preprint, 2022.
Hyperparameter Optimization
In Dive Into Deep Learning, vol. 2 (Chapter 19), 2022.
Memory Efficient Continual Learning with Transformers
Annual Conference on Advances in Neural Information Processing Systems (NeurIPS), 2022.
Memory-efficient Continual Learning for Neural Text Classification
arXiv preprint, 2022.
PASHA: Efficient HPO with Progressive Resource Allocation
Conference on Automated Machine Learning (Late-Breaking Workshop Track), 2022.
Private Synthetic Data for Multitask Learning and Marginal Queries
Annual Conference on Advances in Neural Information Processing Systems (NeurIPS), 2022.
Syne Tune: A Library for Large Scale Hyperparameter Tuning and Reproducible Research
Conference on Automated Machine Learning (Main Track), 2022.
Uncertainty Calibration in Bayesian Neural Networks via Distance-Aware Priors
arXiv preprint, 2022.
A Multi-objective Perspective on Jointly Tuning Hardware and Hyperparameters
ICLR NAS workshop, 2021.
A Resource-efficient Method for Repeated HPO and NAS Problems
ICML AutoML workshop, 2021.
Amazon SageMaker Automatic Model Tuning: Black-box Optimization at Scale
ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2021. Industry track.
BORE: Bayesian Optimization by Density-Ratio Estimation
International Conference on Machine Learning (ICML), 2021. (long presentation)
Diverse Counterfactual Explanations for Anomaly Detection in Time Series
arXiv preprint, 2021.
Dynamic Pruning of a Neural Network via Gradient Signal-to-Noise Ratio
ICML AutoML workshop, 2021.
Fair Bayesian Optimization
AAAI/ACM Conference on Artificial Intelligence, Ethics, and Society (AIES), 2021.
Gradient-matching Coresets for Continual Learning
NeurIPS workshop on Distribution Shifts: Connecting Methods and Applications, 2021.
Hyperparameter Transfer Learning with Adaptive Complexity
International Conference on Artificial Intelligence and Statistics (AISTATS), 2021.
Meta-Forecasting by Combining Global Deep Representations with Local Adaptation
arXiv preprint, 2021.
Multi-objective Asynchronous Successive Halving
arXiv preprint, 2021.
On the Lack of Robustness of Deep Neural Text Classifiers
Annual Meeting of the Association for Computational Linguistics (ACL), 2021. Findings.
Overfitting in Bayesian Optimization: an empirical study and early-stopping solution
ICLR NAS workshop, 2021.
Towards Robust Episodic Meta-Learning
Uncertainty in Artificial Intelligence (UAI), 2021.
Bayesian Optimization by Density Ratio Estimation
NeurIPS workshop on Meta-learning, December 2020. (selected for oral presentation)
Bayesian Optimization with Fairness Constraints
ICML workshop on AutoML, 2020. (best paper award)
Cost-aware Bayesian Optimization
ICML workshop on AutoML, 2020.
LEEP: A New Measure to Evaluate Transferability of Learned Representations
International Conference on Machine Learning (ICML), 2020.
Model-based Asynchronous Hyperparameter and Neural Architecture Search
arXiv preprint, 2020.
Multi-Objective Multi-Fidelity Hyperparameter Optimization with application to Fairness
NeurIPS workshop on Meta-learning, 2020.
Pareto-efficient Acquisition Functions for Cost-Aware Bayesian Optimization
NeurIPS workshop on Meta-learning, 2020.
Constrained Bayesian Optimization with Max-Value Entropy Search
NeurIPS workshop on Meta-learning, 2019.
Learning Search Spaces for Bayesian Optimization: Another View of Hyperparameter Transfer Learning
Annual Conference on Advances in Neural Information Processing Systems (NeurIPS), 2019.
A Simple Transfer Learning Extension of Hyperband
NeurIPS workshop on Meta-Learning, 2018.
Scalable Hyperparameter Transfer Learning
Annual Conference on Advances in Neural Information Processing Systems (NeurIPS), 2018.
An interpretable latent variable model for attribute applicability in the Amazon catalogue
NeurIPS Symposium on Interpretable Machine Learning, 2017.
Bayesian Optimization with Tree-structured Dependencies
International Conference on Machine Learning (ICML), 2017.
Multiple Adaptive Bayesian Linear Regression for Scalable Bayesian Optimization with Warm Start
NeurIPS workshop on Meta-Learning, 2017.
Adaptive Algorithms for Online Convex Optimization with Long-term Constraints
International Conference on Machine Learning (ICML), 2016.
Online Dual Decomposition for Performance and Delivery-based Distributed Ad Allocation
Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), pp. 117-126, 2016.
Online Optimization and Regret Guarantees for Non-additive Long-term Constraints
arXiv preprint, 2016.
Incremental Variational Inference applied to Latent Dirichlet Allocation
NeurIPS workshop on Advances in Approximate Bayesian Inference, 2015.
Incremental Variational Inference for Latent Dirichlet Allocation
arXiv preprint, 2015.
Latent IBP compound Dirichlet Allocation
IEEE transactions in Pattern Analysis and Machine Intelligence (PAMI) 37(2):321-333, 2015.
One-Pass Ranking Models for Low-Latency Product Recommendations
Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), pp. 1789-1798, 2015.
Online Inference for Relation Extraction with a Reduced Feature Set
arXiv preprint, 2015.
Overlapping Trace Norms in Multi-View Learning
arXiv preprint, April 2014.
Towards Crowd-based Customer Service: A Mixed-Initiative Tool for Managing Q&A Sites
Proceedings of the ACM Conference on Human Factors in Computing Systems (CHI), pp. 2725-2734, 2014.
Bringing Representativeness into Social Media Monitoring and Analysis
Proceedings of the 46th Hawaii International Conference on System Sciences (HICSS), pp. 2003-2012, 2013.
Connecting Comments and Tags: Improved Modeling of Social Tagging Systems
In S. Leonardi, A. Panconesi, P. Ferragina, A. Gionis (Eds.), 6th ACM Conference on Web Search and Data Mining (WSDM), pp. 547-556, 2013.
Error Prediction with Partial Feedback
In H. Blockeel, K. Kersting, S. Nijssen, F. Zelezny (Eds.), European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML-PKDD), Lecture Notes in Computer Science (LNCS), 8189:80-94, 2013.
Log-linear Language Models based on Structured Sparsity
Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 233-243, 2013.
Plackett-Luce regression: a new Bayesian model for polychotomous data
In N. de Freitas, K. P. Murphy (Eds.), Uncertainty in Artificial Intelligence (UAI) 28, pp. 84-92, 2012.
Variational Markov chain Monte Carlo for Bayesian smoothing of non-linear diffusions
Computational Statistics 27:1, 149-176, 2012.
Approximate Inference for continuous-time Markov processes
In D. Barber, A. T. Cemgil, and S. Chiappa, Inference and Learning in Dynamic Models, 2011.
Latent IBP compound Dirichlet allocation
NeurIPS 24 workshop on Bayesian nonparametrics: Hope or Hype?, 2011.
Mail2Wiki: low-cost sharing and early curation from email to wikis
In M. Foth, J. Kjeldskov, J. Paay (Eds.), Proceedings of the International Conference on Communities and Technologies (C&T) 5, pp. 98-107, 2011.
Mail2Wiki: posting and curating Wiki content from email
In P. Pu, M. J. Pazzani, E. Andre, D. Riecken (Eds.), Proceedings of the International Conference on Intelligent User Interfaces (IUI), pp 441-442, 2011.
Robust Bayesian Matrix Factorisation
Artificial Intelligence and Statistics (AISTATS) 14. JMLR Workshop and Conference Proceedings 15:425-433, 2011.
Sparse Bayesian multi-task learning
In J. Shawe-Taylor, R. S. Zemel, P. L. Bartlett, F. C. N. Pereira, K. Q. Weinberger (Eds.), Neural Information Processing Systems (NeurIPS) 24, pp. 1755-1763, 2011.
The Sequence Memoizer
Communications of the ACM, 54(2):91-98, 2011.
A Comparison of Variational and Markov Chain Monte Carlo Methods for Inference in Partially Observed Stochastic Dynamic Systems
Journal of Signal Processing Systems, 61(1):51-59, 2010.
Multiple Gaussian process models
NeurIPS 23 workshop on New Directions in Multiple Kernel Learning, 2010.
Prediction of hot spot residues at protein-protein interfaces by combining machine learning and energy-based methods
BMC Bioinformatics, 10: 365-382, 2009.
Sparse Probabilistic Projections
In D. Koller, D. Schuurmans, Y. Bengio, and L. Bottou (Eds.), Neural Information Processing Systems (NeurIPS) 21, pp.17-24, 2009.
Stochastic Memoizer for Sequence Data
In L. Bottou and M. Littman, Proceedings of the 26th International Conference on Machine Learning (ICML), Montreal (Quebec), Canada, June 14-18, 2009, pp. 1129-1136.
Switching Regulatory Models of Cellular Stress Response
Bioinformatics, 25(10): 1280-1286, 2009.
The Variational Gaussian Approximation Revisited
Neural Computation 21(3):786-792, 2009.
Improving the robustness to outliers of mixtures of probabilistic PCAs
In T. Washio, et al. (Eds.), Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD) 12, Lecture notes in Artificial Intelligence (LNAI) 5012:527-535, 2008.
Mixtures of Robust Probabilistic Principal Component Analyzers
Neurocomputing, 71(7-9):1274-1282, 2008.
Using Subspace-Based Template Attacks to Compare and Combine Power and Electromagnetic Information Leakages
In E. Oswald and P. Rohatgi (Eds.), 10th International Workshop on Cryptographic Hardware and Embedded Systems (CHES), Washington, DC, USA, 10-13 August, 2008. Lecture Notes in Computer Science vol. 5154, pp. 411-425.
Variational Inference for Diffusion Processes
In C. Platt, D. Koller, Y. Singer and S. Roweis (Eds.), Neural Information Processing Systems (NeurIPS) 20, pp.17-24, 2008.
Evaluation of Variational and Markov Chain Monte Carlo Methods for Inference in Partially Observed Stochastic Dynamic Systems
Proceedings of the 17th IEEE workshop on Machine Learning for Signal Processing (MLSP), Thessaloniki, Greece, 27-28 August, 2007, pp. 306-311.
Gaussian Process Approximations of Stochastic Differential Equations
Journal of Machine Learning Research Workshop and Conference Proceedings, 1:1-16, 2007.
Mixtures of Robust Probabilistic Principal Component Analyzers
Proceedings of the 15th European Symposium on Artificial Neural Networks (ESANN), Bruges, Belgium, April 25-27, 2007, pp. 229-234.
Robust Bayesian Clustering
Neural Networks, 20:129-138, 2007.
Automatic Adjustment of Discriminant Adaptive Nearest Neighbor
In Y.Y. Tang, P. Wang, G. Lorette and D.S. Yeung (Eds.), Proceedings of the 18th International Conference on Pattern Recognition (ICPR), Hong Kong, P.R.C., 20-24 August, 2006, vol. 2, pp. 525-555.
Robust Probabilistic Projections
In W. W. Cohen and A. Moore (Eds.), Proceedings of the 23rd International Conference on Machine Learning (ICML), Pittsburgh (PA), U.S.A., 25-29 June, 2006, pp. 33-40.
Template Attacks in Principal Subspaces
In L. Goubin and M. Matsui (Eds.), 8th International Workshop on Cryptographic Hardware and Embedded Systems (CHES), Yokohama, Japan, 10-13 October, 2006. Lecture Notes in Computer Science vol. 4249, pp. 1-14.
Towards Security Limits of Side-Channel Attacks
In L. Goubin and M. Matsui (Eds.), 8th International Workshop on Cryptographic Hardware and Embedded Systems (CHES), Yokohama, Japan, 10-13 October, 2006. Lecture Notes in Computer Science vol. 4249, pp. 30-45.
Probabilistic Models in Noisy Environments — And their Application to a Visual Prosthesis for the Blind
Doctoral dissertation, Université catholique de Louvain, Louvain-la-Neuve, Belgium, September 2005.
Local Vector-based Models for Sense Discrimination
In H. Bunt, J. Geertzen and E. Thijsse (Eds.), Proceedings of the 6th International Workshop on Computational Semantics (IWCS), Tilburg, the Netherlands, January 12-14, 2005, pp. 163-174.
Manifold Constrained Finite Gaussian Mixtures
In J. Cabestany, A. Prieto and F. Sandoval Hernández (Eds.), Computational Intelligence and Bioinspired Systems - 8th International Work-Conference on Artificial Neural Networks (IWANN), Vilanova i la Geltrú (Barcelona), Spain, June 8-10, 2005. Lecture Notes in Computer Science, vol. 3512, pp.820-828.
Entropy Minima and Distribution Structural Modifications in Blind Separation of Multi-model Sources
In R. Fisher, R. Preuss and U. von Toussaint, Proceedings of the 24th International Workshop on Bayesian Inference and Maximum Entropy Methods in Science and Engineering (MaxEnt), IPP Garching bei München, Germany, July 25-30, 2004, pp. 589-596.
Flexible and Robust Bayesian Classification by Finite Mixture Models
Proceedings of the 12th European Symposium on Artificial Neural Networks (ESANN), Bruges, Belgium, April 28-30, 2004, pp. 75-80.
Prediction of Visual Perceptions with Artificial Neural Networks in a Visual Prosthesis for the Blind
Artificial Intelligence in Medicine, 32(3):183-194, 2004.
Supervised Nonparametric Information Theoretic Classification
In J. Kittler, M. Petrou and M. Nixon (Eds.), Proceedings of the 17th International Conference on Pattern Recognition (ICPR), Cambridge, U.K., August 23-26, 2004, vol. 3, pp. 414-417.
Towards a Local Separation Performances Estimator using Common ICA contrast Functions?
Proceedings of the 12th European Symposium on Artificial Neural Networks (ESANN), Bruges, Belgium, April 28-30, 2004, pp. 211-216.
Classification of Visual Sensations Generated Electrically in the Visual Field of the Blind
In D. D. Feng and E. R. Carson (Eds.), Proceedings of the 5th IFAC Symposium on Modelling and Control in Biomedical Systems, Melbourne, Australia, August 21-23, 2003, pp. 223-228.
Locally Linear Embedding versus Isotop
Proceedings of the 11th European Symposium on Artificial Neural Networks (ESANN), Bruges, Belgium, April 23-25, 2003, pp. 527-534.
On Convergence Problems of the EM Algorithm for Finite Gaussian Mixtures
Proceedings of the 11th European Symposium on Artificial Neural Networks (ESANN), Bruges, Belgium, April 23-25, 2003, pp. 99-106.
Width Optimization of the Gaussian Kernels in Radial Basis Function Networks
Proceedings of the 10th European Symposium on Artificial Neural Networks (ESANN), Bruges, Belgium, April 24-26, 2002, pp. 425-432.
Phosphene Evaluation in a Visual Prosthesis with Artificial Neural Networks
Proceedings of the 1st European Symposium on Intelligent Technologies, Hybrid Systems and their implementation on Smart Adaptive Systems (EUNITE), Puerto de la Cruz (Tenerife), Spain, December 13-14, 2001, pp. 509-515.