Pieter Delobelle

Building and studying large language models. Pretraining, data, tokenization & AI safety.

dr. ing. ยท Lead AI Scientist at Pleias ยท Guest professor at KU Leuven

Pieter Delobelle at Station F, Paris

What I work on

At Pleias, I work on LLM pretraining and synthetic data — most recently releasing Nemotron-Personas-Belgium together with NVIDIA. At KU Leuven, my research spans tokenization, fairness, and Dutch language models.

I created RobBERT, the Dutch language model family in the top 80 on Hugging Face, and have collaborated on most of the open Dutch LLMs — from Tweety to ChocoLlama. On the safety side, I work on measuring bias, reducing toxicity (with Apple, at ICML), and questioning who decides what “fair” means. I also work on tokenization and developed trans-tokenization, a method to translate LLMs from one language to another.

My work has been published at top AI venues including ICML, NAACL, and EMNLP, including work done at Apple and Aleph Alpha. I've done research visits at MilaNLP (Bocconi), HU Berlin, and the Weizenbaum Institute, and accumulated 1,000+ citations across 32 publications. My research has been covered by WIRED, MIT Technology Review, De Tijd, and VTM Nieuws.

At KU Leuven, I was a member of the GenAI advisory board and sit on the council for research policy. I also contributed as an NLP expert to a EU AI Office workshop on the code of practice for general-purpose AI.

I also frequently give talks on AI and language models. I've spoken at companies like Apple, KBC, VRT, TechWolf, ML6, and Superlinear, and at several events, for instance at Wintercircus. Topics range from LLM pretraining and technical deep dives into LLM inference to AI safety and introductory talks about the Dutch NLP ecosystem.

News

September 05, 2026 I start this semester as guest professor at KU Leuven, teaching advanced NLP in the Master of AI ๐ŸŽ“. September 04, 2026 I was interviewed for De Standaard about OpenAI's GPT-6 and its AGI claims. August 28, 2026 I was interviewed for De Morgen about bias and stereotypes in AI assistants. June 18, 2026 We released Nemotron-Personas-Belgium with NVIDIA. Pleias announcement ๐Ÿ‡ง๐Ÿ‡ช. June 04, 2026 I joined a panel on "AI is changing our research" at SMiLee 2026, our DTAI research workshop. June 04, 2026 Our paper on query-efficient fairness auditing of black-box LLMs got accepted at ACL 2026 Findings ๐ŸŽ‰. May 26, 2026 I gave a talk on what LLMs can and cannot do at Bonus advocaten. May 22, 2026 I led a workshop on synthetic data at the OSFM workshop near Cologne. May 19, 2026 Our preprint on the implications of sunsetting Perspective API was featured in Tagesspiegel Background. May 06, 2026 I gave a lecture on fairness in LLMs at TU Berlin. April 28, 2026 New preprint on arXiv on the implications of Google sunsetting Perspective API . April 23, 2026 I gave a half-day lecture on safety and fairness in LLMs. April 22, 2026 I gave a talk on Belgian LLMs at the Jura Falconis studiedag.
All talks & appearances