Leibniz Open Science Day 2026: Scientific Rigor and Collaborative Research in the Age of AI
17 November 2026
Berlin (Leibniz-Association)

Keynote speaker: Felix Holzmeister, University of Innsbruck
Organized by ZBW – Leibniz Information Centre for Economics, DIW – The German Institute for Economic Research Berlin, WZB Berlin Social Science Center, RWI – Leibniz Institute for Economic Research, and Lab² – Metalab for Better Science, this workshop brings together scholars committed to strengthening the reliability, transparency, and cumulative nature of research in the social sciences.
Programme
8:45 - 9:00
Registration
9:00-9:30
Welcome
Marianne Saam (ZBW – Leibniz-Information Centre for Economics)
Jörg Ankel-Peters (RWI – Leibniz Institute for Economic Research)
Levent Neyse (WZB Berlin Social Science Center, DIW German Institute for Economic Research)
9:30-11:00
Session 1: Open Science – Incentives, Sharing & Publishing
- Data and Code Sharing in the Social Sciences: A Large-Scale Analysis of Trends and Journal Policies
Jan Marcus (Freie Universität Berlin, Germany)
- Incentives, Transparency, and p-Hacking – Evidence from a Diamond Open Access Reform
Oluwatobi Gbadamosi (REGE, Université Savoie Mont Blanc, France)
- Making replications count: Founding and maintaining the Replication Research Journal
Susanne Adler (Ludwig Maximilian University of Munich, Germany)
11:00-11:15
Coffee Break
11:15-12:15
Keynote by Felix Holzmeister(University of Innsbruck, Austria)
The noisy gatekeeper: Measuring the reliability of peer review and testing a way to improve it
Virtually everything science “knows” has passed through peer review. It decides what gets published and – downstream – who gets believed, funded, and hired. Yet while open science has transformed how we report research, the evaluation of research has largely escaped the same scrutiny. So how well does science’s gatekeeper actually work, and could it work better? This keynote presents two complementary attempts at an answer. The first is a diagnosis at unprecedented scale: drawing on 11.7 million reviews of 4.7 million manuscripts at more than 1,900 Elsevier journals, it asks how often reviewers of the same manuscript actually agree. The second is a treatment: a preregistered field experiment within a journal’s editorial process that redesigns what reviewers are asked to produce. Together, the studies put numbers on questions usually argued from anecdote: how much of peer review is signal and how much is lottery, what reviewer disagreement means for the false-positive and false-negative rates of editorial decisions, and which levers could make the gatekeeper better at its job. Some of the answers are reassuring; several are not.
12:15-13:15
Lunch Break
13:15-15:15
Session 2
Room 1: Replication, Reproducibility & Robustness | Room 2: Measurement, Experiments & Generalizability |
Donor Pool Selection in Synthetic Controls: Lessons from Carbon Tax Evaluations | Evaluating and Improving Experimental Designs in Economics |
Revisiting the Origins of Gender Roles: Replication, Measurement, and Robustness in Historical Persistence Research | A meta-analysis of risk elicitation tasks |
Robustness in the Deforestation Literature – An Expert-Enhanced Meta-Reproduction | The (In)Stability of Risk and Time Preferences: Measurement, Context or Genuine Change |
Code–paper discrepancies in economics and finance | Do Student Subject Pools Generalize? Testing Population Heterogeneity in Physicians’ Responses to Payment Incentives |
Same Data, Different Answers: Analytic Multiplicity as Shared Infrastructure for Collaborative Research | When Does an AI-Assisted Replication Replicate? A Multi-Model Reassessment of Corporate Accelerator Typologies |
15:15-15:30
Coffee Break
15:30-17:00
Session 3: AI in Research – Opportunities, Reliability & Epistemic Risks
- Fast, but incomplete and inaccurate? Performance of AI-powered tools for systematic literature reviews
Rene Bekkers (Vrije University Amsterdam, Netherlands)
- Heterogeneity of treatment effects across large language models: evidence from two behavioral cases
Armando Holzknecht (University of Innsbruck, Austria)
- Boosting Samples with Synthetic Respondents? Epistemic Risks of Human and LLM Data Mixtures in Social Research
Amal Labbouz (Karlsruhe Institute of Technology (KIT), Germany)
17:00-17:10
Closing
Marianne Saam (ZBW – Leibniz-Information Centre for Economics)
Jörg Ankel-Peters (RWI – Leibniz Institute for Economic Research)
Levent Neyse (WZB Berlin Social Science Center)
Organizing committee:
- Marianne Saam, ZBW – Leibniz Information Centre for Economics and University of Hamburg
- Doreen Siegfried, ZBW – Leibniz Information Centre for Economics
- Jörg Ankel-Peters, RWI – Leibniz Institute for Economic Research
- Macartan Humphreys, WZB Berlin Social Science Center
- Levent Neyse, DIW – The German Institute for Economic Research Berlin and WZB Berlin Social Science Center
Kontakt
Partner:






