Leibniz Open Science Day 2026: Scientific Rigor and Collaborative Research in the Age of AI

17 November 2026
Berlin (Leibniz-Association)

Keynote speaker: Felix Holzmeister, University of Innsbruck

Organized by ZBW – Leibniz Information Centre for Economics, DIW – The German Institute for Economic Research Berlin, WZB Berlin Social Science Center, RWI – Leibniz Institute for Economic Research, and Lab² – Metalab for Better Science, this workshop brings together scholars committed to strengthening the reliability, transparency, and cumulative nature of research in the social sciences.

Programme

8:45 - 9:00
Registration

9:00-9:30
Welcome 
Marianne Saam (ZBW – Leibniz-Information Centre for Economics)
Jörg Ankel-Peters (RWI – Leibniz Institute for Economic Research)
Levent Neyse (WZB Berlin Social Science Center, DIW German Institute for Economic Research)

9:30-11:00
Session 1: Open Science – Incentives, Sharing & Publishing

  • Data and Code Sharing in the Social Sciences: A Large-Scale Analysis of Trends and Journal Policies
    Jan Marcus (Freie Universität Berlin, Germany)
     
  • Incentives, Transparency, and p-Hacking – Evidence from a Diamond Open Access Reform
    Oluwatobi Gbadamosi (REGE, Université Savoie Mont Blanc, France)
     
  • Making replications count: Founding and maintaining the Replication Research Journal
    Susanne Adler (Ludwig Maximilian University of Munich, Germany)
     

11:00-11:15
Coffee Break

11:15-12:15
Keynote by Felix Holzmeister(University of Innsbruck, Austria)
The noisy gatekeeper: Measuring the reliability of peer review and testing a way to improve it
Virtually everything science “knows” has passed through peer review. It decides what gets published and – downstream – who gets believed, funded, and hired. Yet while open science has transformed how we report research, the evaluation of research has largely escaped the same scrutiny. So how well does science’s gatekeeper actually work, and could it work better? This keynote presents two complementary attempts at an answer. The first is a diagnosis at unprecedented scale: drawing on 11.7 million reviews of 4.7 million manuscripts at more than 1,900 Elsevier journals, it asks how often reviewers of the same manuscript actually agree. The second is a treatment: a preregistered field experiment within a journal’s editorial process that redesigns what reviewers are asked to produce. Together, the studies put numbers on questions usually argued from anecdote: how much of peer review is signal and how much is lottery, what reviewer disagreement means for the false-positive and false-negative rates of editorial decisions, and which levers could make the gatekeeper better at its job. Some of the answers are reassuring; several are not.

12:15-13:15
Lunch Break

13:15-15:15
Session 2
 

Room 1: Replication, Reproducibility & Robustness 

Room 2: Measurement, Experiments & Generalizability 

Donor Pool Selection in Synthetic Controls: Lessons from Carbon Tax Evaluations
Sachintha Fernando (Martin Luther University Halle-Wittenberg, Germany)

Evaluating and Improving Experimental Designs in Economics
Daniel Evans (University of Bonn, Germany)

Revisiting the Origins of Gender Roles: Replication, Measurement, and Robustness in Historical Persistence Research
Yifan Yang (Stockholm School of Economics, Sweden)

A meta-analysis of risk elicitation tasks
Elizaveta Golovanova (Centre for Research on Economic Strategies (CRESE), Université Marie et Louis Pasteur, France)

Robustness in the Deforestation Literature – An Expert-Enhanced Meta-Reproduction
Julian Rose (RWI – Leibniz Institute for Economic Research, Germany)

The (In)Stability of Risk and Time Preferences: Measurement, Context or Genuine Change
Nodir Djanibekov (Leibniz Institute of Agricultural Development in Transition Economies (IAMO), Germany)

Code–paper discrepancies in economics and finance
Christoph Huber (Aalto University, Finland)

Do Student Subject Pools Generalize? Testing Population Heterogeneity in Physicians’ Responses to Payment Incentives 
Julia von Hanxleden (University of Hamburg, Germany)

Same Data, Different Answers: Analytic Multiplicity as Shared Infrastructure for Collaborative Research
Alyssa Columbus (Ludwig Maximilian University of Munich, Germany)

When Does an AI-Assisted Replication Replicate? A Multi-Model Reassessment of Corporate Accelerator Typologies
Marvin Weingärtner (Technical University of Applied Sciences Wildau, Germany)

15:15-15:30
Coffee Break


15:30-17:00
Session 3: AI in Research – Opportunities, Reliability & Epistemic Risks

  • Fast, but incomplete and inaccurate? Performance of AI-powered tools for systematic literature reviews
    Rene Bekkers (Vrije University Amsterdam, Netherlands)
     
  • Heterogeneity of treatment effects across large language models: evidence from two behavioral cases
    Armando Holzknecht (University of Innsbruck, Austria)
     
  • Boosting Samples with Synthetic Respondents? Epistemic Risks of Human and LLM Data Mixtures in Social Research
    Amal Labbouz (Karlsruhe Institute of Technology (KIT), Germany)
     

17:00-17:10
Closing
Marianne Saam (ZBW – Leibniz-Information Centre for Economics)
Jörg Ankel-Peters (RWI – Leibniz Institute for Economic Research)
Levent Neyse (WZB Berlin Social Science Center)

Registration

Register now

Organizing committee:

  • Marianne Saam, ZBW  – Leibniz Information Centre for Economics and University of Hamburg
  • Doreen Siegfried, ZBW – Leibniz Information Centre for Economics
  • Jörg Ankel-Peters, RWI Leibniz Institute for Economic Research
  • Macartan Humphreys, WZB Berlin Social Science Center
  • Levent Neyse, DIW – The German Institute for Economic Research Berlin and WZB Berlin Social Science Center