About · London, UK

Idil Gozel

Researcher in reinforcement learning, AI evaluation and quantum physics.

I study how reinforcement learning algorithms behave when they cannot observe the full state of a system, and how training choices affect evaluations of agent memory. I also work on AI evaluator reliability and reinforcement learning for quantum networks.

Idil Gozel
2023–present

Research experience

  1. 2026–present

    Evaluator Integrity

    Founder · AI evaluation research.

  2. 2025–2026

    University College London

    Reinforcement learning theory and agent memory; quantum networks with BT.

  3. 2024

    University of Rostock

    Diamond-based quantum optics and optical-readout analysis.

  4. 2023–2025

    Bilkent University

    Magnetometry, NV-centre qubit simulation and photodetector characterisation.

Background

About me

I founded Evaluator Integrity, where we test automated AI evaluators and develop reproducible checks for incorrect scores. Our methods can detect evaluator failures even when the correct answers are unknown.

At UCL, I derived closed-form results on reinforcement learning under partial observability and studied how the learning-horizon parameter biases comparisons of agents with and without memory. In a research project with UCL and BT, I developed deep reinforcement learning policies for entanglement distribution in quantum repeater networks.

At the University of Rostock, I contributed to optical-readout analysis for diamond-based quantum systems and co-authored a paper in Nano Letters.

At Bilkent University, I worked on magnetometry in Aybas Lab, avalanche photodetectors at NANOTAM, and NV-centre qubit simulations. My background in physics informs my work in mathematical analysis, numerical simulation and experiments.

MSc Quantum TechnologiesUniversity College London · 2025–2026
BSc PhysicsBilkent University · 2021–2025

Curriculum vitae

Idil Gozel’s curriculum vitae, page 1 of 2
1 +44 7444 659531 London idilgozel@evaluatorintegrity.com Idil Gozel evaluatorintegrity.com GitHub: idilgozel LinkedIn: Idil Gozel ENTERPRISE AND AI ASSURANCE Founder, Evaluator Integrity London August 2026 - present Founded Evaluator Integrity to develop independent methods for detecting and explaining failures in automated AI evaluation. • Developed a system that detects evaluator failures without knowing the correct answers, using contradictions between accepted outputs to produce checkable evidence. • Found scoring failures in MASK (Inspect Evals, developed with the UK AI Security Institute). • Found four specification violations accepted by a code verifier; a repair removed all four while retaining 256 valid solutions. Published evaluation criteria and a framework review; submitted evidence to Ofgem and findings to AISI, OpenAI and Anthropic. Demonstrating how AI evaluators can reproduce their own errors Evaluator Integrity’s first research publication | arXiv:2609.00088 August 2026 • Showed that requiring an AI judge to solve a task before scoring other answers can cause optimisation to reproduce the judge’s mistakes. The defence eliminated incorrect acceptances on one coding task but increased them on another, each across two runs. • Audited eight evaluation frameworks: none of 24 applicable defaults implemented the defence; nine inherited a weaker approach from one prompt. Released records and verification materials. Total experimental spend: approximately £27. EDUCATION MSc Quantum Technologies University College London 2025 - 2026 Expected grade: Distinction. Modules: Advanced Photonic Devices; Techniques in High Performance Computing; Research Software Engineering with Python; Research Software Engineering with C++; Advanced Quantum Theory; Quantum Computation and Communication; Advanced Statistical Mechanics. BSc Physics Bilkent University 2021 - 2025 Overall grade: 89/100. Modules: Numerical Methods in Physics; Quantum Mechanics; Quantum Sensing and Measurement; Condensed Matter Physics; Atomic Molecular Optical Physics; Probability Theory; Statistical Mechanics; Advanced Calculus; Electromagnetic Theory. SELECTED RESEARCH Identifying training-parameter bias in evaluations of agent memory University College London | Manuscript in preparation 2026 • Established how lambda (λ), a standard parameter controlling the learning horizon, distorts the performance of agents without memory. Comparisons intended to measure the benefit of memory can therefore also measure the training recipe. • Derived the exact learning equilibrium across λ settings in a solvable model, revealing a tenfold change in the strength of the agent’s corrective actions. Predicted four thresholds where the distortion disappears and located all four experimentally within 1% of the advance predictions. • Confirmed predicted absence of the effect when actions do not influence the unobserved part of the system, in the test model and public benchmarks. Tested persistence in larger systems and a strongly nonlinear extension. • Made the research independently auditable: froze 25 hypothesis tests before experiments and published every outcome, including 11 that met their criteria and 14 that did not. Downgraded a proposed mechanism under a pre-set rule; released the complete scoring and results.
Idil Gozel’s curriculum vitae, page 2 of 2
2 REINFORCEMENT LEARNING AND QUANTUM RESEARCH Deep reinforcement learning outperforms a quantum-network heuristic Research project | UCL / BT September 2025 - September 2026 • Achieved, to our knowledge, the first demonstration of a deep reinforcement learning policy outperforming the “swap-as-soon-as-possible” heuristic widely used in analytical studies of quantum repeater networks, within the settings studied. • Established end-to-end entanglement in networks of up to seven nodes, corresponding to 140 km, using roughly half as many operations as the heuristic while maintaining or improving the success rate. • Identified how action selection, readout and reward design affected reliability as networks grew. Built the simulator and reproducible CPU/GPU experiments, progressing from Q-learning to deep Q-networks in PyTorch; fixed the evaluation protocol before experiments. Research project: Entanglement Distribution in Quantum Networks: A Reinforcement Learning Approach. Submitted 3 September 2026. Supervisor: Dr Alejandra Beghelli. Reinforcement learning under partial observability University College London | arXiv:2608.07228 2025 - 2026 • Derived theorems explaining how incomplete observations bias learning away from a good, representable policy, separating failure of the learning rule from limits of the available strategies. • Obtained exact expressions for the equilibrium of the expected learning update and its cost. Traced the failure to biased value estimates and identified how changing the learning horizon removes it. In the reference setting, learning settles at a policy costing 35% more than the best available policy. • Tested theoretical predictions with an unmodified deep PPO implementation; released analytical proofs and executable numerical checks against independent solvers. EARLIER RESEARCH Research intern, AG Quantum Technology, University of Rostock (June - September 2024). Improved optical-readout analysis for diamond-based quantum systems; co-author of a Nano Letters paper. NV-centre qubit simulation, Bilkent University (2025). Modelled environmental memory, coherence and gate fidelity in QuTiP; compared echo and dynamical-decoupling protocols and checked physical consistency with approximately 50 invariant tests. Undergraduate researcher, Aybas Lab, Bilkent University (November 2023 - May 2025). Researched methods to improve magnetometer sensitivity and accuracy. Research intern, NANOTAM, Bilkent University (July - August 2023). Characterised avalanche photodetectors, extracted time-series parameters and automated measurements with LabVIEW/Python. PUBLICATIONS Gozel, I. Commit-first LLM judging inherits the judge’s own errors. Preprint, 2026. arXiv:2609.00088. Gozel, I. Learning Suffers More Than the Policy Class Under Partial Observability: A Closed-Form Analysis. Preprint, 2026. arXiv:2608.07228. Paul Weinbrenner, Aina Lopez Benet, İdil Gözel and Friedemann Reinhard. Harnessing the Diamond-Air Interface as an Efficient Photon Antenna for Solid-State Emitters. Nano Letters 26(3), 990-996 (2026). doi:10.1021/acs.nanolett.5c04935. TECHNICAL SKILLS Programming: Python, C++, PyTorch, NumPy, SciPy, QuTiP, MATLAB, Mathematica, Git, LaTeX. Machine learning: DQN, PPO, actor-critic theory, partial observability, sparse/delayed rewards, action masking, curriculum learning, AI evaluator reliability. Mathematics: Monte Carlo, Metropolis/Wang-Landau sampling, parallel tempering, Itô calculus, Fokker-Planck/Langevin dynamics, quantum simulation, signal processing. Engineering: CPU/GPU HPC, Numba, distributed experiments, checkpointing, reproducible pipelines, automated verification, continuous integration, LabVIEW.

Research areas

All research
01 /

Reinforcement learning under partial observability

I derive exact results on learning-rule bias and test how training parameters affect evaluations of agent memory.

Research details
02 /

Reinforcement learning for quantum networks

I develop learning-based control policies for entanglement distribution in quantum repeater networks.

Research details
03 /

AI evaluation & LLM judges

I test whether automated evaluators accept incorrect outputs and whether proposed changes address those failures.

Research details