Research / 2023–2026

Research

My research covers reinforcement learning under partial observability, AI evaluation and quantum technologies.

8 research projects
01 /

Evaluator Integrity
2026

AI evaluation

AI evaluation & LLM judges

I test whether automated evaluators distinguish correct answers from incorrect ones. My commit-first judging study found that asking a judge to solve a task before scoring candidates removed incorrect acceptances on one coding task, but increased them on another when the judge’s own answer was wrong. A census of eight frameworks found that none of 24 applicable default configurations implemented commit-first judging.

At Evaluator Integrity, we also develop methods that can detect evaluator failures without knowing the correct answers, by finding contradictions between accepted outputs. Our latest MASK experiments in Inspect Evals, developed with AISI contributions, reproduced cases where the recorded score contradicted the judge’s own conclusion across GPT-4.1 and GPT-4o. These controlled results have undergone model-assisted review; human review is pending.

02 /

University College London
2026

Manuscript in preparation

Training-parameter bias in evaluations of agent memory

Training parameters can affect a comparison intended to measure the benefit of agent memory. In a solvable model, I derived how the learning-horizon parameter λ changes a memoryless agent’s equilibrium, producing a tenfold change in corrective-action strength. Four predicted thresholds where the distortion disappears were located experimentally within 1% of the predictions.

The study includes 25 tests fixed before the experiments: 11 met their criteria and 14 did not. It also examines when the effect is absent and whether it persists in larger or nonlinear systems.

Manuscript in preparation; a publication link will be added when available.
03 /

Research project · UCL / BT
September 2025–September 2026

Reinforcement learning

Deep reinforcement learning for quantum repeater networks

I built a quantum-network simulator and trained reinforcement learning policies to choose repeater operations. In the settings studied, deep reinforcement learning outperformed the “swap-as-soon-as-possible” heuristic: networks of up to seven nodes, corresponding to 140 km, established end-to-end entanglement using roughly half as many operations while maintaining or improving success rates.

The work progressed from Q-learning to deep Q-networks and examined action selection, readout and reward design as networks grew.

Research project · UCL / BT. Entanglement Distribution in Quantum Networks: A Reinforcement Learning Approach. Research supervisor: Dr Alejandra Beghelli. Public report link forthcoming.
04 /

University College London
2025–2026

Learning theory

Reinforcement learning under partial observability

Under partial observability, a learning rule can converge to a worse policy than the best one the agent can represent. I derived exact expressions for the equilibrium learning update and its cost, tracing the failure to biased value estimates. In the reference setting, learning converges to a policy costing 35% more than the best available policy.

The analysis identifies how changing the learning horizon removes the failure. Theoretical predictions were tested with an unmodified deep PPO implementation, with proofs and numerical checks against independent solvers.

05 /

University of Rostock
June–September 2024

Quantum optics

The diamond–air interface as a photon antenna

During a research internship with AG Quantum Technology, I improved optical-readout analysis for diamond-based quantum systems. This work contributed to a co-authored Nano Letters paper on using the diamond–air interface as an efficient photon antenna for solid-state emitters.

06 /

Bilkent University
2025

Quantum simulation

NV-centre qubits & non-Markovian dynamics

I modelled environmental memory, coherence and gate fidelity in NV-centre qubit simulations using QuTiP. The work compared echo and dynamical-decoupling protocols, with approximately 50 invariant tests checking the physical consistency of the simulations.

07 /

Aybas Lab · Bilkent University
November 2023–May 2025

Quantum sensing

Precision magnetometry

As an undergraduate researcher in Aybas Lab, I investigated methods to improve magnetometer sensitivity and accuracy, including the measurement of weak magnetic signals.

08 /

NANOTAM · Bilkent University
July–August 2023

Device characterisation

Avalanche photodetectors

I characterised avalanche photodetectors, extracted parameters from time-series measurements, and automated measurement workflows with LabVIEW and Python during a research internship at NANOTAM.

Publications

Publications & manuscripts