Research

I’m interested in how language-model systems can improve themselves without changing their weights, and in what that improvement costs. Concretely, that means two threads:

Broader interests: recursive self-improvement · agentic systems & harness engineering · recursive language models (RLMs) · long-horizon reasoning · efficient AI systems.

Both projects below are ongoing. I report results only once the experiments are finished.

Under review · experiments ongoing

Self-Harnessing Recursive Language Models

Submitted to the NeurIPS 2026 Workshop on Meta Agents

Can a frozen recursive language model mine its own failures, rewrite its own harness, and keep working on inputs 8–32× longer than it was optimized on?

• By William Stanford, Mohammed Akram Khan Lodi, Jaloliddin Boymakhammadov, Eliaz Calvar and Simon Coumes
Presented · experiments ongoing

The Carbon Cost of Reasoning

National AI Summit on Industry 5.0, 2026 (Paper ID AIS 052)

Benchmarking the energy and CO₂ cost of Full Fine-Tuning, LoRA, and QLoRA on Gemma 4 for mathematical reasoning, and asking where extra energy stops buying accuracy.

• By Mohammed Akram Khan Lodi