<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
    <title>Akram - fine-tuning</title>
    <subtitle>Mohammed Akram Khan Lodi — CS undergraduate researching recursive language models, self-improving agent harnesses, and the energy cost of efficient AI.</subtitle>
    <link rel="self" type="application/atom+xml" href="http://akramlodi.com/tags/fine-tuning/atom.xml"/>
    <link rel="alternate" type="text/html" href="http://akramlodi.com/"/>
    <generator uri="https://www.getzola.org/">Zola</generator>
    <updated>2026-10-01T00:00:00+00:00</updated>
    <id>http://akramlodi.com/tags/fine-tuning/atom.xml</id>
    <entry xml:lang="en">
        <title>The Carbon Cost of Reasoning</title>
        <published>2026-10-01T00:00:00+00:00</published>
        <updated>2026-10-01T00:00:00+00:00</updated>
        
        <author>
          <name>
            Mohammed Akram Khan Lodi
          </name>
        </author>
        
        <link rel="alternate" type="text/html" href="http://akramlodi.com/research/carbon-cost-of-reasoning/"/>
        <id>http://akramlodi.com/research/carbon-cost-of-reasoning/</id>
        
        <content type="html" xml:base="http://akramlodi.com/research/carbon-cost-of-reasoning/">&lt;dl class=&quot;paper-meta&quot;&gt;
&lt;dt&gt;Full title&lt;&#x2F;dt&gt;
&lt;dd&gt;The Carbon Cost of Reasoning: Benchmarking Energy Efficiency in Fine-Tuning Gemma 4 for Mathematical Logic&lt;&#x2F;dd&gt;
&lt;dt&gt;Author&lt;&#x2F;dt&gt;
&lt;dd&gt;&lt;strong&gt;Mohammed Akram Khan Lodi&lt;&#x2F;strong&gt; (independent research, B.S. Abdur Rahman Crescent Institute of Science and Technology)&lt;&#x2F;dd&gt;
&lt;dt&gt;Presented&lt;&#x2F;dt&gt;
&lt;dd&gt;National AI Summit on Industry 5.0, April 2026, Chennai. Accepted &amp;amp; presented, Paper ID AIS 052&lt;&#x2F;dd&gt;
&lt;dt&gt;Status&lt;&#x2F;dt&gt;
&lt;dd&gt;&lt;span class=&quot;status-pill&quot;&gt;Pipeline built · full experimental runs in progress&lt;&#x2F;span&gt;&lt;&#x2F;dd&gt;
&lt;&#x2F;dl&gt;
&lt;ul class=&quot;link-row&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https:&#x2F;&#x2F;drive.google.com&#x2F;file&#x2F;d&#x2F;1SklCgsuYjd8BWy3lgWytgx1N1_dLAlR6&#x2F;view?usp=share_link&quot;&gt;Proceedings&lt;&#x2F;a&gt;&lt;&#x2F;li&gt;
&lt;li&gt;&lt;span&gt;Code: released with results&lt;&#x2F;span&gt;&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;h2 id=&quot;tl-dr&quot;&gt;TL;DR&lt;&#x2F;h2&gt;
&lt;p&gt;Fine-tuning is the standard way to specialize a base model for multi-step reasoning, and its &lt;em&gt;accuracy&lt;&#x2F;em&gt; is well documented. Its &lt;em&gt;environmental&lt;&#x2F;em&gt; cost (energy drawn and CO₂ emitted) mostly goes unmeasured. Practitioners choose between &lt;strong&gt;Full Fine-Tuning, LoRA, and QLoRA&lt;&#x2F;strong&gt; based on accuracy and available compute, with no standard way to weigh what that choice costs.&lt;&#x2F;p&gt;
&lt;p&gt;This project runs a &lt;strong&gt;controlled, reproducible comparison&lt;&#x2F;strong&gt; of the three methods on a small open-weight model for grade-school math reasoning, measuring accuracy and energy side by side. It introduces two tools for reading the result:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;the &lt;strong&gt;Green Gap&lt;&#x2F;strong&gt;: the point in training where each additional kWh stops buying meaningful accuracy; and&lt;&#x2F;li&gt;
&lt;li&gt;the &lt;strong&gt;Reasoning Efficiency Index (REI)&lt;&#x2F;strong&gt;: accuracy gained over zero-shot, per kWh consumed.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;h2 id=&quot;research-questions&quot;&gt;Research questions&lt;&#x2F;h2&gt;
&lt;ol&gt;
&lt;li&gt;How much &lt;strong&gt;energy and CO₂&lt;&#x2F;strong&gt; do Full Fine-Tuning, LoRA, and QLoRA each consume when specializing a model for multi-step mathematical reasoning?&lt;&#x2F;li&gt;
&lt;li&gt;At what point does additional energy &lt;strong&gt;stop yielding meaningful accuracy gains&lt;&#x2F;strong&gt;, i.e. where is the Green Gap?&lt;&#x2F;li&gt;
&lt;li&gt;Can a standardized metric make &lt;strong&gt;environmental efficiency as comparable across methods as accuracy already is&lt;&#x2F;strong&gt;?&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;h2 id=&quot;methodology&quot;&gt;Methodology&lt;&#x2F;h2&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;&lt;&#x2F;th&gt;&lt;th&gt;&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Base model&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;Gemma 4 E2B-it (Google, ~5.1B stored parameters, Apache 2.0)&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Benchmark&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;&lt;a href=&quot;https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2110.14168&quot;&gt;GSM8K&lt;&#x2F;a&gt;: grade-school math word problems requiring chain-of-thought, scored by exact match on the final numeric answer&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Methods&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;Full Fine-Tuning · LoRA (ranks 8 &#x2F; 16 &#x2F; 32) · QLoRA (4-bit NF4, ranks 8 &#x2F; 16 &#x2F; 32)&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Energy tracking&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;&lt;a href=&quot;https:&#x2F;&#x2F;codecarbon.io&quot;&gt;CodeCarbon&lt;&#x2F;a&gt;: GPU, CPU and RAM energy (kWh) and estimated CO₂e per run, using the cloud region’s grid carbon intensity&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;p&gt;&lt;strong&gt;Full Fine-Tuning is backbone-only.&lt;&#x2F;strong&gt; The embedding and multimodal encoder layers are frozen (about 1.9B trainable parameters). This matches the architecture’s own pretraining conventions &lt;em&gt;and&lt;&#x2F;em&gt; covers the same parts of the network that LoRA and QLoRA can reach, so the comparison is between methods rather than between different amounts of model being trained.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;strong&gt;Controls.&lt;&#x2F;strong&gt; Every condition uses the same fixed prompt template, dataset split, and evaluation protocol. Observed differences should come from the fine-tuning method, not from setup variables.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;metrics&quot;&gt;Metrics&lt;&#x2F;h2&gt;
&lt;h3 id=&quot;reasoning-efficiency-index-rei&quot;&gt;Reasoning Efficiency Index (REI)&lt;&#x2F;h3&gt;
&lt;p&gt;Accuracy gained over the zero-shot baseline, per unit of energy:&lt;&#x2F;p&gt;
&lt;div class=&quot;math&quot;&gt;$$\text{REI} = \frac{\text{Acc}_{\text{fine-tuned}} - \text{Acc}_{\text{zero-shot}}}{E_{\text{training}}\ (\text{kWh})}$$&lt;&#x2F;div&gt;
&lt;p&gt;REI turns “which method is greener?” into a single number that can be compared across methods, ranks, and, if adopted, across papers.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;the-green-gap&quot;&gt;The Green Gap&lt;&#x2F;h3&gt;
&lt;p&gt;Accuracy checkpoints are logged against &lt;strong&gt;cumulative energy&lt;&#x2F;strong&gt; throughout training. That gives an accuracy-vs-kWh curve for each method. The Green Gap is the point on that curve where the marginal gain&lt;&#x2F;p&gt;
&lt;div class=&quot;math&quot;&gt;$$\frac{\Delta\,\text{Accuracy}}{\Delta\,\text{Energy (kWh)}}$$&lt;&#x2F;div&gt;
&lt;p&gt;drops sharply. Past it, you’re mostly paying for energy, not reasoning. The rank ablations show how that point &lt;strong&gt;shifts with adapter capacity&lt;&#x2F;strong&gt;.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;experimental-design&quot;&gt;Experimental design&lt;&#x2F;h2&gt;
&lt;div class=&quot;table-scroll&quot;&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Condition&lt;&#x2F;th&gt;&lt;th&gt;Configuration&lt;&#x2F;th&gt;&lt;th&gt;Repetitions&lt;&#x2F;th&gt;&lt;th&gt;Role&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;Full Fine-Tuning&lt;&#x2F;td&gt;&lt;td&gt;Backbone only, bf16&lt;&#x2F;td&gt;&lt;td&gt;3&lt;&#x2F;td&gt;&lt;td&gt;Core&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;LoRA&lt;&#x2F;td&gt;&lt;td&gt;Rank 16, α = 32&lt;&#x2F;td&gt;&lt;td&gt;3&lt;&#x2F;td&gt;&lt;td&gt;Core&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;QLoRA&lt;&#x2F;td&gt;&lt;td&gt;4-bit NF4, rank 16&lt;&#x2F;td&gt;&lt;td&gt;3&lt;&#x2F;td&gt;&lt;td&gt;Core&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;LoRA&lt;&#x2F;td&gt;&lt;td&gt;Ranks 8 &amp;amp; 32&lt;&#x2F;td&gt;&lt;td&gt;1 each&lt;&#x2F;td&gt;&lt;td&gt;Ablation&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;QLoRA&lt;&#x2F;td&gt;&lt;td&gt;Ranks 8 &amp;amp; 32&lt;&#x2F;td&gt;&lt;td&gt;1 each&lt;&#x2F;td&gt;&lt;td&gt;Ablation&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;&#x2F;div&gt;
&lt;p&gt;The &lt;strong&gt;9 core runs&lt;&#x2F;strong&gt; support a statistically meaningful comparison across methods. The &lt;strong&gt;4 ablation runs&lt;&#x2F;strong&gt; trace how the Green Gap moves with adapter rank. A zero-shot baseline evaluation anchors REI.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;infrastructure&quot;&gt;Infrastructure&lt;&#x2F;h2&gt;
&lt;p&gt;All billed fine-tuning, baseline evaluation, and energy-tracked runs happen on &lt;strong&gt;AWS EC2&lt;&#x2F;strong&gt; (us-east-1). Fixed-spec hardware with per-instance telemetry is what makes the energy numbers credible; managed fine-tuning APIs don’t expose hardware-level energy data.&lt;&#x2F;p&gt;
&lt;div class=&quot;table-scroll&quot;&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Instance&lt;&#x2F;th&gt;&lt;th&gt;GPU&lt;&#x2F;th&gt;&lt;th&gt;Used for&lt;&#x2F;th&gt;&lt;th&gt;Why&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;g5.xlarge&lt;&#x2F;td&gt;&lt;td&gt;1× NVIDIA A10G, 24 GB&lt;&#x2F;td&gt;&lt;td&gt;Zero-shot baseline, LoRA, QLoRA&lt;&#x2F;td&gt;&lt;td&gt;Enough memory for frozen&#x2F;quantized-base training and batched inference&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;g6e.xlarge&lt;&#x2F;td&gt;&lt;td&gt;1× NVIDIA L40S, 48 GB&lt;&#x2F;td&gt;&lt;td&gt;Full Fine-Tuning&lt;&#x2F;td&gt;&lt;td&gt;~1.9B trainable parameters need more headroom than the A10G’s 24 GB&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;&#x2F;div&gt;
&lt;p&gt;To avoid wasting compute, all pipeline development (smoke tests, prompt and evaluation calibration) is done first on a free-tier Colab T4. The full study is budgeted at &lt;strong&gt;about $30–$50&lt;&#x2F;strong&gt;, with an AWS Budget alert at 50% and ~90% of a $50 cap, and instances run only for the duration of a specific experiment.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;plan-status&quot;&gt;Plan &amp;amp; status&lt;&#x2F;h2&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Phase&lt;&#x2F;th&gt;&lt;th&gt;Work&lt;&#x2F;th&gt;&lt;th&gt;Status&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;1. Pipeline validation&lt;&#x2F;td&gt;&lt;td&gt;Smoke tests, prompt&#x2F;eval calibration on Colab&lt;&#x2F;td&gt;&lt;td&gt;&lt;span class=&quot;status-pill done&quot;&gt;Done&lt;&#x2F;span&gt;&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;2. Baseline &amp;amp; core runs&lt;&#x2F;td&gt;&lt;td&gt;Zero-shot baseline; Full FT &#x2F; LoRA &#x2F; QLoRA on AWS&lt;&#x2F;td&gt;&lt;td&gt;&lt;span class=&quot;status-pill&quot;&gt;In progress&lt;&#x2F;span&gt;&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;3. Ablations &amp;amp; analysis&lt;&#x2F;td&gt;&lt;td&gt;Rank ablations, Green Gap curves, REI&lt;&#x2F;td&gt;&lt;td&gt;Planned&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;4. Write-up&lt;&#x2F;td&gt;&lt;td&gt;Results synthesis, full paper&lt;&#x2F;td&gt;&lt;td&gt;Planned&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;blockquote class=&quot;markdown-alert-note&quot;&gt;
&lt;p&gt;The study design and motivation were presented at the National AI Summit on Industry 5.0 (April 2026). The full set of energy-tracked runs is still running, so &lt;strong&gt;no final numbers are reported here yet.&lt;&#x2F;strong&gt; The code, logs, and results will be released together.&lt;&#x2F;p&gt;
&lt;&#x2F;blockquote&gt;
&lt;h2 id=&quot;expected-contributions&quot;&gt;Expected contributions&lt;&#x2F;h2&gt;
&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;reproducible, open-source benchmark&lt;&#x2F;strong&gt; of the environmental cost of three widely used fine-tuning methods on a reasoning task.&lt;&#x2F;li&gt;
&lt;li&gt;The &lt;strong&gt;Green Gap&lt;&#x2F;strong&gt; as a named, measurable phenomenon in fine-tuning cost–benefit analysis.&lt;&#x2F;li&gt;
&lt;li&gt;The &lt;strong&gt;Reasoning Efficiency Index&lt;&#x2F;strong&gt; as a reusable metric others can apply to their own fine-tuning comparisons.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;h2 id=&quot;why-this-matters-to-me&quot;&gt;Why this matters to me&lt;&#x2F;h2&gt;
&lt;p&gt;My other project, &lt;a href=&quot;http:&#x2F;&#x2F;akramlodi.com&#x2F;research&#x2F;self-harnessing-rlms&#x2F;&quot;&gt;Self-Harnessing RLMs&lt;&#x2F;a&gt;, improves a model &lt;em&gt;without&lt;&#x2F;em&gt; touching its weights. This one asks what touching the weights actually costs. Together they’re two sides of one question I keep coming back to: &lt;strong&gt;where should the compute for better reasoning go, and is it worth it?&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
</content>
        
    </entry>
</feed>
