<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://bimarshakhanal.com.np/feed.xml" rel="self" type="application/atom+xml" /><link href="https://bimarshakhanal.com.np/" rel="alternate" type="text/html" /><updated>2026-10-03T18:42:15+00:00</updated><id>https://bimarshakhanal.com.np/feed.xml</id><title type="html">Bimarsha Khanal</title><subtitle>Bimarsha Khanal&apos;s academic portfolio - Machine Learning Engineer working on Multimodal LLMs, AI Agents, and Trustworthy AI</subtitle><author><name>Bimarsha Khanal</name><email>khanalbimarsha@gmail.com</email><uri>https://bimarshakhanal.com.np</uri></author><entry><title type="html">Paper Review: Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models</title><link href="https://bimarshakhanal.com.np/posts/2026/02/paper-review-steering-chain-of-thought/" rel="alternate" type="text/html" title="Paper Review: Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models" /><published>2026-02-01T00:00:00+00:00</published><updated>2026-02-01T00:00:00+00:00</updated><id>https://bimarshakhanal.com.np/posts/2026/02/paper-review-steering-chain-of-thought</id><content type="html" xml:base="https://bimarshakhanal.com.np/posts/2026/02/paper-review-steering-chain-of-thought/"><![CDATA[<p>Review of <a href="https://arxiv.org/pdf/2505.15634">Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models</a> (arXiv:2505.15634).</p>

<p>To illustrate my understanding of the core innovation presented in this paper, I would like to share a story. Some time ago, I watched an interview with a neurosurgeon who explained how diseases like Parkinson’s are treated using deep brain stimulation similar to steering of LLM. Surgeons had the person play a musical instrument, which he was having trouble with due to the disease, to identify the nerve, simulating a similar approach to how VS decomposition is used in this research. The chosen nerve is then stimulated using a digital pacemaker with the required strength, like modifying the original residual activation.</p>

<ol>
  <li>
    <h2 id="surgical-precision-through-vs-decomposition">Surgical Precision through VS Decomposition.</h2>
    <p>The core strength of this work lies in its ability to isolate signals from noise. By calculating the absolute difference between verbal and symbolic feature activations (|αx−αy|), the authors successfully suppress task-irrelevant features like punctuation (e.g., features 19555 and 14602 in DeepSeek-Llama3) while amplifying reasoning-specific ones like “relationship statements involving variables” (feature 462). This subtraction method is mathematically robust; Spearman’s rank correlation tests confirm that subtraction reflects model prediction importance far more accurately than simple addition, which tends to amplify noise.</p>
  </li>
  <li>
    <h2 id="the-sae-free-digital-pacemaker">The SAE-Free “Digital Pacemaker”.</h2>
    <p>Between the two steering approaches, I am particularly impressed by how this research addresses the bottleneck caused by the unavailability of SAEs. The move towards SAE-free steering is a significant leap for “Actionable Interpretability.” By treating the steering challenge as a Rayleigh quotient problem, the paper derives steering directions unit eigenvectors directly from the covariance matrix of activation differences. This allows for a targeted intervention without the prohibitive cost of training layer-specific SAEs. Remarkably, this “raw” eigenvector steering often outperforms SAE-based methods on complex benchmarks like MMLU-high, likely because it derives information from structured reasoning processes rather than general, noisy datasets.</p>
  </li>
</ol>

<p>This approach relies on the orthogonality assumption, which assumes that the most significant feature vectors in a model’s high-dimensional space are approximately pairwise orthogonal, verified through a cosine similarity check. When applied with a recommended steering strength of λ≤0.5, this eigenvector-based approach often outperformed SAE-based steering on benchmarks like MMLU-high. This success is likely because eigenvectors extract information directly from structured reasoning traces rather than the potentially noisy, general datasets used to train traditional SAEs, resulting in more targeted improvements to reasoning depth and attention allocation toward critical mathematical tokens.</p>

<p>As newer, deeper models are released, I’m curious about how the VS decomposition approach will scale. In the very deep architecture, the choice of a reasoning-rich layer may be challenging, and increased feature density may challenge the orthogonality assumption. To address the challenge of choosing a feature extraction layer, I wonder if extracting a more robust ‘global’ reasoning vector via a weighted sum or skip connections could mitigate single-layer noise or late-stage hallucinations.</p>]]></content><author><name>Bimarsha Khanal</name><email>khanalbimarsha@gmail.com</email><uri>https://bimarshakhanal.com.np</uri></author><category term="LLM interpretability" /><category term="chain-of-thought" /><category term="activation steering" /><category term="mechanistic interpretability" /><summary type="html"><![CDATA[Review of Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models (arXiv:2505.15634).]]></summary></entry><entry><title type="html">Math behind backpropagation in Neural Network</title><link href="https://bimarshakhanal.com.np/posts/2022/05/math-behind-backpropagation/" rel="alternate" type="text/html" title="Math behind backpropagation in Neural Network" /><published>2022-05-06T00:00:00+00:00</published><updated>2022-05-06T00:00:00+00:00</updated><id>https://bimarshakhanal.com.np/posts/2022/05/math-behind-backpropagation</id><content type="html" xml:base="https://bimarshakhanal.com.np/posts/2022/05/math-behind-backpropagation/"><![CDATA[<p>Here we will discuss math behind backpropagation in neural network. Here are the prerequisites for this tutorial.</p>

<ul>
  <li>Matrix, Vectors</li>
  <li>Dot product, Element-wise multiplication</li>
  <li>Matrix Calculus(Chain Rule)</li>
  <li>Gradient Descent Algorithm</li>
</ul>

<p>Let us consider a 2 layered Artificial Neural Network as shown in the figure below.
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1651832677384/lNAnxkktG.png" alt="Two layered Neural Network" />
There are two hidden layers(including the output layer). The output activation function is supposed to be a <strong>sigmoid</strong> function. And activation function is any other function (say <strong>g(x)</strong>).</p>

<h3 id="background">Background</h3>
<p>Neural Networks are first initialized with random parameters for every layers. Our objective is to correct these parameters to obtain the best fit. We use gradient descent algorithm which corrects the randomly initialized parameters at each iterations. For this, at each epochs forward propagation is performed which generates intermediate and final  output using the current parameters. Then, cost for that result is calculated. Then, there comes a bad boy, <strong>backpropagation</strong>.
   Backpropagation uses cost to calculate gradients and then comes the gradient descent.</p>

<h3 id="forward-propagation">Forward Propagation</h3>
<p>To start forward propagation, parameter matrix (W) needs to be initialized randomly for all layers. Then, forward propagation proceed as below.
\(Z^{[1]} = W^{[1]}X + b^{[1]}\)
\(A^{[1]} = g^{[1]}(Z^{[1]})\)
\(Z^{[2]} = W^{[2]}A^{[1]} + b^{[2]}\)
\(A^{[2]} = g^{[2]}(Z^{[2]}) = \sigma(Z^{[2]})\)
Note: \(\sigma \) denotes sigmoid function. ie \(\sigma =\frac{1}{1+e^{-z^{[2]}}} \)</p>
<h3 id="back-propagation">Back Propagation</h3>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1651847096444/XWlbyDCTx.png" alt="Computational Graph for back propagation" />
In our two layered Neural Network, we have four parameters \(W^{[1]}, b^{[1]}, W^{[2]}, b^{[2]}\)
Note: Superscript [l] denotes parameters of \(l^{th}\) layer.
So, to update these parameters using gradient descent we need to calculate the respective gradients
\(\frac{\partial J}{\partial W^{[2]}}, \frac{\partial J}{\partial b^{[2]}}, \frac{\partial J}{\partial W^{[1]}}, \frac{\partial J}{\partial b^{[1]}}\)</p>

<p>Chain rule of derivative will come in handy while calculating these gradients.
\(\frac{\partial J}{\partial W^{[2]}} =  \frac{\partial J}{\partial A^{[2]}}  \frac{\partial A^{[2]}}{\partial Z^{[2]}}  \frac{\partial Z^{[2]}}{\partial W^{[2]}}\)
For further simplification, we will calculate each product terms on RHS one by one.
\(\frac{\partial J}{\partial A^{[2]}} = \sigma'(Z) = \frac{d(1+e^{-Z})^{-1}}{dZ}\)
\(= \frac{e^{-Z}}{(1+e^{-z})^{2}} = \frac{1}{1+e^{-Z}}*\frac{1+e^{-Z}-1}{1+e^{-Z}}\)
\(= \sigma(Z)(1- \sigma(Z))\)
Then,
\(\frac{\partial J}{\partial A^{[2]}} = \frac{\partial (-Ylog(A^{[2]}) - (1-Y)log(1-A^{[2]}))}{\partial A^{[2]}}\)
\(= -\frac{Y}{A^{[2]}} - \frac{1-Y}{1-A^{[2]}} = \frac{A^{[2]}-Y}{A^{[2]}(1-A^{[2]})}\)</p>

<p>Similarly,
\(\frac{\partial Z^{[2]}}{\partial W^{[2]}} =  \frac{\partial (W^{[2]}A^{[1]} + b^{[2]})}{\partial W^{[2]}} = A^{[1]}\)
Hence,
\(\frac{\partial J}{\partial W^{[2]}} = (A^{[2]}-Y).A^{[1]}\)
Great! Next we will calculate gradient of \( b^{[2]} \) in similar way.
\(\frac{\partial A}{\partial b^{[2]}} =  \frac{\partial J}{\partial A^{[2]}}  \frac{\partial A^{[2]}}{\partial Z^{[2]}}  \frac{\partial Z^{[2]}}{\partial b^{[2]}}\)</p>

<p>For this, first two terms are already calculated.
\(\frac{\partial Z^{[2]}}{\partial b^{[2]}} = \frac{\partial (W^{[2]}A^{[1]} + b^{[2]})}{\partial b^{[2]}} = 1\)
Hence, \(\frac{\partial Z^{[2]}}{\partial W^{[2]}} =A^{[2]}-Y\)</p>

<p>Next, we will calculate the gradients for first hidden layer \(\frac{\partial J}{\partial W^{[1]}}, \frac{\partial J}{\partial b^{[1]}}\).</p>

<p>\(\frac{\partial J}{\partial W^{[1]}} =  \frac{\partial J}{\partial A^{[2]}}  \frac{\partial A^{[2]}}{\partial Z^{[2]}}  \frac{\partial Z^{[2]}}{\partial A^{[1]}} \frac{\partial A^{[1]}}{\partial Z^{[1]}} \frac{\partial Z^{[1]}}{\partial W^{[1]}}\)
First two terms are already calculated.
\(\frac{\partial Z^{[2]}}{\partial A^{[1]}} = W^{[2]}  \hspace{0.3cm}and \hspace{0.3cm}  \frac{\partial Z^{[1]}}{\partial W^{[1]}}= X\)</p>

<p>Hence, \( \frac{\partial J}{\partial W^{[1]}} =  (A^{[2]}-Y).W^{[2]}* g^{[1]}{‘}(Z^{[1]}).X \)
Similarly, we can calculate gradient for bias term b for first layer.</p>

<p>\( \frac{\partial J}{\partial b^{[1]}} =  (A^{[2]}-Y).W^{[2]}* g^{[1]}{‘}(Z^{[1]}) \)</p>

<p>Finally we can compute updated weight using gradient descent as below.</p>

\[W^{[2]}-= \alpha * \frac{\partial J}{\partial W^{[2]}}\]

\[b^{[2]}-= \alpha * \frac{\partial J}{\partial b^{[2]}}\]

\[W^{[1]}-= \alpha * \frac{\partial J}{\partial W^{[1]}}\]

\[b^{[1]}-= \alpha * \frac{\partial J}{\partial b^{[1]}}\]

<p>Above explanation is only the intuition to understand math behind back propagation in Neural Network. Exact coding implementation differ slightly with use of numpy function like dot product and sums. Also, in real implementation there are multiple training examples, so these formulae needs to be adjusted accordingly.</p>

<p><em>Originally published on <a href="https://bimarshak.com.np/math-behind-backpropagation-in-neural-network">bimarshak.com.np</a> on May 6, 2022.</em></p>]]></content><author><name>Bimarsha Khanal</name><email>khanalbimarsha@gmail.com</email><uri>https://bimarshakhanal.com.np</uri></author><category term="deep learning" /><category term="neural networks" /><category term="backpropagation" /><summary type="html"><![CDATA[Here we will discuss math behind backpropagation in neural network. Here are the prerequisites for this tutorial.]]></summary></entry></feed>