geofrey.dev

Geofrey Kamau

Research notes on mechanistic interpretability, LLM experiments, and building AI agent systems from scratch. Based in Nairobi.

mechanistic-interpretability·Aug 15, 2026·11m read

The Head That Tracks Names: IOI Circuit Universality from GPT-2 to Llama

The canonical GPT-2 IOI circuit — the mechanism that tracks indirect objects across names — exists in Llama-3.2-1B at the same relative depth. One head, L11·H31, fires universally across 8 experiments. It works for Kamau and Wanjiru as well as Mary and John.

mechanistic-interpretabilityioitransformerlensllamacircuitlenscircuit-universalityattention
mechanistic-interpretability·Aug 13, 2026·8m read

I Jailbroke a Language Model and It Called Itself Out

A single jailbreak instruction flips France→Paris to Berlin at 83.3%. Attribution analysis reveals the model isn't updating its beliefs — it's copying the token from the instruction. And in one run, it explicitly said so.

mechanistic-interpretabilityjailbreaktransformerlensllamacircuitlensattribution
mechanistic-interpretability·Aug 12, 2026·7m read

I Found the Attention Head That Knows Paris

OPT-1.3b stores the France→Paris fact in a single attention head — Layer 21, Head 0. Llama 3.2-1B distributes the same knowledge redundantly across dozens. The contrast reveals something fundamental about how different architectures store facts.

mechanistic-interpretabilitytransformerlensoptllamacircuitlensattention
llm-experiments·Aug 10, 2026·6m read

I Fine-Tuned a 1.2B Model to 98.7% Accuracy, Then Scrapped It

How I QLoRA-tuned LFM2.5-1.2B for skill trigger classification, hit near-perfect accuracy, then realized I had solved the wrong problem entirely.

qlorafine-tuninglorallmkroniqocolab