The Builder’s Diary: Confessions of a Claude Code Addict (And Why You Need to Master Context Engineering)
· Mohammad Syed
Note (September 2026): I wrote this in early 2026, when Opus 4.5 and Claude Code 2.0 were brand new. The tools have moved a long way since — I'm on Opus 5.5 now — but I'm leaving this as it was. The model names have aged; the lessons about context and sub-agents mostly haven't. A follow-up is coming.
Key takeaways
- Opus 4.5 in Claude Code 2.0 changed how I build. I direct agents now instead of writing every line.
- Context engineering matters most. Models lose the plot as the context window fills, so I hand off to a fresh session at 50–60% and keep the plan in a todo list.
- Sub-agents like Explore do the token-heavy searching and hand back a short summary, which keeps the main model sharp. Agent Skills load expertise only when it’s needed.
- My workflow: build shared context, let Claude write a throwaway first draft on a branch, audit its blind spots, run a sharper second pass, then have a second model review the diff.
I’m going to be completely honest: I used to treat AI coding agents like glorified autocomplete.
For a long time, I was bouncing between OpenAI’s Codex, Cursor, and earlier versions of Claude. It felt like I was constantly chasing the next benchmark. GPT-5.1-Codex was my daily driver for a while—it was sterile, it was cold, but it got the job done.
But then Anthropic dropped Opus 4.5 and baked it into Claude Code 2.0. And man, everything changed.
If you are building in AI right now and you haven't fundamentally shifted how you interact with these models, you are getting left behind. We are no longer hands-on coders. We are directors of intelligent agents. Andrej Karpathy recently joked about "crashing out" over the sheer speed of AI progress, describing these models as "little spirits/ghosts that live on your computer."
He’s right. And if you don't learn how to manage these "ghosts," they will haunt your codebase.
Here is my unfiltered builder’s diary: the raw wins, the frustrating failures, and the exact workflows I use to actually build shit with Claude Code 2.0.
The Paradigm Shift: Why Opus 4.5 Actually Has a "Soul"
Let’s stop talking about benchmarks for a second. In the real world, building software is a messy, iterative, conversational process.
What makes Opus 4.5 so violently different isn’t just its raw intelligence; it’s the Developer Experience (DX). It doesn’t act like a robotic code dispenser. It acts like that senior pair programmer who actually listens to you. It explains its reasoning beautifully. It has this uncanny "intent detection" where it just gets what I’m trying to do. Some devs on Twitter are half-joking that Anthropic post-trained a "soul" into it. I’m starting to believe them.
The QoL Upgrades I Actually Use
The Claude Code UI is a masterclass. I live in the CLI now, and here is what actually matters:
The Esc+Esc Panic Button (/rewind): Building with AI is 90% experimentation. When an experiment goes totally sideways and Claude starts hallucinating garbage into my codebase, I just hit Esc+Esc. It instantly rolls back the code and the conversation state. It’s a literal time machine.
The Token Paranoia (/context): Agents are absolute token guzzlers. I constantly monitor my context usage. If you don't, your model will suffer from "context rot" and start forgetting its own plan.
Spamming /ultrathink: When I hit a wall with a complex architectural decision, I force the model to think harder. The self-review it does before writing a single line of code has saved me from shipping so many catastrophic bugs.
The Secret Weapon: Sub-Agents and the "Attention Budget"
Here was my biggest "aha" moment: Stop forcing one model to do everything.
In the old days, we made the LLM hold the entire world in its head. Claude Code 2.0 blew my mind because it dynamically spawns Sub-agents.
If I ask Claude to find where a specific component is rendered, it doesn't guess. It quietly spins up the Explore agent—a fast, read-only specialist (usually running on a cheaper model like Haiku) that uses glob and grep to rip through my codebase.
Why is this a big deal? The Attention Budget.
The Explore agent does the dirty work and passes a concise summary back to the main Opus 4.5 agent. By offloading the messy, token-heavy search, my main model only processes high-signal information. It doesn't get lost in the weeds. It stays sharp.
I’ve even started running agents in the background. If I have a long-running script, I just spin up a background agent to monitor the logs while I keep building with the main agent. It feels like I have a whole engineering team living in my terminal.
The Dark Art of Context Engineering
If you take one thing away from this diary, let it be this: You must master Context Engineering.
LLMs are stateless. They only know what is in their current context window. Every tool call, every file read, every output—it all permanently eats up tokens. As that window fills up, the model's brain turns to mush. That’s context rot.
Here is how I fight back:
Ruthless Compaction: I am merciless about killing sessions. When my context hits 50-60%, I use a custom /handoff command. I force Claude to write a summary of what we did and what’s next, kill the session, and start fresh.
The todo.md Hack: Have you noticed agents love making todo.md files? It’s not just a cute UI feature. By constantly rewriting the todo list, the agent is reciting its objectives into the end of the context window. It forces the global plan back into the model's recent attention span. It’s a literal hack to manipulate the model's attention.
Agent Skills ("I Know Kung Fu"): Stop bloating your system prompt (CLAUDE.md). The meta now is Agent Skills. You create modular folders with a SKILL.md file. The model reads the metadata, and if it realises it needs that specific domain expertise, it loads the skill on-demand. Just like Neo in The Matrix.
My Exact Workflow: The "Throw-Away First Draft"
I don’t just theorise; I ship. Here is the exact, raw workflow I use to build complex features right now:
1. The Vibe Check (Exploration)
I start by just talking to Opus 4.5. I ask high-level questions. I make it explain the architecture to me. I make it draw ASCII diagrams. I don't write code; I just build shared context.
2. The Throw-Away Draft
Once we agree on the plan, I create a new branch. I tell Claude to write the feature end-to-end. I sit back, sip my coffee, and watch it work. I fully expect this draft to be hot garbage.
3. The Reality Check (Audit)
I look at what it built and compare it to my mental model. Where did it hallucinate? What context did it clearly miss? This step exposes the model's blind spots.
4. The Real Run
Armed with the hindsight of the failed draft, I run a second iteration. This time, my prompts are violent, precise, and highly targeted to fix the exact blind spots I found.
5. The Codex Review
Claude is a beast for execution, but for finding bugs? I still trust OpenAI. I take the final diff, pass it to GPT-5.2-Codex, and tell it to act as a senior reviewer. It’s the ultimate dynamic duo.
The Bottom Line
The AI space is moving so fast it makes my head spin. But crashing out isn't an option. The only way forward is to augment yourself.
Upskill in your domain so you have the taste and judgment to actually guide these models. Play with new tools. Break things.
The builders who win this year won't be the ones who type the fastest. They will be the ones who know how to engineer context, manage sub-agents, and orchestrate complex workflows.
Stop treating your AI like a search engine. Start treating it like a team.