AI agents
Working with AI coding agents that run real commands: context, permissions, review, safety rails, and where they should run.
53 posts · page 2 of 2
Refactoring AI-Generated Code: Cleaning Up Without Breaking Things
Refactoring improves code's structure without changing what it does. When to refactor an AI-built app, how to do it safely with tests and small steps, the most valuable clean-ups, and prompts that stop the AI from rewriting everything.
Prompt Injection in Coding Agents: A Threat Model
A coding agent with a shell, credentials, and network access reads text written by strangers all day. A threat model — sources, capabilities, sinks — why detection-based defences fail, and the architectural controls that actually bound the damage.
Preview Environments for Every Branch: How They Work and What They Cost
A preview environment gives every branch or pull request its own live URL, so changes are reviewed running rather than read as diffs. How they work, the hard part (databases), the ways to get one, and why they matter more when an AI agent is writing the code.
How to Plan Your First App Before You Ask AI to Build It
Thirty minutes of planning saves days of AI going in circles. How to define the one problem your app solves, cut it down to a first version, describe your users' journeys and your data, and turn it all into a brief an AI tool can build from.
Running Claude Code Agents in Parallel With Git Worktrees
Two agents in one checkout will overwrite each other's work. Git worktrees give each Claude Code session its own files and branch on the same repository. How to set it up, and the parts nobody warns you about: ports, databases, and dependencies.
How to Keep Claude Code Costs Down (Without Making It Worse)
Claude Code usage is driven less by how much you ask and more by how much context every request carries. Where the tokens actually go, how to see them, and the habits that cut usage — clearing between tasks, picking the right model, trimming CLAUDE.md, and planning before building.
How to Write a CLAUDE.md That Actually Changes What Claude Does
Most CLAUDE.md files are either empty or a wall of generic advice Claude would have followed anyway. What to put in one, what to leave out, how the files load, and how to tell whether yours is working.
How to Read a Diff: Reviewing What Your AI Tool Changed
A diff shows exactly what changed in your code: red lines removed, green lines added. How to read unified and side-by-side diffs, what the @@ lines mean, where to look them up in Git, VS Code, and GitHub, and a quick review routine for AI-generated changes.
How to Write Good Prompts for AI Coding Tools
The difference between an AI tool that builds what you want and one that goes in circles is usually the prompt. A simple structure for asking — goal, context, constraints, and how you'll know it's done — with before-and-after examples you can copy.
Evals for Coding Agents: Measuring Whether Your Agent Setup Actually Works
Changing a CLAUDE.md, model, skill, or MCP server changes agent behaviour, usually untested. How to build an eval suite for coding-agent workflows: task selection, hermetic environments, graders, pass@k vs pass^k, cost and trajectory metrics, and running headless in CI.
Debugging for Beginners: A Calm, Repeatable Way to Find Bugs
Debugging isn't guessing until it works. A simple five-step method — reproduce, read, locate, hypothesise, verify — plus console.log, breakpoints, git bisect, rubber-ducking, and how to debug alongside an AI tool without getting stuck in a loop.
What Are Database Migrations? A Plain-English Guide
Migrations are how an app's database changes shape over time without losing data. What they are, why AI-built apps get them wrong, how to make a risky change safely, and the rules that stop a schema change from becoming a data-loss incident.
Claude Pro vs Max vs API Key for Claude Code: Which Should You Pay For?
Claude Code works with a Pro subscription, a Max subscription, or pay-as-you-go API billing. They're metered differently and suit different ways of working. How to choose, with a simple way to check your own usage.
Claude Code vs the Claude Chat App: Which Should You Use for Coding?
Both run on Claude models and come with the same subscription. The chat app is a conversation; Claude Code is an agent that works inside your project. What each does well, where each falls short, how they share usage limits, and a simple rule for when to use which.
Claude Code Subagents: When to Split the Work (and When Not To)
Subagents give Claude Code a second context window: a helper that does a noisy job — running tests, searching a codebase, reading logs — and hands back only the summary. How they work, how to write your own, and the tasks where they help versus the ones where they just add cost.
Claude Code Skills: Turn the Prompts You Keep Retyping Into Commands
If you've pasted the same instructions into Claude Code more than twice, it should be a skill. How skills work, how they differ from CLAUDE.md, how to write one with arguments and live context, and four skills worth having in almost any project.
Claude Code Plan Mode: Make It Think Before It Types
The most expensive Claude Code mistakes happen in the first five minutes, when it confidently heads in the wrong direction. Plan mode makes it research and propose before touching anything. How to use it, how to review a plan properly, and when it's overkill.
Connecting MCP Servers to Claude Code: Setup, Scopes, and the Security Question
MCP servers let Claude Code talk to your issue tracker, database, error monitoring, and docs. Adding one is a single command. Choosing which ones to trust is the harder part. How to connect them, where the configuration lives, and how to keep a useful integration from becoming a data leak.
Claude Code Hooks: Rules the Agent Can't Forget
CLAUDE.md tells Claude Code what you'd like. Hooks make it happen, every time — blocking edits to protected files, formatting after every change, refusing to stop while tests fail. Five practical hooks, how they work, and the honest limits of what a hook can enforce.
Running Claude Code Headless: claude -p in Scripts, Cron, and CI
Add -p and Claude Code stops being a chat and becomes a command: pipe input in, get text or JSON out, exit with a status code. How print mode works, how to control what it's allowed to do and spend, authentication for CI, and where headless runs still need somewhere to live.
How to Use Claude Code From Your Phone: Every Option Compared
You can check on — and steer — a Claude Code session from your phone in at least four different ways. They differ in where the code runs, what happens when your laptop sleeps, and how much you can actually do on a small screen.
Claude Code --dangerously-skip-permissions: When It's Safe, and What to Use Instead
An agent that stops to ask permission can't work while you're away. An agent that never asks can delete things. How Claude Code's permission modes and allow/deny rules work, when skipping prompts is reasonable, and what has to be true of the environment first.
Claude Code Checkpoints: How /rewind Works and What It Can't Undo
Press Esc twice and Claude Code rewinds its file edits to any earlier prompt. It's a genuine safety net — with holes shaped exactly like the mistakes that hurt most: shell commands, database changes, and anything outside the working tree. What's covered, what isn't, and what to use for the rest.
AI Hallucinated a Package: Fake Libraries, Made-Up APIs, and Slopsquatting
AI coding tools sometimes invent packages, functions, and settings that don't exist — and attackers now register the fake package names. Why it happens, how to spot hallucinated imports and APIs, what 'slopsquatting' is, and the habits that keep invented code out of your app.
Egress Control for AI Agents: Designing an Allowlist Proxy
Restricting where an agent can send data is the most reliable defence against exfiltration, and easy to get subtly wrong. Network-layer enforcement, SNI vs TLS interception, DNS as a covert channel, allowlisted domains as leak paths, credential brokering, and testing.
The Sandbox Is the Wrong Abstraction for AI Coding Agents
The industry settled on ephemeral sandboxes for AI agents — isolated, disposable, destroyed after each task. That's exactly right for running untrusted code and exactly wrong for building software. Here's the distinction that matters.
Why Your AI Agent Keeps Forgetting (And Why Bigger Context Windows Won't Fix It)
Your coding agent works brilliantly for twenty minutes, then forgets the architecture you explained. That's not a bug you can prompt your way out of — and million-token windows won't save you. What actually helps.
Why AI Coding Agents Need Persistent Workspaces
Stateless sandboxes make AI agents repeat themselves, lose context, and guess at results they could have measured. Persistent workspaces fix the feedback loop — here's the mechanism, and what it costs to build.
What Is an Autonomous Development Platform?
An Autonomous Development Platform gives AI agents and developers a persistent, managed environment to build in — not just a place to generate code. Here's what separates the category from AI code generators, PaaS, and cloud IDEs.