AI agents
Working with AI coding agents that run real commands: context, permissions, review, safety rails, and where they should run.
69 posts · page 3 of 3
How to Write Good Prompts for AI Coding Tools
The difference between an AI tool that builds what you want and one that goes in circles is usually the prompt. A simple structure for asking — goal, context, constraints, and how you'll know it's done — with before-and-after examples you can copy.
Evals for Coding Agents: Measuring Whether Your Agent Setup Actually Works
Changing a CLAUDE.md, model, skill, or MCP server changes agent behaviour, usually untested. How to build an eval suite for coding-agent workflows: task selection, hermetic environments, graders, pass@k vs pass^k, cost and trajectory metrics, and running headless in CI.
Debugging for Beginners: A Calm, Repeatable Way to Find Bugs
Debugging isn't guessing until it works. A simple five-step method — reproduce, read, locate, hypothesise, verify — plus console.log, breakpoints, git bisect, rubber-ducking, and how to debug alongside an AI tool without getting stuck in a loop.
What Are Database Migrations? A Plain-English Guide
Migrations are how an app's database changes shape over time without losing data. What they are, why AI-built apps get them wrong, how to make a risky change safely, and the rules that stop a schema change from becoming a data-loss incident.
Claude Pro vs Max vs API Key for Claude Code: Which Should You Pay For?
Claude Code works with a Pro subscription, a Max subscription, or pay-as-you-go API billing. They're metered differently and suit different ways of working. How to choose, with a simple way to check your own usage.
Claude Code vs the Claude Chat App: Which Should You Use for Coding?
Both run on Claude models and come with the same subscription. The chat app is a conversation; Claude Code is an agent that works inside your project. What each does well, where each falls short, how they share usage limits, and a simple rule for when to use which.
Claude Code Subagents: When to Split the Work (and When Not To)
Subagents give Claude Code a second context window: a helper that does a noisy job — running tests, searching a codebase, reading logs — and hands back only the summary. How they work, how to write your own, and the tasks where they help versus the ones where they just add cost.
Claude Code Skills: Turn the Prompts You Keep Retyping Into Commands
If you've pasted the same instructions into Claude Code more than twice, it should be a skill. How skills work, how they differ from CLAUDE.md, how to write one with arguments and live context, and four skills worth having in almost any project.
Claude Code Plan Mode: Make It Think Before It Types
The most expensive Claude Code mistakes happen in the first five minutes, when it confidently heads in the wrong direction. Plan mode makes it research and propose before touching anything. How to use it, how to review a plan properly, and when it's overkill.
Connecting MCP Servers to Claude Code: Setup, Scopes, and the Security Question
MCP servers let Claude Code talk to your issue tracker, database, error monitoring, and docs. Adding one is a single command. Choosing which ones to trust is the harder part. How to connect them, where the configuration lives, and how to keep a useful integration from becoming a data leak.
Claude Code Hooks: Rules the Agent Can't Forget
CLAUDE.md tells Claude Code what you'd like. Hooks make it happen, every time — blocking edits to protected files, formatting after every change, refusing to stop while tests fail. Five practical hooks, how they work, and the honest limits of what a hook can enforce.
Running Claude Code Headless: claude -p in Scripts, Cron, and CI
Add -p and Claude Code stops being a chat and becomes a command: pipe input in, get text or JSON out, exit with a status code. How print mode works, how to control what it's allowed to do and spend, authentication for CI, and where headless runs still need somewhere to live.
How to Use Claude Code From Your Phone: Every Option Compared
You can check on — and steer — a Claude Code session from your phone in at least four different ways. They differ in where the code runs, what happens when your laptop sleeps, and how much you can actually do on a small screen.
Claude Code --dangerously-skip-permissions: When It's Safe, and What to Use Instead
An agent that stops to ask permission can't work while you're away. An agent that never asks can delete things. How Claude Code's permission modes and allow/deny rules work, when skipping prompts is reasonable, and what has to be true of the environment first.
Claude Code Checkpoints: How /rewind Works and What It Can't Undo
Press Esc twice and Claude Code rewinds its file edits to any earlier prompt. It's a genuine safety net — with holes shaped exactly like the mistakes that hurt most: shell commands, database changes, and anything outside the working tree. What's covered, what isn't, and what to use for the rest.
AI Hallucinated a Package: Fake Libraries, Made-Up APIs, and Slopsquatting
AI coding tools sometimes invent packages, functions, and settings that don't exist — and attackers now register the fake package names. Why it happens, how to spot hallucinated imports and APIs, what 'slopsquatting' is, and the habits that keep invented code out of your app.
Egress Control for AI Agents: Designing an Allowlist Proxy
Restricting where an agent can send data is the most reliable defence against exfiltration, and easy to get subtly wrong. Network-layer enforcement, SNI vs TLS interception, DNS as a covert channel, allowlisted domains as leak paths, credential brokering, and testing.
The Sandbox Is the Wrong Abstraction for AI Coding Agents
The industry settled on ephemeral sandboxes for AI agents — isolated, disposable, destroyed after each task. That's exactly right for running untrusted code and exactly wrong for building software. Here's the distinction that matters.
Why Your AI Agent Keeps Forgetting (And Why Bigger Context Windows Won't Fix It)
Your coding agent works brilliantly for twenty minutes, then forgets the architecture you explained. That's not a bug you can prompt your way out of — and million-token windows won't save you. What actually helps.
Why AI Coding Agents Need Persistent Workspaces
Stateless sandboxes make AI agents repeat themselves, lose context, and guess at results they could have measured. Persistent workspaces fix the feedback loop — here's the mechanism, and what it costs to build.
What Is an Autonomous Development Platform?
An Autonomous Development Platform gives AI agents and developers a persistent, managed environment to build in — not just a place to generate code. Here's what separates the category from AI code generators, PaaS, and cloud IDEs.