AI features
Building AI into your own app: calling a model safely, streaming replies, tool calling, structured output, embeddings and RAG — and keeping the bill under control.
21 posts
Why Does AI Hallucinate? And How to Reduce It in Your App
AI models sometimes state false things with complete confidence. Why it happens — they predict plausible text rather than look up facts — the kinds of hallucination you'll meet in coding and apps, and practical ways to reduce it: grounding, tools, structure, checks and room to say 'I don't know'.
What Is a System Prompt? How to Give an AI Its Instructions
The system prompt is the standing instruction that shapes every answer an AI model gives in your app — its role, rules, tone and format. How it differs from user messages, what to put in one, a template you can adapt, and why it's not a security boundary.
Streaming LLM Responses to the Browser: SSE, Fetch Streams, and Gotchas
Streaming makes AI features feel fast: words appear as they're generated instead of after a ten-second wait. How to stream from the Claude API on your server, forward it to the browser, read it in React, and fix the proxies and timeouts that buffer or cut off streams.
Prompt Caching Explained: Cut LLM Costs and Latency on Repeated Prompts
If every request repeats the same long system prompt, documents or tool definitions, prompt caching lets the provider reuse that work for a fraction of the price and time. How it works, structuring prompts for cache hits, Claude's cache_control, OpenAI's automatic caching, and verifying it works.
OpenAI API vs Claude API: Which Should Your App Use?
Both APIs let your app send text (and images) to a model and get an answer back. How they differ in request shape, models, tool calling, structured output, caching and pricing — and why many apps keep the choice swappable instead of picking forever.
LLM Temperature Explained: Why the Same Prompt Gives Different Answers
Temperature controls how random an AI model's word choices are. Low temperature gives consistent, predictable answers; high temperature gives varied, creative ones. How it works, what values to use for which tasks, top-p, and why temperature 0 still isn't fully deterministic.
Fine-Tuning vs RAG: How to Teach an AI About Your Data
You want an AI model to know about your products, documents or customers. RAG looks things up and adds them to the prompt; fine-tuning retrains the model. What each actually does, what each is good at, costs and pitfalls, and why most apps should start with neither.
Claude Opus vs Sonnet vs Haiku: Which Model Should You Use?
Anthropic's Claude models come in tiers — Opus, Sonnet and Haiku, plus Fable at the top. What each is good at, what they cost on the API, and a simple rule for choosing for coding, chat and features inside your own app.
Claude API Pricing Explained: What Your AI Feature Will Actually Cost
How Claude API billing works: per-token input and output prices for each model, prompt caching, the 50% Batch discount, and worked examples for a chatbot and a summariser — so you can estimate your bill before you ship.
What Is RAG? Retrieval-Augmented Generation for Beginners
RAG lets an AI answer questions about your own documents by finding the relevant pieces first and handing them to the model. How retrieval-augmented generation works, step by step, when you need it, when you don't, and the mistakes that make it give bad answers.
What Is Function Calling (Tool Calling) in AI? Explained With Examples
Function calling — also called tool calling or tool use — lets an AI model ask your code to do things: look up an order, check the weather, query a database. How it works step by step, a Claude example, why the model never runs your code itself, and how to keep it safe.
What Is an SDK? SDK vs API, Explained for Beginners
An SDK is a ready-made toolkit for using a service from your code. What an SDK contains, how it differs from an API, real examples from Stripe and Anthropic, how to install one, and why AI tools sometimes invent SDK functions that don't exist.
What Is an API Key? A Plain-English Guide (With Claude and ChatGPT Examples)
An API key is a password for software. What API keys are, how they're different from your login, how to get one for Claude or OpenAI, where to keep it, why it must never be in your frontend code, and what to do if one leaks.
What Are Embeddings? How AI Turns Meaning Into Numbers
An embedding is a list of numbers that captures what a piece of text means, so a computer can find similar things. How embeddings work, what they're used for (search, RAG, recommendations), how to store them, and practical tips on models, dimensions and cost.
Structured Output From LLMs: Getting Reliable JSON Every Time
How to get a language model to return JSON your code can trust: why "respond in JSON" isn't enough, schema-constrained structured outputs with Claude and Zod, strict tool use, validation, handling refusals and truncation, and designing schemas models fill well.
Server-Sent Events vs WebSockets: Which for Real-Time Features and AI Streaming?
SSE streams updates from server to browser over plain HTTP; WebSockets open a two-way channel. How each works, code for both, why AI chat responses use SSE-style streaming, proxy buffering and connection-limit gotchas, and a decision guide including plain polling.
pgvector Tutorial: Vector Search in Postgres for RAG and Semantic Search
Add semantic search and RAG to your app without a separate vector database. A hands-on pgvector guide: install the extension, store embeddings, query by cosine distance, add HNSW indexes, filter results correctly, choose dimensions and halfvec, and know when you've outgrown it.
How to Add an AI Chatbot to Your App (Without Leaking Your API Key)
A beginner-friendly guide to adding an AI chat feature: the safe architecture, a working Next.js backend route that streams Claude's replies, the frontend code to display them, and the limits, costs and abuse protection you need before launch.
What Is an LLM? Large Language Models Explained Without the Hype
Claude, GPT, and Gemini are large language models. What an LLM actually is, how it's trained, why it's good at code, why it confidently makes things up, what 'model', 'prompt', and 'temperature' mean, and what that means for building apps with AI.
What Are Tokens in AI? Why Your Usage Is Counted in Pieces of Words
AI models read and write in tokens, not words — and pricing, limits, and context windows are all measured in them. What a token is, roughly how many words it equals, input vs output tokens, why code uses more of them, and practical ways to use fewer.
How to Stop Bots From Running Up Your AI App's Bill
If your app calls an AI model on a user's behalf, every request costs you money — and a bot, a scraper, or one determined user can make thousands of them overnight. Rate limits, usage caps, provider spending limits, and the architecture that keeps a surprise bill from happening.