<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Production Backend Patterns]]></title><description><![CDATA[Production Backend Patterns]]></description><link>https://younggao.hashnode.dev</link><generator>RSS for Node</generator><lastBuildDate>Wed, 07 Oct 2026 13:49:50 GMT</lastBuildDate><atom:link href="https://younggao.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[I Built an AI That Reviews Every PR for Security Bugs — Here's How (2026)]]></title><description><![CDATA[What if every pull request got a security review before merge?
Not a linter check. Not a regex-based scanner. An actual review — the kind a senior security engineer would do — pointing out SQL injecti]]></description><link>https://younggao.hashnode.dev/i-built-an-ai-that-reviews-every-pr-for-security-bugs-here-s-how-2026</link><guid isPermaLink="true">https://younggao.hashnode.dev/i-built-an-ai-that-reviews-every-pr-for-security-bugs-here-s-how-2026</guid><category><![CDATA[AI]]></category><category><![CDATA[Security]]></category><category><![CDATA[GitHub]]></category><category><![CDATA[TypeScript]]></category><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 15:25:00 GMT</pubDate><content:encoded><![CDATA[<p>What if every pull request got a security review before merge?</p>
<p>Not a linter check. Not a regex-based scanner. An actual <em>review</em> — the kind a senior security engineer would do — pointing out SQL injection, hardcoded secrets, command injection, and path traversal bugs, with inline comments on the exact lines that are broken.</p>
<p>I built that. It took a weekend. It costs about $0.003 per review. And it runs on Cloudflare Workers with zero servers to manage.</p>
<p>Let me show you how.</p>
<h2>The Problem Nobody Talks About</h2>
<p>Here's a dirty secret about most engineering teams: <strong>security reviews don't happen.</strong></p>
<p>Oh sure, there's a quarterly pen test. Maybe a SAST tool that generates 400 findings nobody reads. But per-PR, inline, "hey this specific line has a command injection" review? That requires a security engineer looking at every diff. And most teams don't have one. Even the ones that do — they can't review every PR.</p>
<p>Meanwhile, 80%+ of breaches start with a vulnerability that was visible in the code at commit time.</p>
<p>I got tired of this gap. So I built <strong>CodeGuardAI</strong> — a GitHub App that hooks into your PRs and posts security-focused reviews using Claude as the analysis engine.</p>
<h2>Architecture: Embarrassingly Simple</h2>
<p>The entire thing is <strong>four TypeScript files</strong> running on a single Cloudflare Worker:</p>
<pre><code class="language-plaintext">src/
├── index.ts      # Webhook handler + routing (Hono)
├── github.ts     # GitHub API client (auth, fetch diffs, post reviews)
├── reviewer.ts   # AI review engine (Claude API + diff parsing)
├── types.ts      # TypeScript interfaces
└── (that's it)
</code></pre>
<p>The flow:</p>
<ol>
<li><p>GitHub sends a webhook when a PR is opened or updated</p>
</li>
<li><p>Worker verifies the HMAC signature</p>
</li>
<li><p>Fetches the PR diff via GitHub API</p>
</li>
<li><p>Sends the diff to Claude with a security-focused system prompt</p>
</li>
<li><p>Parses the structured JSON response</p>
</li>
<li><p>Posts inline review comments back on the PR</p>
</li>
</ol>
<p>No queues. No containers. No Redis. Just a Worker that wakes up, does the job, and goes back to sleep. Cloudflare's <code>waitUntil()</code> handles the async processing so the webhook returns <code>202 Accepted</code> immediately.</p>
<p>Data layer? A single D1 (SQLite) database tracking installations, reviews, and usage. Three tables. That's the whole backend.</p>
<h2>The Interesting Parts</h2>
<p>Let me walk through the pieces that actually matter.</p>
<h3>1. Webhook Security</h3>
<p>Every GitHub webhook comes with an HMAC-SHA256 signature. You <strong>must</strong> verify it, and you must do it with constant-time comparison (or you're vulnerable to timing attacks):</p>
<pre><code class="language-typescript">export async function verifyWebhookSignature(
  payload: string,
  signature: string | null,
  secret: string
): Promise&lt;boolean&gt; {
  if (!signature) return false;

  const sig = signature.startsWith("sha256=")
    ? signature.slice(7)
    : signature;

  const encoder = new TextEncoder();
  const key = await crypto.subtle.importKey(
    "raw",
    encoder.encode(secret),
    { name: "HMAC", hash: "SHA-256" },
    false,
    ["sign"]
  );

  const signed = await crypto.subtle.sign(
    "HMAC", key, encoder.encode(payload)
  );
  const expected = Array.from(new Uint8Array(signed))
    .map((b) =&gt; b.toString(16).padStart(2, "0"))
    .join("");

  // Constant-time comparison
  if (sig.length !== expected.length) return false;
  let result = 0;
  for (let i = 0; i &lt; sig.length; i++) {
    result |= sig.charCodeAt(i) ^ expected.charCodeAt(i);
  }
  return result === 0;
}
</code></pre>
<p>This runs on the Web Crypto API — no Node.js <code>crypto</code> module needed. Works perfectly in Workers.</p>
<h3>2. The System Prompt (Where the Magic Lives)</h3>
<p>This is the part I iterated on the most. The system prompt turns a general-purpose LLM into a focused security reviewer:</p>
<pre><code class="language-typescript">const SYSTEM_PROMPT = `You are CodeGuardAI, an expert security-focused 
code reviewer. You analyze pull request diffs and identify issues.

Focus areas (in priority order):
1. Security vulnerabilities: SQL injection, XSS, SSRF, path traversal, 
   command injection, prototype pollution, ReDoS
2. Hardcoded secrets: API keys, passwords, tokens, private keys
3. Authentication/Authorization flaws: Missing auth checks, broken 
   access control
4. Race conditions: TOCTOU bugs, unprotected shared state
5. Error handling: Information leakage, missing validation
6. Performance anti-patterns: N+1 queries, unbounded loops

Rules:
- Only comment on ADDED or MODIFIED lines (lines starting with +)
- Be specific — reference the exact code pattern
- Provide a fix suggestion when possible
- Don't flag style/formatting issues
- If the diff looks clean, say so briefly`;
</code></pre>
<p>Key decisions:</p>
<ul>
<li><p><strong>Priority ordering matters.</strong> Claude respects the hierarchy — it won't waste comments on style when there's a SQL injection.</p>
</li>
<li><p><strong>"Only comment on added lines"</strong> prevents noise from reviewing unchanged context.</p>
</li>
<li><p><strong>Structured JSON output</strong> makes parsing deterministic. No regex extraction of natural language.</p>
</li>
<li><p><strong>"If clean, say so briefly"</strong> prevents the AI from inventing problems to justify its existence.</p>
</li>
</ul>
<h3>3. Diff Context Building</h3>
<p>You can't just dump the entire repo into the context window. I build a focused diff payload with smart filtering:</p>
<pre><code class="language-typescript">function buildDiffContext(
  files: FileChange[], 
  maxChars: number = 80000
): string {
  let context = "";
  let truncated = false;

  for (const file of files) {
    if (isIgnoredFile(file.filename)) continue;

    const fileBlock = `\n--- \({file.filename} (\){file.status}) ---\n` +
                      `${file.patch}\n`;

    if (context.length + fileBlock.length &gt; maxChars) {
      truncated = true;
      break;
    }
    context += fileBlock;
  }

  if (truncated) {
    context += "\n[... additional files truncated ...]\n";
  }
  return context;
}
</code></pre>
<p>The <code>isIgnoredFile()</code> function skips lockfiles, sourcemaps, images, build artifacts, and vendored dependencies. No point burning tokens reviewing <code>package-lock.json</code>.</p>
<p>For large PRs (5000+ changed lines), I truncate to the first ~3000 lines and leave a comment telling the author to break up the PR. This is both a cost safeguard and genuinely good advice.</p>
<h3>4. Posting Inline Reviews</h3>
<p>GitHub's review API is powerful but finicky. You post a "review" with inline comments attached to specific lines:</p>
<pre><code class="language-typescript">export async function postReview(
  token: string,
  owner: string,
  repo: string,
  prNumber: number,
  commitSha: string,
  comments: ReviewComment[],
  summary: string
): Promise&lt;void&gt; {
  const body: Record&lt;string, unknown&gt; = {
    commit_id: commitSha,
    body: summary,
    event: "COMMENT",
  };

  if (comments.length &gt; 0) {
    body.comments = comments;
  }

  const resp = await fetch(
    `https://api.github.com/repos/\({owner}/\){repo}/pulls/${prNumber}/reviews`,
    {
      method: "POST",
      headers: {
        Authorization: `token ${token}`,
        Accept: "application/vnd.github+json",
        "User-Agent": "CodeGuardAI/1.0",
      },
      body: JSON.stringify(body),
    }
  );

  // If inline comments fail (line not in diff), retry summary-only
  if (!resp.ok &amp;&amp; comments.length &gt; 0 &amp;&amp; resp.status === 422) {
    await postReview(token, owner, repo, prNumber, commitSha, [], summary);
    return;
  }
}
</code></pre>
<p>The fallback logic is important — GitHub returns 422 if a comment references a line that's not in the diff context. Rather than losing the entire review, we retry with just the summary.</p>
<p>Each comment gets a severity emoji and a formatted body:</p>
<pre><code class="language-typescript">const severityEmoji =
  comment.severity === "critical" ? "🚨"
  : comment.severity === "warning" ? "⚠️"
  : "💡";

reviewComments.push({
  path: comment.file,
  line: comment.line,
  side: "RIGHT",
  body: `\({severityEmoji} **\){comment.severity.toUpperCase()}**\n\n${comment.message}`,
});
</code></pre>
<h2>What It Actually Finds</h2>
<p>I tested this against a deliberately vulnerable PR with common security anti-patterns. Here's what CodeGuardAI flagged:</p>
<p><strong>🚨 CRITICAL — Command Injection:</strong></p>
<pre><code class="language-javascript">// The PR had this:
const output = execSync(`git log --author=${req.query.author}`);
// CodeGuardAI flagged it immediately with a fix:
// Use execFileSync with argument array instead
</code></pre>
<p><strong>🚨 CRITICAL — Path Traversal:</strong></p>
<pre><code class="language-javascript">// The PR had this:
const file = fs.readFileSync(`./uploads/${req.params.filename}`);
// CodeGuardAI: "User input in file path without sanitization. 
// An attacker can use ../../../etc/passwd to read arbitrary files."
</code></pre>
<p><strong>⚠️ WARNING — Hardcoded Secret:</strong></p>
<pre><code class="language-javascript">// The PR had this:
const API_KEY = "sk-proj-abc123...";
// CodeGuardAI: "Hardcoded API key. Use environment variables."
</code></pre>
<p>All three findings appeared as inline comments on the exact lines in the PR diff. The summary at the top rated it <strong>CRITICAL</strong> risk.</p>
<h2>The Economics</h2>
<p>Let's talk money, because this is where it gets interesting.</p>
<p>A typical PR diff is 200-500 lines. With Claude Sonnet, that's roughly:</p>
<ul>
<li><p><strong>Input tokens:</strong> ~2,000-4,000 (system prompt + diff)</p>
</li>
<li><p><strong>Output tokens:</strong> ~500-1,000 (JSON response)</p>
</li>
<li><p><strong>Cost per review:</strong> ~$0.01-0.03</p>
</li>
</ul>
<p>With Claude Haiku, it's even cheaper — around <strong>$0.003 per review</strong>.</p>
<p>For a team doing 50 PRs/week, that's about <strong>\(0.60/week</strong> or <strong>\)2.50/month</strong> in API costs. The Cloudflare Worker free tier handles up to 100K requests/day. D1 is free for 5M reads/day.</p>
<p><strong>Total infrastructure cost for a small team: basically zero.</strong></p>
<p>Compare that to the cost of one security incident that could have been caught in code review.</p>
<h2>What I'd Do Differently</h2>
<p>After running this for a few weeks, here's what I've learned:</p>
<ol>
<li><p><strong>Model choice matters.</strong> Sonnet is better at catching subtle logic bugs. Haiku is fine for the obvious stuff (hardcoded secrets, injection patterns). I'm considering a tiered approach — Haiku for initial scan, Sonnet for files touching auth/crypto/networking.</p>
</li>
<li><p><strong>False positives are the enemy.</strong> If the bot cries wolf too often, developers ignore it. The "don't be pedantic" instruction in the system prompt helps, but I'm still tuning.</p>
</li>
<li><p><strong>Context window limits hurt.</strong> For massive PRs (1000+ files), you can't review everything. The truncation strategy works, but ideally you'd prioritize high-risk files (auth handlers, API routes, database queries) over UI components.</p>
</li>
<li><p><strong>GitHub App auth is painful.</strong> JWT generation, installation tokens, the whole dance. Once it works, it works. But expect to spend time debugging RSA key formatting.</p>
</li>
</ol>
<h2>Try It</h2>
<p>CodeGuardAI is live and free for public repos:</p>
<ul>
<li><p><strong>Install:</strong> <a href="https://github.com/apps/codeguard-ai">github.com/apps/codeguard-ai</a></p>
</li>
<li><p><strong>Landing page:</strong> <a href="https://codeguard-ai.nopkt.com">codeguard-ai.nopkt.com</a></p>
</li>
</ul>
<p>Install it, open a PR, and watch it work. The whole thing is ~600 lines of TypeScript running on the edge.</p>
<hr />
<p>If you're building AI-powered developer tools, I'd love to hear what you're working on. The combination of LLMs + GitHub webhooks + edge computing is wildly underexplored. We're just scratching the surface.</p>
<p><em>Built with Claude, Hono, Cloudflare Workers, and D1. Deployed in under 300ms globally.</em></p>
]]></content:encoded></item><item><title><![CDATA[RAG is Not Dead: Advanced Retrieval Patterns That Actually Work in 2026]]></title><description><![CDATA[Every few months, someone declares RAG (Retrieval-Augmented Generation) dead. "Just use a million-token context window," they say. "Fine-tune instead," others suggest.
They're wrong. RAG isn't dead — ]]></description><link>https://younggao.hashnode.dev/rag-is-not-dead-advanced-retrieval-patterns-that-actually-work-in-2026</link><guid isPermaLink="true">https://younggao.hashnode.dev/rag-is-not-dead-advanced-retrieval-patterns-that-actually-work-in-2026</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 08:05:14 GMT</pubDate><content:encoded><![CDATA[<p>Every few months, someone declares RAG (Retrieval-Augmented Generation) dead. "Just use a million-token context window," they say. "Fine-tune instead," others suggest.</p>
<p>They're wrong. RAG isn't dead — naive RAG is dead. The pattern of "chunk documents → embed → cosine similarity → stuff into prompt" was always a prototype, not a production system. In 2026, production RAG looks radically different.</p>
<p>This article covers the patterns that separate toy demos from systems that actually work.</p>
<h2>Why Naive RAG Fails</h2>
<p>The classic RAG pipeline has predictable failure modes:</p>
<ol>
<li><p><strong>Chunking destroys context</strong> — Splitting at 512 tokens breaks paragraphs, separates questions from answers, and loses document structure</p>
</li>
<li><p><strong>Embedding similarity ≠ relevance</strong> — "How do I reset my password?" and "Password reset policy" have high similarity but serve different intents</p>
</li>
<li><p><strong>Top-K retrieval is crude</strong> — The 5 most similar chunks aren't necessarily the 5 most useful</p>
</li>
<li><p><strong>No query understanding</strong> — The raw user query goes straight to vector search with no transformation</p>
</li>
</ol>
<p>Let's fix each of these.</p>
<h2>Pattern 1: Semantic Chunking</h2>
<p>Instead of fixed-size chunks, split at semantic boundaries:</p>
<pre><code class="language-python">from langchain_experimental.text_splitter import SemanticChunker
from langchain_openai import OpenAIEmbeddings

embeddings = OpenAIEmbeddings(model="text-embedding-3-small")

chunker = SemanticChunker(
    embeddings,
    breakpoint_threshold_type="percentile",
    breakpoint_threshold_amount=90,
)

chunks = chunker.split_text(document_text)
</code></pre>
<p>The semantic chunker computes embeddings for each sentence, then splits where the cosine distance between consecutive sentences exceeds a threshold. Sentences about the same topic stay together.</p>
<h3>Contextual Retrieval (Anthropic's Approach)</h3>
<p>Prepend each chunk with context about where it fits in the document:</p>
<pre><code class="language-python">def add_context(chunk: str, full_document: str) -&gt; str:
    """Use an LLM to generate context for each chunk."""
    prompt = f"""Given this document:
{full_document[:2000]}

And this specific chunk:
{chunk}

Write a 2-3 sentence context that explains where this chunk 
fits within the overall document. Be specific."""
    
    context = llm.invoke(prompt)
    return f"CONTEXT: {context}\n\n{chunk}"
</code></pre>
<p>This costs more upfront but dramatically improves retrieval accuracy — Anthropic reported 49% fewer retrieval failures.</p>
<h2>Pattern 2: Hybrid Search</h2>
<p>Vector search alone misses exact matches. BM25 alone misses semantic similarity. Combine them:</p>
<pre><code class="language-python">from langchain.retrievers import EnsembleRetriever
from langchain_community.retrievers import BM25Retriever
from langchain_community.vectorstores import Qdrant

# Vector retriever
vector_store = Qdrant.from_documents(
    documents, embeddings,
    url="http://localhost:6333",
    collection_name="docs"
)
vector_retriever = vector_store.as_retriever(search_kwargs={"k": 10})

# BM25 retriever  
bm25_retriever = BM25Retriever.from_documents(documents)
bm25_retriever.k = 10

# Combine with Reciprocal Rank Fusion
ensemble = EnsembleRetriever(
    retrievers=[vector_retriever, bm25_retriever],
    weights=[0.6, 0.4],  # Favor semantic for most use cases
)
</code></pre>
<p>The <code>EnsembleRetriever</code> uses Reciprocal Rank Fusion (RRF) to merge results. A document ranking #1 in vector search and #3 in BM25 scores higher than one ranking #2 in both.</p>
<h3>Adding Knowledge Graphs</h3>
<p>For structured relationships (org charts, product hierarchies, dependency trees), add a graph layer:</p>
<pre><code class="language-python">from neo4j import GraphDatabase

def graph_enhanced_retrieval(query: str, vector_results: list) -&gt; list:
    """Enrich vector results with graph context."""
    driver = GraphDatabase.driver("bolt://localhost:7687")
    
    enriched = []
    for doc in vector_results:
        # Find related entities in the graph
        with driver.session() as session:
            result = session.run("""
                MATCH (n)-[r]-(related)
                WHERE n.name = $entity
                RETURN related.name, type(r), related.description
                LIMIT 5
            """, entity=doc.metadata.get("entity"))
            
            context = [f"{r['type(r)']}: {r['related.name']}" for r in result]
            doc.page_content += f"\n\nRelated: {', '.join(context)}"
            enriched.append(doc)
    
    return enriched
</code></pre>
<h2>Pattern 3: Re-Ranking</h2>
<p>Initial retrieval casts a wide net. Re-ranking uses a cross-encoder to score each (query, document) pair more accurately:</p>
<pre><code class="language-python">from langchain.retrievers import ContextualCompressionRetriever
from langchain_cohere import CohereRerank

# Retrieve 20 candidates, re-rank to top 5
reranker = CohereRerank(
    model="rerank-english-v3.0",
    top_n=5,
)

compression_retriever = ContextualCompressionRetriever(
    base_compressor=reranker,
    base_retriever=ensemble,  # From hybrid search above
)

results = compression_retriever.invoke("How do I configure SSO?")
</code></pre>
<p>Cross-encoders process the query and document together (unlike bi-encoders which encode them separately), enabling much more nuanced relevance scoring. The tradeoff is speed — which is why we re-rank a pre-filtered set rather than the entire corpus.</p>
<h3>ColBERT v2: Best of Both Worlds</h3>
<p>ColBERT stores per-token embeddings and uses late interaction for scoring, giving near cross-encoder accuracy at near bi-encoder speed:</p>
<pre><code class="language-python">from ragatouille import RAGPretrainedModel

rag = RAGPretrainedModel.from_pretrained("colbert-ir/colbertv2.0")
rag.index(
    collection=[doc.page_content for doc in documents],
    document_metadatas=[doc.metadata for doc in documents],
    index_name="my_index",
)

results = rag.search(query="SSO configuration", k=5)
</code></pre>
<h2>Pattern 4: Query Transformation</h2>
<p>Don't send the raw user query to retrieval. Transform it first.</p>
<h3>HyDE (Hypothetical Document Embeddings)</h3>
<p>Generate a hypothetical answer, then search for documents similar to that answer:</p>
<pre><code class="language-python">def hyde_retrieval(query: str) -&gt; list:
    # Generate hypothetical answer
    hypothetical = llm.invoke(
        f"Write a short paragraph that would answer: {query}"
    )
    
    # Search using the hypothetical document's embedding
    # This often matches better than the question embedding
    return vector_store.similarity_search(hypothetical, k=5)
</code></pre>
<h3>Multi-Query Expansion</h3>
<p>Generate multiple perspectives on the same question:</p>
<pre><code class="language-python">from langchain.retrievers.multi_query import MultiQueryRetriever

multi_retriever = MultiQueryRetriever.from_llm(
    retriever=vector_store.as_retriever(),
    llm=llm,
)

# Internally generates 3+ query variants and deduplicates results
results = multi_retriever.invoke("Why is our API slow?")
# Generates: "API performance issues", "latency root causes", 
# "slow response time debugging"
</code></pre>
<h3>Step-Back Prompting</h3>
<p>For specific questions, first ask a broader question:</p>
<pre><code class="language-python">def step_back_retrieval(query: str) -&gt; list:
    # Generate a more general question
    broader = llm.invoke(
        f"Given this specific question: '{query}'\n"
        f"What is a more general question that would help answer it?"
    )
    
    # Retrieve for both specific and general queries
    specific_docs = vector_store.similarity_search(query, k=3)
    general_docs = vector_store.similarity_search(broader, k=3)
    
    return deduplicate(specific_docs + general_docs)
</code></pre>
<h2>Pattern 5: Agentic RAG</h2>
<p>The biggest evolution: let the LLM decide what to retrieve and when.</p>
<pre><code class="language-python">from langgraph.graph import StateGraph, END
from typing import TypedDict, Annotated

class AgentState(TypedDict):
    question: str
    documents: list
    answer: str
    needs_more_info: bool

def retrieve(state: AgentState) -&gt; AgentState:
    docs = retriever.invoke(state["question"])
    return {"documents": docs}

def grade_documents(state: AgentState) -&gt; AgentState:
    """Let the LLM decide if retrieved docs are sufficient."""
    grade = llm.invoke(
        f"Question: {state['question']}\n"
        f"Documents: {state['documents']}\n"
        f"Are these documents sufficient to answer the question? "
        f"Reply YES or NO with a brief reason."
    )
    return {"needs_more_info": "NO" in grade.upper()}

def generate(state: AgentState) -&gt; AgentState:
    answer = llm.invoke(
        f"Answer based on these documents:\n"
        f"{state['documents']}\n\n"
        f"Question: {state['question']}"
    )
    return {"answer": answer}

def rewrite_query(state: AgentState) -&gt; AgentState:
    better_query = llm.invoke(
        f"The following question didn't get good search results: "
        f"'{state['question']}'. Rewrite it for better retrieval."
    )
    return {"question": better_query}

# Build the graph
workflow = StateGraph(AgentState)
workflow.add_node("retrieve", retrieve)
workflow.add_node("grade", grade_documents)
workflow.add_node("generate", generate)
workflow.add_node("rewrite", rewrite_query)

workflow.set_entry_point("retrieve")
workflow.add_edge("retrieve", "grade")
workflow.add_conditional_edges(
    "grade",
    lambda s: "rewrite" if s["needs_more_info"] else "generate",
)
workflow.add_edge("rewrite", "retrieve")
workflow.add_edge("generate", END)

app = workflow.compile()
</code></pre>
<p>This agent retrieves, evaluates quality, rewrites the query if needed, and only generates when it has sufficient context.</p>
<h2>Pattern 6: Evaluation with RAGAS</h2>
<p>You can't improve what you don't measure:</p>
<pre><code class="language-python">from ragas import evaluate
from ragas.metrics import (
    faithfulness,
    answer_relevancy,
    context_precision,
    context_recall,
)

result = evaluate(
    dataset=eval_dataset,  # Questions + ground truth answers
    metrics=[
        faithfulness,       # Is the answer grounded in retrieved context?
        answer_relevancy,   # Does the answer address the question?
        context_precision,  # Are the retrieved docs relevant?
        context_recall,     # Did we retrieve all necessary info?
    ],
)

print(result)
# {'faithfulness': 0.92, 'answer_relevancy': 0.88, 
#  'context_precision': 0.85, 'context_recall': 0.79}
</code></pre>
<p>Track these metrics over time. When you change chunking strategy, embedding model, or retrieval pipeline, you'll know immediately if it helped.</p>
<h2>RAG vs. Fine-Tuning vs. Long Context</h2>
<p>When to use each:</p>
<table>
<thead>
<tr>
<th>Approach</th>
<th>Best For</th>
<th>Limitations</th>
</tr>
</thead>
<tbody><tr>
<td><strong>RAG</strong></td>
<td>Dynamic data, source attribution, cost control</td>
<td>Retrieval quality ceiling</td>
</tr>
<tr>
<td><strong>Fine-tuning</strong></td>
<td>Teaching style/format, specialized domains</td>
<td>Stale data, no source citation</td>
</tr>
<tr>
<td><strong>Long context</strong></td>
<td>Small corpora (&lt;100 docs), one-shot analysis</td>
<td>Cost at scale, attention degradation</td>
</tr>
</tbody></table>
<p>The sweet spot for most production systems: <strong>RAG + selective fine-tuning</strong>. Fine-tune for domain language and response style. Use RAG for up-to-date facts and source attribution.</p>
<h2>Production Tips</h2>
<ol>
<li><p><strong>Cache aggressively</strong> — Cache embeddings, cache LLM re-ranking calls, cache final answers for repeated queries</p>
</li>
<li><p><strong>Stream the answer</strong> — Start generating as soon as retrieval completes; don't wait for re-ranking if latency matters</p>
</li>
<li><p><strong>Monitor retrieval quality</strong> — Log which chunks were retrieved and whether users found answers helpful</p>
</li>
<li><p><strong>Use metadata filters</strong> — Filter by date, department, document type before vector search to reduce noise</p>
</li>
<li><p><strong>Implement fallback</strong> — If RAG confidence is low, fall back to a direct LLM response with a disclaimer</p>
</li>
</ol>
<p>RAG in 2026 is a pipeline engineering challenge, not a simple API call. But get it right, and you have a system that's accurate, attributable, and cost-effective at scale.</p>
<hr />
<p><em>If this article helped you, consider</em> <a href="https://ko-fi.com/gps949"><em>buying me a coffee on Ko-fi</em></a><em>! Follow me for more AI engineering content.</em></p>
]]></content:encoded></item><item><title><![CDATA[Building Your First MCP Server in TypeScript: Give AI Access to Your Data]]></title><description><![CDATA[If you've been in the AI space in 2026, you've heard about MCP. The Model Context Protocol, created by Anthropic and now an open standard, is being called the "USB-C for AI" — a universal way to conne]]></description><link>https://younggao.hashnode.dev/building-your-first-mcp-server-in-typescript-give-ai-access-to-your-data</link><guid isPermaLink="true">https://younggao.hashnode.dev/building-your-first-mcp-server-in-typescript-give-ai-access-to-your-data</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 08:04:44 GMT</pubDate><content:encoded><![CDATA[<p>If you've been in the AI space in 2026, you've heard about MCP. The Model Context Protocol, created by Anthropic and now an open standard, is being called the "USB-C for AI" — a universal way to connect AI models to external tools, data sources, and services.</p>
<p>Before MCP, every AI integration was a custom snowflake. Want Claude to query your database? Write a custom tool. Want GPT to search your docs? Build another integration. MCP changes this by providing a standard protocol that any AI host can use to connect to any data source.</p>
<p>In this tutorial, we'll build a fully functional MCP server in TypeScript from scratch.</p>
<h2>What Is MCP and Why Should You Care?</h2>
<p>MCP follows a client-server architecture:</p>
<ul>
<li><p><strong>Host</strong>: The application the user interacts with (Claude Desktop, VS Code, your custom app)</p>
</li>
<li><p><strong>Client</strong>: A protocol client inside the host that maintains a 1:1 connection with a server</p>
</li>
<li><p><strong>Server</strong>: Your code that exposes tools, resources, and prompts to the AI</p>
</li>
</ul>
<p>The protocol defines three core primitives:</p>
<ol>
<li><p><strong>Tools</strong> — Functions the AI can call (like API endpoints)</p>
</li>
<li><p><strong>Resources</strong> — Data the AI can read (like files or database records)</p>
</li>
<li><p><strong>Prompts</strong> — Reusable prompt templates</p>
</li>
</ol>
<h2>Project Setup</h2>
<pre><code class="language-bash">mkdir my-mcp-server &amp;&amp; cd my-mcp-server
npm init -y
npm install @modelcontextprotocol/sdk zod
npm install -D typescript @types/node
npx tsc --init
</code></pre>
<p>Update <code>tsconfig.json</code>:</p>
<pre><code class="language-json">{
  "compilerOptions": {
    "target": "ES2022",
    "module": "Node16",
    "moduleResolution": "Node16",
    "outDir": "./dist",
    "rootDir": "./src",
    "strict": true,
    "esModuleInterop": true,
    "declaration": true
  },
  "include": ["src/**/*"]
}
</code></pre>
<p>Add to <code>package.json</code>:</p>
<pre><code class="language-json">{
  "type": "module",
  "bin": {
    "my-mcp-server": "./dist/index.js"
  },
  "scripts": {
    "build": "tsc",
    "start": "node dist/index.js"
  }
}
</code></pre>
<h2>Building the Server</h2>
<p>Create <code>src/index.ts</code>:</p>
<pre><code class="language-typescript">#!/usr/bin/env node

import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js";
import { z } from "zod";

const server = new McpServer({
  name: "my-mcp-server",
  version: "1.0.0",
});

// --- TOOL 1: Search files by content ---
server.tool(
  "search_files",
  "Search for files containing a specific text pattern",
  {
    pattern: z.string().describe("The text pattern to search for"),
    directory: z.string().default(".").describe("Directory to search in"),
    fileType: z.string().optional().describe("File extension filter (e.g., 'ts', 'py')"),
  },
  async ({ pattern, directory, fileType }) =&gt; {
    const { execSync } = await import("child_process");
    
    let cmd = `grep -rl "\({pattern}" "\){directory}" --include="*.${fileType || '*'}" 2&gt;/dev/null | head -20`;
    
    try {
      const result = execSync(cmd, { encoding: "utf-8", timeout: 10000 });
      const files = result.trim().split("\n").filter(Boolean);
      
      return {
        content: [
          {
            type: "text" as const,
            text: files.length &gt; 0
              ? `Found \({files.length} files matching "\){pattern}":\n\({files.map(f =&gt; `  - \){f}`).join("\n")}`
              : `No files found matching "\({pattern}" in \){directory}`,
          },
        ],
      };
    } catch {
      return {
        content: [{ type: "text" as const, text: `No matches found for "${pattern}"` }],
      };
    }
  }
);

// --- TOOL 2: Execute SQL query ---
server.tool(
  "query_database",
  "Execute a read-only SQL query against the SQLite database",
  {
    query: z.string().describe("SQL SELECT query to execute"),
    database: z.string().describe("Path to SQLite database file"),
  },
  async ({ query, database }) =&gt; {
    // Safety: only allow SELECT queries
    const normalized = query.trim().toUpperCase();
    if (!normalized.startsWith("SELECT")) {
      return {
        content: [{ type: "text" as const, text: "Error: Only SELECT queries are allowed for safety." }],
        isError: true,
      };
    }

    const { execSync } = await import("child_process");
    
    try {
      const result = execSync(
        `sqlite3 -header -csv "\({database}" "\){query}"`,
        { encoding: "utf-8", timeout: 30000 }
      );
      
      return {
        content: [{ type: "text" as const, text: result || "Query returned no results." }],
      };
    } catch (error: any) {
      return {
        content: [{ type: "text" as const, text: `SQL Error: ${error.message}` }],
        isError: true,
      };
    }
  }
);

// --- TOOL 3: HTTP request tool ---
server.tool(
  "http_request",
  "Make an HTTP GET request to a URL and return the response",
  {
    url: z.string().url().describe("The URL to fetch"),
    headers: z.record(z.string()).optional().describe("Optional headers"),
  },
  async ({ url, headers }) =&gt; {
    try {
      const response = await fetch(url, {
        headers: headers || {},
        signal: AbortSignal.timeout(15000),
      });
      
      const contentType = response.headers.get("content-type") || "";
      const body = contentType.includes("json")
        ? JSON.stringify(await response.json(), null, 2)
        : await response.text();
      
      return {
        content: [{
          type: "text" as const,
          text: `HTTP \({response.status} \){response.statusText}\n\n${body.slice(0, 5000)}`,
        }],
      };
    } catch (error: any) {
      return {
        content: [{ type: "text" as const, text: `Request failed: ${error.message}` }],
        isError: true,
      };
    }
  }
);

// --- RESOURCE: Expose project configuration ---
server.resource(
  "config",
  "file:///config/app",
  async (uri) =&gt; {
    const fs = await import("fs/promises");
    
    try {
      const pkg = JSON.parse(
        await fs.readFile("package.json", "utf-8")
      );
      
      return {
        contents: [
          {
            uri: uri.href,
            mimeType: "application/json",
            text: JSON.stringify({
              name: pkg.name,
              version: pkg.version,
              dependencies: Object.keys(pkg.dependencies || {}),
              scripts: Object.keys(pkg.scripts || {}),
            }, null, 2),
          },
        ],
      };
    } catch {
      return {
        contents: [{
          uri: uri.href,
          mimeType: "text/plain",
          text: "No package.json found in current directory",
        }],
      };
    }
  }
);

// --- Start the server ---
async function main() {
  const transport = new StdioServerTransport();
  await server.connect(transport);
  console.error("MCP server running on stdio");
}

main().catch(console.error);
</code></pre>
<h2>Connecting to Claude Desktop</h2>
<p>Build your server:</p>
<pre><code class="language-bash">npm run build
</code></pre>
<p>Edit <code>~/Library/Application Support/Claude/claude_desktop_config.json</code> (macOS) or <code>%APPDATA%\Claude\claude_desktop_config.json</code> (Windows):</p>
<pre><code class="language-json">{
  "mcpServers": {
    "my-server": {
      "command": "node",
      "args": ["/absolute/path/to/my-mcp-server/dist/index.js"],
      "env": {}
    }
  }
}
</code></pre>
<p>Restart Claude Desktop. You should see a hammer icon indicating available tools.</p>
<h2>Testing Your Server</h2>
<p>You can test without Claude Desktop using the MCP Inspector:</p>
<pre><code class="language-bash">npx @modelcontextprotocol/inspector node dist/index.js
</code></pre>
<p>This opens a web UI where you can:</p>
<ul>
<li><p>List all available tools and resources</p>
</li>
<li><p>Call tools with custom arguments</p>
</li>
<li><p>View the raw JSON-RPC messages</p>
</li>
</ul>
<p>For programmatic testing:</p>
<pre><code class="language-typescript">import { Client } from "@modelcontextprotocol/sdk/client/index.js";
import { StdioClientTransport } from "@modelcontextprotocol/sdk/client/stdio.js";

const transport = new StdioClientTransport({
  command: "node",
  args: ["dist/index.js"],
});

const client = new Client({ name: "test-client", version: "1.0.0" });
await client.connect(transport);

// List tools
const tools = await client.listTools();
console.log("Available tools:", tools);

// Call a tool
const result = await client.callTool({
  name: "search_files",
  arguments: { pattern: "TODO", directory: ".", fileType: "ts" },
});
console.log("Result:", result);
</code></pre>
<h2>Production Considerations</h2>
<h3>Authentication</h3>
<p>MCP doesn't include authentication in the base protocol. For remote servers, wrap your transport:</p>
<pre><code class="language-typescript">// Validate API key from environment
const apiKey = process.env.MCP_API_KEY;
if (!apiKey) {
  throw new Error("MCP_API_KEY required");
}

// For HTTP transport (SSE), add auth middleware
server.tool("protected_tool", "...", {}, async (args, extra) =&gt; {
  // The 2026 MCP roadmap includes OAuth 2.1 support
  // For now, use environment-based auth
  // ...
});
</code></pre>
<h3>Error Handling</h3>
<p>Always return structured errors:</p>
<pre><code class="language-typescript">server.tool("risky_operation", "...", { input: z.string() }, async ({ input }) =&gt; {
  try {
    const result = await doSomethingRisky(input);
    return { content: [{ type: "text", text: result }] };
  } catch (error: any) {
    return {
      content: [{ type: "text", text: `Operation failed: ${error.message}` }],
      isError: true, // This tells the AI the tool call failed
    };
  }
});
</code></pre>
<h3>Rate Limiting</h3>
<p>Protect expensive operations:</p>
<pre><code class="language-typescript">const callCounts = new Map&lt;string, { count: number; resetAt: number }&gt;();

function rateLimit(toolName: string, maxPerMinute: number): boolean {
  const now = Date.now();
  const entry = callCounts.get(toolName);
  
  if (!entry || now &gt; entry.resetAt) {
    callCounts.set(toolName, { count: 1, resetAt: now + 60000 });
    return true;
  }
  
  if (entry.count &gt;= maxPerMinute) return false;
  entry.count++;
  return true;
}
</code></pre>
<h2>What's Next for MCP?</h2>
<p>The 2026 MCP roadmap includes exciting developments:</p>
<ul>
<li><p><strong>OAuth 2.1 integration</strong> for standardized auth</p>
</li>
<li><p><strong>Streamable HTTP transport</strong> replacing SSE for better reliability</p>
</li>
<li><p><strong>Agent-to-Agent (A2A) protocol</strong> interoperability</p>
</li>
<li><p><strong>Elicitation</strong> — servers can ask users for input mid-operation</p>
</li>
<li><p><strong>Annotations</strong> — rich metadata on tool responses</p>
</li>
</ul>
<p>MCP is rapidly becoming the standard way AI interacts with the world. Building MCP servers now positions you at the forefront of the agentic AI ecosystem.</p>
<p>The complete source code is available as a template you can fork and customize for your own data sources.</p>
<hr />
<p><em>If this article helped you, consider</em> <a href="https://ko-fi.com/gps949"><em>buying me a coffee on Ko-fi</em></a><em>! Follow me for more AI engineering content.</em></p>
]]></content:encoded></item><item><title><![CDATA[Database Migration Strategies That Wont Take Down Production]]></title><description><![CDATA[Every backend engineer has a migration horror story. Maybe it was the ALTER TABLE that locked a 200-million-row table for 47 minutes. Maybe it was the column rename that brought down every API server ]]></description><link>https://younggao.hashnode.dev/database-migration-strategies-that-wont-take-down-production</link><guid isPermaLink="true">https://younggao.hashnode.dev/database-migration-strategies-that-wont-take-down-production</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 07:57:48 GMT</pubDate><content:encoded><![CDATA[<p>Every backend engineer has a migration horror story. Maybe it was the <code>ALTER TABLE</code> that locked a 200-million-row table for 47 minutes. Maybe it was the column rename that brought down every API server simultaneously. Or maybe it was the "quick fix" migration that corrupted data in a way that took three days to untangle.</p>
<p>Database migrations are the most dangerous routine operation in backend engineering. Your code deploys are (hopefully) stateless and reversible. Your database changes are neither. They mutate persistent state that your entire system depends on, and getting them wrong can mean downtime, data loss, or both.</p>
<p>This article covers the patterns, tooling, and discipline required to run migrations in production without breaking things.</p>
<h2>Why Migrations Are Scary</h2>
<p>The fundamental problem is simple: <strong>schema changes and code changes don't deploy atomically</strong>. There is always a window where your running application code and your database schema are out of sync. During a typical rolling deployment:</p>
<ol>
<li><p>Migration runs, changing the schema</p>
</li>
<li><p>Old application code is still running against the new schema</p>
</li>
<li><p>New application code starts rolling out</p>
</li>
<li><p>Both old and new code run simultaneously against the new schema</p>
</li>
<li><p>Old code finishes draining</p>
</li>
</ol>
<p>Steps 2-4 are where things break. If your migration renames a column, old code looking for the old name throws errors. If it drops a column, same problem. If it adds a <code>NOT NULL</code> column without a default, old code inserting rows fails.</p>
<p>The second problem is <strong>locking</strong>. Most <code>ALTER TABLE</code> operations in traditional databases acquire locks that block reads, writes, or both. On a table with millions of rows, a lock held for even a few seconds can cascade into connection pool exhaustion, request timeouts, and a full outage.</p>
<p>The third problem is <strong>irreversibility</strong>. You can roll back a code deploy in seconds. Rolling back a migration that dropped a column means restoring from backup ‚Äî if you even have a recent one that's consistent.</p>
<h2>Zero-Downtime Migration: The Core Principle</h2>
<p>The rule is straightforward: <strong>at every point during the migration and deployment process, all running code must be compatible with the current database schema</strong>.</p>
<p>This means:</p>
<ul>
<li><p>Never rename a column in a single step</p>
</li>
<li><p>Never drop a column that running code still references</p>
</li>
<li><p>Never add a <code>NOT NULL</code> constraint without a default</p>
</li>
<li><p>Never assume the migration and code deploy happen simultaneously</p>
</li>
</ul>
<p>Every migration must be <strong>backward compatible</strong> with the currently deployed code and <strong>forward compatible</strong> with the code about to be deployed.</p>
<h2>The Expand-Contract Pattern</h2>
<p>This is the most important pattern in zero-downtime migrations. Instead of making a breaking change in one step, you split it into three phases:</p>
<h3>Phase 1: Expand</h3>
<p>Add the new structure alongside the old one. Both old and new code work.</p>
<pre><code class="language-sql">-- Migration: Add new column (nullable, so old code can still INSERT)
ALTER TABLE orders ADD COLUMN status_v2 VARCHAR(50);
</code></pre>
<h3>Phase 2: Migrate</h3>
<p>Deploy code that writes to both old and new structures. Backfill existing data.</p>
<pre><code class="language-sql">-- Backfill in batches (more on this later)
UPDATE orders SET status_v2 = status WHERE status_v2 IS NULL
  AND id BETWEEN \(start AND \)end;
</code></pre>
<h3>Phase 3: Contract</h3>
<p>Once all code uses the new structure and all data is migrated, remove the old one.</p>
<pre><code class="language-sql">-- Only after all application code stops reading `status`
ALTER TABLE orders DROP COLUMN status;
ALTER TABLE orders RENAME COLUMN status_v2 TO status;
</code></pre>
<p>Each phase is a separate migration tied to a separate code deploy. Phase 1 goes out with or before the code that starts using the new column. Phase 3 goes out only after you've verified all services have been updated and the old column is truly unused.</p>
<p><strong>Real-world example: renaming a column</strong></p>
<p>Renaming <code>user_name</code> to <code>display_name</code> on a users table with 10M rows:</p>
<pre><code class="language-plaintext">Deploy 1: Migration adds `display_name` column
Deploy 2: Code writes to both columns, reads from `display_name` with fallback to `user_name`
Deploy 3: Backfill script copies remaining data
Deploy 4: Code reads only from `display_name`
Deploy 5: Migration drops `user_name`
</code></pre>
<p>Yes, that's five deploys for a column rename. That's the cost of zero downtime. In practice, deploys 1-2 often ship together, and 4-5 ship together, so it's usually three deploy cycles.</p>
<h2>Handling Large Table Alterations</h2>
<p>On PostgreSQL, many <code>ALTER TABLE</code> operations are fast because they only update catalog metadata:</p>
<pre><code class="language-sql">-- These are ~instant in PostgreSQL, regardless of table size
ALTER TABLE orders ADD COLUMN notes TEXT;
ALTER TABLE orders ALTER COLUMN notes SET DEFAULT '';
ALTER TABLE orders DROP COLUMN old_field;
</code></pre>
<p>But some operations are not:</p>
<pre><code class="language-sql">-- These rewrite the entire table or scan all rows
ALTER TABLE orders ALTER COLUMN amount TYPE BIGINT;  -- full rewrite
ALTER TABLE orders ADD COLUMN verified BOOLEAN NOT NULL DEFAULT true;
-- (instant in PG 11+ but rewrites in older versions)
CREATE INDEX ON orders (customer_id);  -- full table scan
</code></pre>
<p>For type changes on large tables, use the expand-contract pattern: add a new column with the desired type, backfill, switch reads, drop the old column.</p>
<p>For index creation, always use <code>CONCURRENTLY</code>:</p>
<pre><code class="language-sql">-- Blocks writes:
CREATE INDEX idx_orders_customer ON orders (customer_id);

-- Does NOT block writes (takes longer, but safe):
CREATE INDEX CONCURRENTLY idx_orders_customer ON orders (customer_id);
</code></pre>
<p><strong>Important caveat</strong>: <code>CREATE INDEX CONCURRENTLY</code> cannot run inside a transaction. Most migration tools wrap each migration in a transaction by default. You need to either disable that behavior for this specific migration or use a tool that handles it.</p>
<p>In golang-migrate:</p>
<pre><code class="language-sql">-- +migrate: no-transaction
CREATE INDEX CONCURRENTLY idx_orders_customer ON orders (customer_id);
</code></pre>
<p>In Flyway, you would set <code>executeInTransaction=false</code> on the migration.</p>
<h3>Advisory Locks and Migration Safety</h3>
<p>When running migrations in a horizontally scaled environment, you need to ensure only one instance runs migrations at a time. Most tools handle this with advisory locks:</p>
<pre><code class="language-sql">-- golang-migrate uses PostgreSQL advisory locks
SELECT pg_advisory_lock(12345);
-- ... run migrations ...
SELECT pg_advisory_unlock(12345);
</code></pre>
<p>This prevents two pods starting simultaneously from both trying to run the same migration and corrupting the schema history.</p>
<h2>Migration Tooling</h2>
<h3>golang-migrate</h3>
<p>Straightforward, language-agnostic, uses plain SQL files:</p>
<pre><code class="language-plaintext">migrations/
  000001_create_users.up.sql
  000001_create_users.down.sql
  000002_add_orders.up.sql
  000002_add_orders.down.sql
</code></pre>
<p>Run from CLI or embedded in your Go application:</p>
<pre><code class="language-go">import "github.com/golang-migrate/migrate/v4"

m, err := migrate.New(
    "file://migrations",
    "postgres://localhost:5432/mydb?sslmode=disable",
)
if err != nil {
    log.Fatal(err)
}

if err := m.Up(); err != nil &amp;&amp; err != migrate.ErrNoChange {
    log.Fatal(err)
}
</code></pre>
<p>Strengths: simple, no DSL to learn, works with any language's deployment pipeline. Weaknesses: no built-in support for non-transactional migrations, limited state tracking.</p>
<h3>Flyway</h3>
<p>JVM-based, more opinionated, supports versioned and repeatable migrations:</p>
<pre><code class="language-sql">-- V1__Create_users.sql
CREATE TABLE users (
    id BIGSERIAL PRIMARY KEY,
    email VARCHAR(255) NOT NULL UNIQUE,
    created_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);

-- V2__Add_orders.sql
CREATE TABLE orders (
    id BIGSERIAL PRIMARY KEY,
    user_id BIGINT NOT NULL REFERENCES users(id),
    total_cents BIGINT NOT NULL,
    created_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
</code></pre>
<p>Flyway tracks state in a <code>flyway_schema_history</code> table and supports callbacks, placeholders, and Java-based migrations for complex logic. It's the standard in Java/Kotlin ecosystems.</p>
<h3>Other Notable Tools</h3>
<ul>
<li><p><strong>Alembic</strong> (Python/SQLAlchemy): generates migrations from model diffs, good for Python shops</p>
</li>
<li><p><strong>Sqitch</strong>: dependency-based rather than version-ordered, powerful but steeper learning curve</p>
</li>
<li><p><strong>Atlas</strong> (by Ariga): declarative schema management, computes diffs automatically, gaining traction in Go ecosystems</p>
</li>
<li><p><strong>pg_partman / pgloader</strong>: for partition management and bulk data loading, respectively</p>
</li>
</ul>
<h3>Choosing a Tool</h3>
<p>Pick based on your ecosystem and complexity:</p>
<table>
<thead>
<tr>
<th>Need</th>
<th>Tool</th>
</tr>
</thead>
<tbody><tr>
<td>Simple SQL migrations, any language</td>
<td>golang-migrate</td>
</tr>
<tr>
<td>Java/Kotlin ecosystem</td>
<td>Flyway</td>
</tr>
<tr>
<td>Python/SQLAlchemy</td>
<td>Alembic</td>
</tr>
<tr>
<td>Declarative schema-as-code</td>
<td>Atlas</td>
</tr>
<tr>
<td>Complex dependency graphs</td>
<td>Sqitch</td>
</tr>
</tbody></table>
<h2>Data Backfills Done Right</h2>
<p>Backfilling data across millions of rows is where most migration disasters happen. The naive approach:</p>
<pre><code class="language-sql">-- DO NOT DO THIS
UPDATE orders SET status_v2 = compute_new_status(status);
</code></pre>
<p>This acquires a lock on every row, generates enormous WAL (write-ahead log) volume, and can take hours while blocking other writes.</p>
<p>Instead, backfill in batches with throttling:</p>
<pre><code class="language-python">import time
import psycopg2

BATCH_SIZE = 5000
SLEEP_SECONDS = 0.1  # throttle to reduce replication lag

conn = psycopg2.connect(dsn)
conn.autocommit = True

cursor = conn.cursor()
cursor.execute("SELECT MIN(id), MAX(id) FROM orders")
min_id, max_id = cursor.fetchone()

current = min_id
while current &lt;= max_id:
    cursor.execute("""
        UPDATE orders
        SET status_v2 = compute_new_status(status)
        WHERE id &gt;= %s AND id &lt; %s
          AND status_v2 IS NULL
    """, (current, current + BATCH_SIZE))

    updated = cursor.rowcount
    print(f"Batch {current}-{current + BATCH_SIZE}: updated {updated} rows")

    current += BATCH_SIZE
    time.sleep(SLEEP_SECONDS)
</code></pre>
<p>Key practices:</p>
<ul>
<li><p><strong>Batch by primary key range</strong>, not <code>LIMIT/OFFSET</code> (which gets slower as offset grows)</p>
</li>
<li><p><strong>Sleep between batches</strong> to let replicas catch up and avoid saturating I/O</p>
</li>
<li><p><strong>Make it idempotent</strong> (<code>WHERE status_v2 IS NULL</code>) so you can restart safely if it fails midway</p>
</li>
<li><p><strong>Monitor replication lag</strong> during the backfill and pause if it exceeds your threshold</p>
</li>
<li><p><strong>Run during low-traffic windows</strong> when possible, even if it's technically safe at peak</p>
</li>
</ul>
<p>For truly massive tables (billions of rows), consider doing the backfill at the application level: update the new column whenever a row is naturally read or written, and run the batch backfill for the long tail of untouched rows.</p>
<h2>Blue-Green Database Deployments</h2>
<p>Blue-green deployments for databases are harder than for stateless application servers, but the pattern exists and works well for certain scenarios.</p>
<p>The idea: maintain two database schemas (or databases) ‚Äî blue (current) and green (next). Migrate green, point traffic to it, keep blue as a rollback target.</p>
<h3>Schema-Level Blue-Green</h3>
<pre><code class="language-sql">-- Blue schema (current)
CREATE SCHEMA blue;
-- ... tables in blue schema

-- Green schema (next version)
CREATE SCHEMA green;
-- ... migrated tables in green schema

-- Application config points to schema
SET search_path = 'blue';  -- current
SET search_path = 'green'; -- after cutover
</code></pre>
<p>This works when your migration is a large structural change that's hard to do incrementally. You build the new schema, backfill it from the old one, then switch the application's <code>search_path</code>.</p>
<h3>Limitations</h3>
<p>True blue-green for databases requires either:</p>
<ul>
<li><p>A brief write-freeze during cutover (seconds, not minutes)</p>
</li>
<li><p>Dual-write during the transition period</p>
</li>
<li><p>Logical replication between old and new schemas</p>
</li>
</ul>
<p>For most teams, the expand-contract pattern is more practical than blue-green for databases. Reserve blue-green for major version upgrades (e.g., PostgreSQL 14 to 16) where you set up logical replication between the old and new clusters.</p>
<h2>Rollback Strategies</h2>
<h3>Down Migrations</h3>
<p>Every migration tool supports down migrations. Write them:</p>
<pre><code class="language-sql">-- 000005_add_verified_column.up.sql
ALTER TABLE users ADD COLUMN verified BOOLEAN DEFAULT false;

-- 000005_add_verified_column.down.sql
ALTER TABLE users DROP COLUMN verified;
</code></pre>
<p>But understand their limits. A down migration that drops a column <strong>destroys data</strong>. If you added the column, backfilled it over three days, and then need to roll back, that data is gone.</p>
<h3>Forward-Only Rollbacks</h3>
<p>A safer pattern: instead of rolling back the migration, roll forward with a new migration that undoes the change:</p>
<pre><code class="language-plaintext">000005_add_verified_column.up.sql    -- adds column
000006_remove_verified_column.up.sql -- removes it (if needed)
</code></pre>
<p>This keeps the migration history linear and auditable. Many teams mandate forward-only migrations in production while keeping down migrations for local development.</p>
<h3>Point-In-Time Recovery</h3>
<p>For catastrophic failures, you need PITR:</p>
<pre><code class="language-bash"># PostgreSQL continuous archiving
archive_mode = on
archive_command = 'cp %p /archive/%f'

# Restore to a specific timestamp
recovery_target_time = '2025-03-15 14:30:00 UTC'
</code></pre>
<p>Test your PITR process regularly. The worst time to discover your backups don't work is during an incident.</p>
<h3>The Migration Rollback Matrix</h3>
<table>
<thead>
<tr>
<th>Scenario</th>
<th>Strategy</th>
</tr>
</thead>
<tbody><tr>
<td>Added a column, code not yet deployed</td>
<td>Down migration (safe, no data loss)</td>
</tr>
<tr>
<td>Changed column type, data transformed</td>
<td>Forward migration to revert; may lose precision</td>
</tr>
<tr>
<td>Dropped a column</td>
<td>PITR or restore from backup</td>
</tr>
<tr>
<td>Added an index</td>
<td>Drop it (fast, safe)</td>
</tr>
<tr>
<td>Data backfill went wrong</td>
<td>Forward migration to fix; depends on whether original data is preserved</td>
</tr>
</tbody></table>
<h2>Testing Migrations</h2>
<h3>Against Production-Like Data</h3>
<p>Your test database with 50 rows will not reveal the problems that show up with 50 million rows. Test against a copy of production data:</p>
<pre><code class="language-bash"># Snapshot production (use your cloud provider's snapshot feature)
# Restore to a test instance
# Run the migration
# Measure: time, locks held, WAL generated, replication lag
</code></pre>
<h3>Schema Diffing</h3>
<p>After running migrations on a staging environment, diff the resulting schema against what you expect:</p>
<pre><code class="language-bash"># Using pg_dump to compare schemas
pg_dump --schema-only production_db &gt; prod_schema.sql
pg_dump --schema-only staging_db &gt; staging_schema.sql
diff prod_schema.sql staging_schema.sql
</code></pre>
<p>Atlas has this built in:</p>
<pre><code class="language-bash">atlas schema diff \
  --from "postgres://localhost/production" \
  --to "postgres://localhost/staging"
</code></pre>
<h3>CI Pipeline Integration</h3>
<p>Run migrations in CI against a fresh database and a database with the previous version's schema:</p>
<pre><code class="language-yaml"># .github/workflows/migrations.yml
migration-test:
  services:
    postgres:
      image: postgres:16
      env:
        POSTGRES_DB: test
        POSTGRES_PASSWORD: test
  steps:
    - uses: actions/checkout@v4
    - name: Run all migrations from scratch
      run: migrate -path ./migrations -database "$DB_URL" up
    - name: Verify schema matches expectations
      run: ./scripts/verify-schema.sh
    - name: Test rollback of latest migration
      run: migrate -path ./migrations -database "$DB_URL" down 1
    - name: Re-apply latest migration
      run: migrate -path ./migrations -database "$DB_URL" up
</code></pre>
<h3>Statement-Level Analysis</h3>
<p>Before running a migration in production, analyze what it will actually do:</p>
<pre><code class="language-sql">-- Check if an ALTER TABLE will rewrite the table
-- (PostgreSQL-specific: check pg_catalog after the change)

-- Estimate lock duration using pg_stat_activity during staging test
SELECT pid, wait_event_type, wait_event, query
FROM pg_stat_activity
WHERE wait_event_type = 'Lock';
</code></pre>
<h2>A Migration Checklist</h2>
<p>Before every production migration:</p>
<ol>
<li><p><strong>Is it backward compatible?</strong> Can the currently deployed code work with the new schema?</p>
</li>
<li><p><strong>Is it forward compatible?</strong> Can the about-to-be-deployed code work with the old schema (in case you need to roll back the code deploy)?</p>
</li>
<li><p><strong>Does it hold locks?</strong> If so, for how long? On how many rows?</p>
</li>
<li><p><strong>Has it been tested against production-sized data?</strong></p>
</li>
<li><p><strong>Is there a rollback plan?</strong> Down migration, forward fix, or PITR?</p>
</li>
<li><p><strong>Is the backfill batched and idempotent?</strong></p>
</li>
<li><p><strong>Is the migration wrapped in a transaction?</strong> Should it be? (<code>CREATE INDEX CONCURRENTLY</code> cannot be.)</p>
</li>
<li><p><strong>Has someone else reviewed it?</strong> Migration PRs deserve as much scrutiny as application code.</p>
</li>
</ol>
<h2>Summary</h2>
<p>Database migrations will always carry risk. The goal is not to eliminate risk but to reduce it to a level where deploying a migration is routine rather than terrifying. The core practices:</p>
<ul>
<li><p>Use expand-contract for every breaking schema change</p>
</li>
<li><p>Batch and throttle data backfills</p>
</li>
<li><p>Test against production-scale data</p>
</li>
<li><p>Write rollback plans before you need them</p>
</li>
<li><p>Use <code>CONCURRENTLY</code> for index operations</p>
</li>
<li><p>Keep migrations small, frequent, and reviewable</p>
</li>
</ul>
<p>The teams that migrate databases confidently are not the ones with the best tools. They are the ones with the most disciplined processes.</p>
<hr />
<p><em>Building resilient backend systems? This is article #17 in the</em> <a href="https://dev.to/gps949/series/production-backend-patterns"><em>Production Backend Patterns</em></a> <em>series. Follow for the next one.</em></p>
<hr />
<p><em>If you found this useful, consider supporting my work:</em></p>
<p><a href="https://ko-fi.com/gps949"><img src="https://ko-fi.com/img/githubbutton_sm.svg" alt="Ko-fi" style="display:block;margin:0 auto" /></a></p>
]]></content:encoded></item><item><title><![CDATA[Building Production AI Agents with LangGraph: Beyond the Toy Examples]]></title><description><![CDATA[Building Production AI Agents with LangGraph: Beyond the Toy Examples
Every AI tutorial shows you a chatbot that answers questions. That's not an agent. An agent decides what to do, takes action, obse]]></description><link>https://younggao.hashnode.dev/building-production-ai-agents-with-langgraph-beyond-the-toy-examples</link><guid isPermaLink="true">https://younggao.hashnode.dev/building-production-ai-agents-with-langgraph-beyond-the-toy-examples</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 07:47:54 GMT</pubDate><content:encoded><![CDATA[<h1>Building Production AI Agents with LangGraph: Beyond the Toy Examples</h1>
<p>Every AI tutorial shows you a chatbot that answers questions. That's not an agent. An agent <em>decides what to do</em>, <em>takes action</em>, <em>observes the result</em>, and <em>adapts</em>. In production, it does all of that reliably, with audit trails, error recovery, and human oversight.</p>
<p>LangGraph ‚Äî the graph-based orchestration layer from LangChain ‚Äî has quietly become the framework of choice for teams shipping real agents. Uber routes support workflows through it. LinkedIn uses it for internal knowledge agents. Klarna runs customer-facing agents on it at scale.</p>
<p>This article is the guide I wish I had when I moved from prototype to production. We'll build a Research Assistant agent end-to-end, covering every pattern that matters when uptime counts.</p>
<h2>When to Use Agents (and When Not To)</h2>
<p>Before writing a single line of agent code, ask yourself: <strong>does this task require dynamic decision-making?</strong></p>
<p><strong>Use agents when:</strong></p>
<ul>
<li><p>The number of steps is unknown at design time</p>
</li>
<li><p>The task requires selecting from multiple tools based on context</p>
</li>
<li><p>Intermediate results change the execution path</p>
</li>
<li><p>You need autonomous error recovery</p>
</li>
</ul>
<p><strong>Don't use agents when:</strong></p>
<ul>
<li><p>A fixed pipeline (prompt ‚Üí LLM ‚Üí output) solves the problem</p>
</li>
<li><p>You can enumerate all paths in advance (use a simple chain)</p>
</li>
<li><p>Latency budget is under 2 seconds (agents loop; loops are slow)</p>
</li>
<li><p>The cost of a wrong autonomous action is high and you can't add human checkpoints</p>
</li>
</ul>
<p>Agents add complexity. A well-designed chain with structured outputs will outperform a poorly-designed agent every time. Start with the simplest approach that works, then graduate to agents when you hit the wall.</p>
<h2>LangGraph Core Concepts</h2>
<p>LangGraph models agent logic as a <strong>directed graph</strong> where:</p>
<ul>
<li><p><strong>State</strong> is a typed dictionary that flows through the graph</p>
</li>
<li><p><strong>Nodes</strong> are functions that read and write state</p>
</li>
<li><p><strong>Edges</strong> connect nodes (static or conditional)</p>
</li>
<li><p><strong>Conditional edges</strong> inspect state and route to different nodes</p>
</li>
</ul>
<p>Here's the minimal mental model:</p>
<pre><code class="language-python">from langgraph.graph import StateGraph, START, END
from typing import TypedDict, Annotated
from operator import add

class AgentState(TypedDict):
    messages: Annotated[list, add]  # append-only message list
    step_count: int

def process(state: AgentState) -&gt; dict:
    return {"messages": ["processed"], "step_count": state["step_count"] + 1}

def should_continue(state: AgentState) -&gt; str:
    return "end" if state["step_count"] &gt;= 3 else "process"

graph = StateGraph(AgentState)
graph.add_node("process", process)
graph.add_conditional_edges(START, should_continue, {"process": "process", "end": END})
graph.add_conditional_edges("process", should_continue, {"process": "process", "end": END})

app = graph.compile()
result = app.invoke({"messages": [], "step_count": 0})
</code></pre>
<p>The <code>Annotated[list, add]</code> is critical ‚Äî it tells LangGraph to <strong>merge</strong> list returns instead of overwriting. Without it, each node would clobber the previous messages.</p>
<h2>Building the Research Assistant</h2>
<p>Let's build something real: an agent that takes a research question, searches the web, reads and summarizes relevant pages, and produces a structured report. This is the kind of agent companies actually deploy.</p>
<h3>Step 1: Define the State</h3>
<pre><code class="language-python">from typing import TypedDict, Annotated, Literal
from operator import add
from pydantic import BaseModel

class Source(BaseModel):
    url: str
    title: str
    summary: str
    relevance_score: float

class ResearchState(TypedDict):
    question: str
    search_queries: list[str]
    sources: Annotated[list[Source], add]
    draft_report: str
    critique: str
    final_report: str
    iteration: int
    status: str
</code></pre>
<p>I'm using Pydantic models for <code>Source</code> ‚Äî this gives you validation and serialization for free, which matters when you're persisting state to a database.</p>
<h3>Step 2: Define the Nodes</h3>
<pre><code class="language-python">from langchain_openai import ChatOpenAI
from langchain_core.messages import SystemMessage, HumanMessage
from langchain_community.tools.tavily_search import TavilySearchResults

llm = ChatOpenAI(model="gpt-4o", temperature=0)
search_tool = TavilySearchResults(max_results=5)

async def generate_queries(state: ResearchState) -&gt; dict:
    """Turn the research question into targeted search queries."""
    response = await llm.ainvoke([
        SystemMessage(content="Generate 3 specific search queries to research this topic. Return only the queries, one per line."),
        HumanMessage(content=state["question"])
    ])
    queries = [q.strip() for q in response.content.strip().split("\n") if q.strip()]
    return {"search_queries": queries, "status": "searching"}

async def search_web(state: ResearchState) -&gt; dict:
    """Execute searches and collect sources."""
    all_sources = []
    for query in state["search_queries"]:
        results = await search_tool.ainvoke({"query": query})
        for r in results:
            source = Source(
                url=r["url"],
                title=r.get("title", ""),
                summary=r["content"][:500],
                relevance_score=0.0  # scored in next step
            )
            all_sources.append(source)
    return {"sources": all_sources, "status": "analyzing"}

async def write_report(state: ResearchState) -&gt; dict:
    """Synthesize sources into a structured report."""
    source_text = "\n\n".join(
        f"[{s.title}]({s.url})\n{s.summary}" for s in state["sources"]
    )
    response = await llm.ainvoke([
        SystemMessage(content="""Write a detailed research report based on these sources.
Structure: Executive Summary, Key Findings (numbered), Analysis, Conclusion.
Cite sources inline as [1], [2], etc."""),
        HumanMessage(content=f"Question: {state['question']}\n\nSources:\n{source_text}")
    ])
    return {"draft_report": response.content, "status": "reviewing"}

async def critique_report(state: ResearchState) -&gt; dict:
    """Self-critique the draft for gaps and improvements."""
    response = await llm.ainvoke([
        SystemMessage(content="""Review this research report critically. Identify:
1. Factual gaps or unsupported claims
2. Missing perspectives
3. Areas needing more depth
Be specific and actionable. If the report is solid, say "APPROVED"."""),
        HumanMessage(content=state["draft_report"])
    ])
    return {
        "critique": response.content,
        "iteration": state["iteration"] + 1,
        "status": "critiqued"
    }

async def revise_report(state: ResearchState) -&gt; dict:
    """Revise the report based on critique."""
    response = await llm.ainvoke([
        SystemMessage(content="Revise this report to address the critique. Maintain the same structure."),
        HumanMessage(content=f"Report:\n{state['draft_report']}\n\nCritique:\n{state['critique']}")
    ])
    return {"draft_report": response.content, "status": "revised"}

async def finalize(state: ResearchState) -&gt; dict:
    return {"final_report": state["draft_report"], "status": "complete"}
</code></pre>
<h3>Step 3: Wire the Graph</h3>
<pre><code class="language-python">from langgraph.graph import StateGraph, START, END

def route_after_critique(state: ResearchState) -&gt; Literal["revise", "finalize"]:
    if "APPROVED" in state["critique"] or state["iteration"] &gt;= 3:
        return "finalize"
    return "revise"

builder = StateGraph(ResearchState)

# Add nodes
builder.add_node("generate_queries", generate_queries)
builder.add_node("search_web", search_web)
builder.add_node("write_report", write_report)
builder.add_node("critique_report", critique_report)
builder.add_node("revise_report", revise_report)
builder.add_node("finalize", finalize)

# Add edges
builder.add_edge(START, "generate_queries")
builder.add_edge("generate_queries", "search_web")
builder.add_edge("search_web", "write_report")
builder.add_edge("write_report", "critique_report")
builder.add_conditional_edges("critique_report", route_after_critique)
builder.add_edge("revise_report", "critique_report")  # loop back
builder.add_edge("finalize", END)

research_agent = builder.compile()
</code></pre>
<p>Run it:</p>
<pre><code class="language-python">result = await research_agent.ainvoke({
    "question": "What are the most effective strategies for reducing LLM hallucinations in production systems?",
    "search_queries": [],
    "sources": [],
    "draft_report": "",
    "critique": "",
    "final_report": "",
    "iteration": 0,
    "status": "starting"
})
print(result["final_report"])
</code></pre>
<h2>State Management and Persistence</h2>
<p>In production, agents crash. Servers restart. Users close browsers. You need <strong>checkpointing</strong>.</p>
<p>LangGraph has built-in support for persisting state at every step via checkpointers:</p>
<pre><code class="language-python">from langgraph.checkpoint.postgres.aio import AsyncPostgresSaver

DB_URI = "postgresql://user:pass@localhost:5432/agents"

async with AsyncPostgresSaver.from_conn_string(DB_URI) as checkpointer:
    await checkpointer.setup()  # creates tables on first run

    research_agent = builder.compile(checkpointer=checkpointer)

    # Every invocation now saves state after each node
    config = {"configurable": {"thread_id": "research-001"}}
    result = await research_agent.ainvoke(initial_state, config)
</code></pre>
<p>If the process dies mid-execution, restart with the same <code>thread_id</code> and it picks up exactly where it left off:</p>
<pre><code class="language-python"># Resume from last checkpoint
result = await research_agent.ainvoke(None, config)
</code></pre>
<p><strong>Production tip:</strong> Use <code>thread_id</code> as your correlation ID across logging, tracing, and customer support. When a user reports a problem, you can replay the exact state transitions.</p>
<p>For high-throughput systems, the Postgres checkpointer supports connection pooling. For simpler setups, <code>SqliteSaver</code> works fine. For serverless, use the <code>MemorySaver</code> during development but always switch to a durable store before deploying.</p>
<h2>Human-in-the-Loop Patterns</h2>
<p>Fully autonomous agents are a liability in production. The most reliable pattern is <strong>human-on-the-loop</strong>: the agent runs autonomously but pauses at critical decision points.</p>
<p>LangGraph supports this natively with <code>interrupt</code>:</p>
<pre><code class="language-python">from langgraph.types import interrupt, Command

async def write_report(state: ResearchState) -&gt; dict:
    # ... generate draft ...

    # Pause and wait for human approval
    approval = interrupt({
        "question": "Review this draft report. Reply 'approved' or provide feedback.",
        "draft": draft_content
    })

    if approval.lower() != "approved":
        # Human provided feedback ‚Äî use it as critique
        return {"draft_report": draft_content, "critique": approval, "status": "human_feedback"}

    return {"draft_report": draft_content, "status": "approved"}
</code></pre>
<p>On the calling side, you handle the interrupt:</p>
<pre><code class="language-python">config = {"configurable": {"thread_id": "research-001"}}

# First invocation runs until interrupt
result = await research_agent.ainvoke(initial_state, config)

# Agent is now paused. Show draft to user via your UI.
# When user responds:
result = await research_agent.ainvoke(
    Command(resume="approved"),  # or resume="Add more detail about X"
    config
)
</code></pre>
<p>This pattern maps cleanly to web UIs (show a review screen), Slack bots (send a message and wait for reply), or email workflows.</p>
<p><strong>Advanced pattern ‚Äî tiered autonomy:</strong></p>
<pre><code class="language-python">def route_by_confidence(state: ResearchState) -&gt; str:
    confidence = state.get("confidence_score", 0)
    if confidence &gt; 0.9:
        return "auto_approve"     # agent proceeds
    elif confidence &gt; 0.7:
        return "notify_human"     # agent proceeds but flags for review
    else:
        return "require_approval" # agent pauses
</code></pre>
<p>This lets low-risk actions flow through while escalating uncertain ones ‚Äî the sweet spot for production throughput.</p>
<h2>Tool Calling Best Practices</h2>
<p>Tools are how agents interact with the real world. Get this wrong and you get agents that burn API credits, leak data, or take destructive actions.</p>
<h3>Structured tool definitions</h3>
<pre><code class="language-python">from langchain_core.tools import tool
from pydantic import Field

@tool
def search_knowledge_base(
    query: str = Field(description="Natural language search query"),
    filters: dict | None = Field(default=None, description="Optional metadata filters: {department: str, date_range: str}"),
    max_results: int = Field(default=10, ge=1, le=50, description="Number of results to return")
) -&gt; list[dict]:
    """Search the internal knowledge base for documents matching the query.
    Use this for company-specific information. For general web information, use web_search instead."""
    # implementation
    ...
</code></pre>
<p><strong>Key practices:</strong></p>
<ol>
<li><p><strong>Rich descriptions matter more than you think.</strong> The LLM reads the docstring and field descriptions to decide when and how to call the tool. Vague descriptions lead to wrong tool selection.</p>
</li>
<li><p><strong>Constrain inputs.</strong> Use <code>ge</code>, <code>le</code>, enums, and Pydantic validators. An agent that can pass <code>max_results=10000</code> <em>will</em> eventually do it.</p>
</li>
<li><p><strong>Separate read and write tools.</strong> Never have a single <code>database_tool</code> that can both query and delete. Give the agent <code>db_query</code> and <code>db_delete</code> separately, and only bind <code>db_delete</code> when you've added human approval.</p>
</li>
<li><p><strong>Tool result formatting.</strong> Return structured data, not free text. The LLM processes structured results more reliably:</p>
</li>
</ol>
<pre><code class="language-python">@tool
def get_order_status(order_id: str) -&gt; dict:
    """Look up the status of a customer order."""
    order = db.get_order(order_id)
    return {
        "order_id": order.id,
        "status": order.status,
        "items_count": len(order.items),
        "estimated_delivery": order.eta.isoformat(),
        "action_available": ["cancel"] if order.status == "processing" else []
    }
</code></pre>
<ol>
<li><strong>Bind tools selectively per node.</strong> Not every node needs every tool:</li>
</ol>
<pre><code class="language-python">research_llm = llm.bind_tools([search_tool, scrape_tool])
writing_llm = llm.bind_tools([])  # no tools during writing
</code></pre>
<h2>Error Handling and Retry Strategies</h2>
<p>Production agents face three categories of failures:</p>
<h3>1. Transient failures (API timeouts, rate limits)</h3>
<p>Use LangGraph's built-in retry policy:</p>
<pre><code class="language-python">from langgraph.pregel import RetryPolicy

builder.add_node(
    "search_web",
    search_web,
    retry=RetryPolicy(
        max_attempts=3,
        initial_interval=1.0,  # seconds
        backoff_factor=2.0,
        retry_on=(TimeoutError, RateLimitError)
    )
)
</code></pre>
<h3>2. LLM failures (malformed output, hallucinated tool calls)</h3>
<p>Wrap tool execution with validation:</p>
<pre><code class="language-python">async def safe_tool_executor(state: AgentState) -&gt; dict:
    last_message = state["messages"][-1]

    for tool_call in last_message.tool_calls:
        try:
            # Validate tool exists
            tool = tool_map.get(tool_call["name"])
            if not tool:
                return {"messages": [ToolMessage(
                    content=f"Tool '{tool_call['name']}' does not exist. Available: {list(tool_map.keys())}",
                    tool_call_id=tool_call["id"]
                )]}

            # Execute with timeout
            result = await asyncio.wait_for(
                tool.ainvoke(tool_call["args"]),
                timeout=30.0
            )
            return {"messages": [ToolMessage(content=str(result), tool_call_id=tool_call["id"])]}

        except ValidationError as e:
            return {"messages": [ToolMessage(
                content=f"Invalid arguments: {e}. Please fix and retry.",
                tool_call_id=tool_call["id"]
            )]}
</code></pre>
<p>The agent sees the error message and self-corrects on the next iteration. This works surprisingly well ‚Äî LLMs are good at fixing their own mistakes when given clear error messages.</p>
<h3>3. Logical failures (infinite loops, stuck states)</h3>
<p>Guard against these at the graph level:</p>
<pre><code class="language-python">def route_after_critique(state: ResearchState) -&gt; str:
    # Hard cap on iterations
    if state["iteration"] &gt;= 3:
        return "finalize"

    # Detect stuck state: same critique twice
    if state.get("prev_critique") == state["critique"]:
        return "finalize"

    return "revise"
</code></pre>
<p>Also set a global timeout on the entire graph execution:</p>
<pre><code class="language-python">result = await asyncio.wait_for(
    research_agent.ainvoke(initial_state, config),
    timeout=300.0  # 5 minute hard limit
)
</code></pre>
<h2>Observability with LangSmith</h2>
<p>You cannot operate what you cannot see. LangSmith is the observability layer for LangGraph ‚Äî think Datadog for agent workflows.</p>
<p>Setup is two environment variables:</p>
<pre><code class="language-bash">export LANGCHAIN_TRACING_V2=true
export LANGCHAIN_API_KEY=lsv2_...
</code></pre>
<p>Every node execution, tool call, LLM invocation, and state transition is now traced automatically. No code changes required.</p>
<p><strong>What to monitor in production:</strong></p>
<pre><code class="language-python"># Custom metadata for filtering traces
config = {
    "configurable": {"thread_id": "research-001"},
    "metadata": {
        "user_id": "u_12345",
        "environment": "production",
        "agent_version": "2.1.0"
    },
    "tags": ["research", "priority-high"]
}
</code></pre>
<p><strong>Key metrics to track:</strong></p>
<ul>
<li><p><strong>Tokens per task</strong>: Set budgets. A research agent shouldn't exceed 50k tokens per run. Alert if it does.</p>
</li>
<li><p><strong>Iterations per completion</strong>: If your average is climbing, your prompts or critique logic are degrading.</p>
</li>
<li><p><strong>Tool call success rate</strong>: Below 95%? Your tool descriptions need work.</p>
</li>
<li><p><strong>Time to completion</strong>: Set SLOs. p50 under 30s, p99 under 120s.</p>
</li>
<li><p><strong>Human intervention rate</strong>: Track how often agents escalate. Trending up = model or prompt regression. Trending down = your agent is learning (or your thresholds are too loose).</p>
</li>
</ul>
<p>LangSmith also supports <strong>evaluation datasets</strong> ‚Äî curated input/output pairs that you run nightly to catch regressions:</p>
<pre><code class="language-python">from langsmith import Client

client = Client()

# Create a dataset of expected research outputs
dataset = client.create_dataset("research-agent-evals")
client.create_example(
    inputs={"question": "What is retrieval augmented generation?"},
    outputs={"expected_sections": ["Executive Summary", "Key Findings"]},
    dataset_id=dataset.id
)
</code></pre>
<h2>LangGraph vs. CrewAI vs. AutoGen</h2>
<p>The framework landscape has matured significantly. Here's when to use what:</p>
<table>
<thead>
<tr>
<th>Aspect</th>
<th>LangGraph</th>
<th>CrewAI</th>
<th>AutoGen</th>
</tr>
</thead>
<tbody><tr>
<td><strong>Architecture</strong></td>
<td>Graph-based, explicit control flow</td>
<td>Role-based multi-agent</td>
<td>Conversation-based multi-agent</td>
</tr>
<tr>
<td><strong>Best for</strong></td>
<td>Complex workflows, production systems</td>
<td>Team simulation, parallel task delegation</td>
<td>Research, multi-agent debate</td>
</tr>
<tr>
<td><strong>State management</strong></td>
<td>Built-in, typed, persistent</td>
<td>Limited, via shared memory</td>
<td>Conversation history</td>
</tr>
<tr>
<td><strong>Human-in-the-loop</strong></td>
<td>First-class (<code>interrupt</code>)</td>
<td>Basic approval flows</td>
<td>Chat-based intervention</td>
</tr>
<tr>
<td><strong>Observability</strong></td>
<td>LangSmith native</td>
<td>Basic logging</td>
<td>AutoGen Studio</td>
</tr>
<tr>
<td><strong>Learning curve</strong></td>
<td>Moderate (graph concepts)</td>
<td>Low (intuitive role metaphor)</td>
<td>Low-moderate</td>
</tr>
<tr>
<td><strong>Production readiness</strong></td>
<td>High</td>
<td>Medium</td>
<td>Medium</td>
</tr>
</tbody></table>
<p><strong>Choose LangGraph when:</strong></p>
<ul>
<li><p>You need fine-grained control over execution flow</p>
</li>
<li><p>Persistence and checkpointing are requirements</p>
</li>
<li><p>You're building a single agent with complex routing</p>
</li>
<li><p>You need production-grade observability</p>
</li>
</ul>
<p><strong>Choose CrewAI when:</strong></p>
<ul>
<li><p>Your problem naturally decomposes into roles (researcher, writer, reviewer)</p>
</li>
<li><p>You want rapid prototyping of multi-agent systems</p>
</li>
<li><p>Team-based delegation is the core pattern</p>
</li>
</ul>
<p><strong>Choose AutoGen when:</strong></p>
<ul>
<li><p>You're building conversational multi-agent systems</p>
</li>
<li><p>Agents need to debate or negotiate</p>
</li>
<li><p>Research and experimentation are the primary goals</p>
</li>
</ul>
<p><strong>Hybrid approach (what I recommend):</strong> Use LangGraph as the orchestration layer and implement individual "agents" within it as specialized nodes. You get the reliability of graph-based control flow with the flexibility to swap implementations.</p>
<h2>Production Deployment Tips</h2>
<h3>1. Use LangGraph Platform for managed deployment</h3>
<pre><code class="language-json">// langgraph.json
{
    "graphs": {
        "research_agent": "./agent.py:research_agent"
    },
    "dependencies": ["langchain-openai", "tavily-python"],
    "env": ".env"
}
</code></pre>
<pre><code class="language-bash">langgraph dev     # local development server with hot reload
langgraph build   # Docker image for deployment
langgraph deploy  # deploy to LangGraph Cloud
</code></pre>
<p>The platform gives you a REST API, WebSocket streaming, cron triggers, and a built-in task queue ‚Äî eliminating significant infrastructure work.</p>
<h3>2. Streaming for UX</h3>
<p>Never make users stare at a spinner. Stream intermediate state:</p>
<pre><code class="language-python">async for event in research_agent.astream_events(initial_state, config, version="v2"):
    if event["event"] == "on_chat_model_stream":
        # Token-level streaming for the writing step
        print(event["data"]["chunk"].content, end="", flush=True)
    elif event["event"] == "on_chain_end":
        # Node completion events
        node_name = event.get("name", "")
        print(f"\n[Completed: {node_name}]")
</code></pre>
<h3>3. Rate limiting and cost controls</h3>
<pre><code class="language-python">import tiktoken

class TokenBudget:
    def __init__(self, max_tokens: int = 50_000):
        self.max_tokens = max_tokens
        self.used = 0
        self.encoder = tiktoken.encoding_for_model("gpt-4o")

    def check(self, text: str) -&gt; bool:
        tokens = len(self.encoder.encode(text))
        self.used += tokens
        if self.used &gt; self.max_tokens:
            raise TokenBudgetExceeded(f"Used {self.used}/{self.max_tokens} tokens")
        return True
</code></pre>
<p>Wire this into your LLM callbacks. When an agent hits its budget, force it to the finalize step with whatever it has.</p>
<h3>4. Version your prompts</h3>
<p>Never hardcode prompts in your node functions. Use a prompt registry:</p>
<pre><code class="language-python">from langsmith import Client

client = Client()

# Pull versioned prompts from LangSmith Hub
system_prompt = client.pull_prompt("research-agent/critique:v3")
</code></pre>
<p>This lets you A/B test prompts, roll back bad deployments, and track which prompt version produced which outputs.</p>
<h3>5. Graceful degradation</h3>
<p>Build fallback paths into your graph:</p>
<pre><code class="language-python">def route_search_results(state: ResearchState) -&gt; str:
    if not state["sources"]:
        return "fallback_generate"  # LLM generates from knowledge
    if len(state["sources"]) &lt; 3:
        return "search_again"       # try different queries
    return "write_report"           # proceed normally
</code></pre>
<p>An agent that returns a partial result is infinitely more useful than one that throws a 500.</p>
<h2>Wrapping Up</h2>
<p>The gap between an agent demo and a production agent is the same gap between a script and a service ‚Äî error handling, observability, persistence, and operational controls.</p>
<p>LangGraph gives you the primitives to bridge that gap: typed state, persistent checkpoints, conditional routing, human-in-the-loop interrupts, and native observability. It's opinionated enough to prevent common mistakes but flexible enough to model real workflows.</p>
<p>Start with the simplest graph that solves your problem. Add checkpointing on day one ‚Äî you'll thank yourself the first time a process crashes mid-run. Add human approval gates before any destructive action. Monitor token usage religiously. And version everything: prompts, tools, graph topology.</p>
<p>The agents that succeed in production aren't the cleverest ones ‚Äî they're the most predictable ones.</p>
<hr />
<p><em>If this article helped you, consider</em> <a href="https://ko-fi.com/gps949"><em>buying me a coffee on Ko-fi</em></a><em>! Follow me for more AI engineering content.</em></p>
]]></content:encoded></item><item><title><![CDATA[Database Migration Strategies That Won t Take Down Production]]></title><description><![CDATA[Every backend engineer has a migration horror story. Maybe it was the ALTER TABLE that locked a 200-million-row table for 47 minutes. Maybe it was the column rename that brought down every API server ]]></description><link>https://younggao.hashnode.dev/database-migration-strategies-that-won-t-take-down-production</link><guid isPermaLink="true">https://younggao.hashnode.dev/database-migration-strategies-that-won-t-take-down-production</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 07:38:38 GMT</pubDate><content:encoded><![CDATA[<p>Every backend engineer has a migration horror story. Maybe it was the <code>ALTER TABLE</code> that locked a 200-million-row table for 47 minutes. Maybe it was the column rename that brought down every API server simultaneously. Or maybe it was the "quick fix" migration that corrupted data in a way that took three days to untangle.</p>
<p>Database migrations are the most dangerous routine operation in backend engineering. Your code deploys are (hopefully) stateless and reversible. Your database changes are neither. They mutate persistent state that your entire system depends on, and getting them wrong can mean downtime, data loss, or both.</p>
<p>This article covers the patterns, tooling, and discipline required to run migrations in production without breaking things.</p>
<h2>Why Migrations Are Scary</h2>
<p>The fundamental problem is simple: <strong>schema changes and code changes don't deploy atomically</strong>. There is always a window where your running application code and your database schema are out of sync. During a typical rolling deployment:</p>
<ol>
<li><p>Migration runs, changing the schema</p>
</li>
<li><p>Old application code is still running against the new schema</p>
</li>
<li><p>New application code starts rolling out</p>
</li>
<li><p>Both old and new code run simultaneously against the new schema</p>
</li>
<li><p>Old code finishes draining</p>
</li>
</ol>
<p>Steps 2-4 are where things break. If your migration renames a column, old code looking for the old name throws errors. If it drops a column, same problem. If it adds a <code>NOT NULL</code> column without a default, old code inserting rows fails.</p>
<p>The second problem is <strong>locking</strong>. Most <code>ALTER TABLE</code> operations in traditional databases acquire locks that block reads, writes, or both. On a table with millions of rows, a lock held for even a few seconds can cascade into connection pool exhaustion, request timeouts, and a full outage.</p>
<p>The third problem is <strong>irreversibility</strong>. You can roll back a code deploy in seconds. Rolling back a migration that dropped a column means restoring from backup ‚Äî if you even have a recent one that's consistent.</p>
<h2>Zero-Downtime Migration: The Core Principle</h2>
<p>The rule is straightforward: <strong>at every point during the migration and deployment process, all running code must be compatible with the current database schema</strong>.</p>
<p>This means:</p>
<ul>
<li><p>Never rename a column in a single step</p>
</li>
<li><p>Never drop a column that running code still references</p>
</li>
<li><p>Never add a <code>NOT NULL</code> constraint without a default</p>
</li>
<li><p>Never assume the migration and code deploy happen simultaneously</p>
</li>
</ul>
<p>Every migration must be <strong>backward compatible</strong> with the currently deployed code and <strong>forward compatible</strong> with the code about to be deployed.</p>
<h2>The Expand-Contract Pattern</h2>
<p>This is the most important pattern in zero-downtime migrations. Instead of making a breaking change in one step, you split it into three phases:</p>
<h3>Phase 1: Expand</h3>
<p>Add the new structure alongside the old one. Both old and new code work.</p>
<pre><code class="language-sql">-- Migration: Add new column (nullable, so old code can still INSERT)
ALTER TABLE orders ADD COLUMN status_v2 VARCHAR(50);
</code></pre>
<h3>Phase 2: Migrate</h3>
<p>Deploy code that writes to both old and new structures. Backfill existing data.</p>
<pre><code class="language-sql">-- Backfill in batches (more on this later)
UPDATE orders SET status_v2 = status WHERE status_v2 IS NULL
  AND id BETWEEN \(start AND \)end;
</code></pre>
<h3>Phase 3: Contract</h3>
<p>Once all code uses the new structure and all data is migrated, remove the old one.</p>
<pre><code class="language-sql">-- Only after all application code stops reading `status`
ALTER TABLE orders DROP COLUMN status;
ALTER TABLE orders RENAME COLUMN status_v2 TO status;
</code></pre>
<p>Each phase is a separate migration tied to a separate code deploy. Phase 1 goes out with or before the code that starts using the new column. Phase 3 goes out only after you've verified all services have been updated and the old column is truly unused.</p>
<p><strong>Real-world example: renaming a column</strong></p>
<p>Renaming <code>user_name</code> to <code>display_name</code> on a users table with 10M rows:</p>
<pre><code class="language-plaintext">Deploy 1: Migration adds `display_name` column
Deploy 2: Code writes to both columns, reads from `display_name` with fallback to `user_name`
Deploy 3: Backfill script copies remaining data
Deploy 4: Code reads only from `display_name`
Deploy 5: Migration drops `user_name`
</code></pre>
<p>Yes, that's five deploys for a column rename. That's the cost of zero downtime. In practice, deploys 1-2 often ship together, and 4-5 ship together, so it's usually three deploy cycles.</p>
<h2>Handling Large Table Alterations</h2>
<p>On PostgreSQL, many <code>ALTER TABLE</code> operations are fast because they only update catalog metadata:</p>
<pre><code class="language-sql">-- These are ~instant in PostgreSQL, regardless of table size
ALTER TABLE orders ADD COLUMN notes TEXT;
ALTER TABLE orders ALTER COLUMN notes SET DEFAULT '';
ALTER TABLE orders DROP COLUMN old_field;
</code></pre>
<p>But some operations are not:</p>
<pre><code class="language-sql">-- These rewrite the entire table or scan all rows
ALTER TABLE orders ALTER COLUMN amount TYPE BIGINT;  -- full rewrite
ALTER TABLE orders ADD COLUMN verified BOOLEAN NOT NULL DEFAULT true;
-- (instant in PG 11+ but rewrites in older versions)
CREATE INDEX ON orders (customer_id);  -- full table scan
</code></pre>
<p>For type changes on large tables, use the expand-contract pattern: add a new column with the desired type, backfill, switch reads, drop the old column.</p>
<p>For index creation, always use <code>CONCURRENTLY</code>:</p>
<pre><code class="language-sql">-- Blocks writes:
CREATE INDEX idx_orders_customer ON orders (customer_id);

-- Does NOT block writes (takes longer, but safe):
CREATE INDEX CONCURRENTLY idx_orders_customer ON orders (customer_id);
</code></pre>
<p><strong>Important caveat</strong>: <code>CREATE INDEX CONCURRENTLY</code> cannot run inside a transaction. Most migration tools wrap each migration in a transaction by default. You need to either disable that behavior for this specific migration or use a tool that handles it.</p>
<p>In golang-migrate:</p>
<pre><code class="language-sql">-- +migrate: no-transaction
CREATE INDEX CONCURRENTLY idx_orders_customer ON orders (customer_id);
</code></pre>
<p>In Flyway, you would set <code>executeInTransaction=false</code> on the migration.</p>
<h3>Advisory Locks and Migration Safety</h3>
<p>When running migrations in a horizontally scaled environment, you need to ensure only one instance runs migrations at a time. Most tools handle this with advisory locks:</p>
<pre><code class="language-sql">-- golang-migrate uses PostgreSQL advisory locks
SELECT pg_advisory_lock(12345);
-- ... run migrations ...
SELECT pg_advisory_unlock(12345);
</code></pre>
<p>This prevents two pods starting simultaneously from both trying to run the same migration and corrupting the schema history.</p>
<h2>Migration Tooling</h2>
<h3>golang-migrate</h3>
<p>Straightforward, language-agnostic, uses plain SQL files:</p>
<pre><code class="language-plaintext">migrations/
  000001_create_users.up.sql
  000001_create_users.down.sql
  000002_add_orders.up.sql
  000002_add_orders.down.sql
</code></pre>
<p>Run from CLI or embedded in your Go application:</p>
<pre><code class="language-go">import "github.com/golang-migrate/migrate/v4"

m, err := migrate.New(
    "file://migrations",
    "postgres://localhost:5432/mydb?sslmode=disable",
)
if err != nil {
    log.Fatal(err)
}

if err := m.Up(); err != nil &amp;&amp; err != migrate.ErrNoChange {
    log.Fatal(err)
}
</code></pre>
<p>Strengths: simple, no DSL to learn, works with any language's deployment pipeline. Weaknesses: no built-in support for non-transactional migrations, limited state tracking.</p>
<h3>Flyway</h3>
<p>JVM-based, more opinionated, supports versioned and repeatable migrations:</p>
<pre><code class="language-sql">-- V1__Create_users.sql
CREATE TABLE users (
    id BIGSERIAL PRIMARY KEY,
    email VARCHAR(255) NOT NULL UNIQUE,
    created_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);

-- V2__Add_orders.sql
CREATE TABLE orders (
    id BIGSERIAL PRIMARY KEY,
    user_id BIGINT NOT NULL REFERENCES users(id),
    total_cents BIGINT NOT NULL,
    created_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
</code></pre>
<p>Flyway tracks state in a <code>flyway_schema_history</code> table and supports callbacks, placeholders, and Java-based migrations for complex logic. It's the standard in Java/Kotlin ecosystems.</p>
<h3>Other Notable Tools</h3>
<ul>
<li><p><strong>Alembic</strong> (Python/SQLAlchemy): generates migrations from model diffs, good for Python shops</p>
</li>
<li><p><strong>Sqitch</strong>: dependency-based rather than version-ordered, powerful but steeper learning curve</p>
</li>
<li><p><strong>Atlas</strong> (by Ariga): declarative schema management, computes diffs automatically, gaining traction in Go ecosystems</p>
</li>
<li><p><strong>pg_partman / pgloader</strong>: for partition management and bulk data loading, respectively</p>
</li>
</ul>
<h3>Choosing a Tool</h3>
<p>Pick based on your ecosystem and complexity:</p>
<table>
<thead>
<tr>
<th>Need</th>
<th>Tool</th>
</tr>
</thead>
<tbody><tr>
<td>Simple SQL migrations, any language</td>
<td>golang-migrate</td>
</tr>
<tr>
<td>Java/Kotlin ecosystem</td>
<td>Flyway</td>
</tr>
<tr>
<td>Python/SQLAlchemy</td>
<td>Alembic</td>
</tr>
<tr>
<td>Declarative schema-as-code</td>
<td>Atlas</td>
</tr>
<tr>
<td>Complex dependency graphs</td>
<td>Sqitch</td>
</tr>
</tbody></table>
<h2>Data Backfills Done Right</h2>
<p>Backfilling data across millions of rows is where most migration disasters happen. The naive approach:</p>
<pre><code class="language-sql">-- DO NOT DO THIS
UPDATE orders SET status_v2 = compute_new_status(status);
</code></pre>
<p>This acquires a lock on every row, generates enormous WAL (write-ahead log) volume, and can take hours while blocking other writes.</p>
<p>Instead, backfill in batches with throttling:</p>
<pre><code class="language-python">import time
import psycopg2

BATCH_SIZE = 5000
SLEEP_SECONDS = 0.1  # throttle to reduce replication lag

conn = psycopg2.connect(dsn)
conn.autocommit = True

cursor = conn.cursor()
cursor.execute("SELECT MIN(id), MAX(id) FROM orders")
min_id, max_id = cursor.fetchone()

current = min_id
while current &lt;= max_id:
    cursor.execute("""
        UPDATE orders
        SET status_v2 = compute_new_status(status)
        WHERE id &gt;= %s AND id &lt; %s
          AND status_v2 IS NULL
    """, (current, current + BATCH_SIZE))

    updated = cursor.rowcount
    print(f"Batch {current}-{current + BATCH_SIZE}: updated {updated} rows")

    current += BATCH_SIZE
    time.sleep(SLEEP_SECONDS)
</code></pre>
<p>Key practices:</p>
<ul>
<li><p><strong>Batch by primary key range</strong>, not <code>LIMIT/OFFSET</code> (which gets slower as offset grows)</p>
</li>
<li><p><strong>Sleep between batches</strong> to let replicas catch up and avoid saturating I/O</p>
</li>
<li><p><strong>Make it idempotent</strong> (<code>WHERE status_v2 IS NULL</code>) so you can restart safely if it fails midway</p>
</li>
<li><p><strong>Monitor replication lag</strong> during the backfill and pause if it exceeds your threshold</p>
</li>
<li><p><strong>Run during low-traffic windows</strong> when possible, even if it's technically safe at peak</p>
</li>
</ul>
<p>For truly massive tables (billions of rows), consider doing the backfill at the application level: update the new column whenever a row is naturally read or written, and run the batch backfill for the long tail of untouched rows.</p>
<h2>Blue-Green Database Deployments</h2>
<p>Blue-green deployments for databases are harder than for stateless application servers, but the pattern exists and works well for certain scenarios.</p>
<p>The idea: maintain two database schemas (or databases) ‚Äî blue (current) and green (next). Migrate green, point traffic to it, keep blue as a rollback target.</p>
<h3>Schema-Level Blue-Green</h3>
<pre><code class="language-sql">-- Blue schema (current)
CREATE SCHEMA blue;
-- ... tables in blue schema

-- Green schema (next version)
CREATE SCHEMA green;
-- ... migrated tables in green schema

-- Application config points to schema
SET search_path = 'blue';  -- current
SET search_path = 'green'; -- after cutover
</code></pre>
<p>This works when your migration is a large structural change that's hard to do incrementally. You build the new schema, backfill it from the old one, then switch the application's <code>search_path</code>.</p>
<h3>Limitations</h3>
<p>True blue-green for databases requires either:</p>
<ul>
<li><p>A brief write-freeze during cutover (seconds, not minutes)</p>
</li>
<li><p>Dual-write during the transition period</p>
</li>
<li><p>Logical replication between old and new schemas</p>
</li>
</ul>
<p>For most teams, the expand-contract pattern is more practical than blue-green for databases. Reserve blue-green for major version upgrades (e.g., PostgreSQL 14 to 16) where you set up logical replication between the old and new clusters.</p>
<h2>Rollback Strategies</h2>
<h3>Down Migrations</h3>
<p>Every migration tool supports down migrations. Write them:</p>
<pre><code class="language-sql">-- 000005_add_verified_column.up.sql
ALTER TABLE users ADD COLUMN verified BOOLEAN DEFAULT false;

-- 000005_add_verified_column.down.sql
ALTER TABLE users DROP COLUMN verified;
</code></pre>
<p>But understand their limits. A down migration that drops a column <strong>destroys data</strong>. If you added the column, backfilled it over three days, and then need to roll back, that data is gone.</p>
<h3>Forward-Only Rollbacks</h3>
<p>A safer pattern: instead of rolling back the migration, roll forward with a new migration that undoes the change:</p>
<pre><code class="language-plaintext">000005_add_verified_column.up.sql    -- adds column
000006_remove_verified_column.up.sql -- removes it (if needed)
</code></pre>
<p>This keeps the migration history linear and auditable. Many teams mandate forward-only migrations in production while keeping down migrations for local development.</p>
<h3>Point-In-Time Recovery</h3>
<p>For catastrophic failures, you need PITR:</p>
<pre><code class="language-bash"># PostgreSQL continuous archiving
archive_mode = on
archive_command = 'cp %p /archive/%f'

# Restore to a specific timestamp
recovery_target_time = '2025-03-15 14:30:00 UTC'
</code></pre>
<p>Test your PITR process regularly. The worst time to discover your backups don't work is during an incident.</p>
<h3>The Migration Rollback Matrix</h3>
<table>
<thead>
<tr>
<th>Scenario</th>
<th>Strategy</th>
</tr>
</thead>
<tbody><tr>
<td>Added a column, code not yet deployed</td>
<td>Down migration (safe, no data loss)</td>
</tr>
<tr>
<td>Changed column type, data transformed</td>
<td>Forward migration to revert; may lose precision</td>
</tr>
<tr>
<td>Dropped a column</td>
<td>PITR or restore from backup</td>
</tr>
<tr>
<td>Added an index</td>
<td>Drop it (fast, safe)</td>
</tr>
<tr>
<td>Data backfill went wrong</td>
<td>Forward migration to fix; depends on whether original data is preserved</td>
</tr>
</tbody></table>
<h2>Testing Migrations</h2>
<h3>Against Production-Like Data</h3>
<p>Your test database with 50 rows will not reveal the problems that show up with 50 million rows. Test against a copy of production data:</p>
<pre><code class="language-bash"># Snapshot production (use your cloud provider's snapshot feature)
# Restore to a test instance
# Run the migration
# Measure: time, locks held, WAL generated, replication lag
</code></pre>
<h3>Schema Diffing</h3>
<p>After running migrations on a staging environment, diff the resulting schema against what you expect:</p>
<pre><code class="language-bash"># Using pg_dump to compare schemas
pg_dump --schema-only production_db &gt; prod_schema.sql
pg_dump --schema-only staging_db &gt; staging_schema.sql
diff prod_schema.sql staging_schema.sql
</code></pre>
<p>Atlas has this built in:</p>
<pre><code class="language-bash">atlas schema diff \
  --from "postgres://localhost/production" \
  --to "postgres://localhost/staging"
</code></pre>
<h3>CI Pipeline Integration</h3>
<p>Run migrations in CI against a fresh database and a database with the previous version's schema:</p>
<pre><code class="language-yaml"># .github/workflows/migrations.yml
migration-test:
  services:
    postgres:
      image: postgres:16
      env:
        POSTGRES_DB: test
        POSTGRES_PASSWORD: test
  steps:
    - uses: actions/checkout@v4
    - name: Run all migrations from scratch
      run: migrate -path ./migrations -database "$DB_URL" up
    - name: Verify schema matches expectations
      run: ./scripts/verify-schema.sh
    - name: Test rollback of latest migration
      run: migrate -path ./migrations -database "$DB_URL" down 1
    - name: Re-apply latest migration
      run: migrate -path ./migrations -database "$DB_URL" up
</code></pre>
<h3>Statement-Level Analysis</h3>
<p>Before running a migration in production, analyze what it will actually do:</p>
<pre><code class="language-sql">-- Check if an ALTER TABLE will rewrite the table
-- (PostgreSQL-specific: check pg_catalog after the change)

-- Estimate lock duration using pg_stat_activity during staging test
SELECT pid, wait_event_type, wait_event, query
FROM pg_stat_activity
WHERE wait_event_type = 'Lock';
</code></pre>
<h2>A Migration Checklist</h2>
<p>Before every production migration:</p>
<ol>
<li><p><strong>Is it backward compatible?</strong> Can the currently deployed code work with the new schema?</p>
</li>
<li><p><strong>Is it forward compatible?</strong> Can the about-to-be-deployed code work with the old schema (in case you need to roll back the code deploy)?</p>
</li>
<li><p><strong>Does it hold locks?</strong> If so, for how long? On how many rows?</p>
</li>
<li><p><strong>Has it been tested against production-sized data?</strong></p>
</li>
<li><p><strong>Is there a rollback plan?</strong> Down migration, forward fix, or PITR?</p>
</li>
<li><p><strong>Is the backfill batched and idempotent?</strong></p>
</li>
<li><p><strong>Is the migration wrapped in a transaction?</strong> Should it be? (<code>CREATE INDEX CONCURRENTLY</code> cannot be.)</p>
</li>
<li><p><strong>Has someone else reviewed it?</strong> Migration PRs deserve as much scrutiny as application code.</p>
</li>
</ol>
<h2>Summary</h2>
<p>Database migrations will always carry risk. The goal is not to eliminate risk but to reduce it to a level where deploying a migration is routine rather than terrifying. The core practices:</p>
<ul>
<li><p>Use expand-contract for every breaking schema change</p>
</li>
<li><p>Batch and throttle data backfills</p>
</li>
<li><p>Test against production-scale data</p>
</li>
<li><p>Write rollback plans before you need them</p>
</li>
<li><p>Use <code>CONCURRENTLY</code> for index operations</p>
</li>
<li><p>Keep migrations small, frequent, and reviewable</p>
</li>
</ul>
<p>The teams that migrate databases confidently are not the ones with the best tools. They are the ones with the most disciplined processes.</p>
<hr />
<p><em>Building resilient backend systems? This is article #17 in the</em> <a href="https://dev.to/gps949/series/production-backend-patterns"><em>Production Backend Patterns</em></a> <em>series. Follow for the next one.</em></p>
<hr />
<p><em>If you found this useful, consider supporting my work:</em></p>
<p><a href="https://ko-fi.com/gps949"><img src="https://ko-fi.com/img/githubbutton_sm.svg" alt="Ko-fi" style="display:block;margin:0 auto" /></a></p>
]]></content:encoded></item><item><title><![CDATA[Building a Production-Ready Message Queue Consumer in Go]]></title><description><![CDATA[In the previous article, we explored distributed tracing with OpenTelemetry. Today we tackle one of the most critical ‚Äî and most frequently botched ‚Äî pieces of backend infrastructure: the message ]]></description><link>https://younggao.hashnode.dev/building-a-production-ready-message-queue-consumer-in-go</link><guid isPermaLink="true">https://younggao.hashnode.dev/building-a-production-ready-message-queue-consumer-in-go</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 07:33:45 GMT</pubDate><content:encoded><![CDATA[<p>In the previous article, we explored distributed tracing with OpenTelemetry. Today we tackle one of the most critical ‚Äî and most frequently botched ‚Äî pieces of backend infrastructure: the message queue consumer.</p>
<p>Every production system eventually outgrows synchronous request-response. Order processing, email delivery, image resizing, webhook fanout ‚Äî these all belong in a queue. But the consumer side is where things get ugly. Messages arrive out of order. Workers crash mid-processing. Duplicates sneak through. Your "simple consumer" becomes a tangle of retries, panics, and lost messages.</p>
<p>This article builds a production-grade NATS JetStream consumer in Go, piece by piece, covering the patterns that keep it alive at 3 AM.</p>
<h2>Why Message Queues Matter</h2>
<p>Synchronous architectures hit a wall when:</p>
<ul>
<li><p><strong>Latency budgets are tight.</strong> Your API shouldn't block for 30 seconds while a PDF renders.</p>
</li>
<li><p><strong>Failure domains bleed.</strong> A downstream service outage shouldn't take your entire API down.</p>
</li>
<li><p><strong>Load is bursty.</strong> Black Friday traffic shouldn't require Black Friday compute 365 days a year.</p>
</li>
</ul>
<p>Message queues decouple producers from consumers. The producer writes a message and moves on. The consumer processes it at its own pace, retries on failure, and scales independently. This is table stakes for any system that handles real traffic.</p>
<p>We'll use <strong>NATS JetStream</strong> as our concrete example. It's operationally simpler than RabbitMQ or Kafka, supports persistent streams with at-least-once delivery, and has an excellent Go client. The patterns apply universally.</p>
<h2>The Architecture</h2>
<p>Here's what we're building:</p>
<pre><code class="language-plaintext">Producer ‚Üí NATS JetStream Stream ‚Üí Consumer Group
                                      ‚îú‚îÄ‚îÄ Worker Pool (N goroutines)
                                      ‚îú‚îÄ‚îÄ Retry with Exponential Backoff
                                      ‚îú‚îÄ‚îÄ Dead Letter Queue
                                      ‚îú‚îÄ‚îÄ Idempotency Check
                                      ‚îî‚îÄ‚îÄ Metrics + Tracing
</code></pre>
<p>Let's define the core interfaces first:</p>
<pre><code class="language-go">package consumer

import (
    "context"
    "time"
)

// Message represents a queue message with metadata.
type Message struct {
    ID        string
    Subject   string
    Data      []byte
    Headers   map[string]string
    Timestamp time.Time
    Attempt   int
}

// Handler processes a single message. Return an error to trigger retry.
type Handler func(ctx context.Context, msg Message) error

// IdempotencyStore tracks processed message IDs.
type IdempotencyStore interface {
    Exists(ctx context.Context, messageID string) (bool, error)
    Mark(ctx context.Context, messageID string, ttl time.Duration) error
}
</code></pre>
<h2>Designing Idempotent Consumers</h2>
<p>At-least-once delivery means your handler <strong>will</strong> receive duplicates. Network blips, consumer restarts, rebalancing ‚Äî all cause redelivery. Your handler must be idempotent: processing the same message twice produces the same result.</p>
<p>Two strategies work in practice:</p>
<p><strong>1. Deduplication at the consumer level</strong> ‚Äî check a store before processing:</p>
<pre><code class="language-go">type RedisIdempotencyStore struct {
    client *redis.Client
}

func (s *RedisIdempotencyStore) Exists(ctx context.Context, id string) (bool, error) {
    val, err := s.client.Exists(ctx, "idem:"+id).Result()
    if err != nil {
        return false, fmt.Errorf("idempotency check failed: %w", err)
    }
    return val &gt; 0, nil
}

func (s *RedisIdempotencyStore) Mark(ctx context.Context, id string, ttl time.Duration) error {
    return s.client.Set(ctx, "idem:"+id, "1", ttl).Err()
}
</code></pre>
<p><strong>2. Idempotent operations</strong> ‚Äî design your writes so repeats are safe. Use <code>INSERT ... ON CONFLICT DO NOTHING</code>, conditional updates with version checks, or deterministic IDs derived from message content.</p>
<p>The first approach is simpler to bolt on. The second is more robust. Production systems usually combine both: a deduplication layer as a fast path, with idempotent database operations as the safety net.</p>
<p>The TTL on the idempotency key matters. Set it too short, and late redeliveries slip through. Set it too long, and your store grows unbounded. Match it to your stream's max redelivery window ‚Äî typically 24‚Äì72 hours.</p>
<h2>The Consumer Core</h2>
<p>Now let's build the consumer. The <code>Config</code> struct captures every tunable:</p>
<pre><code class="language-go">type Config struct {
    // NATS connection
    NATSUrl    string
    StreamName string
    Subject    string
    Durable    string // consumer group name

    // Concurrency
    WorkerCount int

    // Retry
    MaxRetries     int
    InitialBackoff time.Duration
    MaxBackoff     time.Duration
    BackoffFactor  float64

    // Idempotency
    IdempotencyTTL time.Duration

    // Observability
    ServiceName string
}

func DefaultConfig() Config {
    return Config{
        WorkerCount:    10,
        MaxRetries:     5,
        InitialBackoff: 1 * time.Second,
        MaxBackoff:     60 * time.Second,
        BackoffFactor:  2.0,
        IdempotencyTTL: 24 * time.Hour,
        ServiceName:    "queue-consumer",
    }
}
</code></pre>
<p>The consumer struct ties everything together:</p>
<pre><code class="language-go">type Consumer struct {
    cfg        Config
    js         nats.JetStreamContext
    handler    Handler
    idemStore  IdempotencyStore
    metrics    *Metrics
    tracer     trace.Tracer
    logger     *slog.Logger
    sub        *nats.Subscription
    wg         sync.WaitGroup
    msgCh      chan *nats.Msg
    shutdownCh chan struct{}
}

func New(nc *nats.Conn, cfg Config, handler Handler, idemStore IdempotencyStore) (*Consumer, error) {
    js, err := nc.JetStream()
    if err != nil {
        return nil, fmt.Errorf("jetstream init: %w", err)
    }

    tp := otel.GetTracerProvider()
    meter := otel.GetMeterProvider().Meter(cfg.ServiceName)

    return &amp;Consumer{
        cfg:        cfg,
        js:         js,
        handler:    handler,
        idemStore:  idemStore,
        metrics:    NewMetrics(meter),
        tracer:     tp.Tracer(cfg.ServiceName),
        logger:     slog.Default(),
        msgCh:      make(chan *nats.Msg, cfg.WorkerCount*2),
        shutdownCh: make(chan struct{}),
    }, nil
}
</code></pre>
<h2>Retry with Exponential Backoff</h2>
<p>The retry logic lives in the message processing path. NATS JetStream supports server-side <code>NakWithDelay</code>, which tells the server to redeliver after a specified duration. This is vastly superior to client-side retry loops because it survives consumer restarts:</p>
<pre><code class="language-go">func (c *Consumer) processMessage(natsMsg *nats.Msg) {
    ctx, span := c.tracer.Start(context.Background(), "process_message",
        trace.WithAttributes(
            attribute.String("subject", natsMsg.Subject),
        ),
    )
    defer span.End()

    meta, err := natsMsg.Metadata()
    if err != nil {
        c.logger.Error("failed to read metadata", "error", err)
        natsMsg.Nak()
        return
    }

    attempt := int(meta.NumDelivered)
    msgID := natsMsg.Header.Get("Nats-Msg-Id")
    if msgID == "" {
        msgID = fmt.Sprintf("%s-%d", meta.Stream, meta.Sequence.Stream)
    }

    span.SetAttributes(
        attribute.String("message_id", msgID),
        attribute.Int("attempt", attempt),
    )

    // Idempotency check
    if exists, err := c.idemStore.Exists(ctx, msgID); err != nil {
        c.logger.Error("idempotency check failed", "error", err, "msg_id", msgID)
        span.RecordError(err)
        // Retry ‚Äî the store might be temporarily down
        c.nakWithBackoff(natsMsg, attempt)
        return
    } else if exists {
        c.logger.Debug("duplicate message, skipping", "msg_id", msgID)
        c.metrics.duplicatesSkipped.Add(ctx, 1)
        natsMsg.Ack()
        return
    }

    msg := Message{
        ID:        msgID,
        Subject:   natsMsg.Subject,
        Data:      natsMsg.Data,
        Headers:   flattenHeaders(natsMsg.Header),
        Timestamp: meta.Timestamp,
        Attempt:   attempt,
    }

    // Process
    start := time.Now()
    err = c.handler(ctx, msg)
    duration := time.Since(start)

    c.metrics.processingDuration.Record(ctx, duration.Seconds(),
        metric.WithAttributes(attribute.Bool("success", err == nil)),
    )

    if err != nil {
        span.RecordError(err)
        span.SetStatus(codes.Error, err.Error())
        c.metrics.processingErrors.Add(ctx, 1)

        if attempt &gt;= c.cfg.MaxRetries {
            c.logger.Error("max retries exceeded, sending to DLQ",
                "msg_id", msgID, "error", err, "attempts", attempt,
            )
            c.sendToDLQ(ctx, natsMsg, err)
            natsMsg.Ack() // Ack original to stop redelivery
            return
        }

        c.logger.Warn("processing failed, scheduling retry",
            "msg_id", msgID, "error", err, "attempt", attempt,
        )
        c.nakWithBackoff(natsMsg, attempt)
        return
    }

    // Success ‚Äî mark as processed and ack
    if err := c.idemStore.Mark(ctx, msgID, c.cfg.IdempotencyTTL); err != nil {
        c.logger.Error("failed to mark idempotency", "error", err, "msg_id", msgID)
        // Don't fail the message ‚Äî it was processed successfully.
        // Worst case: a duplicate gets processed again idempotently.
    }

    c.metrics.messagesProcessed.Add(ctx, 1)
    natsMsg.Ack()
}

func (c *Consumer) nakWithBackoff(msg *nats.Msg, attempt int) {
    delay := c.cfg.InitialBackoff
    for i := 1; i &lt; attempt; i++ {
        delay = time.Duration(float64(delay) * c.cfg.BackoffFactor)
        if delay &gt; c.cfg.MaxBackoff {
            delay = c.cfg.MaxBackoff
            break
        }
    }
    msg.NakWithDelay(delay)
}
</code></pre>
<p>Key decisions here:</p>
<ul>
<li><p><strong>Server-side retry via</strong> <code>NakWithDelay</code> ‚Äî the consumer doesn't hold the message during backoff. If this consumer dies, another picks it up.</p>
</li>
<li><p><strong>Ack after DLQ publish</strong> ‚Äî we ack the original message to stop the redelivery loop. The DLQ is the new source of truth.</p>
</li>
<li><p><strong>Idempotency mark happens after success</strong> ‚Äî if the mark fails, the message might be reprocessed. That's fine because operations are idempotent.</p>
</li>
</ul>
<h2>Dead Letter Queue</h2>
<p>Messages that fail all retries go to a separate stream. Your ops team reviews them, fixes the bug, and replays:</p>
<pre><code class="language-go">func (c *Consumer) sendToDLQ(ctx context.Context, original *nats.Msg, processErr error) {
    _, span := c.tracer.Start(ctx, "send_to_dlq")
    defer span.End()

    headers := nats.Header{}
    // Preserve original headers
    for k, v := range original.Header {
        headers[k] = v
    }
    headers.Set("X-DLQ-Error", processErr.Error())
    headers.Set("X-DLQ-Timestamp", time.Now().UTC().Format(time.RFC3339))
    headers.Set("X-Original-Subject", original.Subject)

    dlqMsg := &amp;nats.Msg{
        Subject: fmt.Sprintf("dlq.%s", c.cfg.Subject),
        Data:    original.Data,
        Header:  headers,
    }

    if _, err := c.js.PublishMsg(dlqMsg); err != nil {
        c.logger.Error("failed to publish to DLQ", "error", err)
        span.RecordError(err)
        c.metrics.dlqFailures.Add(ctx, 1)
        // This is bad. The message will be lost after ack.
        // In practice, also log the full message body for manual recovery.
        c.logger.Error("LOST MESSAGE ‚Äî manual recovery required",
            "subject", original.Subject,
            "data", string(original.Data),
        )
    }
    c.metrics.dlqMessages.Add(ctx, 1)
}
</code></pre>
<p>Create the DLQ stream during setup:</p>
<pre><code class="language-go">func EnsureStreams(js nats.JetStreamContext, streamName, subject string) error {
    // Main stream
    _, err := js.AddStream(&amp;nats.StreamConfig{
        Name:      streamName,
        Subjects:  []string{subject},
        Retention: nats.WorkQueuePolicy,
        MaxAge:    72 * time.Hour,
    })
    if err != nil {
        return fmt.Errorf("create main stream: %w", err)
    }

    // DLQ stream
    _, err = js.AddStream(&amp;nats.StreamConfig{
        Name:      streamName + "_dlq",
        Subjects:  []string{"dlq." + subject},
        Retention: nats.LimitsPolicy,
        MaxAge:    30 * 24 * time.Hour, // Keep DLQ messages for 30 days
    })
    if err != nil {
        return fmt.Errorf("create DLQ stream: %w", err)
    }
    return nil
}
</code></pre>
<h2>Concurrency Control with Worker Pools</h2>
<p>A single goroutine consuming messages wastes most of its time waiting on I/O. A worker pool lets you process N messages concurrently while keeping backpressure under control:</p>
<pre><code class="language-go">func (c *Consumer) Start(ctx context.Context) error {
    sub, err := c.js.PullSubscribe(
        c.cfg.Subject,
        c.cfg.Durable,
        nats.ManualAck(),
        nats.AckWait(30*time.Second),
        nats.MaxDeliver(c.cfg.MaxRetries+1),
    )
    if err != nil {
        return fmt.Errorf("subscribe: %w", err)
    }
    c.sub = sub

    // Start workers
    for i := 0; i &lt; c.cfg.WorkerCount; i++ {
        c.wg.Add(1)
        go c.worker(i)
    }

    // Fetch loop ‚Äî pulls messages from NATS and fans out to workers
    c.wg.Add(1)
    go c.fetchLoop(ctx)

    c.logger.Info("consumer started",
        "workers", c.cfg.WorkerCount,
        "subject", c.cfg.Subject,
        "durable", c.cfg.Durable,
    )
    return nil
}

func (c *Consumer) fetchLoop(ctx context.Context) {
    defer c.wg.Done()
    defer close(c.msgCh)

    for {
        select {
        case &lt;-ctx.Done():
            return
        case &lt;-c.shutdownCh:
            return
        default:
        }

        msgs, err := c.sub.Fetch(c.cfg.WorkerCount, nats.MaxWait(5*time.Second))
        if err != nil {
            if errors.Is(err, nats.ErrTimeout) {
                continue // No messages available, poll again
            }
            c.logger.Error("fetch error", "error", err)
            time.Sleep(time.Second) // Back off on unexpected errors
            continue
        }

        for _, msg := range msgs {
            select {
            case c.msgCh &lt;- msg:
            case &lt;-ctx.Done():
                return
            case &lt;-c.shutdownCh:
                return
            }
        }
    }
}

func (c *Consumer) worker(id int) {
    defer c.wg.Done()

    c.logger.Debug("worker started", "worker_id", id)
    for msg := range c.msgCh {
        c.processMessage(msg)
    }
    c.logger.Debug("worker stopped", "worker_id", id)
}
</code></pre>
<p>The <code>Fetch</code> batch size matches the worker count. This keeps the channel buffer from growing unbounded while ensuring every worker stays fed. The <code>MaxWait</code> on <code>Fetch</code> prevents tight-looping when the stream is empty.</p>
<h2>Graceful Shutdown Integration</h2>
<p>Following the patterns from article #14, we integrate with OS signals to drain in-flight work before exiting:</p>
<pre><code class="language-go">func (c *Consumer) Shutdown(ctx context.Context) error {
    c.logger.Info("shutting down consumer...")

    // Signal the fetch loop to stop
    close(c.shutdownCh)

    // Drain the subscription ‚Äî no new messages will be delivered
    if c.sub != nil {
        if err := c.sub.Drain(); err != nil {
            c.logger.Error("subscription drain failed", "error", err)
        }
    }

    // Wait for in-flight messages to complete (or context deadline)
    done := make(chan struct{})
    go func() {
        c.wg.Wait()
        close(done)
    }()

    select {
    case &lt;-done:
        c.logger.Info("consumer shutdown complete")
        return nil
    case &lt;-ctx.Done():
        c.logger.Warn("shutdown timed out, some messages may be redelivered")
        return ctx.Err()
    }
}
</code></pre>
<p>Wire it into <code>main</code>:</p>
<pre><code class="language-go">func main() {
    ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
    defer stop()

    nc, _ := nats.Connect(nats.DefaultURL)
    defer nc.Close()

    cfg := consumer.DefaultConfig()
    cfg.NATSUrl = nats.DefaultURL
    cfg.StreamName = "ORDERS"
    cfg.Subject = "orders.process"
    cfg.Durable = "order-processor"

    store := &amp;RedisIdempotencyStore{client: redis.NewClient(&amp;redis.Options{Addr: "localhost:6379"})}

    c, err := consumer.New(nc, cfg, handleOrder, store)
    if err != nil {
        log.Fatal(err)
    }

    if err := c.Start(ctx); err != nil {
        log.Fatal(err)
    }

    &lt;-ctx.Done()

    shutdownCtx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()
    c.Shutdown(shutdownCtx)
}

func handleOrder(ctx context.Context, msg consumer.Message) error {
    var order Order
    if err := json.Unmarshal(msg.Data, &amp;order); err != nil {
        return fmt.Errorf("unmarshal order: %w", err) // Will retry, but won't help ‚Äî consider a non-retryable error type
    }
    // Process the order...
    return nil
}
</code></pre>
<p>A subtle point: the <code>handleOrder</code> comment hints at an important production refinement. Some errors are <strong>non-retryable</strong> ‚Äî bad JSON, invalid business data, schema mismatches. Retrying won't help. A production consumer should distinguish these:</p>
<pre><code class="language-go">type PermanentError struct {
    Err error
}

func (e *PermanentError) Error() string { return e.Err.Error() }
func (e *PermanentError) Unwrap() error { return e.Err }

// In processMessage, before retry logic:
var permErr *PermanentError
if errors.As(err, &amp;permErr) {
    c.logger.Error("permanent error, sending to DLQ immediately",
        "msg_id", msgID, "error", err,
    )
    c.sendToDLQ(ctx, natsMsg, err)
    natsMsg.Ack()
    return
}
</code></pre>
<h2>Observability: Metrics and Tracing</h2>
<p>The metrics struct captures the four golden signals for a queue consumer:</p>
<pre><code class="language-go">type Metrics struct {
    messagesProcessed metric.Int64Counter
    processingErrors  metric.Int64Counter
    processingDuration metric.Float64Histogram
    duplicatesSkipped metric.Int64Counter
    dlqMessages       metric.Int64Counter
    dlqFailures       metric.Int64Counter
    inflightMessages  metric.Int64UpDownCounter
}

func NewMetrics(meter metric.Meter) *Metrics {
    m := &amp;Metrics{}
    m.messagesProcessed, _ = meter.Int64Counter("consumer.messages.processed",
        metric.WithDescription("Total messages successfully processed"))
    m.processingErrors, _ = meter.Int64Counter("consumer.messages.errors",
        metric.WithDescription("Total processing errors"))
    m.processingDuration, _ = meter.Float64Histogram("consumer.messages.duration_seconds",
        metric.WithDescription("Message processing duration"),
        metric.WithExplicitBucketBoundaries(0.01, 0.05, 0.1, 0.5, 1, 5, 10, 30))
    m.duplicatesSkipped, _ = meter.Int64Counter("consumer.messages.duplicates",
        metric.WithDescription("Duplicate messages skipped"))
    m.dlqMessages, _ = meter.Int64Counter("consumer.dlq.sent",
        metric.WithDescription("Messages sent to DLQ"))
    m.dlqFailures, _ = meter.Int64Counter("consumer.dlq.failures",
        metric.WithDescription("Failed DLQ publishes ‚Äî potential message loss"))
    m.inflightMessages, _ = meter.Int64UpDownCounter("consumer.messages.inflight",
        metric.WithDescription("Currently processing messages"))
    return m
}
</code></pre>
<p>The <code>dlq.failures</code> counter deserves an alert. If this fires, you're losing messages. Set a PagerDuty threshold at &gt; 0.</p>
<p>For tracing, the spans we set in <code>processMessage</code> create a trace per message. To connect producer and consumer traces, propagate the trace context through NATS headers:</p>
<pre><code class="language-go">// Producer side
func publishWithTrace(ctx context.Context, js nats.JetStreamContext, subject string, data []byte) error {
    msg := &amp;nats.Msg{
        Subject: subject,
        Data:    data,
        Header:  nats.Header{},
    }
    otel.GetTextMapPropagator().Inject(ctx, propagation.HeaderCarrier(msg.Header))
    _, err := js.PublishMsg(msg)
    return err
}

// Consumer side ‚Äî in processMessage, replace context.Background():
ctx = otel.GetTextMapPropagator().Extract(ctx, propagation.HeaderCarrier(natsMsg.Header))
ctx, span := c.tracer.Start(ctx, "process_message", ...)
</code></pre>
<p>Now your traces show the full journey: API request ‚Üí publish ‚Üí consume ‚Üí process ‚Üí downstream calls. This is indispensable for debugging latency in async pipelines.</p>
<h2>Testing Strategies</h2>
<p>Queue consumers are notoriously hard to test. Here's a layered approach.</p>
<p><strong>Unit test the handler in isolation</strong> ‚Äî no queue, no infrastructure:</p>
<pre><code class="language-go">func TestHandleOrder_Success(t *testing.T) {
    msg := consumer.Message{
        ID:   "test-1",
        Data: []byte(`{"id":"order-123","amount":99.99}`),
    }

    err := handleOrder(context.Background(), msg)
    assert.NoError(t, err)
    // Assert side effects: database writes, API calls, etc.
}

func TestHandleOrder_InvalidJSON(t *testing.T) {
    msg := consumer.Message{
        ID:   "test-2",
        Data: []byte(`not json`),
    }

    err := handleOrder(context.Background(), msg)
    assert.Error(t, err)

    var permErr *consumer.PermanentError
    assert.True(t, errors.As(err, &amp;permErr), "bad JSON should be a permanent error")
}
</code></pre>
<p><strong>Integration test with an embedded NATS server</strong> ‚Äî tests the full consumer lifecycle:</p>
<pre><code class="language-go">func TestConsumer_ProcessAndAck(t *testing.T) {
    // Start embedded NATS
    srv, _ := server.NewServer(&amp;server.Options{
        Port:      -1,
        JetStream: true,
        StoreDir:  t.TempDir(),
    })
    go srv.Start()
    defer srv.Shutdown()
    srv.ReadyForConnections(5 * time.Second)

    nc, _ := nats.Connect(srv.ClientURL())
    defer nc.Close()

    js, _ := nc.JetStream()
    consumer.EnsureStreams(js, "TEST", "test.subject")

    processed := make(chan string, 1)
    handler := func(ctx context.Context, msg consumer.Message) error {
        processed &lt;- msg.ID
        return nil
    }

    store := &amp;InMemoryIdempotencyStore{seen: map[string]bool{}}
    cfg := consumer.DefaultConfig()
    cfg.StreamName = "TEST"
    cfg.Subject = "test.subject"
    cfg.Durable = "test-consumer"
    cfg.WorkerCount = 1

    c, _ := consumer.New(nc, cfg, handler, store)

    ctx, cancel := context.WithCancel(context.Background())
    defer cancel()
    c.Start(ctx)

    // Publish a message
    js.Publish("test.subject", []byte(`{"key":"value"}`),
        nats.MsgId("msg-001"))

    select {
    case id := &lt;-processed:
        assert.Equal(t, "msg-001", id)
    case &lt;-time.After(5 * time.Second):
        t.Fatal("message not processed within timeout")
    }
}

func TestConsumer_DuplicateSkipped(t *testing.T) {
    // Same setup as above, but pre-mark the message ID
    store := &amp;InMemoryIdempotencyStore{seen: map[string]bool{"msg-001": true}}
    // ... publish msg-001, assert it gets acked but handler is NOT called
}

func TestConsumer_RetryThenDLQ(t *testing.T) {
    // Handler returns error every time.
    // Assert message appears in DLQ stream after MaxRetries.
}
</code></pre>
<p><strong>Chaos test</strong> ‚Äî kill the consumer mid-processing, restart, and verify no messages are lost:</p>
<pre><code class="language-go">func TestConsumer_CrashRecovery(t *testing.T) {
    // 1. Publish 100 messages
    // 2. Start consumer, let it process ~50
    // 3. Cancel context (simulates crash)
    // 4. Start a new consumer with the same durable name
    // 5. Wait for all 100 to be processed (some may be processed twice)
    // 6. Assert: handler was called &gt;= 100 times, all message IDs covered
}
</code></pre>
<p>The in-memory idempotency store for tests:</p>
<pre><code class="language-go">type InMemoryIdempotencyStore struct {
    mu   sync.Mutex
    seen map[string]bool
}

func (s *InMemoryIdempotencyStore) Exists(_ context.Context, id string) (bool, error) {
    s.mu.Lock()
    defer s.mu.Unlock()
    return s.seen[id], nil
}

func (s *InMemoryIdempotencyStore) Mark(_ context.Context, id string, _ time.Duration) error {
    s.mu.Lock()
    defer s.mu.Unlock()
    s.seen[id] = true
    return nil
}
</code></pre>
<h2>Production Checklist</h2>
<p>Before shipping your consumer:</p>
<ul>
<li><p>[ ] <strong>Idempotency</strong> ‚Äî handler is safe to run twice on the same input</p>
</li>
<li><p>[ ] <strong>Permanent vs transient errors</strong> ‚Äî bad data goes straight to DLQ, not through the retry loop</p>
</li>
<li><p>[ ] <strong>Backoff bounds</strong> ‚Äî <code>MaxBackoff</code> prevents retry storms; jitter is even better</p>
</li>
<li><p>[ ] <strong>DLQ monitoring</strong> ‚Äî alert on <code>dlq.sent &gt; 0</code> and page on <code>dlq.failures &gt; 0</code></p>
</li>
<li><p>[ ] <strong>Graceful shutdown</strong> ‚Äî drain subscription, wait for in-flight, then exit</p>
</li>
<li><p>[ ] <strong>Health check</strong> ‚Äî expose <code>/healthz</code> that checks NATS and idempotency store connectivity</p>
</li>
<li><p>[ ] <strong>Consumer lag metric</strong> ‚Äî monitor <code>consumer.messages.pending</code> via NATS admin API</p>
</li>
<li><p>[ ] <strong>Resource limits</strong> ‚Äî bound worker count, channel buffer, and message size</p>
</li>
<li><p>[ ] <strong>Trace propagation</strong> ‚Äî connect producer and consumer spans for end-to-end visibility</p>
</li>
</ul>
<h2>Wrapping Up</h2>
<p>A message queue consumer is deceptively simple to prototype and deceptively hard to get right in production. The patterns in this article ‚Äî idempotent processing, server-side retry with backoff, dead letter queues, bounded worker pools, graceful shutdown, and full observability ‚Äî form a foundation that handles the failures real systems encounter.</p>
<p>The complete consumer is around 300 lines of focused Go code. No frameworks, no magic. Just clear concurrency patterns and deliberate error handling. That's the kind of code you want running at 3 AM.</p>
<p>Next in the series, we'll look at building a production-ready rate limiter using Redis and sliding window counters. Stay tuned.</p>
<hr />
<p><em>If this article helped you, consider</em> <a href="https://ko-fi.com/gps949"><em>buying me a coffee on Ko-fi</em></a><em>! Follow me for more production backend patterns.</em></p>
]]></content:encoded></item><item><title><![CDATA[Distributed Tracing with OpenTelemetry: A Practical Guide for Go Services]]></title><description><![CDATA[You have logs. You have metrics. A request enters your system through the API gateway, hops across five services, and fails somewhere deep in the order processing pipeline. You open Kibana, grep throu]]></description><link>https://younggao.hashnode.dev/distributed-tracing-with-opentelemetry-a-practical-guide-for-go-services</link><guid isPermaLink="true">https://younggao.hashnode.dev/distributed-tracing-with-opentelemetry-a-practical-guide-for-go-services</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 07:15:32 GMT</pubDate><content:encoded><![CDATA[<p>You have logs. You have metrics. A request enters your system through the API gateway, hops across five services, and fails somewhere deep in the order processing pipeline. You open Kibana, grep through thousands of log lines, and spend forty minutes correlating timestamps by hand.</p>
<p>Distributed tracing eliminates that pain. It gives you a single, end-to-end view of a request as it flows through every service in your architecture. And with OpenTelemetry becoming the industry standard, there has never been a better time to wire it in.</p>
<p>This article walks through instrumenting Go services with OpenTelemetry from scratch. No toy examples ‚Äî everything here is production-grade code you can drop into a real system.</p>
<h2>Why Distributed Tracing Matters</h2>
<p>In a monolith, a stack trace tells you everything. In a distributed system, a single user action might touch an API gateway, an auth service, an order service, a payment provider, a notification queue, and a database. When something goes wrong ‚Äî or just gets slow ‚Äî you need to answer: <em>which service, which call, which dependency?</em></p>
<p>Distributed tracing answers this by assigning a <strong>trace ID</strong> to each incoming request and propagating it through every downstream call. Each unit of work within a service becomes a <strong>span</strong>, and spans nest to form a tree that represents the full request lifecycle.</p>
<p>The payoff is immediate:</p>
<ul>
<li><p><strong>Latency diagnosis</strong>: See exactly which service or database call is the bottleneck.</p>
</li>
<li><p><strong>Error attribution</strong>: Know that the 500 came from the payment service's connection pool, not your code.</p>
</li>
<li><p><strong>Dependency mapping</strong>: Visualize how services actually communicate at runtime, not how the architecture diagram says they should.</p>
</li>
<li><p><strong>SLA tracking</strong>: Measure per-endpoint latency distributions across the entire call chain.</p>
</li>
</ul>
<h2>OpenTelemetry in 60 Seconds</h2>
<p>OpenTelemetry (OTel) is a CNCF project that provides a vendor-neutral API, SDK, and set of tools for generating telemetry data. For tracing, the key concepts are:</p>
<ul>
<li><p><strong>TracerProvider</strong>: Factory that creates tracers and manages span processors.</p>
</li>
<li><p><strong>Tracer</strong>: Creates spans within a specific instrumentation scope (usually one per package).</p>
</li>
<li><p><strong>Span</strong>: Represents a unit of work. Has a name, start/end time, attributes, events, and status.</p>
</li>
<li><p><strong>SpanProcessor</strong>: Handles completed spans (batching, exporting).</p>
</li>
<li><p><strong>Exporter</strong>: Sends span data to a backend (Jaeger, Zipkin, OTLP collector).</p>
</li>
<li><p><strong>Propagator</strong>: Serializes/deserializes trace context across process boundaries (HTTP headers, message queues).</p>
</li>
</ul>
<h2>Setting Up the OTel SDK</h2>
<p>Start with the dependencies. We will use OTLP over gRPC as the export protocol, which works with Jaeger, Tempo, Datadog, and any OTLP-compatible backend.</p>
<pre><code class="language-bash">go get go.opentelemetry.io/otel \
       go.opentelemetry.io/otel/sdk \
       go.opentelemetry.io/otel/sdk/trace \
       go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc \
       go.opentelemetry.io/otel/propagation \
       go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp
</code></pre>
<p>Now build the tracer provider. This is your application's tracing backbone ‚Äî initialize it once at startup and shut it down cleanly on exit.</p>
<pre><code class="language-go">package telemetry

import (
    "context"
    "fmt"
    "time"

    "go.opentelemetry.io/otel"
    "go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc"
    "go.opentelemetry.io/otel/propagation"
    "go.opentelemetry.io/otel/sdk/resource"
    sdktrace "go.opentelemetry.io/otel/sdk/trace"
    semconv "go.opentelemetry.io/otel/semconv/v1.24.0"
)

type Config struct {
    ServiceName    string
    ServiceVersion string
    Environment    string
    OTLPEndpoint   string // e.g., "otel-collector:4317"
    SampleRate     float64
}

func InitTracer(ctx context.Context, cfg Config) (shutdown func(context.Context) error, err error) {
    exporter, err := otlptracegrpc.New(ctx,
        otlptracegrpc.WithEndpoint(cfg.OTLPEndpoint),
        otlptracegrpc.WithInsecure(), // Use WithTLSCredentials in production
    )
    if err != nil {
        return nil, fmt.Errorf("creating OTLP exporter: %w", err)
    }

    res, err := resource.Merge(
        resource.Default(),
        resource.NewWithAttributes(
            semconv.SchemaURL,
            semconv.ServiceName(cfg.ServiceName),
            semconv.ServiceVersion(cfg.ServiceVersion),
            semconv.DeploymentEnvironment(cfg.Environment),
        ),
    )
    if err != nil {
        return nil, fmt.Errorf("creating resource: %w", err)
    }

    sampler := sdktrace.ParentBased(
        sdktrace.TraceIDRatioBased(cfg.SampleRate),
    )

    tp := sdktrace.NewTracerProvider(
        sdktrace.WithBatcher(exporter,
            sdktrace.WithMaxQueueSize(2048),
            sdktrace.WithMaxExportBatchSize(512),
            sdktrace.WithBatchTimeout(5*time.Second),
        ),
        sdktrace.WithResource(res),
        sdktrace.WithSampler(sampler),
    )

    otel.SetTracerProvider(tp)
    otel.SetTextMapPropagator(propagation.NewCompositeTextMapPropagator(
        propagation.TraceContext{},
        propagation.Baggage{},
    ))

    return tp.Shutdown, nil
}
</code></pre>
<p>A few deliberate decisions here worth explaining:</p>
<p><code>ParentBased</code> <strong>sampler</strong>: If an incoming request already carries a sampling decision (from an upstream service), we honor it. This prevents broken traces where a parent is sampled but a child is not. The <code>TraceIDRatioBased</code> sampler only applies to root spans ‚Äî requests that originate at this service.</p>
<p><strong>Batch processor tuning</strong>: The defaults are conservative. For high-throughput services (10k+ requests/sec), increase <code>MaxQueueSize</code> and <code>MaxExportBatchSize</code>. The <code>BatchTimeout</code> controls the maximum delay before a batch is flushed, even if it is not full.</p>
<p><strong>Resource attributes</strong>: These tag every span with service identity. Backend UIs use them for filtering and grouping. Always include service name, version, and environment at minimum.</p>
<h2>Wiring It Into main()</h2>
<pre><code class="language-go">func main() {
    ctx := context.Background()

    shutdown, err := telemetry.InitTracer(ctx, telemetry.Config{
        ServiceName:    "order-service",
        ServiceVersion: "1.4.2",
        Environment:    os.Getenv("APP_ENV"),
        OTLPEndpoint:   os.Getenv("OTEL_EXPORTER_OTLP_ENDPOINT"),
        SampleRate:     0.1, // Sample 10% of new traces
    })
    if err != nil {
        log.Fatalf("init tracer: %v", err)
    }
    defer func() {
        // Give the exporter 10 seconds to flush remaining spans
        ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
        defer cancel()
        if err := shutdown(ctx); err != nil {
            log.Printf("tracer shutdown: %v", err)
        }
    }()

    // ... start HTTP server
}
</code></pre>
<p>The deferred shutdown is critical. Without it, spans buffered in the batch processor are lost when the process exits. This pairs well with the graceful shutdown pattern from article #14 in this series.</p>
<h2>Instrumenting HTTP Handlers</h2>
<p>OpenTelemetry provides <code>otelhttp</code>, a middleware that automatically creates spans for incoming HTTP requests and extracts trace context from headers.</p>
<pre><code class="language-go">package main

import (
    "net/http"

    "go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp"
)

func newRouter() http.Handler {
    mux := http.NewServeMux()
    mux.HandleFunc("/orders", handleListOrders)
    mux.HandleFunc("/orders/create", handleCreateOrder)

    // Wrap the entire mux with OTel instrumentation
    return otelhttp.NewHandler(mux, "http-server",
        otelhttp.WithMessageEvents(otelhttp.ReadEvents, otelhttp.WriteEvents),
    )
}
</code></pre>
<p>Every incoming request now generates a span named after the HTTP route, with attributes for method, status code, URL scheme, and request/response size. The middleware also extracts <code>traceparent</code> and <code>tracestate</code> headers automatically ‚Äî this is how trace context arrives from upstream services.</p>
<p>For more granular route naming (avoiding high-cardinality span names like <code>/orders/abc123</code>), use a custom span name formatter:</p>
<pre><code class="language-go">otelhttp.NewHandler(mux, "http-server",
    otelhttp.WithSpanNameFormatter(func(operation string, r *http.Request) string {
        // Use the route pattern, not the actual URL
        return fmt.Sprintf("%s %s", r.Method, r.Pattern)
    }),
)
</code></pre>
<h2>Propagating Context Across Services</h2>
<p>Trace context propagation is where distributed tracing earns its name. When service A calls service B, it must inject the current trace context into the outgoing request headers. Service B then extracts it and continues the same trace.</p>
<p>Wrap your HTTP client with <code>otelhttp.NewTransport</code>:</p>
<pre><code class="language-go">package httpclient

import (
    "net/http"
    "time"

    "go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp"
)

// New returns an HTTP client that propagates trace context
// and creates client spans for every outgoing request.
func New() *http.Client {
    return &amp;http.Client{
        Timeout: 30 * time.Second,
        Transport: otelhttp.NewTransport(
            &amp;http.Transport{
                MaxIdleConns:        100,
                MaxIdleConnsPerHost: 10,
                IdleConnTimeout:     90 * time.Second,
            },
        ),
    }
}
</code></pre>
<p>Now every outgoing HTTP call automatically injects <code>traceparent</code> headers and creates a client-side span. Here is what it looks like in a handler that calls a downstream service:</p>
<pre><code class="language-go">func handleCreateOrder(w http.ResponseWriter, r *http.Request) {
    ctx := r.Context()

    // This span is a child of the HTTP server span created by otelhttp
    order, err := processOrder(ctx, r)
    if err != nil {
        http.Error(w, "failed to process order", http.StatusInternalServerError)
        return
    }

    // The client automatically propagates trace context
    client := httpclient.New()
    req, _ := http.NewRequestWithContext(ctx, "POST",
        "http://payment-service/charge", orderToJSON(order))
    resp, err := client.Do(req)
    // ...
}
</code></pre>
<p>The key detail: <strong>always pass</strong> <code>ctx</code> <strong>through</strong>. The context carries the current span. If you use <code>context.Background()</code> instead of the request context, you break the trace chain.</p>
<h2>Custom Spans and Attributes</h2>
<p>Auto-instrumentation covers HTTP boundaries, but the most valuable tracing data comes from custom spans around your business logic.</p>
<pre><code class="language-go">package order

import (
    "context"
    "fmt"

    "go.opentelemetry.io/otel"
    "go.opentelemetry.io/otel/attribute"
    "go.opentelemetry.io/otel/codes"
    "go.opentelemetry.io/otel/trace"
)

var tracer = otel.Tracer("github.com/yourorg/order-service/internal/order")

func ProcessOrder(ctx context.Context, req CreateOrderRequest) (*Order, error) {
    ctx, span := tracer.Start(ctx, "order.Process",
        trace.WithAttributes(
            attribute.String("order.customer_id", req.CustomerID),
            attribute.Int("order.item_count", len(req.Items)),
            attribute.String("order.currency", req.Currency),
        ),
    )
    defer span.End()

    // Validate inventory
    if err := validateInventory(ctx, req.Items); err != nil {
        span.RecordError(err)
        span.SetStatus(codes.Error, "inventory validation failed")
        return nil, fmt.Errorf("validate inventory: %w", err)
    }

    // Calculate pricing
    total, err := calculateTotal(ctx, req.Items, req.Currency)
    if err != nil {
        span.RecordError(err)
        span.SetStatus(codes.Error, "pricing calculation failed")
        return nil, fmt.Errorf("calculate total: %w", err)
    }
    span.SetAttributes(attribute.Float64("order.total", total))

    // Persist
    order, err := saveOrder(ctx, req, total)
    if err != nil {
        span.RecordError(err)
        span.SetStatus(codes.Error, "persistence failed")
        return nil, fmt.Errorf("save order: %w", err)
    }

    span.SetAttributes(attribute.String("order.id", order.ID))
    return order, nil
}
</code></pre>
<p>Several patterns at work here:</p>
<ol>
<li><p><strong>Tracer per package</strong>: <code>otel.Tracer("...")</code> creates an instrumentation scope. Use the fully qualified package path. This shows up in tracing UIs and helps identify which code produced a span.</p>
</li>
<li><p><strong>Record errors AND set status</strong>: <code>RecordError</code> adds an exception event to the span (with stack trace if available). <code>SetStatus</code> marks the span as failed. Do both ‚Äî some backends use one, some the other.</p>
</li>
<li><p><strong>Add attributes progressively</strong>: You do not need to know all attributes upfront. Add them as the function progresses. The <code>order.id</code> is only available after the database write, so we set it then.</p>
</li>
<li><p><strong>Return the enriched ctx</strong>: <code>tracer.Start</code> returns a new context with the span. Pass this to downstream functions so their spans nest correctly.</p>
</li>
</ol>
<h3>Instrumenting Database Calls</h3>
<p>Database calls are almost always where latency hides. Create spans around them:</p>
<pre><code class="language-go">func saveOrder(ctx context.Context, req CreateOrderRequest, total float64) (*Order, error) {
    ctx, span := tracer.Start(ctx, "order.SaveToDB",
        trace.WithAttributes(
            attribute.String("db.system", "postgresql"),
            attribute.String("db.operation", "INSERT"),
            attribute.String("db.sql.table", "orders"),
        ),
    )
    defer span.End()

    query := `INSERT INTO orders (customer_id, total, currency, status)
              VALUES (\(1, \)2, $3, 'pending') RETURNING id, created_at`

    var order Order
    err := db.QueryRowContext(ctx, query, req.CustomerID, total, req.Currency).
        Scan(&amp;order.ID, &amp;order.CreatedAt)
    if err != nil {
        span.RecordError(err)
        span.SetStatus(codes.Error, "db insert failed")
        return nil, err
    }

    return &amp;order, nil
}
</code></pre>
<p>Use the <a href="https://opentelemetry.io/docs/specs/semconv/database/">OpenTelemetry semantic conventions for databases</a> (<code>db.system</code>, <code>db.operation</code>, <code>db.sql.table</code>) so your tracing backend can render database-specific views. Do <strong>not</strong> put the full SQL query in a span attribute in production ‚Äî it may contain PII.</p>
<h2>Connecting to Jaeger or Zipkin</h2>
<h3>Option 1: OTLP-native (Recommended)</h3>
<p>Modern Jaeger (v1.35+) natively supports OTLP over gRPC on port 4317. Our setup already works ‚Äî just point <code>OTEL_EXPORTER_OTLP_ENDPOINT</code> at <code>jaeger:4317</code>.</p>
<p>For Grafana Tempo, it is the same: OTLP on port 4317.</p>
<h3>Option 2: OpenTelemetry Collector</h3>
<p>For production, run an OTel Collector as a sidecar or daemonset. This decouples your services from the tracing backend and adds capabilities like tail-based sampling, attribute filtering, and multi-backend fan-out.</p>
<pre><code class="language-yaml"># otel-collector-config.yaml
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317

processors:
  batch:
    timeout: 5s
    send_batch_size: 512
  memory_limiter:
    check_interval: 1s
    limit_mib: 512
    spike_limit_mib: 128

exporters:
  otlp/jaeger:
    endpoint: jaeger:4317
    tls:
      insecure: true
  otlp/tempo:
    endpoint: tempo:4317
    tls:
      insecure: true

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [otlp/jaeger, otlp/tempo]
</code></pre>
<p>The collector's <code>memory_limiter</code> processor is essential. Without it, a traffic spike can OOM the collector and you lose all buffered spans.</p>
<h3>Option 3: Zipkin Exporter</h3>
<p>If you are locked into Zipkin, use the dedicated exporter:</p>
<pre><code class="language-go">import "go.opentelemetry.io/otel/exporters/zipkin"

exporter, err := zipkin.New("http://zipkin:9411/api/v2/spans")
</code></pre>
<h2>Trace Context in Async Workflows</h2>
<p>HTTP propagation covers synchronous calls, but many production systems use message queues. You need to manually inject and extract trace context.</p>
<p><strong>Producer (publishing to a queue):</strong></p>
<pre><code class="language-go">func publishOrderEvent(ctx context.Context, event OrderCreatedEvent) error {
    carrier := make(propagation.MapCarrier)
    otel.GetTextMapPropagator().Inject(ctx, carrier)

    // Attach trace context as message headers
    headers := make(map[string]string)
    for _, key := range []string{"traceparent", "tracestate"} {
        if val := carrier.Get(key); val != "" {
            headers[key] = val
        }
    }

    msg := &amp;queue.Message{
        Body:    encodeEvent(event),
        Headers: headers,
    }
    return queue.Publish(ctx, "order.created", msg)
}
</code></pre>
<p><strong>Consumer (reading from a queue):</strong></p>
<pre><code class="language-go">func handleOrderCreatedMessage(msg *queue.Message) error {
    carrier := propagation.MapCarrier(msg.Headers)
    ctx := otel.GetTextMapPropagator().Extract(context.Background(), carrier)

    ctx, span := tracer.Start(ctx, "queue.ProcessOrderCreated",
        trace.WithSpanKind(trace.SpanKindConsumer),
        trace.WithAttributes(
            attribute.String("messaging.system", "rabbitmq"),
            attribute.String("messaging.operation", "process"),
        ),
    )
    defer span.End()

    // Process with the restored trace context
    return processOrderCreated(ctx, msg.Body)
}
</code></pre>
<p>The consumer's span becomes a child of the producer's span, even though they run in different processes and potentially on different machines. This is the real power of distributed tracing.</p>
<h2>Production Best Practices</h2>
<h3>1. Sample Intelligently</h3>
<p>Tracing 100% of traffic is expensive. Use <code>ParentBased(TraceIDRatioBased(0.1))</code> to sample 10% of new traces while always honoring upstream sampling decisions. For error investigation, implement tail-based sampling at the collector level ‚Äî it captures all traces that contain an error span.</p>
<h3>2. Control Span Cardinality</h3>
<p>Every unique span name creates a series in your backend. Avoid dynamic values in span names:</p>
<pre><code class="language-go">// BAD: creates thousands of unique span names
tracer.Start(ctx, fmt.Sprintf("getUser-%s", userID))

// GOOD: use attributes for dynamic values
tracer.Start(ctx, "user.Get",
    trace.WithAttributes(attribute.String("user.id", userID)),
)
</code></pre>
<h3>3. Keep Attribute Values Bounded</h3>
<p>High-cardinality attributes (full URLs, raw SQL, request bodies) bloat storage and slow down queries. Stick to IDs, enums, and short strings. Use span events for detailed debugging data that you only need occasionally.</p>
<h3>4. Set Span Kind Correctly</h3>
<p>Span kind tells the backend how to render the span:</p>
<ul>
<li><p><code>SpanKindServer</code> ‚Äî incoming RPC/HTTP (set by <code>otelhttp</code> middleware)</p>
</li>
<li><p><code>SpanKindClient</code> ‚Äî outgoing RPC/HTTP (set by <code>otelhttp</code> transport)</p>
</li>
<li><p><code>SpanKindProducer</code> ‚Äî message queue publish</p>
</li>
<li><p><code>SpanKindConsumer</code> ‚Äî message queue consume</p>
</li>
<li><p><code>SpanKindInternal</code> ‚Äî default, local operations</p>
</li>
</ul>
<h3>5. Use Span Links for Fan-Out</h3>
<p>When one request triggers multiple async operations, use span links instead of parent-child relationships. A batch job that processes 1000 messages should link to each producer span, not parent them ‚Äî otherwise you get an unreadable trace tree.</p>
<pre><code class="language-go">links := make([]trace.Link, len(messages))
for i, msg := range messages {
    carrier := propagation.MapCarrier(msg.Headers)
    remoteCtx := otel.GetTextMapPropagator().Extract(context.Background(), carrier)
    links[i] = trace.LinkFromContext(remoteCtx)
}

ctx, span := tracer.Start(ctx, "batch.ProcessMessages",
    trace.WithLinks(links...),
)
</code></pre>
<h3>6. Correlate Traces with Logs</h3>
<p>Inject the trace ID into your structured logger so log lines are searchable by trace:</p>
<pre><code class="language-go">func loggerFromContext(ctx context.Context) *slog.Logger {
    spanCtx := trace.SpanContextFromContext(ctx)
    return slog.Default().With(
        slog.String("trace_id", spanCtx.TraceID().String()),
        slog.String("span_id", spanCtx.SpanID().String()),
    )
}
</code></pre>
<p>Grafana, Datadog, and most observability platforms can jump from a log line to the full trace when trace_id is present.</p>
<h3>7. Flush on Shutdown</h3>
<p>We covered this in the setup, but it bears repeating: <code>TracerProvider.Shutdown()</code> must be called with a generous timeout. The batch processor may have hundreds of spans queued. Pair it with your graceful shutdown handler.</p>
<h2>What a Full Trace Looks Like</h2>
<p>After instrumenting two services, here is what a trace looks like in Jaeger:</p>
<pre><code class="language-plaintext">order-service: POST /orders/create         [============]  320ms
  ‚îú‚îÄ order.Process                         [==========]    280ms
  ‚îÇ   ‚îú‚îÄ order.ValidateInventory           [===]            85ms
  ‚îÇ   ‚îú‚îÄ order.CalculateTotal              [=]              12ms
  ‚îÇ   ‚îî‚îÄ order.SaveToDB                    [==]             45ms
  ‚îî‚îÄ HTTP POST payment-service/charge      [=====]        130ms
      payment-service: POST /charge        [====]         125ms
        ‚îú‚îÄ payment.ValidateCard            [=]              15ms
        ‚îî‚îÄ payment.ProcessCharge           [===]           105ms
</code></pre>
<p>At a glance: 320ms total, and the payment service accounts for 130ms of that. The database insert is 45ms ‚Äî reasonable for a write with index updates. If this endpoint starts breaching its latency SLO, you know exactly where to look.</p>
<h2>Wrapping Up</h2>
<p>The initial investment in distributed tracing pays for itself the first time you debug a cross-service latency issue without manually correlating logs. The setup is:</p>
<ol>
<li><p>Initialize a <code>TracerProvider</code> with OTLP export and sensible batching.</p>
</li>
<li><p>Wrap HTTP servers with <code>otelhttp.NewHandler</code> for automatic span creation.</p>
</li>
<li><p>Wrap HTTP clients with <code>otelhttp.NewTransport</code> for automatic context propagation.</p>
</li>
<li><p>Add custom spans around business logic and database calls.</p>
</li>
<li><p>Manually propagate context through message queues.</p>
</li>
<li><p>Sample, control cardinality, correlate with logs.</p>
</li>
</ol>
<p>All the code in this article uses the stable OpenTelemetry Go SDK. It works with any OTLP-compatible backend ‚Äî Jaeger, Tempo, Datadog, Honeycomb, or the OTel Collector. Pick one, point the exporter at it, and start seeing your system for what it really is.</p>
<hr />
<p><em>If this article helped you, consider</em> <a href="https://ko-fi.com/gps949"><em>buying me a coffee on Ko-fi</em></a><em>! Follow me for more production backend patterns.</em></p>
<hr />
<p>I was unable to save this to a file because both Write and Bash permissions were denied. The complete article is above -- you can copy it or grant file write permission so I can save it to <code>/Users/miniclaw/workspace/article-15-distributed-tracing-otel.md</code>.</p>
]]></content:encoded></item><item><title><![CDATA[Graceful Shutdown in Go: Patterns Every Production Service Needs]]></title><description><![CDATA[Your service just got a SIGTERM. You have roughly 30 seconds before Kubernetes sends SIGKILL. In that window, you need to finish in-flight requests, flush buffered data, close database connections, an]]></description><link>https://younggao.hashnode.dev/graceful-shutdown-in-go-patterns-every-production-service-needs</link><guid isPermaLink="true">https://younggao.hashnode.dev/graceful-shutdown-in-go-patterns-every-production-service-needs</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 06:57:54 GMT</pubDate><content:encoded><![CDATA[<p>Your service just got a <code>SIGTERM</code>. You have roughly 30 seconds before Kubernetes sends <code>SIGKILL</code>. In that window, you need to finish in-flight requests, flush buffered data, close database connections, and deregister from service discovery — all without dropping a single user request.</p>
<p>Get it wrong and you get data loss, broken client connections, and 502s during every deploy. Get it right and your deployments become invisible to users.</p>
<p>This article walks through the patterns that make graceful shutdown reliable in production Go services.</p>
<h2>Why Graceful Shutdown Matters</h2>
<p>Three things go wrong when a service dies abruptly:</p>
<p><strong>Data loss.</strong> Buffered writes — log batches, metrics, queue messages — vanish. Database transactions in progress get rolled back (best case) or leave inconsistent state (worst case).</p>
<p><strong>Connection drops.</strong> In-flight HTTP requests get TCP RST. gRPC streams break mid-message. WebSocket clients see unexpected disconnections. Depending on client retry logic, this can cascade.</p>
<p><strong>Load balancer draining failures.</strong> Kubernetes removes the pod from the Service endpoints, but if your process exits before the kubelet propagates that change, the load balancer still sends traffic to a dead pod. The standard fix is to keep serving for a few seconds after receiving <code>SIGTERM</code> — but only if your shutdown sequence is correct.</p>
<h2>The Foundation: os.Signal + context.Context</h2>
<p>Every graceful shutdown in Go starts with the same two primitives: signal notification and context cancellation.</p>
<pre><code class="language-go">func main() {
    // Create a context that cancels on SIGINT or SIGTERM
    ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
    defer stop()

    // Pass ctx to everything that needs to know about shutdown
    if err := run(ctx); err != nil {
        log.Fatalf("service exited with error: %v", err)
    }
}
</code></pre>
<p><code>signal.NotifyContext</code> (added in Go 1.16) is the cleanest way to wire OS signals into the context tree. When the signal arrives, the context's <code>Done()</code> channel closes, and every goroutine watching that context knows it's time to wrap up.</p>
<p>The key principle: <strong>propagate context everywhere.</strong> Every HTTP handler, every database query, every background worker should accept a <code>context.Context</code>. This is how you make shutdown cooperative rather than forceful.</p>
<h2>HTTP Server Graceful Shutdown</h2>
<p><code>http.Server.Shutdown()</code> does exactly what you want: it stops accepting new connections, waits for in-flight requests to complete, and then returns.</p>
<pre><code class="language-go">func runHTTPServer(ctx context.Context) error {
    mux := http.NewServeMux()
    mux.HandleFunc("/api/data", handleData)

    srv := &amp;http.Server{
        Addr:    ":8080",
        Handler: mux,
    }

    // Start server in a goroutine
    errCh := make(chan error, 1)
    go func() {
        if err := srv.ListenAndServe(); err != http.ErrServerClosed {
            errCh &lt;- err
        }
        close(errCh)
    }()

    // Wait for shutdown signal
    &lt;-ctx.Done()
    log.Println("shutting down HTTP server...")

    // Give in-flight requests a deadline to finish
    shutdownCtx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
    defer cancel()

    if err := srv.Shutdown(shutdownCtx); err != nil {
        return fmt.Errorf("http server shutdown: %w", err)
    }

    // Check if server goroutine hit an error before shutdown
    return &lt;-errCh
}
</code></pre>
<p>Notice the shutdown context uses <code>context.Background()</code>, <strong>not</strong> the already-cancelled parent context. This is a common mistake — if you derive the shutdown timeout from the cancelled context, the timeout is already expired and <code>Shutdown()</code> returns immediately, dropping in-flight requests.</p>
<h2>gRPC Graceful Stop</h2>
<p>gRPC has its own graceful shutdown method that mirrors the HTTP pattern:</p>
<pre><code class="language-go">func runGRPCServer(ctx context.Context) error {
    lis, err := net.Listen("tcp", ":9090")
    if err != nil {
        return fmt.Errorf("listen: %w", err)
    }

    grpcServer := grpc.NewServer()
    pb.RegisterMyServiceServer(grpcServer, &amp;myService{})

    // Enable reflection for debugging
    reflection.Register(grpcServer)

    errCh := make(chan error, 1)
    go func() {
        if err := grpcServer.Serve(lis); err != nil {
            errCh &lt;- err
        }
        close(errCh)
    }()

    &lt;-ctx.Done()
    log.Println("shutting down gRPC server...")

    // GracefulStop waits for active RPCs to finish
    // but we add a hard deadline as a safety net
    stopped := make(chan struct{})
    go func() {
        grpcServer.GracefulStop()
        close(stopped)
    }()

    select {
    case &lt;-stopped:
        log.Println("gRPC server stopped gracefully")
    case &lt;-time.After(10 * time.Second):
        log.Println("gRPC server force stop (timeout)")
        grpcServer.Stop()
    }

    return &lt;-errCh
}
</code></pre>
<p><code>GracefulStop()</code> has no built-in timeout, so you must wrap it with a deadline yourself. Without this, a single slow-draining stream can hang your shutdown indefinitely.</p>
<h2>Worker Pool Draining with sync.WaitGroup</h2>
<p>Background workers — queue consumers, cron jobs, batch processors — need a different pattern. The worker checks <code>ctx.Done()</code> between units of work, and a <code>WaitGroup</code> lets the main goroutine wait for all workers to finish their current task.</p>
<pre><code class="language-go">func runWorkerPool(ctx context.Context, queue &lt;-chan Job) error {
    const numWorkers = 8
    var wg sync.WaitGroup

    for i := 0; i &lt; numWorkers; i++ {
        wg.Add(1)
        go func(id int) {
            defer wg.Done()
            for {
                select {
                case &lt;-ctx.Done():
                    log.Printf("worker %d: shutting down", id)
                    return
                case job, ok := &lt;-queue:
                    if !ok {
                        return // channel closed
                    }
                    if err := processJob(ctx, job); err != nil {
                        log.Printf("worker %d: job failed: %v", id, err)
                    }
                }
            }
        }(i)
    }

    // Wait for all workers to finish their current job
    wg.Wait()
    return nil
}
</code></pre>
<p>The critical detail: <code>processJob</code> receives the context so it can abandon long-running work. If a job takes 5 minutes and your shutdown budget is 30 seconds, the job needs to respect context cancellation or you'll hit <code>SIGKILL</code>.</p>
<h2>Database Connection Cleanup</h2>
<p>Database connections are the easiest to get right and the most painful to get wrong. Leaked connections exhaust the connection pool on the database side, causing failures across all services.</p>
<pre><code class="language-go">func setupDatabase(ctx context.Context) (*sql.DB, error) {
    db, err := sql.Open("postgres", os.Getenv("DATABASE_URL"))
    if err != nil {
        return nil, err
    }

    db.SetMaxOpenConns(25)
    db.SetMaxIdleConns(5)
    db.SetConnMaxLifetime(5 * time.Minute)

    if err := db.PingContext(ctx); err != nil {
        db.Close()
        return nil, fmt.Errorf("database ping: %w", err)
    }

    return db, nil
}
</code></pre>
<p>The cleanup is simple — <code>db.Close()</code> — but the <strong>ordering</strong> matters. Close the database <strong>after</strong> servers and workers have stopped, since they may still be executing queries during their drain period.</p>
<h2>Composing Everything with errgroup</h2>
<p>Now let's stitch it all together. <code>errgroup</code> from <code>golang.org/x/sync</code> gives us structured concurrency: run components in parallel, shut them all down if any one fails, and collect errors.</p>
<pre><code class="language-go">func run(ctx context.Context) error {
    db, err := setupDatabase(ctx)
    if err != nil {
        return fmt.Errorf("database: %w", err)
    }
    defer db.Close() // Always last to close

    jobQueue := make(chan Job, 100)

    g, ctx := errgroup.WithContext(ctx)

    // HTTP server
    g.Go(func() error {
        return runHTTPServer(ctx)
    })

    // gRPC server
    g.Go(func() error {
        return runGRPCServer(ctx)
    })

    // Worker pool
    g.Go(func() error {
        return runWorkerPool(ctx, jobQueue)
    })

    // Health monitor (optional: triggers shutdown on critical failures)
    g.Go(func() error {
        return runHealthMonitor(ctx, db)
    })

    log.Println("service started")
    err = g.Wait()
    log.Println("all components stopped")

    return err
}
</code></pre>
<p><code>errgroup.WithContext</code> creates a derived context that cancels when <strong>any</strong> goroutine in the group returns an error. This means if the HTTP server fails to bind its port, the context cancels, and the gRPC server and workers also begin shutting down. Unified lifecycle management.</p>
<p>Note that <code>db.Close()</code> runs via <code>defer</code> — it executes after <code>g.Wait()</code> returns, guaranteeing all servers and workers have fully drained before we close database connections. <strong>Dependency ordering through deferred cleanup</strong> is the simplest correct pattern.</p>
<h2>Health Check Integration with Kubernetes</h2>
<p>Kubernetes needs to know when your pod is ready to receive traffic and when it should stop sending traffic. The readiness probe is your coordination point.</p>
<pre><code class="language-go">type healthHandler struct {
    ready atomic.Bool
}

func (h *healthHandler) ServeHTTP(w http.ResponseWriter, r *http.Request) {
    if h.ready.Load() {
        w.WriteHeader(http.StatusOK)
        w.Write([]byte("ok"))
    } else {
        w.WriteHeader(http.StatusServiceUnavailable)
        w.Write([]byte("shutting down"))
    }
}

func runHTTPServer(ctx context.Context) error {
    health := &amp;healthHandler{}
    health.ready.Store(true)

    mux := http.NewServeMux()
    mux.Handle("/healthz", health)
    mux.HandleFunc("/api/data", handleData)

    srv := &amp;http.Server{Addr: ":8080", Handler: mux}

    errCh := make(chan error, 1)
    go func() {
        if err := srv.ListenAndServe(); err != http.ErrServerClosed {
            errCh &lt;- err
        }
        close(errCh)
    }()

    &lt;-ctx.Done()

    // Step 1: Mark as not ready — Kubernetes stops sending traffic
    health.ready.Store(false)

    // Step 2: Wait for load balancer to propagate the change
    time.Sleep(5 * time.Second)

    // Step 3: Now shut down the server
    shutdownCtx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
    defer cancel()

    if err := srv.Shutdown(shutdownCtx); err != nil {
        return fmt.Errorf("http server shutdown: %w", err)
    }

    return &lt;-errCh
}
</code></pre>
<p>That <code>time.Sleep(5 * time.Second)</code> looks wrong, but it's essential. Kubernetes endpoint propagation is <strong>eventually consistent</strong>. After your readiness probe starts failing, it takes a few seconds for kube-proxy/iptables rules to update across nodes. During that window, traffic still arrives. The sleep keeps your server alive to handle it.</p>
<p>In your pod spec:</p>
<pre><code class="language-yaml">readinessProbe:
  httpGet:
    path: /healthz
    port: 8080
  periodSeconds: 2
  failureThreshold: 1
terminationGracePeriodSeconds: 45
</code></pre>
<p>Set <code>terminationGracePeriodSeconds</code> to comfortably exceed your total shutdown budget (sleep + drain timeout + cleanup).</p>
<h2>Common Mistakes</h2>
<p><strong>1. Deriving shutdown context from the cancelled parent.</strong></p>
<pre><code class="language-go">// WRONG: ctx is already cancelled, timeout has no effect
shutdownCtx, cancel := context.WithTimeout(ctx, 10*time.Second)

// RIGHT: fresh context with its own deadline
shutdownCtx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
</code></pre>
<p><strong>2. No shutdown timeout at all.</strong> <code>GracefulStop()</code> and custom drain loops can block forever. Always wrap with a deadline and fall back to a forced stop.</p>
<p><strong>3. Wrong cleanup ordering.</strong> If you close the database before the HTTP server finishes draining, in-flight request handlers get connection errors. Rule of thumb: shut down in reverse order of initialization. Servers first, then workers, then infrastructure (DB, caches, message brokers).</p>
<p><strong>4. Forgetting to stop timers and tickers.</strong></p>
<pre><code class="language-go">ticker := time.NewTicker(30 * time.Second)
defer ticker.Stop() // Don't leak this

for {
    select {
    case &lt;-ctx.Done():
        return nil
    case &lt;-ticker.C:
        doPeriodicWork()
    }
}
</code></pre>
<p><strong>5. Not testing shutdown.</strong> Send <code>SIGTERM</code> to your service in integration tests. Verify that in-flight requests complete, no errors appear in logs, and the process exits with code 0.</p>
<h2>Putting It All Together</h2>
<p>The complete pattern:</p>
<ol>
<li><p>Trap <code>SIGINT</code>/<code>SIGTERM</code> with <code>signal.NotifyContext</code></p>
</li>
<li><p>Start all components under an <code>errgroup</code> with the signal-aware context</p>
</li>
<li><p>Each component watches <code>ctx.Done()</code> and begins draining</p>
</li>
<li><p>Mark readiness probe as unhealthy, sleep for LB propagation</p>
</li>
<li><p>Shut down servers (HTTP, gRPC) with a timeout</p>
</li>
<li><p>Wait for workers to finish current tasks</p>
</li>
<li><p>Close infrastructure connections (DB, caches) last via <code>defer</code></p>
</li>
</ol>
<p>This sequence handles the normal deploy case (Kubernetes rolling update), the crash case (a component fails and triggers group shutdown), and the manual stop case (<code>Ctrl+C</code> in development).</p>
<p>Graceful shutdown isn't glamorous, but it's the difference between deploys that are invisible to users and deploys that generate a spike of errors in your dashboard. Build it in from the start — retrofitting it later is always harder.</p>
<hr />
<p><em>This is part of the</em> <em><strong>Production Backend Patterns</strong></em> <em>series. Next up: structured logging and distributed tracing in Go.</em></p>
<hr />
<p><em>If this article helped you, consider</em> <a href="https://ko-fi.com/gps949"><em>buying me a coffee on Ko-fi</em></a><em>! Follow me for more production backend patterns.</em></p>
]]></content:encoded></item><item><title><![CDATA[Building a High-Performance Cache Layer in Go]]></title><description><![CDATA[atgsvcs]]></description><link>https://younggao.hashnode.dev/building-a-high-performance-cache-layer-in-go</link><guid isPermaLink="true">https://younggao.hashnode.dev/building-a-high-performance-cache-layer-in-go</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 06:36:38 GMT</pubDate><content:encoded><![CDATA[<pre><code class="language-plaintext">atgsvcs
</code></pre>
]]></content:encoded></item><item><title><![CDATA[Pragmatic Error Handling in Go]]></title><description><![CDATA[Pragmatic Error Handling in Go
Beyond if err != nil - techniques that work.
Always Wrap
return fmt.Errorf("get order %s: %w", id, err)

Sentinels
var ErrNotFound = errors.New("not found")

HTTP Patter]]></description><link>https://younggao.hashnode.dev/pragmatic-error-handling-in-go</link><guid isPermaLink="true">https://younggao.hashnode.dev/pragmatic-error-handling-in-go</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 06:32:00 GMT</pubDate><content:encoded><![CDATA[<h1>Pragmatic Error Handling in Go</h1>
<p>Beyond if err != nil - techniques that work.</p>
<h2>Always Wrap</h2>
<pre><code class="language-go">return fmt.Errorf("get order %s: %w", id, err)
</code></pre>
<h2>Sentinels</h2>
<pre><code class="language-go">var ErrNotFound = errors.New("not found")
</code></pre>
<h2>HTTP Pattern</h2>
<pre><code class="language-go">type HandlerFunc func(w http.ResponseWriter, r *http.Request) error
</code></pre>
<h2>errgroup</h2>
<pre><code class="language-go">g, ctx := errgroup.WithContext(ctx)
g.Go(func() error { return fn() })
</code></pre>
<hr />
<p><em>If this was helpful, you can support my work at</em> <a href="https://ko-fi.com/nopkt"><em>ko-fi.com/nopkt</em></a></p>
]]></content:encoded></item><item><title><![CDATA[PostgreSQL Performance: 10 Queries You Are Writing Wrong]]></title><description><![CDATA[PostgreSQL is incredibly powerful, but even experienced developers write queries that silently destroy performance. Here are the 10 most common mistakes and how to fix each one.

1. Using SELECT * Ins]]></description><link>https://younggao.hashnode.dev/postgresql-performance-10-queries-you-are-writing-wrong</link><guid isPermaLink="true">https://younggao.hashnode.dev/postgresql-performance-10-queries-you-are-writing-wrong</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 06:31:21 GMT</pubDate><content:encoded><![CDATA[<p>PostgreSQL is incredibly powerful, but even experienced developers write queries that silently destroy performance. Here are the 10 most common mistakes and how to fix each one.</p>
<hr />
<h2>1. Using SELECT * Instead of Specific Columns</h2>
<p><strong>Bad SQL:</strong></p>
<pre><code class="language-sql">SELECT * FROM orders WHERE customer_id = 42;
</code></pre>
<p><strong>Good SQL:</strong></p>
<pre><code class="language-sql">SELECT order_id, total_amount, created_at
FROM orders
WHERE customer_id = 42;
</code></pre>
<p><strong>Why:</strong> SELECT * bloats I/O, defeats index-only scans, and wastes bandwidth. Selecting specific columns lets PostgreSQL use an index-only scan without touching the heap.</p>
<p><strong>Typical speedup: 2-10x</strong></p>
<hr />
<h2>2. Missing Indexes on WHERE and JOIN Columns</h2>
<p><strong>Bad SQL:</strong></p>
<pre><code class="language-sql">-- No index on orders.customer_id
SELECT order_id, total_amount
FROM orders
WHERE customer_id = 42;
</code></pre>
<p><strong>Good SQL:</strong></p>
<pre><code class="language-sql">CREATE INDEX idx_orders_customer_id ON orders (customer_id);

SELECT order_id, total_amount
FROM orders
WHERE customer_id = 42;
</code></pre>
<p><strong>Why:</strong> Without an index, PostgreSQL scans every row in the table. On a 10M-row table, that means scanning all rows to find maybe 50 matches. An index turns this into a B-tree lookup touching a handful of pages.</p>
<p><strong>Typical speedup: 100-10,000x</strong> on large tables. The single biggest performance win.</p>
<hr />
<h2>3. OFFSET Pagination vs Keyset Pagination</h2>
<p><strong>Bad SQL:</strong></p>
<pre><code class="language-sql">SELECT id, title, created_at
FROM articles
ORDER BY created_at DESC
LIMIT 20 OFFSET 9980;
</code></pre>
<p><strong>Good SQL:</strong></p>
<pre><code class="language-sql">SELECT id, title, created_at
FROM articles
WHERE created_at &lt; '2024-01-15 10:30:00'
ORDER BY created_at DESC
LIMIT 20;
</code></pre>
<p><strong>Why:</strong> OFFSET forces PostgreSQL to fetch and discard 9,980 rows before returning your 20. It is O(n) where n is the offset. Keyset pagination uses a WHERE clause to jump directly to the right position - O(1) regardless of depth.</p>
<p><strong>Typical speedup: 10-1000x</strong> on deep pages</p>
<hr />
<h2>4. N+1 Queries vs JOINs</h2>
<p><strong>Bad SQL (application code pattern):</strong></p>
<pre><code class="language-sql">SELECT id, name FROM customers LIMIT 100;
-- Then 100 separate queries:
SELECT SUM(total_amount) FROM orders WHERE customer_id = 1;
SELECT SUM(total_amount) FROM orders WHERE customer_id = 2;
-- ... 98 more
</code></pre>
<p><strong>Good SQL:</strong></p>
<pre><code class="language-sql">SELECT c.id, c.name, COALESCE(SUM(o.total_amount), 0) AS total_spent
FROM customers c
LEFT JOIN orders o ON o.customer_id = c.id
GROUP BY c.id, c.name
LIMIT 100;
</code></pre>
<p><strong>Why:</strong> The N+1 pattern sends 101 separate queries with individual round-trips. A single JOIN lets PostgreSQL optimize the entire operation with hash joins, merge joins, or nested loops.</p>
<p><strong>Typical speedup: 5-50x</strong></p>
<hr />
<h2>5. Not Using EXPLAIN ANALYZE</h2>
<p><strong>Bad approach:</strong></p>
<pre><code class="language-sql">-- Guessing which indexes to add
CREATE INDEX idx_whatever ON orders (status);
CREATE INDEX idx_whatever2 ON orders (status, customer_id);
</code></pre>
<p><strong>Good approach:</strong></p>
<pre><code class="language-sql">EXPLAIN (ANALYZE, BUFFERS, FORMAT TEXT)
SELECT o.order_id, c.name
FROM orders o
JOIN customers c ON c.id = o.customer_id
WHERE o.status = 'pending'
  AND o.created_at &gt; NOW() - INTERVAL '7 days';
</code></pre>
<p><strong>Why:</strong> EXPLAIN ANALYZE executes the query and shows exactly what PostgreSQL did: which indexes it used, row estimate accuracy, and where time was spent. Look for Seq Scans on large tables, estimate vs actual row divergence, and expensive sorts. Stop guessing, start measuring.</p>
<p><strong>Typical speedup:</strong> Tells you exactly what to fix.</p>
<hr />
<h2>6. LIKE with Leading Wildcards vs pg_trgm</h2>
<p><strong>Bad SQL:</strong></p>
<pre><code class="language-sql">SELECT id, name FROM products
WHERE name LIKE '%%wireless%%';
</code></pre>
<p><strong>Good SQL:</strong></p>
<pre><code class="language-sql">CREATE EXTENSION IF NOT EXISTS pg_trgm;
CREATE INDEX idx_products_name_trgm
  ON products USING gin (name gin_trgm_ops);

SELECT id, name FROM products
WHERE name LIKE '%%wireless%%';
</code></pre>
<p><strong>Why:</strong> B-tree indexes only help with prefix matching (LIKE 'wireless%%'). A leading wildcard forces a sequential scan. The pg_trgm extension builds a GIN index over trigrams, enabling index usage for substring, ILIKE, and regex matching.</p>
<p><strong>Typical speedup: 50-500x</strong> on large text columns</p>
<hr />
<h2>7. COUNT(*) for Existence Checks vs EXISTS</h2>
<p><strong>Bad SQL:</strong></p>
<pre><code class="language-sql">SELECT CASE WHEN COUNT(*) &gt; 0 THEN true ELSE false END
FROM orders
WHERE customer_id = 42 AND status = 'pending';
</code></pre>
<p><strong>Good SQL:</strong></p>
<pre><code class="language-sql">SELECT EXISTS (
  SELECT 1 FROM orders
  WHERE customer_id = 42 AND status = 'pending'
);
</code></pre>
<p><strong>Why:</strong> COUNT(*) scans and counts every matching row. If there are 10,000 matches, it counts all 10,000 just to check if any exist. EXISTS returns true on the first match and stops - a short-circuit operation.</p>
<p><strong>Typical speedup: 10-10,000x</strong> depending on match count</p>
<hr />
<h2>8. Single-Row INSERTs vs Batching / COPY</h2>
<p><strong>Bad SQL:</strong></p>
<pre><code class="language-sql">INSERT INTO events (type, data, created_at) VALUES ('click', '{}', NOW());
INSERT INTO events (type, data, created_at) VALUES ('view', '{}', NOW());
-- 9,998 more individual statements
</code></pre>
<p><strong>Good SQL:</strong></p>
<pre><code class="language-sql">-- Multi-value INSERT
INSERT INTO events (type, data, created_at) VALUES
  ('click', '{}', NOW()),
  ('view', '{}', NOW()),
  ('scroll', '{}', NOW());

-- Or COPY for bulk loads
COPY events (type, data, created_at)
FROM STDIN WITH (FORMAT csv);
</code></pre>
<p><strong>Why:</strong> Each INSERT has overhead: network round-trip, SQL parsing, query planning, WAL logging, and transaction commit. Batching and COPY amortize this dramatically.</p>
<p><strong>Typical speedup: 10-100x</strong> for batching, <strong>100-1000x</strong> for COPY</p>
<hr />
<h2>9. No Connection Pooling (Use PgBouncer)</h2>
<p><strong>Bad approach:</strong> Opening a new database connection per request. Each connection forks a new PostgreSQL backend process and allocates ~10MB of memory.</p>
<p><strong>Good approach:</strong> Use PgBouncer with transaction-mode pooling:</p>
<pre><code class="language-ini">[databases]
mydb = host=localhost port=5432 dbname=mydb

[pgbouncer]
listen_port = 6432
pool_mode = transaction
max_client_conn = 1000
default_pool_size = 20
</code></pre>
<p><strong>Why:</strong> Connection creation is expensive. Under load, hundreds of connections create memory pressure and context-switching overhead. PgBouncer maintains a small pool of actual connections serving thousands of clients, and protects against connection storms during traffic spikes.</p>
<p><strong>Typical speedup: 2-20x</strong> under concurrent load</p>
<hr />
<h2>10. Not Using Partial Indexes</h2>
<p><strong>Bad SQL:</strong></p>
<pre><code class="language-sql">CREATE INDEX idx_orders_status ON orders (status);
SELECT * FROM orders WHERE status = 'failed';
</code></pre>
<p><strong>Good SQL:</strong></p>
<pre><code class="language-sql">CREATE INDEX idx_orders_failed ON orders (created_at)
WHERE status = 'failed';

SELECT * FROM orders
WHERE status = 'failed'
ORDER BY created_at DESC LIMIT 10;
</code></pre>
<p><strong>Why:</strong> If 95%% of orders are completed and only 1%% are failed, a full index on status is massive but mostly useless. A partial index with WHERE status = 'failed' indexes only the 1%% you query. It is dramatically smaller, fits in cache, and is faster to scan and maintain.</p>
<p>Use partial indexes for: status flags, soft deletes, queue tables, and any skewed distribution column.</p>
<p><strong>Typical speedup: 3-10x</strong> with significantly less disk space and write overhead</p>
<hr />
<h2>Summary</h2>
<table>
<thead>
<tr>
<th>#</th>
<th>Mistake</th>
<th>Fix</th>
<th>Speedup</th>
</tr>
</thead>
<tbody><tr>
<td>1</td>
<td>SELECT *</td>
<td>Specific columns</td>
<td>2-10x</td>
</tr>
<tr>
<td>2</td>
<td>Missing indexes</td>
<td>Targeted indexes</td>
<td>100-10,000x</td>
</tr>
<tr>
<td>3</td>
<td>OFFSET pagination</td>
<td>Keyset pagination</td>
<td>10-1,000x</td>
</tr>
<tr>
<td>4</td>
<td>N+1 queries</td>
<td>JOINs</td>
<td>5-50x</td>
</tr>
<tr>
<td>5</td>
<td>Blind optimization</td>
<td>EXPLAIN ANALYZE</td>
<td>Enables all fixes</td>
</tr>
<tr>
<td>6</td>
<td>LIKE leading wildcard</td>
<td>pg_trgm GIN index</td>
<td>50-500x</td>
</tr>
<tr>
<td>7</td>
<td>COUNT(*) existence</td>
<td>EXISTS</td>
<td>10-10,000x</td>
</tr>
<tr>
<td>8</td>
<td>Single-row INSERTs</td>
<td>Batch / COPY</td>
<td>10-1,000x</td>
</tr>
<tr>
<td>9</td>
<td>No pooling</td>
<td>PgBouncer</td>
<td>2-20x</td>
</tr>
<tr>
<td>10</td>
<td>Full indexes</td>
<td>Partial indexes</td>
<td>3-10x</td>
</tr>
</tbody></table>
<p>Every optimization follows one principle: <strong>give PostgreSQL less work to do</strong>. Start with EXPLAIN ANALYZE on your slowest queries today.</p>
]]></content:encoded></item><item><title><![CDATA[Building a CLI Tool in Rust: From Zero to Published on crates.io]]></title><description><![CDATA[Building a CLI Tool in Rust: From Zero to Published on crates.io
Why Build a CLI Tool in Rust?
Rust is an excellent language for CLI tools. It compiles to a single static binary, starts instantly, and]]></description><link>https://younggao.hashnode.dev/building-a-cli-tool-in-rust-from-zero-to-published-on-crates-io</link><guid isPermaLink="true">https://younggao.hashnode.dev/building-a-cli-tool-in-rust-from-zero-to-published-on-crates-io</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 06:31:00 GMT</pubDate><content:encoded><![CDATA[<h1>Building a CLI Tool in Rust: From Zero to Published on <a href="http://crates.io">crates.io</a></h1>
<h2>Why Build a CLI Tool in Rust?</h2>
<p>Rust is an excellent language for CLI tools. It compiles to a single static binary, starts instantly, and gives you memory safety without a garbage collector.</p>
<p>Many beloved tools like <code>ripgrep</code>, <code>fd</code>, <code>bat</code>, and <code>exa</code> are written in Rust.</p>
<p>In this tutorial, we will build <strong>rsearch</strong> -- a fast file content searcher with colored output. By the end, you will have a tool published on crates.io that others can install with <code>cargo install</code>.</p>
<hr />
<h2>Step 1: Project Setup</h2>
<pre><code class="language-bash">cargo new rsearch
cd rsearch
</code></pre>
<p>This gives you:</p>
<pre><code class="language-plaintext">rsearch/
  Cargo.toml
  src/
    main.rs
</code></pre>
<h2>Step 2: Dependencies in Cargo.toml</h2>
<p>Open <code>Cargo.toml</code> and add the crates we need:</p>
<pre><code class="language-toml">[package]
name = "rsearch"
version = "0.1.0"
edition = "2021"
description = "A fast file content searcher with colored output"
license = "MIT"

[dependencies]
clap = { version = "4", features = ["derive"] }
regex = "1"
colored = "2"
anyhow = "1"
ignore = "0.4"

[dev-dependencies]
assert_cmd = "2"
predicates = "3"
tempfile = "3"

[profile.release]
opt-level = "z"
lto = true
codegen-units = 1
strip = true
panic = "abort"
</code></pre>
<p>Here is what each crate does:</p>
<table>
<thead>
<tr>
<th>Crate</th>
<th>Purpose</th>
</tr>
</thead>
<tbody><tr>
<td><code>clap</code></td>
<td>Argument parsing with derive macros</td>
</tr>
<tr>
<td><code>regex</code></td>
<td>Regular expression matching</td>
</tr>
<tr>
<td><code>colored</code></td>
<td>Terminal color output</td>
</tr>
<tr>
<td><code>anyhow</code></td>
<td>Ergonomic error handling</td>
</tr>
<tr>
<td><code>ignore</code></td>
<td>Respects <code>.gitignore</code> while walking directories</td>
</tr>
</tbody></table>
<p>The <code>[profile.release]</code> section enables aggressive binary size optimization -- we will revisit this at the end.</p>
<hr />
<h2>Step 3: Argument Parsing with Clap Derive API</h2>
<p>Clap derive API lets you define your CLI interface as a struct. Replace the contents of <code>src/main.rs</code>:</p>
<pre><code class="language-rust">use anyhow::{Context, Result};
use clap::Parser;
use colored::Colorize;
use ignore::WalkBuilder;
use regex::Regex;
use std::fs;
use std::path::PathBuf;

/// rsearch - A fast file content searcher
#[derive(Parser, Debug)]
#[command(name = "rsearch", version, about)]
struct Args {
    /// The regex pattern to search for
    pattern: String,

    /// Directory to search in
    #[arg(default_value = ".")]
    path: PathBuf,

    /// File extension filter
    #[arg(short, long)]
    ext: Option&lt;String&gt;,

    /// Show line numbers
    #[arg(short = 'n', long, default_value_t = true)]
    line_numbers: bool,

    /// Maximum search depth
    #[arg(short, long)]
    depth: Option&lt;usize&gt;,

    /// Include hidden files
    #[arg(long)]
    hidden: bool,

    /// Case-insensitive search
    #[arg(short = 'i', long)]
    ignore_case: bool,
}
</code></pre>
<p>Every field becomes a CLI flag. Clap auto-generates <code>--help</code> and <code>--version</code> for free.</p>
<p>Running <code>rsearch --help</code> will output:</p>
<pre><code class="language-console">rsearch - A fast file content searcher

Usage: rsearch [OPTIONS] &lt;PATTERN&gt; [PATH]

Arguments:
  &lt;PATTERN&gt;  The regex pattern to search for
  [PATH]     Directory to search in [default: .]

Options:
  -e, --ext &lt;EXT&gt;      File extension filter
  -n, --line-numbers   Show line numbers
  -d, --depth &lt;DEPTH&gt;  Maximum search depth
      --hidden         Include hidden files
  -i, --ignore-case    Case-insensitive search
  -h, --help           Print help
  -V, --version        Print version
</code></pre>
<hr />
<h2>Step 4: File Walking with the ignore Crate</h2>
<p>The <code>ignore</code> crate (from the ripgrep family) walks directories while automatically respecting <code>.gitignore</code> rules.</p>
<pre><code class="language-rust">fn walk_files(args: &amp;Args) -&gt; impl Iterator&lt;Item = PathBuf&gt; + '_ {
    let mut builder = WalkBuilder::new(&amp;args.path);
    builder.hidden(!args.hidden);

    if let Some(depth) = args.depth {
        builder.max_depth(Some(depth));
    }

    let ext_filter = args.ext.clone();

    builder.build().filter_map(move |entry| {
        let entry = entry.ok()?;
        let path = entry.path();

        if !path.is_file() {
            return None;
        }

        if let Some(ref ext) = ext_filter {
            if path.extension()?.to_str()? != ext.as_str() {
                return None;
            }
        }

        Some(path.to_path_buf())
    })
}
</code></pre>
<p>Key points:</p>
<ul>
<li><p><code>WalkBuilder</code> handles <code>.gitignore</code> parsing automatically</p>
</li>
<li><p>We filter by extension if the user passed <code>--ext</code></p>
</li>
<li><p>Hidden files are skipped by default unless <code>--hidden</code> is set</p>
</li>
<li><p>The iterator is lazy -- files are yielded on demand</p>
</li>
</ul>
<hr />
<h2>Step 5: Regex Search with Colored Output</h2>
<p>Now the core search logic. For each file, we read its contents, search line by line, and print matches with color highlighting:</p>
<pre><code class="language-rust">fn search_file(
    path: &amp;PathBuf,
    re: &amp;Regex,
    show_line_nums: bool,
) -&gt; Result&lt;Vec&lt;String&gt;&gt; {
    let content = fs::read_to_string(path)
        .with_context(|| format!("Failed to read {}", path.display()))?;

    let mut matches = Vec::new();

    for (line_num, line) in content.lines().enumerate() {
        if re.is_match(line) {
            let highlighted = re.replace_all(
                line,
                |caps: &amp;regex::Captures| {
                    caps[0].red().bold().to_string()
                },
            );

            let formatted = if show_line_nums {
                format!(
                    "{}:{} {}",
                    path.display().to_string().green(),
                    (line_num + 1).to_string().yellow(),
                    highlighted
                )
            } else {
                format!(
                    "{}: {}",
                    path.display().to_string().green(),
                    highlighted
                )
            };

            matches.push(formatted);
        }
    }

    Ok(matches)
}
</code></pre>
<p>The output looks like: <code>src/main.rs:42 fn search_file(...)</code> where the filename is green, the line number is yellow, and the matched text is bold red.</p>
<hr />
<h2>Step 6</h2>
<p>test</p>
<p>Main fn uses anyhow Result for error handling.</p>
<hr />
<h2>Step 7: Unit Tests</h2>
<p>Add a tests module at the bottom of main.rs:</p>
<pre><code class="language-rust">#[cfg(test)]
mod tests {
    use super::*;
    use std::io::Write;
    use tempfile::NamedTempFile;

    #[test]
    fn test_search_finds_match() {
        let mut file = NamedTempFile::new().unwrap();
        writeln!(file, "hello world").unwrap();
        writeln!(file, "foo bar").unwrap();
        writeln!(file, "hello rust").unwrap();
        let re = Regex::new("hello").unwrap();
        let path = file.path().to_path_buf();
        let results = search_file(&amp;path, &amp;re, false).unwrap();
        assert_eq!(results.len(), 2);
    }

    #[test]
    fn test_search_no_match() {
        let mut file = NamedTempFile::new().unwrap();
        writeln!(file, "nothing here").unwrap();
        let re = Regex::new("xyz").unwrap();
        let results = search_file(&amp;file.path().to_path_buf(), &amp;re, false).unwrap();
        assert!(results.is_empty());
    }

    #[test]
    fn test_case_insensitive() {
        let mut file = NamedTempFile::new().unwrap();
        writeln!(file, "Hello World").unwrap();
        let re = Regex::new("(?i)hello").unwrap();
        let results = search_file(&amp;file.path().to_path_buf(), &amp;re, false).unwrap();
        assert_eq!(results.len(), 1);
    }
}
</code></pre>
<p>Run them with <code>cargo test</code>. Each test creates a temp file and verifies search behavior.</p>
<hr />
<h2>Step 8: Integration Tests with assert_cmd</h2>
<p>Create <code>tests/integration.rs</code> for end-to-end testing:</p>
<pre><code class="language-rust">use assert_cmd::Command;
use predicates::prelude::*;
use std::fs;
use tempfile::TempDir;

#[test]
fn test_search_in_directory() {
    let dir = TempDir::new().unwrap();
    fs::write(dir.path().join("test.txt"), "hello world
foo bar
hello rust
").unwrap();
    Command::cargo_bin("rsearch").unwrap()
        .args(&amp;["hello", dir.path().to_str().unwrap()])
        .assert().success()
        .stdout(predicate::str::contains("hello"));
}
</code></pre>
<p>assert_cmd tests the real compiled binary.</p>
<hr />
<h2>Step 9: Publishing to crates.io</h2>
<h3>Binary Size Optimization</h3>
<table>
<thead>
<tr>
<th>Option</th>
<th>Effect</th>
</tr>
</thead>
<tbody><tr>
<td>opt-level z</td>
<td>Optimize for size</td>
</tr>
<tr>
<td>lto true</td>
<td>Link-Time Optimization</td>
</tr>
<tr>
<td>codegen-units 1</td>
<td>Better optimization</td>
</tr>
<tr>
<td>strip true</td>
<td>Remove debug symbols</td>
</tr>
<tr>
<td>panic abort</td>
<td>Smaller binary</td>
</tr>
</tbody></table>
<p>A typical Rust CLI drops from 4MB to 1.5MB with these settings.</p>
<h3>Publishing Steps</h3>
<ol>
<li><p>Create account on crates.io (GitHub login)</p>
</li>
<li><p><code>cargo login your-api-token</code></p>
</li>
<li><p><code>cargo publish --dry-run</code></p>
</li>
<li><p><code>cargo publish</code></p>
</li>
</ol>
<p>Now anyone can install: <code>cargo install rsearch</code></p>
<hr />
<h2>Wrapping Up</h2>
<p>We built a real CLI tool in Rust from scratch:</p>
<ul>
<li><p><strong>Clap derive</strong> gave us a polished CLI interface</p>
</li>
<li><p><strong>ignore</strong> handles .gitignore rules automatically</p>
</li>
<li><p><strong>regex + colored</strong> produce scannable colored output</p>
</li>
<li><p><strong>anyhow</strong> made error handling clean</p>
</li>
<li><p><strong>assert_cmd</strong> let us write integration tests</p>
</li>
<li><p><strong>Release profile</strong> cut binary size significantly</p>
</li>
</ul>
<p>The resulting binary starts in under 5ms and can be installed with a single command. Happy hacking!</p>
]]></content:encoded></item><item><title><![CDATA[Zero-Downtime Deployments on Kubernetes: Rolling Updates, Blue-Green, and Canary]]></title><description><![CDATA[Zero-Downtime Deployments on Kubernetes: Rolling Updates, Blue-Green, and Canary
Deploying without downtime isn't just about setting strategy: RollingUpdate. It's about health checks that actually ver]]></description><link>https://younggao.hashnode.dev/zero-downtime-deployments-on-kubernetes-rolling-updates-blue-green-and-canary</link><guid isPermaLink="true">https://younggao.hashnode.dev/zero-downtime-deployments-on-kubernetes-rolling-updates-blue-green-and-canary</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 06:30:39 GMT</pubDate><content:encoded><![CDATA[<h1>Zero-Downtime Deployments on Kubernetes: Rolling Updates, Blue-Green, and Canary</h1>
<p>Deploying without downtime isn't just about setting <code>strategy: RollingUpdate</code>. It's about health checks that actually verify readiness, connection draining that doesn't drop requests, and rollback triggers that catch problems before users do.</p>
<p>Here's how to set up each deployment strategy correctly on Kubernetes.</p>
<h2>The Basics: Why Deployments Fail</h2>
<p>Most "zero-downtime" deployments still drop requests because of three mistakes:</p>
<ol>
<li><p><strong>Readiness probes that lie</strong> — returning 200 before the app can serve traffic</p>
</li>
<li><p><strong>No graceful shutdown</strong> — pods killed mid-request</p>
</li>
<li><p><strong>Missing preStop hooks</strong> — pod removed from service before in-flight requests complete</p>
</li>
</ol>
<h2>Rolling Update (Default)</h2>
<p>Rolling updates replace pods incrementally. Kubernetes creates new pods before terminating old ones.</p>
<pre><code class="language-yaml"># deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: api-server
spec:
  replicas: 4
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1        # Create 1 extra pod during update
      maxUnavailable: 0   # Never reduce below desired count
  selector:
    matchLabels:
      app: api-server
  template:
    metadata:
      labels:
        app: api-server
    spec:
      terminationGracePeriodSeconds: 60
      containers:
        - name: api
          image: myregistry/api-server:v2.1.0
          ports:
            - containerPort: 8080

          # Readiness: "Can this pod serve traffic?"
          readinessProbe:
            httpGet:
              path: /ready
              port: 8080
            initialDelaySeconds: 5
            periodSeconds: 5
            failureThreshold: 3

          # Liveness: "Is this pod stuck/deadlocked?"
          livenessProbe:
            httpGet:
              path: /health
              port: 8080
            initialDelaySeconds: 15
            periodSeconds: 10
            failureThreshold: 3

          # Startup: "Is the app still starting up?"
          startupProbe:
            httpGet:
              path: /health
              port: 8080
            failureThreshold: 30
            periodSeconds: 2
            # Gives app up to 60s to start before liveness kicks in

          lifecycle:
            preStop:
              exec:
                command: ["sh", "-c", "sleep 10"]
                # Wait for endpoints controller to remove pod from Service
</code></pre>
<h3>Why <code>preStop: sleep 10</code>?</h3>
<p>When Kubernetes terminates a pod, two things happen <strong>concurrently</strong>:</p>
<ol>
<li><p>The Endpoints controller removes the pod from the Service</p>
</li>
<li><p>The container receives SIGTERM</p>
</li>
</ol>
<p>If SIGTERM arrives before the Endpoints update propagates, the load balancer still sends traffic to a pod that's shutting down. The <code>preStop</code> sleep gives time for the Endpoints change to propagate.</p>
<h3>Readiness vs Liveness vs Startup</h3>
<pre><code class="language-go">// In your Go server:
func main() {
    ready := false

    // Startup: load config, warm caches, connect to DB
    go func() {
        connectDB()
        warmCache()
        ready = true // Only now accept traffic
    }()

    http.HandleFunc("/health", func(w http.ResponseWriter, r *http.Request) {
        // Liveness: am I alive and not deadlocked?
        w.WriteHeader(http.StatusOK)
    })

    http.HandleFunc("/ready", func(w http.ResponseWriter, r *http.Request) {
        // Readiness: can I serve production traffic right now?
        if !ready {
            w.WriteHeader(http.StatusServiceUnavailable)
            return
        }
        // Optionally check DB connection
        if err := db.Ping(); err != nil {
            w.WriteHeader(http.StatusServiceUnavailable)
            return
        }
        w.WriteHeader(http.StatusOK)
    })

    http.ListenAndServe(":8080", nil)
}
</code></pre>
<p><strong>Never make liveness probes depend on external services.</strong> If your database goes down, liveness fails, Kubernetes restarts your pod, the new pod also can't reach the database, it restarts again — crash loop. Use readiness to stop traffic; use liveness only for detecting internal deadlocks.</p>
<h2>Blue-Green Deployment</h2>
<p>Run two identical environments. Switch traffic from blue (current) to green (new) atomically.</p>
<pre><code class="language-yaml"># blue-deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: api-server-blue
spec:
  replicas: 4
  selector:
    matchLabels:
      app: api-server
      version: blue
  template:
    metadata:
      labels:
        app: api-server
        version: blue
    spec:
      containers:
        - name: api
          image: myregistry/api-server:v2.0.0
---
# green-deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: api-server-green
spec:
  replicas: 4
  selector:
    matchLabels:
      app: api-server
      version: green
  template:
    metadata:
      labels:
        app: api-server
        version: green
    spec:
      containers:
        - name: api
          image: myregistry/api-server:v2.1.0
---
# service.yaml — switch by changing selector
apiVersion: v1
kind: Service
metadata:
  name: api-server
spec:
  selector:
    app: api-server
    version: blue   # Change to "green" to switch
  ports:
    - port: 80
      targetPort: 8080
</code></pre>
<p>Switch traffic:</p>
<pre><code class="language-bash"># Deploy green with new version
kubectl apply -f green-deployment.yaml

# Wait for all green pods to be ready
kubectl rollout status deployment/api-server-green

# Switch traffic (atomic — one API call)
kubectl patch service api-server -p '{"spec":{"selector":{"version":"green"}}}'

# Verify, then scale down blue
kubectl scale deployment api-server-blue --replicas=0
</code></pre>
<p><strong>Advantage:</strong> Instant rollback — just switch the Service selector back to "blue".</p>
<p><strong>Disadvantage:</strong> Requires 2x resources during deployment.</p>
<h2>Canary Deployment</h2>
<p>Route a small percentage of traffic to the new version. If metrics look good, gradually increase.</p>
<h3>Simple Canary with Replica Ratios</h3>
<pre><code class="language-yaml"># Stable: 9 replicas of v2.0.0
apiVersion: apps/v1
kind: Deployment
metadata:
  name: api-server-stable
spec:
  replicas: 9
  selector:
    matchLabels:
      app: api-server
      track: stable
  template:
    metadata:
      labels:
        app: api-server
        track: stable
    spec:
      containers:
        - name: api
          image: myregistry/api-server:v2.0.0
---
# Canary: 1 replica of v2.1.0 (gets ~10% of traffic)
apiVersion: apps/v1
kind: Deployment
metadata:
  name: api-server-canary
spec:
  replicas: 1
  selector:
    matchLabels:
      app: api-server
      track: canary
  template:
    metadata:
      labels:
        app: api-server
        track: canary
    spec:
      containers:
        - name: api
          image: myregistry/api-server:v2.1.0
---
# Service selects both — traffic split by replica count
apiVersion: v1
kind: Service
metadata:
  name: api-server
spec:
  selector:
    app: api-server    # Matches both stable and canary
  ports:
    - port: 80
      targetPort: 8080
</code></pre>
<p>Scale up canary gradually:</p>
<pre><code class="language-bash"># Start: 10% canary
kubectl scale deployment api-server-canary --replicas=1
kubectl scale deployment api-server-stable --replicas=9

# 50% canary
kubectl scale deployment api-server-canary --replicas=5
kubectl scale deployment api-server-stable --replicas=5

# 100% canary (promote)
kubectl scale deployment api-server-canary --replicas=10
kubectl scale deployment api-server-stable --replicas=0
</code></pre>
<h3>Automated Canary with Metrics</h3>
<p>Use a shell script (or Flagger/Argo Rollouts in production):</p>
<pre><code class="language-bash">#!/bin/bash
# canary-promote.sh

CANARY_DEPLOY="api-server-canary"
STABLE_DEPLOY="api-server-stable"
TOTAL_REPLICAS=10
ERROR_THRESHOLD=1  # percent

for pct in 10 25 50 75 100; do
    canary_replicas=$((TOTAL_REPLICAS * pct / 100))
    stable_replicas=$((TOTAL_REPLICAS - canary_replicas))

    echo "Setting canary to \({pct}% (\){canary_replicas} replicas)"
    kubectl scale deployment \(CANARY_DEPLOY --replicas=\)canary_replicas
    kubectl scale deployment \(STABLE_DEPLOY --replicas=\)stable_replicas

    # Wait and check error rate
    sleep 60

    # Query Prometheus for error rate (adjust query for your setup)
    error_rate=$(curl -s "http://prometheus:9090/api/v1/query?query=\
        rate(http_requests_total{deployment=\"${CANARY_DEPLOY}\",code=~\"5..\"}[1m])\
        /rate(http_requests_total{deployment=\"${CANARY_DEPLOY}\"}[1m])*100" \
        | jq '.data.result[0].value[1] // "0"' -r)

    if (( \((echo "\)error_rate &gt; $ERROR_THRESHOLD" | bc -l) )); then
        echo "Error rate ${error_rate}% exceeds threshold. Rolling back."
        kubectl scale deployment $CANARY_DEPLOY --replicas=0
        kubectl scale deployment \(STABLE_DEPLOY --replicas=\)TOTAL_REPLICAS
        exit 1
    fi

    echo "Error rate ${error_rate}% — looks good"
done

echo "Canary promoted to 100%"
</code></pre>
<h2>Graceful Shutdown</h2>
<p>Your application must handle SIGTERM correctly:</p>
<pre><code class="language-go">func main() {
    srv := &amp;http.Server{Addr: ":8080", Handler: mux}

    // Start server
    go func() {
        if err := srv.ListenAndServe(); err != http.ErrServerClosed {
            log.Fatal(err)
        }
    }()

    // Wait for SIGTERM
    quit := make(chan os.Signal, 1)
    signal.Notify(quit, syscall.SIGTERM, syscall.SIGINT)
    &lt;-quit

    log.Println("Shutting down — finishing in-flight requests...")

    // Give in-flight requests up to 30s to complete
    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()

    if err := srv.Shutdown(ctx); err != nil {
        log.Printf("Forced shutdown: %v", err)
    }

    log.Println("Server stopped")
}
</code></pre>
<p>The shutdown sequence:</p>
<ol>
<li><p>Kubernetes sends SIGTERM</p>
</li>
<li><p><code>preStop</code> hook runs (sleep 10 — lets Endpoints update propagate)</p>
</li>
<li><p>App receives SIGTERM, stops accepting new connections</p>
</li>
<li><p>App finishes in-flight requests (up to 30s)</p>
</li>
<li><p>App exits cleanly</p>
</li>
<li><p>Kubernetes waits up to <code>terminationGracePeriodSeconds</code> (60s), then SIGKILL</p>
</li>
</ol>
<h2>Pod Disruption Budgets</h2>
<p>Prevent too many pods from going down simultaneously (especially during node maintenance):</p>
<pre><code class="language-yaml">apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: api-server-pdb
spec:
  minAvailable: 3    # Always keep at least 3 pods running
  selector:
    matchLabels:
      app: api-server
</code></pre>
<h2>Quick Reference</h2>
<table>
<thead>
<tr>
<th>Strategy</th>
<th>Downtime</th>
<th>Rollback Speed</th>
<th>Resource Cost</th>
<th>Complexity</th>
</tr>
</thead>
<tbody><tr>
<td>Rolling</td>
<td>Zero</td>
<td>Minutes</td>
<td>1.25x</td>
<td>Low</td>
</tr>
<tr>
<td>Blue-Green</td>
<td>Zero</td>
<td>Seconds</td>
<td>2x</td>
<td>Medium</td>
</tr>
<tr>
<td>Canary</td>
<td>Zero</td>
<td>Seconds</td>
<td>1.1x–2x</td>
<td>High</td>
</tr>
</tbody></table>
<p><strong>Start with rolling updates</strong> (with proper probes and preStop hooks). Move to canary when you have metrics/monitoring in place. Use blue-green for databases or stateful services where you need instant rollback.</p>
<h2>Conclusion</h2>
<p>Zero-downtime deployments require getting three things right: readiness probes that genuinely verify readiness, graceful shutdown that drains connections, and preStop hooks that account for Endpoint propagation delay. The deployment strategy (rolling, blue-green, canary) is secondary — if your health checks lie, every strategy will drop requests.</p>
<hr />
<p><em>If this was helpful, you can support my work at</em> <a href="https://ko-fi.com/nopkt"><em>ko-fi.com/nopkt</em></a> ☕</p>
]]></content:encoded></item><item><title><![CDATA[Building a Production Rate Limiter from Scratch in Go]]></title><description><![CDATA[Building a Production Rate Limiter from Scratch in Go
Every API needs rate limiting. Most developers reach for a library or cloud service, but understanding the algorithms lets you pick the right one ]]></description><link>https://younggao.hashnode.dev/building-a-production-rate-limiter-from-scratch-in-go</link><guid isPermaLink="true">https://younggao.hashnode.dev/building-a-production-rate-limiter-from-scratch-in-go</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 06:30:19 GMT</pubDate><content:encoded><![CDATA[<h1>Building a Production Rate Limiter from Scratch in Go</h1>
<p>Every API needs rate limiting. Most developers reach for a library or cloud service, but understanding the algorithms lets you pick the right one and debug it when things go wrong. Here's how to build four different rate limiters in Go — from simple to production-grade.</p>
<h2>The Four Algorithms</h2>
<table>
<thead>
<tr>
<th>Algorithm</th>
<th>Behavior</th>
<th>Best For</th>
</tr>
</thead>
<tbody><tr>
<td>Fixed Window</td>
<td>Count requests per time window</td>
<td>Simple API limits</td>
</tr>
<tr>
<td>Sliding Window</td>
<td>Weighted count across windows</td>
<td>Smoother traffic shaping</td>
</tr>
<tr>
<td>Token Bucket</td>
<td>Tokens replenish at fixed rate</td>
<td>Burst-tolerant APIs</td>
</tr>
<tr>
<td>Sliding Window Log</td>
<td>Exact per-request tracking</td>
<td>Strict compliance</td>
</tr>
</tbody></table>
<h2>Shared Interface</h2>
<pre><code class="language-go">// limiter.go
package ratelimit

import "time"

type Result struct {
    Allowed   bool
    Remaining int
    ResetAt   time.Time
    RetryAt   time.Time // Zero if allowed
}

type Limiter interface {
    Allow(key string) Result
}
</code></pre>
<h2>1. Fixed Window Counter</h2>
<p>Simplest approach: count requests in fixed time intervals (e.g., 100 requests per minute). At the boundary of each window, the count resets.</p>
<pre><code class="language-go">// fixed_window.go
package ratelimit

import (
    "sync"
    "time"
)

type FixedWindow struct {
    limit    int
    window   time.Duration
    counters map[string]*counter
    mu       sync.Mutex
}

type counter struct {
    count    int
    windowStart time.Time
}

func NewFixedWindow(limit int, window time.Duration) *FixedWindow {
    return &amp;FixedWindow{
        limit:    limit,
        window:   window,
        counters: make(map[string]*counter),
    }
}

func (fw *FixedWindow) Allow(key string) Result {
    fw.mu.Lock()
    defer fw.mu.Unlock()

    now := time.Now()
    windowStart := now.Truncate(fw.window)

    c, exists := fw.counters[key]
    if !exists || c.windowStart != windowStart {
        c = &amp;counter{count: 0, windowStart: windowStart}
        fw.counters[key] = c
    }

    resetAt := windowStart.Add(fw.window)

    if c.count &gt;= fw.limit {
        return Result{
            Allowed:   false,
            Remaining: 0,
            ResetAt:   resetAt,
            RetryAt:   resetAt,
        }
    }

    c.count++
    return Result{
        Allowed:   true,
        Remaining: fw.limit - c.count,
        ResetAt:   resetAt,
    }
}
</code></pre>
<p><strong>Problem:</strong> At the boundary between two windows, a client can send <code>limit * 2</code> requests (e.g., 100 at 12:00:59 and 100 at 12:01:00). The sliding window fixes this.</p>
<h2>2. Sliding Window Counter</h2>
<p>Approximates a sliding window by weighting the previous window's count based on how far into the current window we are.</p>
<pre><code class="language-go">// sliding_window.go
package ratelimit

import (
    "sync"
    "time"
)

type SlidingWindow struct {
    limit    int
    window   time.Duration
    counters map[string]*windowPair
    mu       sync.Mutex
}

type windowPair struct {
    prevCount    int
    prevStart    time.Time
    currentCount int
    currentStart time.Time
}

func NewSlidingWindow(limit int, window time.Duration) *SlidingWindow {
    return &amp;SlidingWindow{
        limit:    limit,
        window:   window,
        counters: make(map[string]*windowPair),
    }
}

func (sw *SlidingWindow) Allow(key string) Result {
    sw.mu.Lock()
    defer sw.mu.Unlock()

    now := time.Now()
    currentWindowStart := now.Truncate(sw.window)

    wp, exists := sw.counters[key]
    if !exists {
        wp = &amp;windowPair{currentStart: currentWindowStart}
        sw.counters[key] = wp
    }

    // Rotate windows if needed
    if currentWindowStart != wp.currentStart {
        if currentWindowStart.Sub(wp.currentStart) == sw.window {
            wp.prevCount = wp.currentCount
            wp.prevStart = wp.currentStart
        } else {
            wp.prevCount = 0
        }
        wp.currentCount = 0
        wp.currentStart = currentWindowStart
    }

    // Weight: how far through the current window are we?
    elapsed := now.Sub(currentWindowStart)
    weight := float64(sw.window-elapsed) / float64(sw.window)
    estimated := int(float64(wp.prevCount)*weight) + wp.currentCount

    resetAt := currentWindowStart.Add(sw.window)

    if estimated &gt;= sw.limit {
        return Result{
            Allowed:   false,
            Remaining: 0,
            ResetAt:   resetAt,
            RetryAt:   now.Add(time.Duration(float64(sw.window) * (1 - weight))),
        }
    }

    wp.currentCount++
    return Result{
        Allowed:   true,
        Remaining: sw.limit - estimated - 1,
        ResetAt:   resetAt,
    }
}
</code></pre>
<p>This uses only two counters per key — O(1) memory regardless of request volume.</p>
<h2>3. Token Bucket</h2>
<p>Tokens accumulate at a fixed rate up to a maximum. Each request consumes one token. Allows bursts up to the bucket size while maintaining a long-term average rate.</p>
<pre><code class="language-go">// token_bucket.go
package ratelimit

import (
    "sync"
    "time"
)

type TokenBucket struct {
    rate       float64 // tokens per second
    bucketSize int
    buckets    map[string]*bucket
    mu         sync.Mutex
}

type bucket struct {
    tokens   float64
    lastFill time.Time
}

func NewTokenBucket(ratePerSecond float64, bucketSize int) *TokenBucket {
    return &amp;TokenBucket{
        rate:       ratePerSecond,
        bucketSize: bucketSize,
        buckets:    make(map[string]*bucket),
    }
}

func (tb *TokenBucket) Allow(key string) Result {
    tb.mu.Lock()
    defer tb.mu.Unlock()

    now := time.Now()

    b, exists := tb.buckets[key]
    if !exists {
        b = &amp;bucket{tokens: float64(tb.bucketSize), lastFill: now}
        tb.buckets[key] = b
    }

    // Add tokens based on elapsed time
    elapsed := now.Sub(b.lastFill).Seconds()
    b.tokens += elapsed * tb.rate
    if b.tokens &gt; float64(tb.bucketSize) {
        b.tokens = float64(tb.bucketSize)
    }
    b.lastFill = now

    if b.tokens &lt; 1 {
        // Calculate when next token arrives
        waitSeconds := (1 - b.tokens) / tb.rate
        return Result{
            Allowed:   false,
            Remaining: 0,
            RetryAt:   now.Add(time.Duration(waitSeconds * float64(time.Second))),
        }
    }

    b.tokens--
    return Result{
        Allowed:   true,
        Remaining: int(b.tokens),
    }
}
</code></pre>
<p>Token bucket is what most cloud APIs use (AWS, Stripe, GitHub). It naturally handles bursty traffic.</p>
<h2>4. HTTP Middleware</h2>
<pre><code class="language-go">// middleware.go
package ratelimit

import (
    "encoding/json"
    "fmt"
    "net/http"
    "strconv"
    "time"
)

func Middleware(limiter Limiter, keyFunc func(r *http.Request) string) func(http.Handler) http.Handler {
    return func(next http.Handler) http.Handler {
        return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
            key := keyFunc(r)
            result := limiter.Allow(key)

            // Always set rate limit headers (draft RFC 7231 extension)
            w.Header().Set("X-RateLimit-Remaining", strconv.Itoa(result.Remaining))
            if !result.ResetAt.IsZero() {
                w.Header().Set("X-RateLimit-Reset", strconv.FormatInt(result.ResetAt.Unix(), 10))
            }

            if !result.Allowed {
                w.Header().Set("Retry-After", fmt.Sprintf("%.0f", time.Until(result.RetryAt).Seconds()))
                w.Header().Set("Content-Type", "application/json")
                w.WriteHeader(http.StatusTooManyRequests)
                json.NewEncoder(w).Encode(map[string]any{
                    "error":      "rate limit exceeded",
                    "retry_after": time.Until(result.RetryAt).Seconds(),
                })
                return
            }

            next.ServeHTTP(w, r)
        })
    }
}

// Common key functions
func ByIP(r *http.Request) string {
    // Check X-Forwarded-For first (when behind a reverse proxy)
    if xff := r.Header.Get("X-Forwarded-For"); xff != "" {
        return "ip:" + xff
    }
    return "ip:" + r.RemoteAddr
}

func ByAPIKey(r *http.Request) string {
    key := r.Header.Get("X-API-Key")
    if key == "" {
        key = r.URL.Query().Get("api_key")
    }
    if key == "" {
        return ByIP(r) // Fallback to IP
    }
    return "key:" + key
}
</code></pre>
<h2>Usage</h2>
<pre><code class="language-go">// main.go
package main

import (
    "fmt"
    "log"
    "net/http"

    "yourmod/ratelimit"
)

func main() {
    // 100 requests per minute per IP
    limiter := ratelimit.NewSlidingWindow(100, time.Minute)

    mux := http.NewServeMux()
    mux.HandleFunc("/api/data", func(w http.ResponseWriter, r *http.Request) {
        fmt.Fprintf(w, `{"status": "ok"}`)
    })

    // Apply rate limiting
    handler := ratelimit.Middleware(limiter, ratelimit.ByIP)(mux)

    log.Println("Server starting on :8080")
    log.Fatal(http.ListenAndServe(":8080", handler))
}
</code></pre>
<h2>Distributed Rate Limiting with Redis</h2>
<p>The in-memory implementations above work for single servers. For distributed systems, use Redis:</p>
<pre><code class="language-go">// redis_limiter.go
package ratelimit

import (
    "context"
    "time"

    "github.com/redis/go-redis/v9"
)

type RedisLimiter struct {
    client *redis.Client
    limit  int
    window time.Duration
}

func NewRedisLimiter(client *redis.Client, limit int, window time.Duration) *RedisLimiter {
    return &amp;RedisLimiter{client: client, limit: limit, window: window}
}

// Lua script ensures atomic increment + expiry check
var slidingWindowScript = redis.NewScript(`
    local key = KEYS[1]
    local limit = tonumber(ARGV[1])
    local window = tonumber(ARGV[2])
    local now = tonumber(ARGV[3])

    -- Remove expired entries
    redis.call('ZREMRANGEBYSCORE', key, 0, now - window)

    -- Count current entries
    local count = redis.call('ZCARD', key)

    if count &lt; limit then
        -- Add this request
        redis.call('ZADD', key, now, now .. '-' .. math.random(1000000))
        redis.call('PEXPIRE', key, window)
        return {1, limit - count - 1}
    end

    return {0, 0}
`)

func (rl *RedisLimiter) Allow(key string) Result {
    ctx := context.Background()
    now := time.Now()

    result, err := slidingWindowScript.Run(ctx, rl.client,
        []string{"ratelimit:" + key},
        rl.limit,
        rl.window.Milliseconds(),
        now.UnixMilli(),
    ).Int64Slice()

    if err != nil {
        // On Redis failure, allow the request (fail open)
        return Result{Allowed: true, Remaining: rl.limit}
    }

    return Result{
        Allowed:   result[0] == 1,
        Remaining: int(result[1]),
        RetryAt:   now.Add(rl.window),
    }
}
</code></pre>
<p>Using a Lua script makes the check-and-increment atomic — no race conditions even under high concurrency.</p>
<h2>Choosing the Right Algorithm</h2>
<ul>
<li><p><strong>Fixed Window</strong>: Simplest. Use when exact precision at boundaries doesn't matter (internal APIs, development).</p>
</li>
<li><p><strong>Sliding Window Counter</strong>: Best default. Low memory, smooth limiting, good enough for most production APIs.</p>
</li>
<li><p><strong>Token Bucket</strong>: When you want to allow bursts. Good for user-facing APIs where occasional spikes are expected.</p>
</li>
<li><p><strong>Sliding Window Log</strong> (Redis sorted set): When you need exact counting. Higher memory usage but zero approximation error.</p>
</li>
</ul>
<h2>Memory Cleanup</h2>
<p>For in-memory limiters, run periodic cleanup:</p>
<pre><code class="language-go">func (fw *FixedWindow) Cleanup() {
    fw.mu.Lock()
    defer fw.mu.Unlock()

    cutoff := time.Now().Add(-fw.window * 2)
    for key, c := range fw.counters {
        if c.windowStart.Before(cutoff) {
            delete(fw.counters, key)
        }
    }
}

// Run every minute
go func() {
    ticker := time.NewTicker(time.Minute)
    for range ticker.C {
        limiter.Cleanup()
    }
}()
</code></pre>
<h2>Conclusion</h2>
<p>Rate limiting isn't complex, but picking the wrong algorithm causes either false rejections (lost revenue) or ineffective protection (outages). Start with sliding window counters for single servers, move to Redis when you scale horizontally. Always set <code>Retry-After</code> headers — good clients will back off automatically.</p>
<hr />
<p><em>If this was helpful, you can support my work at</em> <a href="https://ko-fi.com/nopkt"><em>ko-fi.com/nopkt</em></a> ☕</p>
]]></content:encoded></item><item><title><![CDATA[Server-Sent Events: The Underrated Alternative to WebSockets for Real-Time Notifications]]></title><description><![CDATA[Server-Sent Events: The Underrated Alternative to WebSockets for Real-Time Notifications
WebSockets are the default answer for real-time features. But for most notification systems — activity feeds, l]]></description><link>https://younggao.hashnode.dev/server-sent-events-the-underrated-alternative-to-websockets-for-real-time-notifications</link><guid isPermaLink="true">https://younggao.hashnode.dev/server-sent-events-the-underrated-alternative-to-websockets-for-real-time-notifications</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 06:29:58 GMT</pubDate><content:encoded><![CDATA[<h1>Server-Sent Events: The Underrated Alternative to WebSockets for Real-Time Notifications</h1>
<p>WebSockets are the default answer for real-time features. But for most notification systems — activity feeds, live dashboards, deployment status updates — they're overkill. Server-Sent Events (SSE) give you real-time push over plain HTTP with zero client-side libraries, automatic reconnection, and drastically simpler server code.</p>
<p>Here's how to build a production notification system using SSE with Node.js and TypeScript.</p>
<h2>Why SSE Over WebSockets</h2>
<table>
<thead>
<tr>
<th>Feature</th>
<th>SSE</th>
<th>WebSocket</th>
</tr>
</thead>
<tbody><tr>
<td>Direction</td>
<td>Server → Client</td>
<td>Bidirectional</td>
</tr>
<tr>
<td>Protocol</td>
<td>HTTP</td>
<td>WS (separate protocol)</td>
</tr>
<tr>
<td>Reconnection</td>
<td>Built-in</td>
<td>Manual implementation</td>
</tr>
<tr>
<td>Auth</td>
<td>Standard HTTP headers/cookies</td>
<td>Custom handshake</td>
</tr>
<tr>
<td>Load balancers</td>
<td>Works everywhere</td>
<td>Needs sticky sessions or upgrade support</td>
</tr>
<tr>
<td>Client code</td>
<td><code>new EventSource(url)</code></td>
<td><code>new WebSocket(url)</code> + message framing</td>
</tr>
<tr>
<td>Browser support</td>
<td>All modern browsers</td>
<td>All modern browsers</td>
</tr>
</tbody></table>
<p><strong>If your client only needs to receive updates</strong> (notifications, live feeds, progress bars), SSE is the simpler, more reliable choice.</p>
<h2>The Protocol in 30 Seconds</h2>
<p>SSE is just an HTTP response with <code>Content-Type: text/event-stream</code> that stays open. The server writes lines in this format:</p>
<pre><code class="language-yaml">event: notification
data: {"id": 1, "message": "Deploy succeeded"}
id: 1

event: notification
data: {"id": 2, "message": "New comment on PR #42"}
id: 2
</code></pre>
<p>Each message is separated by a blank line. That's it. The browser handles parsing, reconnection, and event dispatching.</p>
<h2>Project Setup</h2>
<pre><code class="language-bash">mkdir sse-notifications &amp;&amp; cd sse-notifications
npm init -y
npm install express cors
npm install -D typescript @types/express @types/cors tsx
</code></pre>
<pre><code class="language-json">// tsconfig.json
{
  "compilerOptions": {
    "target": "ES2022",
    "module": "NodeNext",
    "moduleResolution": "NodeNext",
    "strict": true,
    "esModuleInterop": true,
    "outDir": "dist"
  }
}
</code></pre>
<h2>Step 1: The SSE Connection Manager</h2>
<pre><code class="language-typescript">// src/sse.ts
import { Response } from 'express';

interface Client {
  id: string;
  userId: string;
  res: Response;
  connectedAt: Date;
  lastEventId: number;
}

class SSEManager {
  private clients = new Map&lt;string, Client&gt;();
  private eventCounter = 0;
  private eventBuffer: Array&lt;{
    id: number;
    event: string;
    data: string;
    userId: string | null; // null = broadcast
    timestamp: number;
  }&gt; = [];
  private readonly BUFFER_SIZE = 1000;
  private readonly BUFFER_TTL = 5 * 60 * 1000; // 5 minutes

  addClient(userId: string, res: Response, lastEventId?: string): string {
    const clientId = `\({userId}-\){Date.now()}-${Math.random().toString(36).slice(2, 8)}`;

    // Set SSE headers
    res.writeHead(200, {
      'Content-Type': 'text/event-stream',
      'Cache-Control': 'no-cache',
      Connection: 'keep-alive',
      'X-Accel-Buffering': 'no', // Disable nginx buffering
    });

    // Send initial connection event
    res.write(`event: connected\ndata: ${JSON.stringify({ clientId })}\n\n`);

    const client: Client = {
      id: clientId,
      userId,
      res,
      connectedAt: new Date(),
      lastEventId: lastEventId ? parseInt(lastEventId) : 0,
    };

    this.clients.set(clientId, client);

    // Replay missed events if client reconnected
    if (client.lastEventId &gt; 0) {
      this.replayEvents(client);
    }

    // Send heartbeat every 30s to keep connection alive
    const heartbeat = setInterval(() =&gt; {
      try {
        res.write(': heartbeat\n\n');
      } catch {
        clearInterval(heartbeat);
      }
    }, 30000);

    // Clean up on disconnect
    res.on('close', () =&gt; {
      clearInterval(heartbeat);
      this.clients.delete(clientId);
    });

    return clientId;
  }

  // Send event to a specific user (all their connected devices)
  sendToUser(userId: string, event: string, data: unknown): void {
    const id = ++this.eventCounter;
    const payload = JSON.stringify(data);

    this.bufferEvent(id, event, payload, userId);

    for (const client of this.clients.values()) {
      if (client.userId === userId) {
        this.writeEvent(client.res, id, event, payload);
      }
    }
  }

  // Broadcast to all connected clients
  broadcast(event: string, data: unknown): void {
    const id = ++this.eventCounter;
    const payload = JSON.stringify(data);

    this.bufferEvent(id, event, payload, null);

    for (const client of this.clients.values()) {
      this.writeEvent(client.res, id, event, payload);
    }
  }

  getStats() {
    const userCounts = new Map&lt;string, number&gt;();
    for (const client of this.clients.values()) {
      userCounts.set(client.userId, (userCounts.get(client.userId) || 0) + 1);
    }
    return {
      totalConnections: this.clients.size,
      uniqueUsers: userCounts.size,
      eventsBuffered: this.eventBuffer.length,
      totalEventsSent: this.eventCounter,
    };
  }

  private writeEvent(res: Response, id: number, event: string, data: string): void {
    try {
      res.write(`id: \({id}\nevent: \){event}\ndata: ${data}\n\n`);
    } catch {
      // Client disconnected — cleanup happens via the 'close' handler
    }
  }

  private bufferEvent(id: number, event: string, data: string, userId: string | null): void {
    this.eventBuffer.push({ id, event, data, userId, timestamp: Date.now() });

    // Prune old events
    const cutoff = Date.now() - this.BUFFER_TTL;
    while (this.eventBuffer.length &gt; this.BUFFER_SIZE ||
           (this.eventBuffer.length &gt; 0 &amp;&amp; this.eventBuffer[0].timestamp &lt; cutoff)) {
      this.eventBuffer.shift();
    }
  }

  private replayEvents(client: Client): void {
    const missed = this.eventBuffer.filter(
      (e) =&gt; e.id &gt; client.lastEventId &amp;&amp; (e.userId === null || e.userId === client.userId)
    );

    for (const event of missed) {
      this.writeEvent(client.res, event.id, event.event, event.data);
    }
  }
}

export const sseManager = new SSEManager();
</code></pre>
<p>The key features:</p>
<ul>
<li><p><strong>Event buffering</strong> — when a client reconnects, <code>Last-Event-ID</code> header tells us where they left off, and we replay missed events</p>
</li>
<li><p><strong>Heartbeats</strong> — keep the connection alive through proxies and load balancers</p>
</li>
<li><p><code>X-Accel-Buffering: no</code> — prevents nginx from buffering the stream (a common gotcha)</p>
</li>
<li><p><strong>Per-user targeting</strong> — send notifications to specific users across all their devices</p>
</li>
</ul>
<h2>Step 2: Notification Service</h2>
<pre><code class="language-typescript">// src/notifications.ts
import { sseManager } from './sse.js';

interface Notification {
  id: string;
  type: 'info' | 'success' | 'warning' | 'error';
  title: string;
  message: string;
  link?: string;
  timestamp: string;
}

// In-memory store (use Redis or a database in production)
const notifications = new Map&lt;string, Notification[]&gt;();

export function createNotification(
  userId: string,
  type: Notification['type'],
  title: string,
  message: string,
  link?: string
): Notification {
  const notification: Notification = {
    id: crypto.randomUUID(),
    type,
    title,
    message,
    link,
    timestamp: new Date().toISOString(),
  };

  // Store
  const userNotifs = notifications.get(userId) || [];
  userNotifs.unshift(notification);
  if (userNotifs.length &gt; 100) userNotifs.pop(); // Keep last 100
  notifications.set(userId, userNotifs);

  // Push via SSE
  sseManager.sendToUser(userId, 'notification', notification);

  return notification;
}

export function getNotifications(userId: string, limit = 20): Notification[] {
  return (notifications.get(userId) || []).slice(0, limit);
}

export function broadcastAnnouncement(title: string, message: string): void {
  sseManager.broadcast('announcement', {
    id: crypto.randomUUID(),
    type: 'info',
    title,
    message,
    timestamp: new Date().toISOString(),
  });
}
</code></pre>
<h2>Step 3: Express Routes</h2>
<pre><code class="language-typescript">// src/server.ts
import express from 'express';
import cors from 'cors';
import { sseManager } from './sse.js';
import { createNotification, getNotifications, broadcastAnnouncement } from './notifications.js';

const app = express();
app.use(cors());
app.use(express.json());

// SSE endpoint — client connects here
app.get('/events/:userId', (req, res) =&gt; {
  const { userId } = req.params;
  const lastEventId = req.headers['last-event-id'] as string | undefined;

  const clientId = sseManager.addClient(userId, res, lastEventId);
  console.log(`Client connected: \({clientId} (user: \){userId})`);
});

// REST: Get notification history
app.get('/api/notifications/:userId', (req, res) =&gt; {
  const limit = parseInt(req.query.limit as string) || 20;
  const notifs = getNotifications(req.params.userId, limit);
  res.json({ notifications: notifs, count: notifs.length });
});

// REST: Send notification to user
app.post('/api/notifications/:userId', (req, res) =&gt; {
  const { type = 'info', title, message, link } = req.body;

  if (!title || !message) {
    return res.status(400).json({ error: 'title and message are required' });
  }

  const notification = createNotification(req.params.userId, type, title, message, link);
  res.status(201).json(notification);
});

// REST: Broadcast to all
app.post('/api/broadcast', (req, res) =&gt; {
  const { title, message } = req.body;

  if (!title || !message) {
    return res.status(400).json({ error: 'title and message are required' });
  }

  broadcastAnnouncement(title, message);
  res.json({ sent: true, connections: sseManager.getStats().totalConnections });
});

// Stats
app.get('/api/stats', (_, res) =&gt; {
  res.json(sseManager.getStats());
});

// Demo page
app.get('/', (_, res) =&gt; {
  res.send(`&lt;!DOCTYPE html&gt;
&lt;html&gt;
&lt;head&gt;&lt;title&gt;SSE Notifications Demo&lt;/title&gt;&lt;/head&gt;
&lt;body&gt;
  &lt;h1&gt;Notifications&lt;/h1&gt;
  &lt;div id="notifs"&gt;&lt;/div&gt;
  &lt;script&gt;
    const userId = 'demo-user';
    const events = new EventSource('/events/' + userId);

    events.addEventListener('connected', (e) =&gt; {
      console.log('Connected:', JSON.parse(e.data));
    });

    events.addEventListener('notification', (e) =&gt; {
      const n = JSON.parse(e.data);
      const div = document.createElement('div');
      div.style.cssText = 'padding:12px;margin:8px 0;border-radius:8px;background:#f0f4ff;border-left:4px solid #4f46e5';
      div.innerHTML = '&lt;strong&gt;' + n.title + '&lt;/strong&gt;&lt;br&gt;' + n.message + '&lt;br&gt;&lt;small&gt;' + n.timestamp + '&lt;/small&gt;';
      document.getElementById('notifs').prepend(div);
    });

    events.addEventListener('announcement', (e) =&gt; {
      const n = JSON.parse(e.data);
      const div = document.createElement('div');
      div.style.cssText = 'padding:12px;margin:8px 0;border-radius:8px;background:#fef3c7;border-left:4px solid #f59e0b';
      div.innerHTML = '📢 &lt;strong&gt;' + n.title + '&lt;/strong&gt;&lt;br&gt;' + n.message;
      document.getElementById('notifs').prepend(div);
    });

    events.onerror = () =&gt; console.log('Reconnecting...');
  &lt;/script&gt;
&lt;/body&gt;
&lt;/html&gt;`);
});

const PORT = process.env.PORT || 3000;
app.listen(PORT, () =&gt; console.log(`Server running on http://localhost:${PORT}`));
</code></pre>
<h2>Step 4: Testing It</h2>
<pre><code class="language-bash">npx tsx src/server.ts
</code></pre>
<p>In another terminal:</p>
<pre><code class="language-bash"># Send a notification
curl -X POST http://localhost:3000/api/notifications/demo-user \
  -H "Content-Type: application/json" \
  -d '{"type": "success", "title": "Deploy Complete", "message": "v2.3.1 deployed to production"}'

# Broadcast to all
curl -X POST http://localhost:3000/api/broadcast \
  -H "Content-Type: application/json" \
  -d '{"title": "Scheduled Maintenance", "message": "Database maintenance at 2am UTC"}'

# Check stats
curl http://localhost:3000/api/stats
</code></pre>
<p>Open <code>http://localhost:3000</code> in your browser — notifications appear instantly without polling.</p>
<h2>Production Considerations</h2>
<h3>Scaling Beyond One Server</h3>
<p>SSE connections are stateful — each lives on a specific server. For multiple servers, use Redis pub/sub:</p>
<pre><code class="language-typescript">// src/sse-redis.ts
import { createClient } from 'redis';
import { sseManager } from './sse.js';

const pub = createClient({ url: process.env.REDIS_URL });
const sub = pub.duplicate();

await pub.connect();
await sub.connect();

// Publish events to Redis (instead of calling sseManager directly)
export async function publishNotification(userId: string, event: string, data: unknown) {
  await pub.publish('sse:events', JSON.stringify({ userId, event, data }));
}

// Each server subscribes and delivers to its local clients
await sub.subscribe('sse:events', (message) =&gt; {
  const { userId, event, data } = JSON.parse(message);
  if (userId) {
    sseManager.sendToUser(userId, event, data);
  } else {
    sseManager.broadcast(event, data);
  }
});
</code></pre>
<h3>Connection Limits</h3>
<p>Each SSE connection holds an open HTTP connection. Default Node.js limit is ~1000 concurrent sockets. Increase it:</p>
<pre><code class="language-typescript">import http from 'http';
const server = http.createServer(app);
server.maxConnections = 10000;
server.listen(PORT);
</code></pre>
<p>For higher numbers, use <code>uWebSockets.js</code> or put a reverse proxy (nginx with <code>proxy_buffering off</code>) in front.</p>
<h3>Authentication</h3>
<p>SSE uses standard HTTP, so auth works normally:</p>
<pre><code class="language-typescript">app.get('/events/:userId', authenticate, (req, res) =&gt; {
  // authenticate middleware validates JWT from cookie or Authorization header
  if (req.user.id !== req.params.userId) {
    return res.status(403).json({ error: 'Forbidden' });
  }
  sseManager.addClient(req.user.id, res);
});
</code></pre>
<p>The <code>EventSource</code> API doesn't support custom headers, so use cookies or query params for auth:</p>
<pre><code class="language-javascript">// Client-side with auth token
const events = new EventSource('/events/me?token=' + authToken);
</code></pre>
<p>Or use the <code>fetch</code>-based polyfill for custom headers:</p>
<pre><code class="language-javascript">// Using fetch for SSE with headers
const response = await fetch('/events/me', {
  headers: { Authorization: `Bearer ${token}` },
});
const reader = response.body.getReader();
const decoder = new TextDecoder();

while (true) {
  const { done, value } = await reader.read();
  if (done) break;
  const text = decoder.decode(value);
  // Parse SSE format manually
  for (const line of text.split('\n')) {
    if (line.startsWith('data: ')) {
      const data = JSON.parse(line.slice(6));
      handleEvent(data);
    }
  }
}
</code></pre>
<h3>Graceful Shutdown</h3>
<pre><code class="language-typescript">process.on('SIGTERM', () =&gt; {
  // Close all SSE connections — clients will auto-reconnect to another server
  for (const client of sseManager.clients.values()) {
    client.res.end();
  }
  server.close();
});
</code></pre>
<h2>When to Use WebSockets Instead</h2>
<p>SSE is wrong when:</p>
<ul>
<li><p><strong>Clients send frequent messages</strong> (chat, collaborative editing, gaming)</p>
</li>
<li><p><strong>Binary data streaming</strong> (SSE is text-only)</p>
</li>
<li><p><strong>You need sub-10ms latency</strong> (SSE adds HTTP overhead)</p>
</li>
</ul>
<p>SSE is right when:</p>
<ul>
<li><p><strong>Server pushes updates to clients</strong> (notifications, feeds, dashboards)</p>
</li>
<li><p><strong>You want simplicity</strong> (no library needed, standard HTTP)</p>
</li>
<li><p><strong>Reliability matters</strong> (auto-reconnect with event replay)</p>
</li>
</ul>
<h2>Conclusion</h2>
<p>SSE gives you real-time server-to-client push with less code, better reliability, and simpler infrastructure than WebSockets. The built-in reconnection with <code>Last-Event-ID</code> means clients never miss events, even across network drops. For notification systems, live dashboards, and status feeds, it's the pragmatic choice.</p>
<hr />
<p><em>If this was helpful, you can support my work at</em> <a href="https://ko-fi.com/nopkt"><em>ko-fi.com/nopkt</em></a></p>
]]></content:encoded></item><item><title><![CDATA[Testing Strategies for TypeScript APIs That Actually Catch Bugs]]></title><description><![CDATA[Testing Strategies for TypeScript APIs That Actually Catch Bugs
Most API test suites are theater. They test that 200 OK is returned for valid input, but miss the edge cases that break production — rac]]></description><link>https://younggao.hashnode.dev/testing-strategies-for-typescript-apis-that-actually-catch-bugs</link><guid isPermaLink="true">https://younggao.hashnode.dev/testing-strategies-for-typescript-apis-that-actually-catch-bugs</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 06:29:17 GMT</pubDate><content:encoded><![CDATA[<h1>Testing Strategies for TypeScript APIs That Actually Catch Bugs</h1>
<p>Most API test suites are theater. They test that 200 OK is returned for valid input, but miss the edge cases that break production — race conditions, partial failures, invalid state transitions, and subtle serialization bugs.</p>
<p>Here's a testing strategy built from real production incidents, organized by what each test type catches and how much it costs to maintain.</p>
<h2>The Testing Pyramid for APIs</h2>
<pre><code class="language-plaintext">        /  E2E  \        — 5% of tests, catch integration failures
       / Contract \       — 15%, catch API compatibility breaks
      / Integration \     — 30%, catch database + service bugs
     /    Unit Tests  \   — 50%, catch logic bugs fast
</code></pre>
<h2>Project Setup</h2>
<pre><code class="language-bash">mkdir api-testing &amp;&amp; cd api-testing
npm init -y
npm install express zod drizzle-orm better-sqlite3
npm install -D vitest supertest @types/supertest testcontainers msw
</code></pre>
<pre><code class="language-typescript">// vitest.config.ts
import { defineConfig } from 'vitest/config';

export default defineConfig({
  test: {
    globals: true,
    environment: 'node',
    coverage: {
      provider: 'v8',
      reporter: ['text', 'lcov'],
      exclude: ['**/test/**', '**/*.test.ts'],
    },
    // Run tests in sequence for integration tests
    pool: 'forks',
    poolOptions: {
      forks: { singleFork: true },
    },
  },
});
</code></pre>
<h2>Unit Tests: Pure Logic, No I/O</h2>
<p>Unit tests should cover business logic — validation, transformations, calculations. No database, no HTTP, no file system.</p>
<pre><code class="language-typescript">// src/services/pricing.ts
export function calculateDiscount(
  basePrice: number,
  quantity: number,
  customerTier: 'standard' | 'premium' | 'enterprise'
): { price: number; discount: number; total: number } {
  if (basePrice &lt; 0 || quantity &lt; 1) {
    throw new Error('Invalid input');
  }

  const tierMultipliers = { standard: 0, premium: 0.1, enterprise: 0.2 };
  const volumeDiscount = quantity &gt;= 100 ? 0.15 : quantity &gt;= 10 ? 0.05 : 0;
  const discount = Math.min(tierMultipliers[customerTier] + volumeDiscount, 0.35);
  const total = basePrice * quantity * (1 - discount);

  return { price: basePrice, discount, total: Math.round(total * 100) / 100 };
}
</code></pre>
<pre><code class="language-typescript">// src/services/pricing.test.ts
import { describe, it, expect } from 'vitest';
import { calculateDiscount } from './pricing';

describe('calculateDiscount', () =&gt; {
  it('applies no discount for standard tier, small quantity', () =&gt; {
    const result = calculateDiscount(100, 5, 'standard');
    expect(result).toEqual({ price: 100, discount: 0, total: 500 });
  });

  it('caps combined discount at 35%', () =&gt; {
    // Enterprise (20%) + volume 100+ (15%) = 35% cap
    const result = calculateDiscount(100, 100, 'enterprise');
    expect(result.discount).toBe(0.35);
    expect(result.total).toBe(6500);
  });

  it('stacks tier and volume discounts', () =&gt; {
    // Premium (10%) + volume 10+ (5%) = 15%
    const result = calculateDiscount(50, 20, 'premium');
    expect(result.discount).toBe(0.15);
    expect(result.total).toBe(850);
  });

  it('rejects negative prices', () =&gt; {
    expect(() =&gt; calculateDiscount(-10, 1, 'standard')).toThrow('Invalid input');
  });

  it('rejects zero quantity', () =&gt; {
    expect(() =&gt; calculateDiscount(100, 0, 'standard')).toThrow('Invalid input');
  });

  // This test caught a real bug — floating point precision
  it('rounds total to 2 decimal places', () =&gt; {
    const result = calculateDiscount(19.99, 3, 'premium');
    // Without rounding: 19.99 * 3 * 0.9 = 53.973
    expect(result.total).toBe(53.97);
  });
});
</code></pre>
<p><strong>What unit tests catch:</strong> Logic errors, edge cases, floating point bugs, off-by-one errors. They're fast (run in &lt;1ms each) and stable (no external dependencies).</p>
<h2>Integration Tests: Real Database, Real Queries</h2>
<p>Integration tests hit the actual database. Mocking SQL queries is worse than useless — it tests that your mock matches your assumptions, not that your query works.</p>
<pre><code class="language-typescript">// src/repositories/user.integration.test.ts
import { describe, it, expect, beforeAll, afterAll, beforeEach } from 'vitest';
import { drizzle } from 'drizzle-orm/better-sqlite3';
import Database from 'better-sqlite3';
import { migrate } from 'drizzle-orm/better-sqlite3/migrator';
import { users } from '../schema';
import { UserRepository } from './user';

describe('UserRepository', () =&gt; {
  let db: ReturnType&lt;typeof drizzle&gt;;
  let sqlite: Database.Database;
  let repo: UserRepository;

  beforeAll(() =&gt; {
    sqlite = new Database(':memory:');
    db = drizzle(sqlite);
    migrate(db, { migrationsFolder: './drizzle' });
    repo = new UserRepository(db);
  });

  afterAll(() =&gt; {
    sqlite.close();
  });

  beforeEach(() =&gt; {
    // Clean slate for each test
    db.delete(users).run();
  });

  it('creates a user and retrieves by ID', async () =&gt; {
    const created = await repo.create({
      email: 'test@example.com',
      name: 'Test User',
    });

    expect(created.id).toBeDefined();
    expect(created.email).toBe('test@example.com');

    const found = await repo.findById(created.id);
    expect(found).toEqual(created);
  });

  it('enforces unique email constraint', async () =&gt; {
    await repo.create({ email: 'dup@example.com', name: 'User 1' });

    await expect(
      repo.create({ email: 'dup@example.com', name: 'User 2' })
    ).rejects.toThrow(/UNIQUE constraint/);
  });

  it('handles pagination correctly at boundaries', async () =&gt; {
    // Create exactly 25 users
    for (let i = 0; i &lt; 25; i++) {
      await repo.create({ email: `user\({i}@test.com`, name: `User \){i}` });
    }

    const page1 = await repo.list({ page: 1, limit: 10 });
    const page3 = await repo.list({ page: 3, limit: 10 });

    expect(page1.items).toHaveLength(10);
    expect(page3.items).toHaveLength(5);
    expect(page3.totalPages).toBe(3);
  });

  it('soft-deletes without removing from database', async () =&gt; {
    const user = await repo.create({ email: 'del@test.com', name: 'Delete Me' });
    await repo.softDelete(user.id);

    // findById should not return soft-deleted users
    const found = await repo.findById(user.id);
    expect(found).toBeNull();

    // But the row still exists (for audit/recovery)
    const raw = sqlite.prepare('SELECT * FROM users WHERE id = ?').get(user.id);
    expect(raw).toBeDefined();
    expect((raw as any).deleted_at).not.toBeNull();
  });
});
</code></pre>
<p><strong>What integration tests catch:</strong> SQL bugs, constraint violations, migration issues, ORM misuse, query performance regressions. Using SQLite in-memory is fast enough for most tests (sub-second for hundreds of queries).</p>
<h2>API Tests: HTTP Layer + Middleware</h2>
<p>Test the full HTTP stack — routing, middleware, serialization, error handling.</p>
<pre><code class="language-typescript">// src/app.test.ts
import { describe, it, expect, beforeAll, afterAll } from 'vitest';
import request from 'supertest';
import { createApp } from './app';

describe('API endpoints', () =&gt; {
  let app: ReturnType&lt;typeof createApp&gt;;

  beforeAll(() =&gt; {
    app = createApp({ database: ':memory:' });
  });

  describe('POST /api/users', () =&gt; {
    it('creates user with valid input', async () =&gt; {
      const res = await request(app)
        .post('/api/users')
        .send({ email: 'new@test.com', name: 'New User' })
        .expect(201);

      expect(res.body).toMatchObject({
        email: 'new@test.com',
        name: 'New User',
      });
      expect(res.body.id).toBeDefined();
      expect(res.body.createdAt).toBeDefined();
    });

    it('returns 400 with validation details for bad input', async () =&gt; {
      const res = await request(app)
        .post('/api/users')
        .send({ email: 'not-an-email', name: '' })
        .expect(400);

      expect(res.body.error).toBe('Validation failed');
      expect(res.body.details).toHaveLength(2);
      expect(res.body.details[0].path).toContain('email');
      expect(res.body.details[1].path).toContain('name');
    });

    it('returns 409 for duplicate email', async () =&gt; {
      await request(app)
        .post('/api/users')
        .send({ email: 'dup@test.com', name: 'First' });

      const res = await request(app)
        .post('/api/users')
        .send({ email: 'dup@test.com', name: 'Second' })
        .expect(409);

      expect(res.body.error).toContain('already exists');
    });
  });

  describe('GET /api/users', () =&gt; {
    it('returns paginated results with correct headers', async () =&gt; {
      // Seed some data
      for (let i = 0; i &lt; 15; i++) {
        await request(app)
          .post('/api/users')
          .send({ email: `page\({i}@test.com`, name: `User \){i}` });
      }

      const res = await request(app)
        .get('/api/users?page=2&amp;limit=10')
        .expect(200);

      expect(res.body.items).toHaveLength(5);
      expect(res.body.pagination.page).toBe(2);
      expect(res.body.pagination.totalPages).toBe(2);
    });

    it('ignores unknown query parameters', async () =&gt; {
      await request(app)
        .get('/api/users?foo=bar&amp;page=1')
        .expect(200);
    });
  });

  describe('Authentication', () =&gt; {
    it('returns 401 for missing token', async () =&gt; {
      await request(app).get('/api/admin/stats').expect(401);
    });

    it('returns 401 for expired token', async () =&gt; {
      const expiredToken = createExpiredJWT();
      await request(app)
        .get('/api/admin/stats')
        .set('Authorization', `Bearer ${expiredToken}`)
        .expect(401);
    });
  });
});
</code></pre>
<p><strong>What API tests catch:</strong> Routing bugs, middleware ordering issues, status code mistakes, response format errors, authentication bypass.</p>
<h2>Contract Tests: API Compatibility</h2>
<p>When other services depend on your API, contract tests prevent breaking changes.</p>
<pre><code class="language-typescript">// src/contracts/user-api.contract.test.ts
import { describe, it, expect } from 'vitest';
import { z } from 'zod';
import request from 'supertest';
import { createApp } from '../app';

// This schema represents what consumers expect
const UserResponseContract = z.object({
  id: z.string().uuid(),
  email: z.string().email(),
  name: z.string(),
  createdAt: z.string().datetime(),
  updatedAt: z.string().datetime(),
});

const PaginatedUsersContract = z.object({
  items: z.array(UserResponseContract),
  pagination: z.object({
    page: z.number(),
    limit: z.number(),
    total: z.number(),
    totalPages: z.number(),
  }),
});

describe('User API contract', () =&gt; {
  const app = createApp({ database: ':memory:' });

  it('POST /api/users response matches contract', async () =&gt; {
    const res = await request(app)
      .post('/api/users')
      .send({ email: 'contract@test.com', name: 'Contract Test' });

    const parsed = UserResponseContract.safeParse(res.body);
    expect(parsed.success).toBe(true);
  });

  it('GET /api/users response matches paginated contract', async () =&gt; {
    const res = await request(app).get('/api/users');

    const parsed = PaginatedUsersContract.safeParse(res.body);
    expect(parsed.success).toBe(true);
  });

  it('error responses always include error field', async () =&gt; {
    const res = await request(app)
      .post('/api/users')
      .send({ email: 'bad' });

    expect(res.body).toHaveProperty('error');
    expect(typeof res.body.error).toBe('string');
  });
});
</code></pre>
<p><strong>What contract tests catch:</strong> Accidental field renames, type changes, missing fields. They fail when you'd break a consumer.</p>
<h2>Testing External Service Calls with MSW</h2>
<p>Mock Service Worker intercepts HTTP calls at the network level — no mocking fetch or axios.</p>
<pre><code class="language-typescript">// src/services/payment.test.ts
import { describe, it, expect, beforeAll, afterAll, afterEach } from 'vitest';
import { setupServer } from 'msw/node';
import { http, HttpResponse } from 'msw';
import { PaymentService } from './payment';

const server = setupServer();

beforeAll(() =&gt; server.listen());
afterEach(() =&gt; server.resetHandlers());
afterAll(() =&gt; server.close());

describe('PaymentService', () =&gt; {
  const service = new PaymentService('https://api.stripe.test');

  it('creates a payment intent', async () =&gt; {
    server.use(
      http.post('https://api.stripe.test/v1/payment_intents', () =&gt; {
        return HttpResponse.json({
          id: 'pi_test_123',
          status: 'requires_payment_method',
          amount: 1000,
        });
      })
    );

    const result = await service.createPaymentIntent(1000, 'usd');
    expect(result.id).toBe('pi_test_123');
    expect(result.amount).toBe(1000);
  });

  it('retries on 503 then succeeds', async () =&gt; {
    let attempts = 0;
    server.use(
      http.post('https://api.stripe.test/v1/payment_intents', () =&gt; {
        attempts++;
        if (attempts &lt; 3) {
          return new HttpResponse(null, { status: 503 });
        }
        return HttpResponse.json({ id: 'pi_retry', status: 'created', amount: 500 });
      })
    );

    const result = await service.createPaymentIntent(500, 'usd');
    expect(result.id).toBe('pi_retry');
    expect(attempts).toBe(3);
  });

  it('throws after max retries exhausted', async () =&gt; {
    server.use(
      http.post('https://api.stripe.test/v1/payment_intents', () =&gt; {
        return new HttpResponse(null, { status: 503 });
      })
    );

    await expect(service.createPaymentIntent(500, 'usd')).rejects.toThrow(
      'Payment service unavailable'
    );
  });

  it('handles malformed JSON response', async () =&gt; {
    server.use(
      http.post('https://api.stripe.test/v1/payment_intents', () =&gt; {
        return new HttpResponse('not json', {
          headers: { 'Content-Type': 'application/json' },
        });
      })
    );

    await expect(service.createPaymentIntent(500, 'usd')).rejects.toThrow();
  });
});
</code></pre>
<p><strong>What MSW tests catch:</strong> Retry logic bugs, timeout handling, error mapping, response parsing failures.</p>
<h2>Test Organization</h2>
<pre><code class="language-plaintext">src/
├── services/
│   ├── pricing.ts
│   └── pricing.test.ts          # Unit tests: co-located
├── repositories/
│   ├── user.ts
│   └── user.integration.test.ts # Integration tests: co-located
├── app.test.ts                  # API tests: top-level
└── contracts/
    └── user-api.contract.test.ts # Contract tests: separate dir
</code></pre>
<p>Run them separately:</p>
<pre><code class="language-json">{
  "scripts": {
    "test": "vitest run",
    "test:unit": "vitest run --testPathPattern='.test.ts$' --testPathIgnorePatterns='integration|contract'",
    "test:integration": "vitest run --testPathPattern='integration'",
    "test:contract": "vitest run --testPathPattern='contract'",
    "test:watch": "vitest watch"
  }
}
</code></pre>
<h2>What Not to Test</h2>
<ul>
<li><p><strong>Framework behavior:</strong> Don't test that Express returns 404 for unregistered routes. That's Express's job.</p>
</li>
<li><p><strong>Simple getters/setters:</strong> If a function just passes through data, the integration test covers it.</p>
</li>
<li><p><strong>Third-party library internals:</strong> Don't test that Zod validates emails correctly.</p>
</li>
<li><p><strong>Implementation details:</strong> Test the output, not how a function internally computes it.</p>
</li>
</ul>
<h2>Conclusion</h2>
<p>The strategy is straightforward:</p>
<ul>
<li><p><strong>Unit tests</strong> for business logic and calculations (fast, numerous)</p>
</li>
<li><p><strong>Integration tests</strong> for database operations (use real databases, not mocks)</p>
</li>
<li><p><strong>API tests</strong> for HTTP behavior and middleware (use supertest)</p>
</li>
<li><p><strong>Contract tests</strong> for API stability (use Zod schemas)</p>
</li>
<li><p><strong>MSW</strong> for external service interactions (network-level mocking)</p>
</li>
</ul>
<p>Each layer catches different bugs. Skip a layer and those bugs reach production.</p>
<hr />
<p><em>If this was helpful, you can support my work at</em> <a href="https://ko-fi.com/nopkt"><em>ko-fi.com/nopkt</em></a></p>
]]></content:encoded></item><item><title><![CDATA[Docker Multi-Stage Builds for Go: From 1GB to 12MB Production Images]]></title><description><![CDATA[Docker Multi-Stage Builds for Go: From 1GB to 12MB Production Images
Most Go Docker images are built wrong. A typical FROM golang:1.22 image weighs 800MB+. Your compiled Go binary is probably 10-15MB.]]></description><link>https://younggao.hashnode.dev/docker-multi-stage-builds-for-go-from-1gb-to-12mb-production-images</link><guid isPermaLink="true">https://younggao.hashnode.dev/docker-multi-stage-builds-for-go-from-1gb-to-12mb-production-images</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 06:28:55 GMT</pubDate><content:encoded><![CDATA[<h1>Docker Multi-Stage Builds for Go: From 1GB to 12MB Production Images</h1>
<p>Most Go Docker images are built wrong. A typical <code>FROM golang:1.22</code> image weighs 800MB+. Your compiled Go binary is probably 10-15MB. Here's how to use multi-stage builds to ship minimal, secure images — and avoid the common pitfalls.</p>
<h2>The Problem</h2>
<pre><code class="language-dockerfile"># DON'T do this
FROM golang:1.22
WORKDIR /app
COPY . .
RUN go build -o server .
CMD ["./server"]
</code></pre>
<p>This image includes the entire Go toolchain, build cache, and source code. Result: ~850MB image containing ~840MB of stuff you don't need in production.</p>
<h2>The Basic Multi-Stage Fix</h2>
<pre><code class="language-dockerfile"># Stage 1: Build
FROM golang:1.22-alpine AS builder
WORKDIR /app
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN CGO_ENABLED=0 GOOS=linux go build -o server .

# Stage 2: Run
FROM alpine:3.19
RUN apk --no-cache add ca-certificates
COPY --from=builder /app/server /server
CMD ["/server"]
</code></pre>
<p>This gets you to ~20MB. But we can do better, and there are production concerns to address.</p>
<h2>The Production Dockerfile</h2>
<pre><code class="language-dockerfile"># syntax=docker/dockerfile:1

# ============ Stage 1: Dependencies ============
FROM golang:1.22-alpine AS deps
WORKDIR /app
COPY go.mod go.sum ./
RUN --mount=type=cache,target=/go/pkg/mod \
    go mod download &amp;&amp; go mod verify

# ============ Stage 2: Build ============
FROM golang:1.22-alpine AS builder
WORKDIR /app

# Copy cached modules
COPY --from=deps /go/pkg/mod /go/pkg/mod

# Copy source
COPY . .

# Build with optimizations
ARG VERSION=dev
ARG COMMIT=unknown
RUN --mount=type=cache,target=/go/pkg/mod \
    --mount=type=cache,target=/root/.cache/go-build \
    CGO_ENABLED=0 GOOS=linux GOARCH=amd64 \
    go build \
      -ldflags="-s -w -X main.version=\({VERSION} -X main.commit=\){COMMIT}" \
      -trimpath \
      -o /server \
      ./cmd/server

# ============ Stage 3: Production ============
FROM gcr.io/distroless/static-debian12:nonroot

COPY --from=builder /server /server

EXPOSE 8080
USER nonroot:nonroot

ENTRYPOINT ["/server"]
</code></pre>
<p>Let's break down what each piece does and why.</p>
<h2>Key Decisions Explained</h2>
<h3><code>CGO_ENABLED=0</code></h3>
<p>Go can link against C libraries (CGo). Disabling it produces a fully static binary — no libc dependency, no shared library issues, runs on any Linux. If you need CGo (SQLite, certain crypto), you'll need a different base image.</p>
<h3><code>-ldflags="-s -w"</code></h3>
<p><code>-s</code> strips the symbol table, <code>-w</code> strips DWARF debug information. Together they reduce binary size by 20-30%. You lose <code>go tool pprof</code> symbol names, but production binaries shouldn't need them.</p>
<h3><code>-trimpath</code></h3>
<p>Removes local file paths from the binary. Without this, panic stack traces reveal your build directory structure — a minor security concern.</p>
<h3><code>--mount=type=cache</code></h3>
<p>BuildKit cache mounts persist the Go module cache and build cache across builds. First build downloads all modules; subsequent builds reuse the cache. On a project with 200 dependencies, this cuts build time from 90s to 15s.</p>
<h3>Distroless vs Alpine vs Scratch</h3>
<pre><code class="language-dockerfile"># Option 1: scratch (smallest, ~12MB total)
FROM scratch
COPY --from=builder /etc/ssl/certs/ca-certificates.crt /etc/ssl/certs/
COPY --from=builder /server /server

# Option 2: distroless (recommended, ~14MB total)
FROM gcr.io/distroless/static-debian12:nonroot

# Option 3: alpine (~22MB total)
FROM alpine:3.19
RUN apk --no-cache add ca-certificates
</code></pre>
<p><strong>Scratch</strong> has no shell, no tools, no users — just your binary. Debugging is hard (can't <code>exec</code> into the container).</p>
<p><strong>Distroless</strong> adds CA certificates, timezone data, a <code>nonroot</code> user, and <code>/tmp</code>. Still no shell or package manager. This is the sweet spot for most production services.</p>
<p><strong>Alpine</strong> adds a shell, package manager, and musl libc. Use it when you need to debug containers in production or when you have CGo dependencies.</p>
<h3>Non-Root User</h3>
<pre><code class="language-dockerfile">USER nonroot:nonroot
</code></pre>
<p>Never run containers as root. Distroless provides a <code>nonroot</code> user (UID 65532). If using scratch:</p>
<pre><code class="language-dockerfile">FROM scratch
COPY --from=builder /etc/passwd /etc/passwd
USER 65532:65532
</code></pre>
<h2>Handling Static Files and Config</h2>
<p>If your Go binary needs to serve static files or read config:</p>
<pre><code class="language-dockerfile">FROM gcr.io/distroless/static-debian12:nonroot

# Copy binary
COPY --from=builder /server /server

# Copy static files (embedded in binary is better, but sometimes you need files)
COPY --from=builder /app/static /static
COPY --from=builder /app/migrations /migrations

EXPOSE 8080
USER nonroot:nonroot
ENTRYPOINT ["/server"]
</code></pre>
<p>Better approach — embed files in the binary:</p>
<pre><code class="language-go">//go:embed static/*
var staticFiles embed.FS

//go:embed migrations/*.sql
var migrations embed.FS
</code></pre>
<p>Now your binary is self-contained. No files to copy, no paths to get wrong.</p>
<h2>CI/CD Integration</h2>
<p>Build with version info from Git:</p>
<pre><code class="language-bash">docker build \
  --build-arg VERSION=$(git describe --tags --always) \
  --build-arg COMMIT=$(git rev-parse --short HEAD) \
  -t myapp:$(git describe --tags --always) \
  .
</code></pre>
<p>In your main.go:</p>
<pre><code class="language-go">var (
    version = "dev"
    commit  = "unknown"
)

func main() {
    slog.Info("starting server", "version", version, "commit", commit)
    // ...
}
</code></pre>
<h2>Health Check for Orchestrators</h2>
<pre><code class="language-dockerfile"># If using alpine (has wget)
HEALTHCHECK --interval=30s --timeout=3s --start-period=5s \
  CMD wget --no-verbose --tries=1 --spider http://localhost:8080/health || exit 1
</code></pre>
<p>For distroless (no wget), implement the health check in your Go binary:</p>
<pre><code class="language-go">http.HandleFunc("/health", func(w http.ResponseWriter, r *http.Request) {
    w.WriteHeader(http.StatusOK)
    w.Write([]byte("ok"))
})
</code></pre>
<p>And let Kubernetes/ECS probe the endpoint directly.</p>
<h2>Multi-Architecture Builds</h2>
<p>Ship for both AMD64 and ARM64:</p>
<pre><code class="language-bash">docker buildx create --name multiarch --use
docker buildx build \
  --platform linux/amd64,linux/arm64 \
  --build-arg VERSION=$(git describe --tags) \
  -t myregistry/myapp:latest \
  --push \
  .
</code></pre>
<p>Update your Dockerfile to use the build arg:</p>
<pre><code class="language-dockerfile">ARG TARGETOS TARGETARCH
RUN CGO_ENABLED=0 GOOS=\({TARGETOS} GOARCH=\){TARGETARCH} \
    go build -ldflags="-s -w" -trimpath -o /server ./cmd/server
</code></pre>
<p>Docker sets <code>TARGETOS</code> and <code>TARGETARCH</code> automatically during multi-platform builds.</p>
<h2>Security Scanning</h2>
<pre><code class="language-bash"># Scan with trivy
trivy image myapp:latest

# Scan with grype
grype myapp:latest
</code></pre>
<p>Distroless and scratch images typically have zero CVEs because there's nothing to scan — no OS packages, no libraries. Alpine images usually have a few low-severity findings in musl or busybox.</p>
<h2>Size Comparison</h2>
<table>
<thead>
<tr>
<th>Base Image</th>
<th>Go Binary</th>
<th>Total Image</th>
<th>CVEs</th>
</tr>
</thead>
<tbody><tr>
<td>golang:1.22</td>
<td>15MB</td>
<td>~850MB</td>
<td>100+</td>
</tr>
<tr>
<td>alpine:3.19</td>
<td>15MB</td>
<td>~22MB</td>
<td>0-3</td>
</tr>
<tr>
<td>distroless/static</td>
<td>12MB</td>
<td>~14MB</td>
<td>0</td>
</tr>
<tr>
<td>scratch</td>
<td>12MB</td>
<td>~12MB</td>
<td>0</td>
</tr>
</tbody></table>
<p>The binary is smaller with <code>-ldflags="-s -w"</code> (15MB → 12MB).</p>
<h2>Common Mistakes</h2>
<p><strong>1. Copying go.sum but not go.mod</strong> — Docker layer caching works on <code>go.mod</code> and <code>go.sum</code> together. Copy both before <code>go mod download</code>.</p>
<p><strong>2. Not separating dependency download from build</strong> — If you copy all source before downloading modules, every code change invalidates the module cache layer.</p>
<p><strong>3. Forgetting CA certificates</strong> — TLS connections fail silently or with cryptic errors. Distroless includes them. Scratch and alpine need manual handling.</p>
<p><strong>4. Building for the wrong architecture</strong> — Always set <code>GOOS</code> and <code>GOARCH</code> explicitly. Your CI might be ARM but your target is AMD64.</p>
<p><strong>5. Including</strong> <code>.git</code> <strong>in the build context</strong> — Add <code>.git</code> to <code>.dockerignore</code>. A large git history can add hundreds of MB to the build context.</p>
<pre><code class="language-conf"># .dockerignore
.git
*.md
docs/
**/*_test.go
</code></pre>
<h2>Conclusion</h2>
<p>A production Go Docker image should be under 15MB, run as non-root, and contain nothing except your binary and its required certificates. Multi-stage builds make this straightforward — one stage for building, one for running. The build cache flags (<code>--mount=type=cache</code>) keep iteration fast.</p>
<p>Start with distroless for most services. Drop to scratch if you need the absolute minimum. Use alpine only when you need a shell for debugging.</p>
<hr />
<p><em>If this was helpful, you can support my work at</em> <a href="https://ko-fi.com/nopkt"><em>ko-fi.com/nopkt</em></a></p>
]]></content:encoded></item><item><title><![CDATA[Building a Type-Safe REST API with Zod, Express, and TypeScript: From Validation to OpenAPI Docs]]></title><description><![CDATA[Every REST API needs input validation and documentation. Most teams treat these as separate concerns — writing Joi schemas for validation, then manually maintaining OpenAPI specs. They inevitably drif]]></description><link>https://younggao.hashnode.dev/building-a-type-safe-rest-api-with-zod-express-and-typescript-from-validation-to-openapi-docs</link><guid isPermaLink="true">https://younggao.hashnode.dev/building-a-type-safe-rest-api-with-zod-express-and-typescript-from-validation-to-openapi-docs</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 06:28:35 GMT</pubDate><content:encoded><![CDATA[<p>Every REST API needs input validation and documentation. Most teams treat these as separate concerns — writing Joi schemas for validation, then manually maintaining OpenAPI specs. They inevitably drift apart. Here'''s how to use <strong>Zod</strong> as the single source of truth for both.</p>
<h2>What We'''re Building</h2>
<p>A complete REST API for a task management system with:</p>
<ul>
<li><p>Request/response validation using Zod schemas</p>
</li>
<li><p>Automatic OpenAPI 3.1 spec generation from those schemas</p>
</li>
<li><p>Type-safe request handlers (no <code>any</code> types)</p>
</li>
<li><p>Proper error responses with structured validation messages</p>
</li>
</ul>
<h2>Prerequisites</h2>
<ul>
<li><p>Node.js 20+ and npm</p>
</li>
<li><p>Basic TypeScript and Express knowledge</p>
</li>
<li><p>Familiarity with REST API concepts</p>
</li>
</ul>
<h2>Project Setup</h2>
<pre><code class="language-bash">mkdir task-api &amp;&amp; cd task-api
npm init -y
npm install express zod zod-to-openapi uuid
npm install -D typescript @types/express @types/uuid tsx
</code></pre>
<p>Create <code>tsconfig.json</code>:</p>
<pre><code class="language-json">{
  "compilerOptions": {
    "target": "ES2022",
    "module": "NodeNext",
    "moduleResolution": "NodeNext",
    "strict": true,
    "esModuleInterop": true,
    "outDir": "dist",
    "rootDir": "src"
  },
  "include": ["src/**/*"]
}
</code></pre>
<h2>Step 1: Define Schemas as the Source of Truth</h2>
<pre><code class="language-typescript">// src/schemas/task.ts
import { z } from '''zod''';
import { extendZodWithOpenApi } from '''zod-to-openapi''';

extendZodWithOpenApi(z);

// Base task schema — represents what'''s stored
export const TaskSchema = z.object({
  id: z.string().uuid().openapi({ example: '''550e8400-e29b-41d4-a716-446655440000''' }),
  title: z.string().min(1).max(200).openapi({ example: '''Ship v2.0 release''' }),
  description: z.string().max(2000).optional().openapi({ example: '''Final QA pass and deploy''' }),
  status: z.enum(['''todo''', '''in_progress''', '''done''']).openapi({ example: '''todo''' }),
  priority: z.enum(['''low''', '''medium''', '''high''', '''critical''']).openapi({ example: '''high''' }),
  assignee: z.string().email().optional().openapi({ example: '''dev@company.com''' }),
  dueDate: z.string().datetime().optional().openapi({ example: '''2025-06-15T00:00:00Z''' }),
  tags: z.array(z.string().max(50)).max(10).default([]).openapi({ example: ['''backend''', '''release'''] }),
  createdAt: z.string().datetime(),
  updatedAt: z.string().datetime(),
}).openapi('''Task''');

// Create request — subset of fields
export const CreateTaskSchema = TaskSchema.pick({
  title: true,
  description: true,
  priority: true,
  assignee: true,
  dueDate: true,
  tags: true,
}).extend({
  title: z.string().min(1).max(200),
}).openapi('''CreateTask''');

// Update request — all fields optional
export const UpdateTaskSchema = CreateTaskSchema.partial().openapi('''UpdateTask''');

// Query parameters for listing
export const TaskQuerySchema = z.object({
  status: z.enum(['''todo''', '''in_progress''', '''done''']).optional(),
  priority: z.enum(['''low''', '''medium''', '''high''', '''critical''']).optional(),
  assignee: z.string().email().optional(),
  page: z.coerce.number().int().min(1).default(1),
  limit: z.coerce.number().int().min(1).max(100).default(20),
  sortBy: z.enum(['''createdAt''', '''updatedAt''', '''dueDate''', '''priority''']).default('''createdAt'''),
  order: z.enum(['''asc''', '''desc''']).default('''desc'''),
}).openapi('''TaskQuery''');

// Derive TypeScript types directly from schemas
export type Task = z.infer&lt;typeof TaskSchema&gt;;
export type CreateTask = z.infer&lt;typeof CreateTaskSchema&gt;;
export type UpdateTask = z.infer&lt;typeof UpdateTaskSchema&gt;;
export type TaskQuery = z.infer&lt;typeof TaskQuerySchema&gt;;
</code></pre>
<p>Every type, every validation rule, every OpenAPI example lives in one place. Change the schema, and validation, types, and docs all update together.</p>
<h2>Step 2: Type-Safe Validation Middleware</h2>
<pre><code class="language-typescript">// src/middleware/validate.ts
import { Request, Response, NextFunction } from '''express''';
import { z, ZodError, ZodSchema } from '''zod''';

type ValidationTarget = '''body''' | '''query''' | '''params''';

interface ValidationSchema {
  body?: ZodSchema;
  query?: ZodSchema;
  params?: ZodSchema;
}

function formatZodError(error: ZodError) {
  return {
    error: '''Validation failed''',
    details: error.issues.map((issue) =&gt; ({
      path: issue.path.join('''.'''),
      message: issue.message,
      code: issue.code,
    })),
  };
}

export function validate(schemas: ValidationSchema) {
  return (req: Request, res: Response, next: NextFunction) =&gt; {
    const targets: ValidationTarget[] = ['''body''', '''query''', '''params'''];
    const errors: z.ZodIssue[] = [];

    for (const target of targets) {
      const schema = schemas[target];
      if (!schema) continue;

      const result = schema.safeParse(req[target]);
      if (!result.success) {
        for (const issue of result.error.issues) {
          issue.path = [target, ...issue.path];
          errors.push(issue);
        }
      } else {
        (req as any)[target] = result.data;
      }
    }

    if (errors.length &gt; 0) {
      const zodError = new ZodError(errors);
      return res.status(400).json(formatZodError(zodError));
    }

    next();
  };
}
</code></pre>
<p>Two details matter here. First, <code>safeParse</code> returns the <strong>transformed</strong> data (coerced numbers, defaults applied), and we replace <code>req.body</code>/<code>req.query</code> with it. Downstream handlers get clean, typed data. Second, we validate all targets before responding, so the client sees every error at once.</p>
<h2>Step 3: Request Handlers with Full Type Safety</h2>
<pre><code class="language-typescript">// src/routes/tasks.ts
import { Router, Request, Response } from '''express''';
import { v4 as uuidv4 } from '''uuid''';
import { validate } from '''../middleware/validate.js''';
import {
  CreateTaskSchema,
  UpdateTaskSchema,
  TaskQuerySchema,
  type Task,
  type CreateTask,
  type TaskQuery,
} from '''../schemas/task.js''';
import { z } from '''zod''';

const router = Router();
const tasks = new Map&lt;string, Task&gt;();

const IdParamSchema = z.object({ id: z.string().uuid() });

// LIST tasks with filtering, pagination, sorting
router.get('''/''',
  validate({ query: TaskQuerySchema }),
  (req: Request, res: Response) =&gt; {
    const query = req.query as unknown as TaskQuery;
    let results = Array.from(tasks.values());

    if (query.status) results = results.filter((t) =&gt; t.status === query.status);
    if (query.priority) results = results.filter((t) =&gt; t.priority === query.priority);
    if (query.assignee) results = results.filter((t) =&gt; t.assignee === query.assignee);

    results.sort((a, b) =&gt; {
      const aVal = a[query.sortBy] ?? '''''';
      const bVal = b[query.sortBy] ?? '''''';
      const cmp = String(aVal).localeCompare(String(bVal));
      return query.order === '''asc''' ? cmp : -cmp;
    });

    const total = results.length;
    const start = (query.page - 1) * query.limit;
    const items = results.slice(start, start + query.limit);

    res.json({
      items,
      pagination: { page: query.page, limit: query.limit, total, totalPages: Math.ceil(total / query.limit) },
    });
  }
);

// CREATE task
router.post('''/''',
  validate({ body: CreateTaskSchema }),
  (req: Request, res: Response) =&gt; {
    const input = req.body as CreateTask;
    const now = new Date().toISOString();
    const task: Task = {
      id: uuidv4(),
      status: '''todo''',
      ...input,
      tags: input.tags ?? [],
      createdAt: now,
      updatedAt: now,
    };
    tasks.set(task.id, task);
    res.status(201).json(task);
  }
);

// GET single task
router.get('''/:id''',
  validate({ params: IdParamSchema }),
  (req: Request, res: Response) =&gt; {
    const task = tasks.get(req.params.id);
    if (!task) return res.status(404).json({ error: '''Task not found''' });
    res.json(task);
  }
);

// UPDATE task
router.patch('''/:id''',
  validate({ params: IdParamSchema, body: UpdateTaskSchema }),
  (req: Request, res: Response) =&gt; {
    const existing = tasks.get(req.params.id);
    if (!existing) return res.status(404).json({ error: '''Task not found''' });
    const updated: Task = {
      ...existing,
      ...req.body,
      id: existing.id,
      createdAt: existing.createdAt,
      updatedAt: new Date().toISOString(),
    };
    tasks.set(updated.id, updated);
    res.json(updated);
  }
);

// DELETE task
router.delete('''/:id''',
  validate({ params: IdParamSchema }),
  (req: Request, res: Response) =&gt; {
    if (!tasks.delete(req.params.id)) return res.status(404).json({ error: '''Task not found''' });
    res.status(204).send();
  }
);

export default router;
</code></pre>
<p>Notice there are zero manual type assertions beyond the initial cast from <code>req.query</code>. The validation middleware guarantees the shape, and TypeScript carries it forward.</p>
<h2>Step 4: Auto-Generate OpenAPI Spec</h2>
<pre><code class="language-typescript">// src/openapi.ts
import { OpenAPIRegistry, OpenApiGeneratorV31 } from '''zod-to-openapi''';
import { TaskSchema, CreateTaskSchema, UpdateTaskSchema, TaskQuerySchema } from '''./schemas/task.js''';
import { z } from '''zod''';

const registry = new OpenAPIRegistry();

registry.register('''Task''', TaskSchema);
registry.register('''CreateTask''', CreateTaskSchema);
registry.register('''UpdateTask''', UpdateTaskSchema);

const IdParam = registry.registerParameter('''taskId''',
  z.string().uuid().openapi({ param: { name: '''id''', in: '''path''' } })
);

registry.registerPath({
  method: '''get''',
  path: '''/api/tasks''',
  summary: '''List tasks with filtering and pagination''',
  request: { query: TaskQuerySchema },
  responses: {
    200: {
      description: '''Paginated task list''',
      content: {
        '''application/json''': {
          schema: z.object({
            items: z.array(TaskSchema),
            pagination: z.object({
              page: z.number(), limit: z.number(),
              total: z.number(), totalPages: z.number(),
            }),
          }),
        },
      },
    },
  },
});

registry.registerPath({
  method: '''post''',
  path: '''/api/tasks''',
  summary: '''Create a new task''',
  request: { body: { content: { '''application/json''': { schema: CreateTaskSchema } } } },
  responses: {
    201: { description: '''Task created''', content: { '''application/json''': { schema: TaskSchema } } },
    400: { description: '''Validation error''' },
  },
});

registry.registerPath({
  method: '''get''',
  path: '''/api/tasks/{id}''',
  summary: '''Get a task by ID''',
  request: { params: z.object({ id: IdParam }) },
  responses: {
    200: { description: '''Task found''', content: { '''application/json''': { schema: TaskSchema } } },
    404: { description: '''Task not found''' },
  },
});

registry.registerPath({
  method: '''patch''',
  path: '''/api/tasks/{id}''',
  summary: '''Update a task''',
  request: {
    params: z.object({ id: IdParam }),
    body: { content: { '''application/json''': { schema: UpdateTaskSchema } } },
  },
  responses: {
    200: { description: '''Task updated''', content: { '''application/json''': { schema: TaskSchema } } },
    404: { description: '''Task not found''' },
  },
});

registry.registerPath({
  method: '''delete''',
  path: '''/api/tasks/{id}''',
  summary: '''Delete a task''',
  request: { params: z.object({ id: IdParam }) },
  responses: { 204: { description: '''Task deleted''' }, 404: { description: '''Task not found''' } },
});

const generator = new OpenApiGeneratorV31(registry.definitions);
export const openApiSpec = generator.generateDocument({
  openapi: '''3.1.0''',
  info: { title: '''Task Management API''', version: '''1.0.0''', description: '''Type-safe REST API with automatic validation and documentation''' },
  servers: [{ url: '''http://localhost:3000''' }],
});
</code></pre>
<h2>Step 5: Wire Everything Together</h2>
<pre><code class="language-typescript">// src/index.ts
import express from '''express''';
import taskRoutes from '''./routes/tasks.js''';
import { openApiSpec } from '''./openapi.js''';

const app = express();
app.use(express.json());
app.get('''/openapi.json''', (_, res) =&gt; res.json(openApiSpec));
app.use('''/api/tasks''', taskRoutes);

app.use((err: Error, _req: express.Request, res: express.Response, _next: express.NextFunction) =&gt; {
  console.error(err);
  res.status(500).json({ error: '''Internal server error''' });
});

const PORT = process.env.PORT || 3000;
app.listen(PORT, () =&gt; {
  console.log(`Server running on http://localhost:${PORT}`);
  console.log(`OpenAPI spec: http://localhost:${PORT}/openapi.json`);
});
</code></pre>
<h2>Testing It</h2>
<pre><code class="language-bash">npx tsx src/index.ts
</code></pre>
<p>Create a task:</p>
<pre><code class="language-bash">curl -X POST http://localhost:3000/api/tasks \
  -H "Content-Type: application/json" \
  -d '''{"title": "Write API docs", "priority": "high", "tags": ["docs"]}'''
</code></pre>
<p>Hit the validation:</p>
<pre><code class="language-bash">curl -X POST http://localhost:3000/api/tasks \
  -H "Content-Type: application/json" \
  -d '''{"title": "", "priority": "extreme"}'''
</code></pre>
<p>Response:</p>
<pre><code class="language-json">{
  "error": "Validation failed",
  "details": [
    { "path": "body.title", "message": "String must contain at least 1 character(s)", "code": "too_small" },
    { "path": "body.priority", "message": "Invalid enum value", "code": "invalid_enum_value" }
  ]
}
</code></pre>
<h2>Why This Works Better Than Alternatives</h2>
<p><strong>vs. Joi + Swagger JSDoc:</strong> Joi schemas can'''t generate TypeScript types. You maintain types, Joi schemas, and JSDoc annotations separately. Three sources of truth that drift apart.</p>
<p><strong>vs. class-validator + class-transformer:</strong> Decorator-based approaches require classes, which conflicts with functional TypeScript. Zod works with plain objects.</p>
<p><strong>vs. tRPC:</strong> tRPC is excellent for TypeScript-to-TypeScript communication. But if non-TypeScript clients need your API, you need OpenAPI — and <code>zod-to-openapi</code> gives you that.</p>
<h2>Going Further</h2>
<p>Add Swagger UI:</p>
<pre><code class="language-bash">npm install swagger-ui-express @types/swagger-ui-express
</code></pre>
<pre><code class="language-typescript">import swaggerUi from '''swagger-ui-express''';
app.use('''/docs''', swaggerUi.serve, swaggerUi.setup(openApiSpec));
</code></pre>
<p>Now <code>http://localhost:3000/docs</code> serves interactive API documentation generated entirely from your Zod schemas.</p>
<h2>Conclusion</h2>
<p>Zod as the single source of truth eliminates an entire class of bugs — the ones where validation accepts something the type system doesn'''t expect, or the API docs describe a field that the code ignores. Define once, validate everywhere, document automatically.</p>
<p>The full source is about 250 lines of actual logic. The OpenAPI spec it generates is production-ready — import it into Postman, generate client SDKs, or feed it to API gateways.</p>
<hr />
<p><em>If this was helpful, you can support my work at</em> <a href="https://ko-fi.com/nopkt"><em>ko-fi.com/nopkt</em></a></p>
]]></content:encoded></item><item><title><![CDATA[Building a Production API Gateway on Cloudflare Workers with Hono]]></title><description><![CDATA[Modern APIs need rate limiting, authentication, caching, and observability — but running a dedicated gateway server adds cost and complexity. Cloudflare Workers lets you build a full-featured API gate]]></description><link>https://younggao.hashnode.dev/building-a-production-api-gateway-on-cloudflare-workers-with-hono</link><guid isPermaLink="true">https://younggao.hashnode.dev/building-a-production-api-gateway-on-cloudflare-workers-with-hono</guid><dc:creator><![CDATA[Young Gao]]></dc:creator><pubDate>Sat, 21 Mar 2026 06:28:14 GMT</pubDate><content:encoded><![CDATA[<p>Modern APIs need rate limiting, authentication, caching, and observability — but running a dedicated gateway server adds cost and complexity. Cloudflare Workers lets you build a full-featured API gateway at the edge, with zero cold starts and global distribution.</p>
<p>In this guide, we'll build a production-ready API gateway using <strong>Hono</strong> (a lightweight web framework), <strong>Durable Objects</strong> (for distributed rate limiting), and Workers' built-in <strong>Cache API</strong>.</p>
<h2>Prerequisites</h2>
<ul>
<li><p>Node.js 18+ and npm</p>
</li>
<li><p>A Cloudflare account (free tier works)</p>
</li>
<li><p>Basic familiarity with TypeScript and REST APIs</p>
</li>
</ul>
<h2>Project Setup</h2>
<pre><code class="language-bash">npm create cloudflare@latest api-gateway -- --template hono
cd api-gateway
npm install hono jose
</code></pre>
<p>Your <code>wrangler.toml</code> needs Durable Object bindings:</p>
<pre><code class="language-toml">name = "api-gateway"
main = "src/index.ts"
compatibility_date = "2024-01-01"

[durable_objects]
bindings = [
  { name = "RATE_LIMITER", class_name = "RateLimiter" }
]

[[migrations]]
tag = "v1"
new_classes = ["RateLimiter"]

[vars]
UPSTREAM_URL = "https://api.example.com"
RATE_LIMIT_RPM = "60"
</code></pre>
<h2>Gateway Architecture</h2>
<pre><code class="language-plaintext">Client -&gt; [Auth] -&gt; [Rate Limit] -&gt; [Cache Check] -&gt; [Proxy] -&gt; [Log] -&gt; Response
</code></pre>
<p>Each step is a Hono middleware. If any step fails, the request short-circuits with an error.</p>
<h2>Step 1: The Hono Application Shell</h2>
<pre><code class="language-typescript">// src/index.ts
import { Hono } from 'hono';
import { cors } from 'hono/cors';
import { RateLimiter } from './rate-limiter';
import { authMiddleware } from './middleware/auth';
import { rateLimitMiddleware } from './middleware/rate-limit';
import { cacheMiddleware } from './middleware/cache';
import { loggingMiddleware } from './middleware/logging';
import { proxyHandler } from './handlers/proxy';

type Bindings = {
  RATE_LIMITER: DurableObjectNamespace;
  UPSTREAM_URL: string;
  RATE_LIMIT_RPM: string;
  JWT_SECRET: string;
};

const app = new Hono&lt;{ Bindings: Bindings }&gt;();

app.use('*', cors());
app.use('*', loggingMiddleware);

app.get('/health', (c) =&gt; c.json({ status: 'ok', edge: c.req.header('cf-ray') }));

app.use('/api/*', authMiddleware);
app.use('/api/*', rateLimitMiddleware);
app.get('/api/*', cacheMiddleware);
app.all('/api/*', proxyHandler);

export default app;
export { RateLimiter };
</code></pre>
<h2>Step 2: JWT Authentication</h2>
<pre><code class="language-typescript">// src/middleware/auth.ts
import { createMiddleware } from 'hono/factory';
import * as jose from 'jose';

export const authMiddleware = createMiddleware(async (c, next) =&gt; {
  const authHeader = c.req.header('Authorization');
  if (!authHeader?.startsWith('Bearer ')) {
    return c.json({ error: 'Missing or invalid Authorization header' }, 401);
  }

  const token = authHeader.slice(7);
  try {
    const secret = new TextEncoder().encode(c.env.JWT_SECRET);
    const { payload } = await jose.jwtVerify(token, secret, {
      algorithms: ['HS256'],
    });
    c.set('userId', payload.sub as string);
    c.set('scopes', (payload.scopes as string[]) || []);
    await next();
  } catch (err) {
    if (err instanceof jose.errors.JWTExpired) {
      return c.json({ error: 'Token expired' }, 401);
    }
    return c.json({ error: 'Invalid token' }, 401);
  }
});
</code></pre>
<p><strong>Why</strong> <code>jose</code> <strong>over</strong> <code>jsonwebtoken</code><strong>?</strong> <code>jose</code> uses the Web Crypto API natively -- perfect for edge runtimes without Node.js polyfills.</p>
<h2>Step 3: Distributed Rate Limiting with Durable Objects</h2>
<pre><code class="language-typescript">// src/rate-limiter.ts
export class RateLimiter {
  private state: DurableObjectState;
  private requests: number[] = [];

  constructor(state: DurableObjectState) {
    this.state = state;
  }

  async fetch(request: Request): Promise&lt;Response&gt; {
    const url = new URL(request.url);
    const limit = parseInt(url.searchParams.get('limit') || '60');
    const windowMs = parseInt(url.searchParams.get('window') || '60000');
    const now = Date.now();

    const stored = await this.state.storage.get&lt;number[]&gt;('requests');
    if (stored) this.requests = stored;

    this.requests = this.requests.filter((ts) =&gt; now - ts &lt; windowMs);

    if (this.requests.length &gt;= limit) {
      const oldestInWindow = Math.min(...this.requests);
      const retryAfter = Math.ceil((oldestInWindow + windowMs - now) / 1000);
      return new Response(
        JSON.stringify({ error: 'Rate limit exceeded', retryAfter, limit, remaining: 0 }),
        {
          status: 429,
          headers: {
            'Content-Type': 'application/json',
            'Retry-After': retryAfter.toString(),
            'X-RateLimit-Limit': limit.toString(),
            'X-RateLimit-Remaining': '0',
          },
        }
      );
    }

    this.requests.push(now);
    await this.state.storage.put('requests', this.requests);

    const remaining = limit - this.requests.length;
    return new Response(
      JSON.stringify({ allowed: true, remaining, limit }),
      {
        headers: {
          'X-RateLimit-Limit': limit.toString(),
          'X-RateLimit-Remaining': remaining.toString(),
        },
      }
    );
  }
}
</code></pre>
<p>The middleware:</p>
<pre><code class="language-typescript">// src/middleware/rate-limit.ts
import { createMiddleware } from 'hono/factory';

export const rateLimitMiddleware = createMiddleware(async (c, next) =&gt; {
  const userId = c.get('userId') || c.req.header('cf-connecting-ip') || 'anonymous';
  const id = c.env.RATE_LIMITER.idFromName(userId);
  const limiter = c.env.RATE_LIMITER.get(id);

  const limit = parseInt(c.env.RATE_LIMIT_RPM || '60');
  const resp = await limiter.fetch(
    new Request(`https://limiter/?limit=${limit}&amp;window=60000`)
  );

  const result = await resp.json&lt;{ allowed?: boolean; remaining: number; limit: number }&gt;();

  c.header('X-RateLimit-Limit', result.limit.toString());
  c.header('X-RateLimit-Remaining', result.remaining.toString());

  if (!result.allowed) {
    return c.json({ error: 'Rate limit exceeded' }, 429);
  }
  await next();
});
</code></pre>
<p>Each user gets their own Durable Object instance -- rate limits are per-user and globally consistent across all edge locations.</p>
<h2>Step 4: Response Caching</h2>
<pre><code class="language-typescript">// src/middleware/cache.ts
import { createMiddleware } from 'hono/factory';

export const cacheMiddleware = createMiddleware(async (c, next) =&gt; {
  if (c.req.method !== 'GET') { await next(); return; }

  const cache = caches.default;
  const cacheKey = new Request(c.req.url, { method: 'GET' });

  const cached = await cache.match(cacheKey);
  if (cached) {
    c.header('X-Cache', 'HIT');
    const body = await cached.text();
    const headers = Object.fromEntries(cached.headers.entries());
    return c.body(body, 200, headers);
  }

  c.header('X-Cache', 'MISS');
  await next();

  if (c.res.status === 200) {
    const response = c.res.clone();
    const cacheResponse = new Response(response.body, {
      headers: {
        ...Object.fromEntries(response.headers.entries()),
        'Cache-Control': 'public, max-age=60',
      },
    });
    c.executionCtx.waitUntil(cache.put(cacheKey, cacheResponse));
  }
});
</code></pre>
<h2>Step 5: Request Proxying</h2>
<pre><code class="language-typescript">// src/handlers/proxy.ts
import { createMiddleware } from 'hono/factory';

export const proxyHandler = createMiddleware(async (c) =&gt; {
  const upstreamUrl = new URL(c.req.path.replace('/api', ''), c.env.UPSTREAM_URL);

  const requestUrl = new URL(c.req.url);
  requestUrl.searchParams.forEach((value, key) =&gt; {
    upstreamUrl.searchParams.set(key, value);
  });

  const headers = new Headers(c.req.raw.headers);
  headers.delete('Authorization');
  headers.set('X-Forwarded-For', c.req.header('cf-connecting-ip') || '');
  headers.set('X-Request-ID', crypto.randomUUID());

  const upstreamReq = new Request(upstreamUrl.toString(), {
    method: c.req.method,
    headers,
    body: ['GET', 'HEAD'].includes(c.req.method) ? null : c.req.raw.body,
  });

  const startTime = Date.now();
  const response = await fetch(upstreamReq);
  const duration = Date.now() - startTime;

  const responseHeaders = new Headers(response.headers);
  responseHeaders.set('X-Gateway-Duration', `${duration}ms`);
  responseHeaders.set('X-Request-ID', headers.get('X-Request-ID')!);

  return new Response(response.body, {
    status: response.status,
    headers: responseHeaders,
  });
});
</code></pre>
<h2>Step 6: Structured Logging</h2>
<pre><code class="language-typescript">// src/middleware/logging.ts
import { createMiddleware } from 'hono/factory';

export const loggingMiddleware = createMiddleware(async (c, next) =&gt; {
  const requestId = crypto.randomUUID();
  const startTime = Date.now();
  c.header('X-Request-ID', requestId);

  await next();

  const logEntry = {
    timestamp: new Date().toISOString(),
    requestId,
    method: c.req.method,
    path: c.req.path,
    status: c.res.status,
    duration: Date.now() - startTime,
    ip: c.req.header('cf-connecting-ip'),
    userAgent: c.req.header('user-agent'),
    country: c.req.header('cf-ipcountry'),
    userId: c.get('userId') || null,
    cacheStatus: c.res.headers.get('X-Cache') || 'N/A',
  };

  console.log(JSON.stringify(logEntry));
});
</code></pre>
<h2>Performance</h2>
<table>
<thead>
<tr>
<th>Component</th>
<th>Overhead</th>
</tr>
</thead>
<tbody><tr>
<td>JWT verification</td>
<td>~1-2ms</td>
</tr>
<tr>
<td>Rate limit (Durable Object)</td>
<td>~5-15ms</td>
</tr>
<tr>
<td>Cache hit</td>
<td>~1ms</td>
</tr>
<tr>
<td>Cache miss + proxy</td>
<td>Upstream latency + ~2ms</td>
</tr>
<tr>
<td>Logging (async)</td>
<td>0ms</td>
</tr>
</tbody></table>
<p>Total overhead for cached responses: under 5ms. For uncached with rate limiting: 10-20ms.</p>
<h2>Production Hardening</h2>
<pre><code class="language-typescript">app.onError((err, c) =&gt; {
  console.error(JSON.stringify({
    error: err.message,
    stack: err.stack,
    path: c.req.path,
  }));
  return c.json({ error: 'Internal gateway error' }, 500);
});

// Request size limit
app.use('/api/*', async (c, next) =&gt; {
  const contentLength = parseInt(c.req.header('content-length') || '0');
  if (contentLength &gt; 10 * 1024 * 1024) {
    return c.json({ error: 'Request too large' }, 413);
  }
  await next();
});
</code></pre>
<h2>Conclusion</h2>
<p>Under 300 lines of TypeScript gives you authentication, distributed rate limiting, caching, and structured logging at the edge. Key advantages:</p>
<ul>
<li><p><strong>Zero cold starts</strong> and global distribution across 300+ cities</p>
</li>
<li><p><strong>Pay per request</strong> ($0.50/million on paid plan)</p>
</li>
<li><p><strong>Strongly consistent rate limiting</strong> via Durable Objects</p>
</li>
</ul>
<p>Next steps: API key management (KV), request transformation, A/B routing, WebSocket proxying.</p>
<hr />
<p><em>If this was helpful, you can support my work at</em> <a href="https://ko-fi.com/nopkt"><em>ko-fi.com/nopkt</em></a></p>
]]></content:encoded></item></channel></rss>