Skip to main content
The Hard Parts.dev

The Hard Parts

Engineering reference

  • Failure modes.
  • Red flags.
  • Trade-offs.
  • Playbooks.

For senior engineers,
architects, and
engineering leaders.

Issue 02

Software fails
the same way.
Every time.

A field guide for staff engineers, tech leads, architects, and engineering managers dealing with recurring software delivery problems.

Sections

04

Entries

181

Use in

  • Incident reviews.
  • Architecture reviews.
  • AI adoption discussions.
  • Retros & decisions.
Essay 13 May 2026

AI Accelerates Old Failure Modes

AI did not invent the ordinary gaps in software delivery like incomplete specifications, rushed reviews unclear ownership, or architecture that is harder to explain than we would like. It made those gaps easier to carry forward into working code. This is why engineering judgment matters more than ever.

Pick your way in

Start from the question you have

Field guide

Failure Modes: named patterns that recur in software teams.

From Failure Modes

A glimpse of the catalog. Each entry walks through how the pattern starts, how it escalates, what it looks like at early, mid, and late stages, and what good responses look like.

Browse all 37 Failure Modes →

Decision ledger

Tech Decisions: structured trade-off references for consequential engineering calls.

From Tech Decisions

One decision per axis the catalog covers: architecture, delivery, team, quality, and AI systems. Each entry lays out two concrete options with their real conditions, costs, hidden costs, and failure modes when misapplied.

Browse all 46 Tech Decisions →

Signal bulletin

Red Flags: early-warning signals across code, teams, process, leadership, and AI-enabled work.

From Red Flags

One signal per layer the catalog covers: code, team, process, leadership, and AI. Each entry opens with what you would actually notice and walks through what it usually indicates and what to check next.

Browse all 50 Red Flags →

Playbook

Engineering Playbook: practical routines for recurring engineering situations.

From Engineering Playbook

One playbook per situation the catalog covers: delivery, architecture, operations, teamwork, and AI adoption. Each entry lists the owner, the real conditions, and the sequence of moves.

EP-20 Delivery

Run a phased migration

Move from old to new in controlled slices, where each slice has explicit ownership, cutover criteria, rollback, and retirement of the old path.

Owner · tech lead
EP-44 Team

Repair trust after a painful incident

Repair trust by making the event intelligible, changing the conditions that produced it, and demonstrating through behavior that the team is safer, more honest, and more accountable than before.

Owner · engineering manager
EP-19 Architecture

Refactor a dangerous hotspot

Refactor a hotspot by targeting the specific reasons it is dangerous: high churn, poor testability, unclear ownership, or oversized responsibility - and improving it in narrow, repeatable steps.

Owner · maintainer
EP-31 Operations

Run an incident review that actually helps

Turn an incident review into a system-learning exercise that explains what happened, why it made sense at the time, what conditions enabled it, and what changes will reduce recurrence.

Owner · incident lead
EP-01 Ai

Upgrade code review for AI-assisted work

Redesign review so that AI-assisted changes are judged by risk, understanding, and behavioral correctness, not by surface polish or author confidence.

Owner · tech lead
EP-05 Ai

Evaluate an AI feature against real tasks

Evaluate the feature against real user jobs, realistic failure patterns, and operational constraints so the team learns whether the system actually helps, not just whether it performs well on curated examples.

Owner · evaluation owner

Browse all 48 Playbooks →