Error Handling — Middle¶
At middle level, focus on this question:
Where does Error Handling belong in a maintainable component, and which trade-off selects the design?
Use the smallest realistic scenario that exposes the decision and its failure behavior.¶
Core Concepts¶
1. Design an error taxonomy, not just individual error types¶
A service benefits from a small, deliberate set of error categories — e.g. NotFound, InvalidInput, Unauthorized, Unavailable (retryable), Internal — each mapped consistently to an HTTP status, a retry policy, and a log level. Individual error values/types then carry which category they belong to, so a single dispatch point (an HTTP middleware, an RPC interceptor) can handle all of them uniformly instead of every handler reinventing the mapping.
type Kind int
const (
KindNotFound Kind = iota
KindInvalidInput
KindUnauthorized
KindUnavailable
KindInternal
)
type AppError struct {
Kind Kind
Message string
Err error
}
func (e *AppError) Error() string { return e.Message }
func (e *AppError) Unwrap() error { return e.Err }
2. errors.Join combines independent errors (Go 1.20+)¶
errors.Join is for when multiple independent operations can each fail and you need to report all of them (e.g. cleaning up several resources in a defer) — errors.Is/errors.As still work across a joined error, checking each one in turn.
3. Retry only what's actually retryable¶
var ErrUnavailable = errors.New("temporarily unavailable")
func withRetry(fn func() error, attempts int) error {
var err error
for i := 0; i < attempts; i++ {
if err = fn(); err == nil || !errors.Is(err, ErrUnavailable) {
return err
}
time.Sleep(backoff(i))
}
return err
}
Retrying a KindInvalidInput or KindNotFound error is pointless — the input won't become valid by trying again. Only errors explicitly marked retryable (typically transient network/availability failures) should trigger a retry loop; blindly retrying everything hides real bugs and can amplify load during an outage.
4. panic/recover is for programmer errors, not expected failures¶
func mustParse(s string) int {
n, err := strconv.Atoi(s)
if err != nil {
panic(fmt.Sprintf("mustParse: invalid input %q", s))
}
return n
}
panic is reserved for situations that indicate a bug (an invariant violated, a "this should never happen" branch) — not for ordinary, expected failure modes like a missing file or a failed network call, which should always be a returned error. recover belongs at well-defined boundaries (an HTTP middleware, a goroutine's top-level defer), converting an unexpected panic into a 500 response or a logged crash, rather than scattered throughout business logic.
5. Boring error flows beat clever ones¶
An error-handling strategy that's easy to explain in one sentence per function ("if X fails, wrap and return; if Y fails, it's retryable, so bubble up as KindUnavailable") is more valuable long-term than one that's technically elegant but requires tracing through several layers of custom logic to understand what happens on failure. Optimize for "a new engineer can predict what this does when it fails by reading it once."
Code Examples¶
Example 1 — A single error-to-response mapping point¶
func writeError(w http.ResponseWriter, err error) {
var ae *AppError
if !errors.As(err, &ae) {
ae = &AppError{Kind: KindInternal, Message: "internal error", Err: err}
}
status := map[Kind]int{
KindNotFound: 404,
KindInvalidInput: 400,
KindUnauthorized: 401,
KindUnavailable: 503,
KindInternal: 500,
}[ae.Kind]
if ae.Kind == KindInternal {
log.Printf("internal error: %v", ae.Err) // full detail server-side only
}
http.Error(w, ae.Message, status)
}
Every handler just returns an *AppError (or a plain error, mapped to Internal); the HTTP status/logging decision lives in exactly one place.
Example 2 — errors.Join for cleanup¶
func closeAll(closers ...io.Closer) error {
var errs []error
for _, c := range closers {
if err := c.Close(); err != nil {
errs = append(errs, err)
}
}
return errors.Join(errs...)
}
Example 3 — Recover at a goroutine boundary¶
func safeGo(fn func()) {
go func() {
defer func() {
if r := recover(); r != nil {
log.Printf("recovered panic: %v\n%s", r, debug.Stack())
}
}()
fn()
}()
}
Best Practices¶
- Define a small, fixed set of error kinds and map them consistently to status codes/log levels in one place.
- Never retry an error you haven't explicitly classified as transient.
- Reserve
panicfor programmer errors;recoveronly at clear boundaries, never as routine control flow. - Log full error detail once, at the boundary where it's handled — not at every layer it passes through.
Edge Cases & Pitfalls¶
- Recovering a panic and continuing as if nothing happened can leave a goroutine or service in an inconsistent state — recovery should usually still fail the current request/operation, just without crashing the whole process.
- Blanket retry logic applied to every error type can turn a brief downstream blip into a self-inflicted thundering herd.
errors.Join'sUnwrap() []error(plural) meanserrors.Is/errors.Ascheck each joined error, but a naiveerrors.Unwrap(err)(singular) call on a joined error returns nothing useful — useerrors.Is/errors.As, not manual unwrapping, on joined errors.
Common Mistakes¶
| Mistake | Fix |
|---|---|
| Logging the same error at every layer it passes through | Log once, fully, at the boundary |
| Retrying everything "just in case" | Explicitly classify retryable vs. non-retryable |
Using panic for expected failure conditions | Return an error; reserve panic for invariant violations |
Tricky Points¶
- An
AppErrorembedding anotherAppErrorviaErrstill needs its ownUnwrap()method forerrors.Is/errors.Asto see through it — implementingError()alone isn't enough for wrapping to work. recover()only works when called directly inside a deferred function — calling it indirectly (through another function call) returnsnileven during an active panic.
Apply it¶
- Find a real component where Error Handling affects an interface or dependency.
- Write two plausible choices and the constraint that favors each one.
- Make the smallest reversible change at that boundary.
- Exercise the component alone, then exercise the integrated flow.
- Keep the decision note with the evidence that selected the option.
Verify your work¶
- A focused check proves the local behavior.
- An integrated check proves callers and dependencies still agree.
- Logs, traces, compiler output, or benchmarks expose the boundary.
- Reverting the change restores the previous behavior without unrelated edits.
Review questions¶
- Which boundary is most affected by Error Handling?
- What constraint would make you choose the alternative design?
- How would you isolate a local defect from an integration defect?
- What evidence shows that the change remains maintainable?
In this topic
- junior
- middle
- senior
- professional