Skip to content

Non-Functional Requirements — Mistake

When the Quality-Risk Loop earns its cost

  • Use the full loop when a change touches a critical user journey, sensitive data, permissions, an external dependency, background work, a material traffic/data change, or on-call ownership. These are the places where a short ticket can hide a costly failure mode.
  • Use a lighter scan for a small, reversible change with no changed risk. Write the reason it is low risk; “the feature is small” is not evidence that its data or access impact is small.
  • Do not treat the loop as legal, privacy, or security approval. It exposes questions early; the accountable specialist or policy owner still makes the decision that needs their authority.

Common mistakes

  • Writing “fast,” “secure,” or “highly available” as if they were requirements. Nobody can tell if the promise was met, and each person silently supplies a different meaning. Fix: state the flow, condition, measure, threshold, time window, and owner—for example, “90% of eligible exports complete within 10 minutes over 30 days.”
  • Starting with a favourite solution. “Add a queue” or “turn on encryption” can hide the user impact, make alternatives invisible, and create work the feature does not need. Fix: name the risk and desired outcome first; choose a control only after the requirement is clear.
  • Asking only the product manager. Product may know the goal but not the data classification, support history, dependency limits, or on-call burden. Fix: bring in the people who own those facts early, with concrete questions about the changed flow.
  • Using the daily average as the capacity estimate. A healthy daily number can hide the month-end minute when everyone requests an export; it can also hide a slow or failed high-value flow. Fix: calculate requests in the busiest credible window, keep average, peak, and total volume separate, and say where the traffic is measured. Google SRE's case study makes the same distinction when selecting traffic measures and sizing for a retail peak (source).
  • Counting one user action as one request. “600 exports” hides job creation, status polling, source reads, file writes, downloads, retries, and cache misses. Fix: make a small operation table: operation, boundary, peak RPS/QPS, payload, and whether it reads or writes. The CSV example's polling endpoint can reach 120 QPS even though export creation is only 10 RPS.
  • Calculating file size but forgetting retention, copies, and egress. A 50 MiB export sounds small until it is made 1,200 times a day, retained for a week, replicated, backed up, and downloaded during the same busy window. Fix: show items × size × retention × copies, list backups separately, and estimate outgoing bandwidth separately from stored capacity.
  • Presenting a rough estimate as an exact capacity plan. Rounded input, implementation overhead, skewed data, retries, and waiting lines can all change the result. Fix: label every assumption, use the calculation to find the questions and a test target, then validate the actual implementation under realistic load. Google SRE calls back-of-the-envelope math a sanity check and says load testing is still necessary (source).
  • Treating a provider price page as a budget decision. Price depends on the service, region, commitment, operation mix, storage class, and date; replication may multiply more than one line item. Fix: record the calculator inputs and date, show a range if an input is uncertain, and get the product or finance owner to accept the material cost. A documented Azure example separates storage and throughput and multiplies capacity for replicated regions (source).
  • Copying a cloud provider's SLA as the product requirement. A provider agreement covers a defined slice of a service, not your code, configuration, dependencies, or user journey. Fix: treat vendor commitments as an input and define an end-to-end target that the team can measure; Microsoft cautions against using an SLA without understanding its coverage (source).
  • Treating security as a last-minute scan. Access rules, data flows, and abuse cases become much harder to change after the feature shape is fixed. Fix: include protection needs, data-flow review, and at least one abuse case while the ticket is being clarified; OWASP recommends these activities throughout the development lifecycle (source).
  • Calling privacy “security with a different name.” A feature can have strong access control and still collect too much data, retain it too long, or use it beyond the agreed purpose. Fix: trace the data life cycle, policy, legal requirements, and acceptable risk separately; NIST maps those privacy outcomes into planning, design, deployment, and operation (source).
  • Creating an alert without a useful action. A noisy alert trains people to ignore it; an alert with no owner or runbook turns an incident into improvisation. Fix: name the symptom, threshold, recipient, urgency, and first safe action. Keep the human-facing signal simple and tied to a real failure.
  • Adding every control because the feature feels risky. An enormous checklist delays delivery and makes the important risks harder to see. Fix: rank risks by affected flow, impact, likelihood, and policy; use a standard such as ASVS as a guide, not as an excuse to apply every control unchanged.
  • Leaving the requirement in the ticket after release. Usage, policies, dependencies, and support evidence change. Fix: review the promises after rollout, an incident, a material dependency change, or a new data use; update the requirement and its evidence together.

Continue to Best Practise.