What we actually look at in a Well-Architected Review
The framework asks good questions. Most reviews answer them badly
The AWS Well-Architected Framework is a genuinely useful structure. Six pillars, a long list of questions, and a vocabulary that lets two engineers who have never met discuss the same system.
The trouble is what happens when a review becomes a form-filling exercise. You end up with a document that scores every question, flags two hundred items, sorts them by a severity label, and leaves the reader no wiser about what to do on Monday. Nobody acts on it, which means the review changed nothing.
A review is only worth the time if it ends with a short, ordered list that a team can start on. Here is what we chase inside each pillar to get there.
Operational excellence: who finds out, and what do they read?
We do not start with whether monitoring exists. It almost always exists. We start at the moment of failure and work backwards.
When this workload breaks at 2am, what fires? Who does it reach, on what device, and is that person on a roster or did they just happen to see it? When they open their laptop, is there a runbook, and does the runbook describe this system as it is now or as it was two years ago?
The gap we find most often is not missing alerts. It is alerts that reach a shared mailbox, or that fire so frequently that the team has learned to ignore them. An alert nobody acts on is worse than no alert, because it creates the belief that the system is watched.
Security: how far does one leaked credential reach?
We ask the question in that direction deliberately. Asking “is it secure” produces a defence of the design. Asking how far a single compromised credential travels produces a map.
So we trace it. What can each identity read and write? Which identities can assume others? Is any permission chain able to turn a modest-looking write into code execution with higher privileges? Are the account and network boundaries real, or is production one flat space with optimistic security groups?
Then we check whether any of that movement would be noticed. Logging that lands somewhere nothing reads is a common and expensive finding, because it costs money every month and buys nothing.
Reliability: what single thing takes the whole workload with it?
Every system has one. The job is to name it out loud and decide, deliberately, whether to accept it.
We look for the component with no replica, the single availability zone that quietly became load-bearing, the manual step in the recovery path, and the backup that has never been restored. That last one deserves its own sentence: a backup you have not restored is a hypothesis, not a backup.
We also ask what the recovery target actually is, and then whether the architecture can meet it. Those two numbers disagree more often than not, and the disagreement is usually news to somebody senior.
Performance efficiency: what is sized for a load you no longer have?
Workloads accumulate capacity the way roles accumulate permissions. Somebody sized an instance for a launch that happened three years ago, and nothing has questioned it since because it works.
We compare provisioned capacity against observed use, look for the database doing table scans because an index was never added, and check whether anything is paying for provisioned throughput it has never approached. This pillar and the cost pillar overlap heavily, and findings here usually pay for the engineering time several times over.
Cost optimisation: what are you paying for that nothing uses?
We look for the unattached storage volumes, the old snapshots nobody will ever restore, the idle load balancers, the data crossing availability zones for no reason, the logs retained forever by default, and the non-production environments running at full size overnight and over weekends.
None of this is clever. All of it is money, and it is usually the finding that buys the team political permission to do the harder security work.
Sustainability: where does it burn capacity it never uses?
In practice this pillar overlaps almost entirely with the two above it. Right-sizing, shutting down what is idle, and choosing managed services that scale to nothing when nothing is happening all show up here too.
We report it honestly as overlap rather than inventing separate findings to fill the section.
The part that makes a review useful
Everything above produces findings. Findings are not the deliverable. The deliverable is the order.
We rank by consequence, not by a severity label. The question for each finding is what actually happens if this is never fixed, and how likely that is given how the system is really used. A critical-labelled finding on a component with no sensitive data and no availability requirement ranks below a medium-labelled finding sitting on the only path to your customer records.
Then we say what we would do first, second and third, and why. A team can argue with that. A team cannot argue with a list of two hundred items sorted alphabetically, so it does not engage with it at all.
What a review cannot do
It is a point-in-time engineering opinion. It is not an accredited assessment, it is not a certification, and it does not survive six months of deployments unexamined.
It is also not a substitute for fixing anything. The review tells you where to spend the next month. Somebody still has to spend it.
We run these for free on one workload, because it is a better way to meet an engineering team than a capability deck. The details are here, including what access we need and what we hand back.
Ready to Unlock the Full Power of AWS?
Let’s talk about your cloud goals — no pressure, no hard sells. We’ll audit your setup, suggest improvements, and help you scale smarter.





