Most engineering dashboards measure what is easy to count, not what is worth knowing. This is a working guide to the engineering metrics that actually signal a healthy team, how to spot vanity metrics before they set the wrong incentives, and why the human cost of the work belongs in the picture now that AI writes so much of the code.
Engineering metrics are measurements that describe how a software team builds and ships software, and how the work behaves as a system. Good ones answer questions a leader actually has: are changes reaching users quickly and safely, is quality holding, and can the team keep this pace without burning down.
The useful ones share three traits. They track a system or a team rather than a person. They point at an outcome, not just activity. And they are read as signals to investigate, not as verdicts to rank by. A metric that fails those tests can still be interesting, but it should not steer decisions on its own. This guide sits under our wider work on developer productivity metrics, and narrows in on the measurement question itself.
Almost every argument about engineering metrics comes down to confusing three different things. Naming them keeps a program honest.
Measures whether the work achieved a result a user or the business cares about: a feature adopted, an incident avoided, a cycle shortened. Harder to attribute, and the closest thing to what matters.
Counts the work itself: commits, pull requests, story points, deploys. Useful as context and for spotting flow problems, but easy to mistake for progress when it is only motion.
Rises without anything getting better, and changes no decision. Lines of code is the classic example: more of it is often worse, not better. If a number cannot end an argument, it is decoration.
The practical rule is to lead with outcomes, keep a few output measures as supporting context, and retire vanity metrics rather than displaying them out of habit. A number earns its place on a dashboard only if you can name the decision it informs.
Two frameworks show up in nearly every serious conversation about engineering metrics. Neither is a scorecard to copy blindly, but both give a shared vocabulary that beats inventing metrics from scratch.
Developed by the DevOps Research and Assessment program, DORA measures how well a team delivers software through four metrics:
The first two describe speed; the last two describe stability. Read together they resist the trap of shipping fast by breaking things, or being stable by shipping nothing.
The SPACE framework was created to counter single-metric thinking. It argues productivity is multidimensional and names five dimensions to sample across, rather than a fixed set of numbers:
Its central advice is to pick metrics from more than one dimension, so no single number gets gamed at the expense of the others.
Frameworks summarised from their public definitions. No benchmark thresholds are quoted here, because healthy ranges depend on your context, not a universal number.
DORA tells you how the delivery system is performing. SPACE reminds you that delivery is only part of the story, and that satisfaction and flow are legitimate things to measure. That leaves one dimension most tooling still misses, which is where a picture of cognitive load comes in.
The single most common way engineering metrics go wrong is turning them on individual people. Ranking engineers by commits, pull requests or lines of code feels objective and is deeply misleading.
Metrics do their best work at the team and system level, where they surface questions to explore together. The moment a number is used to compare two engineers on a spreadsheet, it has left the domain where it can be trusted. This is closely tied to the wider developer experience, which those individual scorecards tend to quietly degrade.
AI coding tools change what an engineer's day is made of. Less of the time goes on writing code by hand, and more goes on steering the tool, reviewing what it produced, holding context across several open loops, and deciding what to keep. The visible output can look the same, or better, while the work underneath gets heavier.
Delivery metrics and output counts can stay flat or improve while the mental burden on the people doing the work climbs out of view. A metrics program that only watches the code will call that a win.
This is the gap Ascenda Flow was built for. It reads the activity that AI coding tools such as Claude Code, Cursor and Codex already record, and shows the engineer driving them what the work looked like: the retries, the loop steering, the context switching, not just the raw output. It never claims to read how anyone feels.
Adding that dimension keeps the rest of your engineering metrics honest. It is what stops a team from mistaking a more draining process for a healthier one, and it fits naturally beside DORA delivery signals and a SPACE-style view of satisfaction and flow.
The frameworks matter less than the habits around them. A metrics program that helps rather than distorts tends to follow a few principles.
Pick metrics because you have a decision to make, not because a tool can produce them. If no decision changes with the number, do not track it.
Combine delivery signals, quality signals and a human signal so no single number can be gamed at the expense of the others. This is the core SPACE lesson.
Measure systems and teams, never individuals. Use the numbers to open a conversation, not to close one with a ranking.
Review the set regularly and retire anything that has stopped changing a decision. A dashboard is not an archive; every metric on it should still earn its place. Treated this way, engineering metrics become an instrument for improving the work rather than a report card that reshapes it.
Ascenda Flow shows engineers how they work with AI coding agents, the human side behind delivery and output numbers. Get early access and see it on your own history.
Get early access