NewResearch: the same pull request data supports opposite conclusions
Creativity is the new productivity

Every dashboard measures how much more you shipped.That's just more AI slop!

When writing is free, volume stops being an achievement. Behold measures the judgment behind the work: what it costs you to keep the slop, and where a person is still the most valuable thing in the room. All of it in $$$.

Last 12 monthsBy teamBy model
AI spend12 mo
$48,214+12%
JFMAMJJASOND
Cost per merged PRtrailing
$54−34%
Spend by teamcost / PR
Platform
$54
Payments
$71
Growth
$44
Mobile
$118

Reads from the systems you already run

  • GitHub
  • GitLab
  • Claude Code
  • GitHub Copilot
  • Cursor
  • Anthropic
  • OpenAI
  • AWS Bedrock
  • Google Vertex
  • LiteLLM
  • Jira
  • Linear
  • Slack

Same team. Same quarter. Two dashboards.

Both of these are accurate. One of them is what your tooling reports and what gets taken to the board. The other is what the quarter actually cost you.

What every dashboard sees

A breakout quarter.

  • Pull requests merged

    vs. same quarter last year

    2.1×
  • Throughput per engineer

    sustained over 11 weeks

    +41%
  • AI adoption

    of engineers, weekly active

    78%
  • Cost per merged PR

    spend down, output up

    −34%

Nothing here is wrong. Every number is real, and every tool you already run will show you some version of it.

What actually happened

A bill that arrives later.

  • Merged work rewritten within 90 days

    up from 9%

    31%
  • Review time per reviewer

    the cost moved, it did not vanish

    2.4×
  • Changes no author could explain

    sampled at review

    1 in 6
  • Cognitive complexity

    compounding, not transient

    +27%

None of it appears in a throughput chart, because none of it is production. It is the cost of keeping what was produced.

measuring output rewards whoever produces the most of it

When writing was expensive, volume was a reasonable proxy for effort. It is not one any more. The scarce input is no longer production, it is judgment: knowing what to build, what to keep, what to throw away, and where a person still has to be the one deciding.

The industry spent forty years learning not to count lines. Then it started counting again.

For four decades the industry agreed that counting lines was a poor way to measure software. It rewards volume, and the engineers you most want to keep are the ones who remove volume. A rewrite that deletes two thousand lines and a feature that adds two thousand score as opposites, when the first is often worth more. The argument was settled and the metric was retired.

Then AI arrived and the same number came back wearing a new name. Share of code written by AI. Tokens consumed. Suggestions accepted. Pull requests merged. Every one of them counts the act of writing, and that too at the exact moment writing stopped being the expensive part. A metric that was merely gameable when a human had to type it is unbounded when a machine does.

So there is a great deal to account for now, and no single instrument reaches all of it. We use telemetry to establish what happened, and we ask people directly where the answer only exists in their heads. Then the part that actually decides the outcome: every organization adopts this differently enough that a reading taken from one does not transfer to another.

TelemetrySurveysExpertsis far better than any one of them alone

We answer the right questions.

Scientifically curated, and built to measure what actually determines the outcome rather than what happens to be easy to count.

Slop Index

What plausible-looking work costs you after it merges.

Computed from what gets rewritten, what drags review, what compounds as complexity, and what nobody can account for three weeks later. Priced, so it can be argued about in a budget meeting rather than a retro.

61/ 100

Up from 24 before assisted authoring

Priced

$412k

per quarter

  • Rework38%
  • Review burden27%
  • Complexity added21%
  • Unexplained change14%
Judgment Rate

Where a person changed the direction rather than the syntax.

Read from the work itself: the rejections, the redirections, the designs that were thrown away before they cost anything. It is the only one of our numbers that goes up when people think harder, and it cannot be gamed by producing more.

22%

Of human review changed direction, not syntax

Avoided

$780k

per quarter

  • Payments34%
  • Identity28%
  • Ledger19%
  • Growth8%

Growth ships fastest and thinks least. That is the finding, and no throughput chart contains it.

Three products, and the work they read.

Engineering intelligence

Know whether the work was any good.

Throughput tells you a team was busy. It cannot tell you whether anyone exercised judgment, and when producing is free that is the only question left worth asking.

Merged this quarter112
  • Slop16
  • Real judgment11
  • Unremarkable85
Token intelligence

Every AI dollar, and what it actually bought.

Your provider can tell you what you spent. Nobody but you can tell you what it was worth, because the return shows up somewhere the invoice never looks.

Billed in July

$48,214

day 90 · still standing

$31,683

rewritten or abandoned
$31,683 still standing
billed30d60d90d
  • Claude Code$21,480$15,337
  • Cursor$12,905$8,069
  • Copilot$7,640$5,078
  • Agents (API)$4,312$2,101
  • Everything else$1,877$1,100
Optimization

Not every request needs the expensive model.

The hard part is not saving money. It is proving that quality held while you did, which takes measuring the work rather than the latency.

~/acme/webbar 80

Rename the billing props

4 files changed, 61 lines

~/acme/ledgerbar 85

Why is the payments test flaky?

bisected 9 runs, found it

~/acme/platformbar 88

Design the caching layer

write-through, 90s TTL

~/acme/webbar 82

Add a /changelog page from MDX

2 routes, RSS 2.0

~/acme/apibar 86

Write the migration for orders

reversible, 1 new index

~/acme/infrabar 78

Bump the Terraform providers

no plan diff

~/acme/authbar 87

Refactor the session module

11 call sites updated

~/acme/webbar 84

Add tests for the cart reducer

18 cases, 3 edge

~/acme/searchbar 86

Explain why recall dropped

analyzer change, week 31

~/acme/webbar 80

Rename the billing props

4 files changed, 61 lines

~/acme/ledgerbar 85

Why is the payments test flaky?

bisected 9 runs, found it

~/acme/platformbar 88

Design the caching layer

write-through, 90s TTL

~/acme/webbar 82

Add a /changelog page from MDX

2 routes, RSS 2.0

~/acme/apibar 86

Write the migration for orders

reversible, 1 new index

~/acme/infrabar 78

Bump the Terraform providers

no plan diff

~/acme/authbar 87

Refactor the session module

11 call sites updated

~/acme/webbar 84

Add tests for the cart reducer

18 cases, 3 edge

~/acme/searchbar 86

Explain why recall dropped

analyzer change, week 31

Haiku

4.5

scores42$0.80routed

GLM-5.2

Z.ai

scores55$4.40routed

Gemini

3.1 Pro

scores71$12routed

Sonnet

4.6

scores83$15routed

GPT-5.5

Codex

scores87$30routed

Opus

4.8

scores94$75routed

One instrument, pointed at the whole lifecycle.

Engineering analytics

Delivery measured end to end, from the moment work is picked up to the moment a customer can use it, with the time each stage actually consumed.

Last 6 monthsBy teamBy stage
Idea → productionmedian
9.4 days
JFMAMJ
Merged PRs / engineerper month
18.6+22%
JFMAMJ
Where the time goesdays
Review
3.8
Coding
2.1
QA
1.9
Release
1.6

What engineering leaders use it for.

Each of these is a decision already being made, usually without the evidence to settle it.

Align AI Spend With What Matters

See exactly how AI and engineering effort maps to the initiatives you care about, so leaders can confirm investment is flowing to the highest-priority work, not just the loudest.

Show the Real ROI of AI

Tie AI costs and engineering activity directly to shipped outcomes. Know what your AI adoption is actually delivering, and whether it's genuinely accelerating how fast you ship.

Build a Smarter Metrics Practice

Design flexible metrics spanning delivery, investment, and day-to-day workflows, with AI that reads the data for you and highlights what deserves attention next.

Elevate How Teams Use AI

See how engineers actually prompt, iterate, and weave AI into the development lifecycle. Spot the patterns and habits that reliably lead to better results, and spread them across teams.

Remove Friction From Developer Workflows

Capture structured feedback from your engineers to pinpoint what's slowing them down, then act on it to lift productivity and satisfaction.

Simplify Compliance and Financial Reporting

Replace manual time tracking with automated, audit-ready reports for software capitalization and R&D tax credits.

AI adoption in SDLC workflows is near-universal, but most organizations have no way to study how it is impacting their software. That gap is not a tooling problem. Every organization has different constraints, and no single metric can tell you how yours is performing. We are researchers and practitioners with more than a decade in developer productivity and software engineering. Closing that gap is our work.

AI experts from the world's leading organizations

and more…

Our selected work.

Bring evidence to the next budget conversation.

We're working with a small number of engineering organizations to get this right. Connect a repository and see your own numbers.

Behold Labs

Tell us a little more.