Estimating AI-augmented engineering
Last updated: July 2026
Story points, t-shirt sizing, hours estimates — the classical engineering estimation vocabulary calibrated to a specific idea of how fast a senior engineer could ship. When the AI-augmented senior engineer ships five to ten times faster on the parts where AI does the typing, the calibration breaks. A three-point story is now half a day. A "large" t-shirt is a morning. The number stayed the same; the underlying unit shifted. Teams that did not notice keep planning as if a sprint holds twenty story points, when it actually holds sixty. This post is what breaks, what replaces it, what stays hard even with AI, and how to sanity-check a partner's AI-augmented estimate. It is a cluster child of our AI-native agency operating model pillar.
Key takeaways
- Story points broke because they measured typing effort, which AI now does. The unit had to shift from typing hours to editorial hours. Every team's point calibration is now wrong until they recalibrate.
- The parts that stayed hard did not get faster. Design, architecture, integration, incident response, business-rules encoding — all the parts that always were the hard part of engineering. AI helps at the margins; they still take real time.
- The right unit is now a feature or a slice, not a point. Estimation in features + delivery pace = confidence interval. Points are two calibration steps removed from anything a stakeholder cares about.
- Weekly re-estimation replaces sprint-boundary estimation. The velocity is high enough that a two-week sprint estimate is stale by end of week one. The estimate has to live in the same working cycle as the code.
- Sanity-checking a partner's estimate has three specific tests. Their breakdown of what is AI-typing vs judgement time, their weekly re-estimation cadence, and their variance track record.
What broke about story points
Story points were a shorthand. They compressed complexity, effort, and risk into one number that a team could estimate in aggregate. The compression worked when the underlying components were stable — a team that consistently shipped 20 points a sprint had a real signal.
Two things changed in 2026:
The typing effort collapsed. A story that would previously have taken a senior engineer three days of typing now takes a day of editorial time. The point was calibrated to the three days; the shipped work took a day. Points inflated silently.
The editorial effort became the actual constraint. The senior engineer's day is now bounded by how many AI-generated diffs they can meaningfully review, not by how many lines they can type. A point that reflects typing effort does not reflect the actual constraint.
The result: story points that used to have a fuzzy but useful relationship to elapsed time now have no consistent relationship. A three-point story might be half a day (all typing, AI does it) or three days (design decisions plus integration). Teams that keep points as their unit are not lying; they are using a broken instrument.
What replaces points
The unit we settled on: features shipped per week as the top-line measure, with hours of editorial time per feature as the diagnostic.
Features shipped per week. A team ships some number of user-visible or system-user-visible features each week. This is what the stakeholder cares about, and it is what the team can commit to. In an AI-native team the number is typically 3-8 features per week for a five-person delivery unit. In a classical team the same unit shipped 1-2 per week.
Hours of editorial time per feature. Under the features, the diagnostic that catches drift. If a feature normally takes 4-6 hours of senior editorial time and this one is taking 12, something is off — usually design ambiguity or an unclear acceptance criterion. The diagnostic is a signal, not a plan.
The two measures compose into an honest estimate for a stakeholder. "This scope is 15 features. Our team ships 5 per week. Three weeks." When the actual work involves 4 hours of editorial per feature on the routine ones and 20 hours on two tricky ones, the plan holds. When the tricky ones balloon to 40 hours, the plan updates weekly and the stakeholder knows in week one, not week three.
Story points are missing from this vocabulary deliberately. If a team wants to use points internally as a shorthand for editorial hours, that is fine; the stakeholder-facing number is features and elapsed time.
What stayed hard
The parts of engineering that always were the hard parts did not get faster in proportion to the typing collapse.
Design. Deciding what to build. What the entities are, what the flows are, what the user's mental model needs to be. AI helps produce artefacts (specs, diagrams, prototypes) but the decisions are human. Design time did not compress in a 10x way; it compressed maybe 1.5-2x.
Architecture. Service boundaries, data model, integration points. Same shape as design — the AI helps produce, the human decides. Architecture time compressed maybe 2x.
Integration. Getting the new work to talk to the existing systems. Legacy APIs, upstream data quality, undocumented conventions. AI helps with the typing but the discovery is human. Integration time compressed maybe 2-3x.
Incident response. When something breaks in production, judgement about severity and remediation is human work. AI helps with hypothesis generation and diagnostics; it does not shorten the debug loop as much as the marketing suggests.
Business rules encoding. Turning the enterprise's actual rules into code. The rules live in someone's head or in a fragmented spec. Getting them out and coded correctly is human interviewing plus editorial work.
Add up the hard parts of a real engagement and they still take most of the elapsed time. AI-augmented engineering is not "everything is 10x faster." It is "the typing is 10x faster, and the hard parts got a little faster." A partner that estimates as if everything is 10x is going to overrun on the parts that stayed hard.
The weekly re-estimation cadence
Classical estimation was a sprint-boundary event. The team estimated at planning, ran the sprint, retrospected at the end. Two-week feedback loops.
At AI-native velocity, a two-week loop is stale. In two weeks the team ships 10-16 features rather than 2-3, so the estimate at planning is far from the reality at the retrospective. The cadence has to move to weekly:
Monday. Look at what remains in scope. Look at what shipped last week. Update the estimate.
Wednesday. Mid-week check on the current features. Anything that is running long? Reforecast.
Friday. What actually shipped. What did not. Why. Feed the answer into Monday's estimate.
This is a lightweight cadence — the meetings are 20 minutes each — but the discipline is real. A team that runs weekly re-estimation catches drift in week one and can renegotiate scope in week two. A team that runs sprint-boundary estimation catches drift at the end and has to scramble to save the sprint.
Sanity-checking a partner's estimate
The enterprise buying delivery services from a partner in 2026 will get AI-adjusted estimates from every serious partner. Three tests separate the honest partners from the marketing.
Test 1: The AI-typing vs judgement breakdown. Ask the partner to break the estimate into "what AI writes" and "what humans decide." An honest partner has the breakdown ready. The AI-writing portion is typically 30-50% of the elapsed time on greenfield work, less on brownfield. A partner whose estimate says "AI writes 80% of it" is either lying or does not understand their own delivery.
Test 2: The weekly re-estimation cadence. Ask how they will keep the estimate current. An honest partner has a weekly cadence and a specific meeting. A partner who says "we will let you know if it slips" has a sprint-boundary cadence and will notice slippage too late.
Test 3: The variance track record. Ask what percentage of their last ten engagements delivered within 10% of the estimate. An honest partner has the number and it is 70-90%. A partner who says "all of them" is lying. A partner who says "we do not track" is not disciplined enough to trust with a fixed scope.
The three tests are cheap to run in a discovery call and expensive to fake in real delivery. Any partner that fails one of the three has an estimation problem that will show up in the engagement.
What this shifts for the buyer
The enterprise's own estimation practice needs an update too. Two moves:
Recalibrate the internal team's velocity numbers. If the internal engineering team is using AI coding tools, their story-point calibration is now off. A quarterly reset is the minimum. Weekly is better.
Change what the roadmap conversation measures. The classical roadmap conversation was "how many story points fit in a quarter." The new conversation is "how many features can we ship in a quarter, at what confidence." The unit shift is small; the conversation quality it produces is large.
The estimation shift is one of the biggest visible signs of an AI-native operating model. A team still counting points at sprint boundaries is not AI-native, no matter what the marketing says. A team that measures features per week, re-estimates weekly, and tracks variance is doing the work.
let's talk


