How we run an AI-native agency
Last updated: July 2026
"AI-native" has become the marketing claim every services firm makes in 2026. This post is the specific operating model behind the claim, at Twistag, as of July 2026. It is honest about what we automated, what we did not, what our team shape looks like now, how we estimate, how we price, and where we are still figuring things out. It is a template other agencies and internal delivery teams can adapt — not a copy-paste playbook, but a set of shapes that we have found stick. It is the pillar of our cluster on AI-native delivery process (G).
Key takeaways
- AI-native is an operating model, not a tool selection. The tools are the visible layer. The invisible layer — how work gets scoped, staffed, estimated, priced, and reviewed — is where 70% code generation actually comes from.
- We automated the internal ops that produce the biggest time savings: proposal generation, meeting synthesis, spec drafting, code review, delivery reporting. We did not automate the parts that require judgement or relationship — sales, delivery leadership, senior engineering, client conversation.
- Our team shape changed. Senior count stayed. Mid-level count compressed. Junior hiring changed shape entirely. Nobody gets hired into a typing role anymore. Every role has an editorial or judgement component that AI does not do.
- Estimation changed twice. First it got faster and looser, then it got tighter again. The right shape is: we estimate less, we estimate more accurately, we revisit weekly. Fixed-scope engagements got easier; hourly engagements got harder to justify.
- Pricing is moving from time-and-materials toward outcome-based. Not all the way, not overnight, but the direction is clear. When velocity is 5-10x, T&M no longer represents what the client is buying.
What we automated (and what we did not)
The exercise most agencies run wrong is trying to automate the client-facing surface first. That is where relationships live and where AI has the least payback. The internal ops surface — the stuff nobody outside the agency ever sees — is where the automation actually pays off.
Automated internal ops (net time saved is substantial):
- Proposal generation. The Twistag proposal skill takes a discovery-call transcript, a scope note, and produces a first-draft proposal in the correct brand shape. A senior consultant edits it in 30 minutes rather than writing it in 3 hours.
- Meeting synthesis. Every internal and external meeting gets transcribed and summarised automatically. Action items surface without a manual capture step. The synthesis is close enough to right that a two-minute review catches what matters.
- Spec drafting. Product managers use AI to draft feature specs from acceptance criteria. Same pattern as the coding tools — the AI produces the first draft, the human edits.
- Code review layer. The four-layer review from our 70%-AI-generated codebase pillar — author self-check, reviewer agents, senior review, integration tests. The reviewer-agent layer is the biggest single time-saver.
- Delivery reporting. Client status reports draft themselves from the sprint's telemetry and merged PRs. The delivery lead edits, adds narrative, sends.
Not automated (deliberately):
- Sales conversations. The prospect wants a person. AI helps prep, but the conversation is human.
- Delivery leadership. The delivery lead's judgement about pace, quality, and client mood is the point of the role. AI cannot do it.
- Senior engineering. Architecture, hard trade-offs, incident response. The AI helps, but the senior engineer is accountable.
- Client relationship management. The relationship is a person-to-person contract. AI cannot own it.
- Hiring interviews. The interview is a judgement about whether a person will thrive on the team. AI could screen; we do not use it that way.
The pattern: automate everything that is repetitive knowledge work with an editable output. Do not automate anything that trades on trust or judgement.
The team shape change
Our staffing shape at the end of 2025 vs mid-2026 shows the biggest structural change of the year:
Senior engineers: count stayed roughly flat. The role changed (see the 70%-codebase pillar) but the headcount is the same. Each senior is more productive; we absorbed the productivity into taking on more or larger engagements rather than shedding seniors.
Mid-level engineers: count compressed by about a third. The typing work that used to fill a mid-level's day is now AI-generated. Mid-levels who transitioned to editorial / review roles stayed; mid-levels who were doing well but leaned on typing over judgement did not.
Junior engineers: hiring stopped for pure typing roles. We still hire junior engineers — but into apprenticeship-style pairings with a senior, where the junior is learning judgement, not producing typing volume. It is a smaller pipeline and a slower ramp; the juniors we do bring in ramp to editorial productivity in about six months.
Design engineers: new role, small count. The role we describe in the design engineering pipeline post. Two or three across the agency. Very high value.
Product managers, designers: shape held. These roles were already judgement-heavy. AI helped them ship faster; the count did not change.
Delivery leads: count grew slightly. Because we run more engagements per senior engineer, we need slightly more delivery-lead capacity. The role is more relational than before — less spec-writing, more client-facing.
Internal ops: count dropped by half. Proposal ops, meeting ops, delivery ops — the roles that were coordinating the internal surface are largely automated. The people who did them either moved into delivery leadership or moved on.
The visible outcome: a leaner org that ships more per person. The invisible outcome: a much higher requirement on judgement across every remaining role.
Estimation, the second time round
Our estimation practice went through a shape change in 2026 that is worth naming.
First shift (early 2026). Estimation got faster and looser. AI could produce a plausible estimate from a spec in minutes. We ran with the AI estimate. Some engagements went well, some went 20% over. The variance was higher than the classical estimation.
Second shift (mid-2026). We tightened back. The AI estimate is now the first-cut, not the final. A senior engineer reviews it against the specific engagement risks — integrations, data quality, team familiarity. The estimate ships accurate to within 10% on most engagements.
The pattern that produces the tight estimate:
- AI produces the first cut from the spec, using our historical delivery data as reference.
- Senior engineer overrides the AI's assumptions on the specific risks that are not in the historical data.
- We revisit weekly. If the estimate is drifting, the client knows within a week, not at the end of the engagement.
Fixed-scope engagements became easier to sell because we can honestly commit to them. The client knows what they are getting; we know what we are delivering. This is the shape outcome-based pricing sits on top of.
Pricing, still moving
Our pricing model was mostly time-and-materials through 2024, mostly fixed-scope through 2025, and moving toward outcome-based in 2026. The direction is set; the exact shape is still being worked out.
Time-and-materials. Broke because our velocity was hard to describe. When one senior engineer with an AI stack ships in a day what previously took a week, the client is paying for the same outcome at a fraction of the hours. Clients notice, and either negotiate a discount or feel underpaid-for. The model does not survive this asymmetry.
Fixed-scope. Our main model in 2026. We commit to a scope, we commit to a price, we deliver. The AI-native operating model is what makes this work — we can estimate accurately and we can absorb small overruns because our unit cost is lower. The client is paying for the outcome, not the hours.
Outcome-based. Emerging on a subset of engagements. The client pays a base fee for the engagement and a bonus tied to a specific business metric. The pattern only works where the business metric is clean and attributable, which is a smaller set of engagements than the marketing suggests.
The shape we ended up at: fixed-scope for most work, outcome-based where the metric is clean, T&M only for open-ended discovery work where the scope genuinely cannot be nailed down.
Where we are still figuring things out
An honest section, because the AI-native operating model is not fully solved and we do not want to pretend it is.
- Junior engineer development. We stopped hiring for typing roles, but we have not fully worked out the apprenticeship model that gets a junior to editorial competence in six months. The pattern works when it works; it does not scale as cleanly as the classical junior-to-mid pipeline did.
- Estimation on genuinely novel work. Our historical data covers what we have done before. When the engagement is architecturally novel, the AI estimate is worse than the classical estimate. Senior judgement overrides it, but that means the estimation cost is still real.
- Client understanding of the model. Some clients still buy time. When we quote a fixed scope that reflects our AI-native velocity, the client sometimes reads it as a low quote and worries about corners cut. The conversation about why the price is what it is happens on every engagement.
- Internal ops that we have not automated yet. Contract review, some finance operations, some parts of the delivery pipeline still have manual work that has an AI-native shape we have not built yet. Prioritising them against client work is a recurring choice.
What this means if you are running an agency
Three things transfer from our operating model to any agency (or internal delivery team) trying to become AI-native.
Start with internal ops, not client-facing surfaces. The payback lives inside. The client-facing polish comes later — and comes for free — once the internal ops are running.
Shift the team shape deliberately, not by attrition. Deciding what your team should look like at the end of the shift is a Phase 1 conversation. Waiting for the shape to emerge through hiring and departures takes years and produces regret.
Fix the estimation and pricing together. They move together. A T&M business model with AI-native delivery pace is a losing shape. A fixed-scope model with pre-AI estimation is a different losing shape. The internal discipline changes both at once.
The operating model shift is the real work. The tools are the easy part. The agencies that will still exist in 2028 have started the operating-model work in 2026. The ones that are still marketing "AI-native" without having done the internal shift will find that clients notice the gap by 2027.
let's talk


