MaintenanceOS: The Self-Maintaining Tech Stack

AI is making software dramatically easier to create. It isn't making software easier to own.
Every service an agent builds, dependency it adds, database it provisions, API it integrates with and piece of infrastructure it creates becomes something that needs to be maintained for years.
And software doesn't stay still.
Versions reach end of life. APIs get deprecated. Dependencies introduce breaking changes. Cloud providers change defaults. Configuration drifts. Infrastructure becomes unused. Internal standards change. Teams reorganize and ownership gets lost.
None of this is new.
The scale is.
As the cost of creating software collapses, companies will create far more of it. Things that weren't worth building before suddenly will be. Internal tools, services, automations, infrastructure and integrations will multiply.
The maintenance surface grows with them.
The current operating model doesn't.
Software maintenance is still a human system
Modern engineering has automated almost every part of software creation and operation. Infrastructure is programmable. Builds are automated. Deployments are automated. Production is observable. Agents can now write code, run tests, open PRs and increasingly execute complex engineering tasks on their own.
Maintenance still works very differently.
A tool finds something. A dashboard turns red. A vendor announces an EOL. Someone creates a ticket. Someone figures out who owns it. Someone investigates whether it matters. Someone prioritizes it against everything else the team needs to do. Eventually someone makes the change and hopefully verifies that it worked.
For one issue, this is manageable.
At enterprise scale, it happens across thousands of services, cloud resources, databases, runtimes, dependencies and teams. It becomes a permanent tax on engineering: enormous backlogs, upgrade campaigns, stale infrastructure, lost ownership and teams spending cycles rediscovering context that already exists somewhere in the organization.
Adding an agent to the end of this process doesn't fix it.
It just makes the last step faster.
Writing the fix is becoming the easy part
The agentic engineering conversation is heavily focused on execution.
Can an agent upgrade this dependency? Change this Terraform? Migrate this database? Open the PR? Run the tests?
Increasingly, yes.
But before an agent changes anything, it needs to know what should change.
What actually exists? How is it being used? Is it production? What depends on it? Who owns it? Is the issue urgent? What's the right target state? What does company policy allow? What could break? Has this change been safely made before?
And after the change, something needs to determine whether it actually worked.
Autonomous maintenance is a context and governance problem before it is a coding problem.
The execution layer is getting commoditized quickly. The harder layer is continuously understanding the tech stack well enough to make good decisions about it.
From maintenance work to an operating system
The end state is not an agent with production credentials changing whatever it wants, and it's not another maintenance copilot waiting for a ticket.
It's a continuously running maintenance loop with enough context and governance to decide when humans are needed and when they're not.
At Draftt, we think about that loop in six parts.
1. Identify
Continuously understand what actually exists across infrastructure and software.
Not a periodic inventory. The live state of the stack: technologies, versions, resources, configurations, dependencies, ownership and usage.
Maintenance starts with knowing what you actually run.
2. Understand
A PostgreSQL version reaching EOL is a fact. Whether it matters now depends on context.
Where is it running? Is it production? What depends on it? Who owns it? How is it configured? What will break if it changes?
The same issue can be irrelevant in one environment and a major engineering project in another.
A self-maintaining system needs context before action.
3. Prioritize
Finding everything creates a backlog. Maintenance requires deciding what deserves attention.
Urgency, risk, effort, business impact, dependencies and organizational policies all matter. The system has to determine what should happen now, what can wait and what can be ignored.
This is where maintenance becomes governance rather than another scanner.
4. Act
Once the problem and context are understood, the system can determine the right action.
Sometimes that's routing work to the owner with the investigation already done. Sometimes it's generating the change. Sometimes it's opening a PR.
And for changes that are well understood, reversible and within policy, eventually it means fixing them automatically.
The level of autonomy should follow the risk of the action, not the ambition of the AI strategy.
5. Verify
A PR isn't a resolution.
The system needs to know whether the change actually happened, whether the issue disappeared and whether the environment is still healthy afterward.
Maintenance only becomes a closed loop when the result is verified against the real environment.
6. Govern
Every maintenance decision creates context for the next one.
Which changes were approved? Which were rejected? What broke? Which teams have different policies? Which maintenance windows work? Which recommendations repeatedly get ignored?
Over time, the system should get better at maintaining your stack, not an abstract version of one.
How maintenance gets executed
Most autonomy models are a trust slider: the human approves everything, then some things, then nothing.
We think that's the wrong axis.
The real question in an enterprise is who executes the change and what they need to execute it safely. Different teams, environments and risk levels will answer that differently, inside the same company, at the same time.
Draftt works across four execution models.
Follow the plan:
Draftt identifies the maintenance work, prioritizes it, finds the owner and hands them a plan with everything already resolved: what's affected, why it matters here, what the change is and what to watch afterward. The engineer executes. This is where most maintenance work lives today, and it's a very different job when the investigation is already done.
Automatic PR:
For changes Draftt understands well, it generates the change and opens the PR. A human reviews and merges. Draftt then verifies the result against the environment, not against the merge.
Contextual engineering:
Many enterprises already have coding agents doing the work. They don't need another agent. They need theirs to know what Draftt knows: the live state of the stack, dependencies, ownership, policies and blast radius.
Draftt provides that context directly to the customer's agents over MCP. They bring the execution layer. Draftt brings the maintenance context.
Agent-to-agent (A2A) handoff:
The furthest execution model delegates maintenance work directly to the customer's agent fleet. Draftt scopes the task, supplies the context, sets the policy boundaries and verifies the outcome over A2A.
The customer's agents do the engineering. Draftt makes sure it's the right engineering.
-
These models can coexist. A production database upgrade might require a human. A dependency update might be handled through an automatic PR. Another team might already have an internal coding agent they want to execute the work.
The goal isn't maximum autonomy everywhere.
It's the right execution path for each change, with enough context and governance to make it safe.
Today, humans run the maintenance loop and software assists them.
The self-maintaining tech stack reverses that model: software runs the loop continuously, and humans step in where judgment, approval or exception handling is actually required.
MaintenanceOS
This is what we're building at Draftt.
MaintenanceOS is the governance and maintenance layer for the engineering stack. It continuously understands what exists, what needs attention, what matters and what can safely change. Then it closes the loop through the right execution path, whether that's a human, Draftt or the enterprise's own agents.
The goal isn't a better maintenance backlog.
It's to progressively remove the manual maintenance loop itself.
Software will keep aging. Dependencies will keep changing. Infrastructure will keep drifting. Standards will keep moving. AI will dramatically increase the amount of software companies have to own.
Engineering organizations can't solve that by generating tickets faster.
As software creation becomes agentic, maintenance becomes agentic too.
The tech stack will maintain itself.
That's MaintenanceOS.


.png)
