Skip to main content

Command Palette

Search for a command to run...

From Portals to IDEOps: The Age of Agentic Platform Engineering

Updated
•15 min read•View as Markdown

A colleague was working with Claude Code on a project running on IBM Cloud. He kicked off a rebuild and sat back to wait. The agent, noticing the build had stalled, decided on its own to investigate. It accessed the deployment pipeline, re-ran the compilation, and redeployed the service. Nobody asked it to. When the developer looked back at his terminal, the work was already done.

That moment, an AI agent autonomously detecting a stuck process and fixing it without being told, is exactly the shift this article is about. And it should make you both excited and deeply uncomfortable.

The metric that died

I used to measure my productivity the way everyone does: lines written per day, features closed per sprint, PRs merged per week. Editors optimized for that. Frameworks optimized for that. The entire ecosystem revolved around one obsession: produce code faster.

That metric is dead. It didn't fade; it just stopped mattering. The constraint today isn't how fast we write code, it's how well we supervise the agents writing it for us. And that changes everything: how we work individually, how we build teams, and what our platforms need to look like.

The competitive pressure is already real. According to Weave Intelligence's analysis, high-performing engineering organizations are five times more likely to invest structurally in AI platforms rather than simply buying better code assistants. As Luca Galante puts it: "If your competition orchestrates ten agents in parallel while you're still reviewing every line manually, they're not 10% faster. They're on a different curve."

If you're an architect, SRE, or platform engineer still deploying portals with colorful dashboards so developers can click buttons, this article is for you. You probably won't like everything it says.

Levels of autonomy: from manual operator to orchestrator

To understand where we're going, we need a map. The human intervention framework inspired by autonomous driving levels gives us exactly that: a maturity scale describing how the relationship between developer and AI agent evolves.

The important thing here: you can't jump from Level 1 to Level 4 by magic. Each step demands architectural maturity, not just better tooling.

  • Level 0: Human-driven. The developer does everything manually gathers context, writes code, runs CI, validates, deploys. This is the default state for most teams today. There's nothing wrong with admitting that.

  • Level 1: In the loop. Agents act as local assistants inside the IDE: autocomplete, suggest refactors, generate tests. Individual speed goes up, but gains are strictly linear. The bottleneck is still a human approving every action. Without structural change, some teams have actually reported slowdowns from fighting the noise the AI generates.

  • Level 2: On the loop. The human shifts from executor to dispatcher. This is the "Dispatch Work to Agents" pattern: multiple agents generating pull requests in parallel, resolving issues, proposing changes. But the bottleneck moves too: now the problem is validation and review quality. The gate model of Levels 0–1 (human approves before anything moves) gives way to an iterative validation loop. Agents propose, tests run, the human reviews outcomes rather than individual actions. From here on, throughput stops being linear and starts compounding.

  • Level 3: Orchestrating. The system dispatches work based on observability signals: alerts, metrics, incidents. The human no longer initiates work; they govern it. They define policies, set guardrails, and supervise outcomes. The bottleneck here shifts to architectural maturity and token economics: can your platform sustain the cost and complexity of agents operating at scale?

  • Level 4: Outside the loop (outlook). Full autonomy for defined classes of work. Systems self-heal and optimize in the background based on deterministic platform policies, with no direct human intervention. This level remains aspirational for most organizations. The bottleneck is platform quality and the strength of your guardrails. Most teams today are still navigating between Levels 1 and 2.

The question you should be asking yourself: what kind of platform do you need to build to enable Levels 3 & 4?

Spoiler: it's not a web portal with forms.

The critical leap: from Level 2 to Level 3

Of all these transitions, the jump from Level 2 to 3 generates the most friction. And it's not technical friction, it's organizational.

At Level 2, the human is still the starting point: they decide what work gets dispatched and when. At Level 3, that starting point shifts to the system. An observability alert triggers orchestrated actions without anyone pressing a button. For most organizations, that means a deep cultural shift: accepting that a system initiates and orchestrates work on its own. 

This is not science fiction or a three-year roadmap anymore,remember the Claude Code story from the opening? That is exactly this transition. 

The agent stopped being an assistant waiting for instructions and became an operator that detected a stalled situation and acted. To me what makes it fascinating also makes it unsettling, there was no observability system triggering the action, instead the model itself interpreted the wait as unproductive and decided to intervene. 

That is precisely where you need to stop and think. 

The same behavior that resolved a redeploy today without consequences could trigger a destructive action against a production environment tomorrow. Without guardrails, without policies codified in the MCP server, without a security perimeter constraining what the agent can and cannot do, that emergent autonomy goes from advantage to concrete risk. 

The anecdote is compelling, but it's also the best argument for everything we're about to discuss: "giving an agent skills isn't enough", you have to govern how it uses them.

Here's a trap I've seen play out in multiple teams: AI amplifies whatever organizational friction already exists. I watched teams deploy agents on clusters where applying a simple ConfigMap required three manual approvals in a Jira ticket. The result was predictable: agents generated work faster than the organization could approve it, and what they got was a massive bottleneck with extra steps. 

You cannot automate a broken process, you have to fix the process first.

The key to unlocking this transition isn't making the leap all at once. It's designing zones of controlled autonomy:

  • Start with ephemeral and development environments, where the blast radius of a bad decision is minimal. The agent creates, destroys, and redeploys without approval.

  • Scale toward staging with asynchronous supervision. The agent acts, but generates a summary of actions that a human reviews after, not before. The feedback loop inverts.

  • Production only with verified deterministic policies. Agent autonomy is bounded by strict guardrails the platform enforces "not by human vigilance".

But technical trust is only half the equation, the leap to Level 3 demands a sociotechnical shift, and this is exactly where most engineering teams look the other way. AI introduces roles like the AI/ML Platform Engineer that live between data, platform, and application teams. Success depends on breaking silos and creating a model where AI capabilities can be reused and scaled across the organization "not encapsulated in one isolated team".

I've seen CTOs approve platform budgets in Q1 and reassign the team to "priority projects" by Q3. Without permanent dedication, what you build rusts. And it'll probably die.

Beyond the Dashboard: Why the Next Generation of Developer Portals is Invisible

Developer portals were born with a clear promise: centralize tools, reduce cognitive load, and offer a self-service experience through a unified graphical interface.

I know teams with impeccable Red Hat Developer Hub or Backstage instances, complete software catalogs, polished templates, plugins for everything. Their mean time to incident resolution didn't improve by a single minute. When an alert comes in at 3 AM, and believe me, it comes, nobody opens the portal. They open a terminal.

But there's a deeper problem that emerges when coding becomes probabilistic, driven by language models that generate, infer, and decide. In that world, platform engineering can't keep being a storefront of buttons. It needs to become a deterministic harness that gives structure and safety to agent execution.

An AI agent doesn't care how pretty your portal's dashboard is. What it needs are clear APIs, granular permissions, and rich context. It needs skills: well-defined, typed capabilities it can invoke to act on infrastructure. The platform is what must guide models through guardrails and security boundaries on the road to production.

The portal isn't dead. But it's no longer the primary interface. The Internal Developer Platform underneath remains the foundation: the catalog, the pipelines, the policies, the golden paths. What changes is how that platform gets consumed. For humans, the portal still works. For agents, you need something else entirely.

That's what an Agentic Developer Platform (ADP) is: the evolution of your existing IDP, extended with a new consumption layer designed for agents. Not a replacement, but a new surface on top of the same foundation. ADPs don't center on the UI. They center on skills and protocols like MCP.

MCP: the glue between agents and your infrastructure

Here's what happens today in most organizations that have adopted AI coding assistants: the agent lives inside the IDE, it's great at writing code, but it has no idea what's running in your cluster, what your CI pipeline looks like, or who owns the service it just generated a PR for. It's like hiring a brilliant engineer and then locking them in a room with no access to your systems. All that intelligence, completely disconnected from your operational reality.

MCP (Model Context Protocol) solves exactly that. It's an open standard that defines how AI models discover and interact with external tools and data sources. Repositories, CI/CD pipelines, Kubernetes clusters, observability stacks, service catalogs MCP gives agents a structured way to reach all of it without hardcoding integrations for each one.

The shift this enables is fundamental, it changes everything. Platform teams used to spend their energy building UIs on top of complex APIs so humans could click through workflows without needing to understand what's underneath. That job description is changing. We're no longer designing screens for people. We're designing skills for agents well-defined, typed operations that a model can discover, understand, and invoke.

In practice, an MCP server works like this: the agent connects and asks,

"What can I do here?"

The server responds with a catalog of available skills, query service metadata from the Developer Hub catalog, pull logs from a specific pipeline run, trigger a GitOps sync through Argo CD. The agent picks the right tools for the task, chains them together, and executes.

No dashboards. No tab switching. No copy-pasting output between three browser windows.

I've worked with platform teams that spent months building beautiful Backstage plugins for deployment visibility. When we wired the same data through an MCP server, the agent could pull that information in seconds and actually act on it - not just display it. That's the difference. A portal shows you what's happening. An MCP-enabled agent does something about it.

The agents live in the IDE. The platform exposes its capabilities as skills through MCP. The workflow concentrates in the editor.

IDEOps: the real command center

This convergence is what we call IDEOps: code writing, infrastructure operation, and platform consumption; all executed from the integrated development environment.

Instead of constant context-switching "opening a web portal to provision a cluster, another tab to review logs, jumping to Slack to ask who owns the service" the developer simply talks to their agent:

"The payment service pods are restarting on OpenShift. Check the logs, and if it's an OOMKilled after the last deployment, roll back to the previous version via Argo."

The real cost of context-switching

Let's run the same incident through both worlds.

With a traditional portal: You get a Slack alert at 2 AM. You open Grafana to understand the metrics. You search the repo on GitHub to locate the affected service. You go to Backstage to find the owner and dependencies. You open the OpenShift console to dig through pod logs, or check the Tekton pipeline. You finally go back to your IDE to read the code and find the problem. That's roughly 45 minutes of investigation before writing a single line of fix.

With IDEOps and MCP: The observability alert arrives directly in the IDE through the agent. The agent uses MCP skills to pull logs, traces, and service metadata from Developer Hub. It analyzes the context and tells you: "The payment service pods are dying from OOMKilled since the deployment of commit a3f82d1. Want me to roll back to the previous stable version and generate a PR with the memory limits fix?"

Three minutes. No tab switching. Without leaving the IDE.

That difference isn't an incremental improvement. It's an order of magnitude leap in incident response time.

The architecture behind it

The flow breaks down into three layers. The interaction layer "IDE plus local agent" is where the developer states intent in natural language. The context layer "the MCP platform server" is where the agent discovers available tools. This is where platform teams invest their real effort: designing skills that are well-defined, secure, and composable. The execution layer "deterministic infrastructure with guardrails" is where the agent's plan meets reality. It runs on rails: get the failed pipeline ID, read the logs, identify OOMKilled, trigger an Argo sync for rollback.

Security shift-left: the MCP server as Policy Enforcement Point

A natural objection: are we really going to let an LLM execute actions against production?

The answer is that with IDEOps, we don't eliminate security - we shift it left. The MCP server acts as a Policy Enforcement Point: every skill invocation passes through a policy layer before touching real infrastructure.

Take an extreme case not as hypothetical as we'd like. The LLM hallucinates and decides the best way to fix a memory leak is to delete the entire production namespace. Without guardrails, that's catastrophic. In a well-designed ADP, the flow is different:

The agent invokes the namespace deletion method. The MCP server intercepts the call and evaluates it against defined policies - OPA rules or Gatekeeper. The policy detects it's a production namespace, the action is destructive, and the developer lacks admin permissions. The MCP blocks the action before the command reaches the OpenShift API, returning an explicit error: "Action denied."

The platform enforces RBAC and policies at the skill level, not the UI level. Security doesn't depend on a human clicking the right button; it's codified in the protocol. Every skill has its own declarative, auditable, version-controlled security perimeter.

This is Policy as Code applied to agent-infrastructure interaction. The same principles we already use for CI/CD pipelines and Kubernetes deployments, now extended to the execution plane of autonomous agents. If your team already knows how to write OPA policies, 80% of the conceptual work is done. What changes is where they're enforced.

Governing autonomy: trust, audit, and agent observability

Giving skills to an agent isn't enough if you can't answer three questions afterward: what did it do, why did it do it, and who's accountable when something goes wrong?

In an agentic model, every action has an implicit delegation chain. The developer states an intent ("roll back if it's OOM"). The agent interprets that intent, traces a plan, and executes skills. The platform enforces policies and runs the action against infrastructure. When something fails, the question "who did this?" needs a clear answer. For that, the MCP server must generate an immutable audit log for every invocation: what skill was called, with what parameters, what policy was evaluated, what the result was, and which developer "through which agent" initiated it. This isn't optional. It's the equivalent of the audit trail that any enterprise system already requires for compliance.

But auditing alone isn't enough. If agents are the operational workforce, they need the same level of observability we give any production service. It's not enough to know what they did, we need to understand how they reasoned.

That means instrumenting the agent with traces that capture: what context it consumed (which discovery skills it called, what catalog data it read), what plan it traced (what sequence of actions it chose and why it discarded alternatives), what it executed (every skill call with latency, result, and errors), and what it decided (the final decision and the model's reasoning behind it).

These traces are, conceptually, OpenTelemetry spans for agents. They extend the distributed observability model we already know into the AI reasoning plane. A trace of an agent resolving an incident should be viewable in Jaeger or Grafana Tempo with the same naturalness as an HTTP request trace crossing microservices.

Without agent observability, there's no trust. And without trust, there's no viable path to Levels 3 and 4.

Build a better factory, not just faster workers

That line from Luca Galante's Weave Intelligence analysis captures it perfectly. IDEOps is an architecture where the IDE becomes the command center, agents become the operational workforce, and the deterministic platform becomes the nervous system connecting everything with precision.

The question isn't whether your portal needs a UI redesign. The question is: is your platform ready to be consumed by agents?

If the answer is no, you know where to start. And if the answer is "I don't know" that's also a no.

Open your IDE. Connect an agent to your infrastructure. Ask it to do something operational. What happens in the next 30 seconds tells you exactly where you stand.