Cisco AppDynamics · 2023

Designing an inventory-first agent-management workflow

This case study follows the inventory model, the hierarchy changes made after evaluation, and the release trade-offs behind the agent-management workflow.

Timeline
6 months
Product
Cisco AppDynamics
Role
Senior product designer
Scroll to discover↓

At a glance

Install, upgrade and operate agents from one place.

Agent Management brought installation, upgrades, rollbacks and ongoing operations into one workflow around Smart Agent; so ITOps teams could manage thousands of machines without stitching the process together by hand.

≈5×faster measured install17 min → 3 min in the guided flow
21.5%SaaS adoptionof eligible accounts within six months
2 weeks → 2 hourscustomer-reported time to valuefor a representative rollout
“highlight”at Cisco Liveamong a broad launch portfolio

I carried the work from ambiguity to an operable system.

Turn scattered operational pain into one lifecycle product model.

Problem framing + product model

Translate discovery research into product priorities and decisions.

Research planning + synthesis

Standardise installation and configuration across agents.

Working sessions with PM · architecture · engineering

Protect the core workflow through release and implementation constraints.

Prioritisation + design-to-engineering delivery

With product management · engineering · architecture · research · field support · design system · UX writing · IX

let’s understand agent managementthrough a city.

Applications are like cities.

Services are buildings, dependencies are roads, and requests are the traffic moving between them.

The useful signal is at the junction.

Many paths meet here. A slowdown at one intersection can ripple through the system.

Agents are the eyes.

They observe what moves, where it slows, and what needs attention; without becoming the traffic itself.

Skip city story

“It takes weeks.”

95,000 agents at enterprise scale

1 dot = 1,000 agents

Installation and upgrades were assembled through scripts, documentation and coordination across teams. At this scale, even a small inconsistency became operational work.

Why it mattered

A security event turned Christmas into an emergency upgrade window.

Customers' employees gave up planned holiday time to urgently upgrade agents across their environments. The emergency exposed the human cost of a fragmented upgrade workflow; not just the technical risk.

Admin AdnanA bearded man with three eyes and a headset mic works at a laptop while connected servers pulse around him and a gear turns.

Admin Adnan

Responsible for the stability and security of systems; while coordinating changes that can affect thousands of agents.

01Deploy + monitor agentsAcross applications, hosts and teams.
02Optimise configurationsWithout breaking working environments.
03Set up alertsAct before customers feel impact.

Discovery research

Across the research, admins lost context as one task moved across tools, owners and configurations.

What the research exposed

Decision 01

Start with inventory.

Admins needed to know what existed and its state before taking action.

What it took to get an install through.

Identification

  • Target applications
  • Deployment strategy
  • Identify stakeholders

Planning & testing

  • Construct and test scripts
  • Create and test install steps
  • Validate dependencies

Execution

  • Schedule a window
  • Deploy scripts or install manually
  • Verify and restart safely

Where the workflow became difficult to follow

The difficult part started before execution.

Research highlighted the difficulty of identifying the right agents, coordinating ownership and permissions, and deciding which configuration was safe enough to apply.

Identifying targets

Identifying the right applications and agents required collaborative effort across owners and operators.

Coordinating permissions

Installation required admin permission plus application-owner changes to link applications to agents.

Checking configuration

An incompatible configuration could break a working environment, so testing carried high consequence.

Decision 02

Make the target clear before the action.

Show ownership, status and configuration before asking admins to commit.

The predecessor

ZFI simplified one install.

At fleet scale, admins also needed bulk actions, persistent status and ongoing management.

Learn more about ZFI

AppDynamics introduced Zero Friction Installation (ZFI), also known as Agent Installer, to simplify agent installation. The project review identified four limitations:

  1. Installation only. The other phases of the agent lifecycle were not covered. At the time, only four customers had adopted ZFI.
  2. One agent at a time. The installer did not address bulk deployment.
  3. Limited agent support. Functional gaps left some agent needs unsupported.
  4. Stronger competing solutions. The review found competitors offered better solutions than ZFI.
Zero Friction Installation guided Java agent setup showing a step-by-step installation flow
Zero Friction Installation

Learning from ZFI

Decision 03

Keep the guidance. Make it work for fleets.

Bulk actions and persistent status extend the guided model.

Focus on the lifecycle moments with the most leverage.

We prioritised installation, upgradation and configuration: the recurring moments with the most coordination and risk.

See why

These three stages repeated most often and carried the highest cost when the workflow broke.

Technical landscape

from many installs to one orchestrator.

Before

Every agent was installed and managed independently.

Agent-specific setup repeated on every host.

After

One Smart Agent orchestrates the host.

A shared host-level layer can coordinate agent operations.

Before architecture with many manually installed agents After architecture with a Smart Agent on each host
Decision 04

Standardise the steps behind each agent.

Smart Agent could orchestrate the differences while the workflow stayed familiar.

Brainstorming

Make the system visible before making it operable.

I used the early model to connect the recurring decisions that had been scattered across installation, inventory, configuration and ongoing operations.

Hand-drawn brainstorming concepts for agent status, market updates, history, and installing or maintaining agents
Brainstorming explored status and alerts, market updates, historical performance and configuration history, and installing or maintaining agents.
Early layout sketch separating an action bar, main content area and inspector panel

Layout design

Then give repeated actions a stable place to live.

The layout separates global context, selection-dependent actions and detailed configuration without turning each task into a new page.

Low-fidelity validation

Teams from 6 companies evaluated the direction.

3 min

Install flow, down from 17 min

173
Unified table dashboard iteration

Preferred direction

Unified dashboard

Preferred by 4 of 6 teams: “more data at a glance.”

Dynamic card view iteration

Explored

Dynamic view

Grouped by agent type.

What validation changed

We tested the operating model, not just the screens.

Inventory first

Customers wanted to understand agent state and context before choosing an action. That reinforced inventory as the starting point for the product.

Restart with explicit consent

Customers were hesitant about agents automatically restarting applications. We added a restart reminder in the panel so the consequence stayed visible and under admin control.

Evaluation informed a change from a tree hierarchy to a column-based layout.

Three of six teams found it hard to use at scale. Adjacent columns kept host, application, tier and node visible.

BeforeNested tree concept for application hierarchy
Nested tree
Hand-drawn arrow
AfterColumn-based hierarchy mapping redesign
Adjacent columns

The final experience

Admin Adnan installs Java agents for newly deployed applications.

Adnan begins on the application monitoring home and moves into Agent Management when the fleet needs attention.

Decision 05

Keep long-running work visible.

Status and recovery stay available after the install panel closes.

The balancing act

A fixed release date forced prioritisation. We kept the end-to-end jobs intact and ranked secondary capabilities with engineering.

Must have

Keep inventory, fleet actions, guided setup and status connected.

+

Good to have

Rank secondary capabilities by impact and effort.

01Scheduling
02Live logs
03Sidebar interactions
04In-app view
Hand-drawn character balancing two trays

Impact

Adoption across accounts and hosts.*

430+

SaaS customers

Direct adoption of Smart Agents and Agent Management.

680+

SaaS accounts

Using the portal across multiple accounts.

3,000+

SaaS hosts

Powered by Smart Agent within six months.

21.5%

Eligible SaaS accounts

Adopted Agent Management within six months.

* These are direct SaaS adoption metrics. On-premises adoption data is classified and intentionally omitted.

Less time between intent and a working agent.

Customer-reported time to value

2 weeks 2 hours

Reported by a customer for a representative rollout.

Observed guided install

17 min 3 min

Observed benchmark completion time. The shipped workflow reduced the measured installation journey by roughly 5×.

What design influenced

These were the parts of the product model and workflow that design directly shaped.

One workflow

Installation, upgrades, rollbacks and ongoing operations shared one product model.

Fleet operations

Bulk actions and persistent status made the workflow usable across thousands of agents.

Visible state

Inventory, ownership and configuration stayed visible before admins committed changes.

Guided configuration

The guided install model became a repeatable pattern for safer configuration.

Orchestration and execution performance were enabled by the broader product and engineering system.

Release scope and trade-offs

The release decisions centered on what needed to work as a complete workflow and which capabilities could follow later.

Ansible / Chef management preceded the simplified management UI.

  1. This case study

    Simplified management UI

    Cisco AppDynamics

  2. Later work

    Inventory & actions

    Cisco Observability Platform

  3. Later work

    Profiles & templates

    Cisco Observability Platform

  4. Later year

    Auto discovery / auto attach

    Cisco AppDynamics

  5. Subsequent work

    Host machine

    Cisco Observability Platform

This case study covers the simplified management UI. The following stages were separate, later work.

A parallel experience for Cisco Observability Platform.

I stayed involved as a parallel Agent Management experience was created for Cisco Observability Platform. The work explored Smart Agent behavior while keeping inventory, installation, configuration and background operations familiar.

Agent Management on Cisco Observability Platform

Four lessons that stayed with me.