Agents Tracker

AI Agents Status Tracker

Making enterprise AI automation trustworthy through visibility.

Overview

The AI Agents Status Tracker panel provides a unified, real-time view of all AI-driven jobs initiated by the user. It centralizes progress tracking, human-in-the-loop approvals, and completed outcomes across multiple agents. The panel is anchored inside the chat experience, where users naturally initiate work, and serves as a persistent, reliable place to monitor automation activity.

My Role:

Product Designer (concept originator & lead designer)

Company:

IBM|Product name: IBM Concert

Tools:

BOB (IBM Internal AI Prototyping tool)| Figma| Figma Make

Timeline:

~3 months, concept to validated MVP

Team:

Team: Design (me), UX Research, Product Managemer, Engineering
 

Problem

I kept noticing the same gap in conversations with people using AI agents day to day: they'd kick off an agent, step away, and have no way to know what happened next. Did it finish? Did it need something from them? Did it break?

That absence isn't just a missing UI affordance, it's a trust problem. As AI agents take on more complex, multi-step work, people have no centralized way to track progress, approvals, or outcomes across tasks, and if they can't see what's happening, they stop delegating anything they can't sit and babysit, which defeats the point of agentic automation in the first place.

I brought the idea to an internal AI design team. The response was strong enough that I decided not to let it sit as an artifact, I started scoping it as an actual project.

What it is

The AI Agents Status Panel gives users a structured view of automation activity through three tabs: Active, for jobs currently running with step-level progress; Attention, for tasks needing human approval or intervention; and Completed, for finished work with outcomes and follow-up actions. It's accessible directly from chat, where users naturally start agent actions, so they can run multiple jobs at once, step away, and return to one consolidated view instead of re-asking each agent for an update.

Why we need this

Multi-step, multi-agent work often runs in parallel and needs occasional human decisions, but without a centralized status layer, users lose track of what's running, miss approval steps, and have no reliable place to check on finished work. The panel closes that gap with continuous transparency while jobs run, clear checkpoints when a decision is needed, and a persistent history to return to, so every feature ties back to a specific failure mode rather than being added as a nice-to-have.

 

Research and Concept testing

I began ideation in Figma and Figma Make, then built the prototype using IBM's internal AI tool. When the first draft was ready, I partnered with research to run concept testing. That first version was intentionally simple, I just wanted to know whether participants understood the concept. We ran moderated sessions with four Site Reliability Engineers already using AI agents in their daily work.

 

Active Tab - Version 1

Jobs currently running, with step-level progress (e.g., "Step 3 of 5"), progress bars, and short descriptions, updating in real time.

Attention Tab - Version 1

Tasks paused for human approval, showing exactly which step needs review, with Approve/Reject actions and required contextual explanation for the decision.

Completed Tab - Version 1

Finished tasks with timestamps, an agent tag identifying which AI agent performed the work (e.g., Compliance Agent, Resilience Agent), and access to results, logs, and follow-up actions.

 

What validated

Task comprehension: All 4 of 4 participants understood the panel as a centralized operational workspace, and the three-tab structure without prompting; the core information architecture validated on the first pass.

Approval task success: Did not pass in its original form. Approve/reject alone was too binary for participants to decide confidently, which directly drove the Attention tab redesign (risk indicators, Request Changes, Add Note, impact evidence).

Adoption intent: 4 of 4 participants said they would use the tool going forward; conditional on pilot/POC validation, richer approval detail, and reporting, all of which were folded into the next design iteration rather than treated as blockers.

Trust: Validated as real but conditional; participants consistently named visible steps, logs, timestamps, and production proof as the specific things that would convert conceptual trust into operational reliance.

From Findings to Redesign

Active tab: Displays all running jobs, showing step numbers (e.g., Step 3/5), progress bars, and short descriptions that update in real time as steps complete. Expanding a job reveals its full step-by-step plan, with every step marked Processing, Queued, or Complete, replacing the single progress bar with the step-level explainability participants had asked for."

Active tab with expandable job report

Attention tab: Replaced the flat approve/reject action with a risk-level indicator (Low/High) plus a confidence percentage, an expandable "View impact & evidence" section, and two new actions, Request Changes and Add Note, alongside Approve and Reject.

Attention tab with expandable job report

Completed tab: Added an agent/domain tag (e.g., Resilience, Performance, Compliance), final status, and direct links to View in Chat and View Report a first pass at the audit-trail requirement research had surfaced.

Completed tab with expandable job report

Every change in this iteration traced back to a specific finding from the four sessions, which kept the redesign grounded in what users actually asked for rather than my own assumptions about what "richer" should mean.

Pitching the Vision

From Single Product to Cross-Team Opportunity

I put together a project proposal to share progress and get alignment from key stakeholders, including our AI lead PM and AI design managers. I received genuinely positive feedback on the direction. one AI lead PM, who had already flagged the feature as scalable, was in the process of aligning at the leadership level with a sibling AI orchestration team and a partner product team. The PM moved quickly too, agreeing on the spot to add the project to the AI product roadmap. In parallel, I reached out directly to the design leads on those two partner teams myself, both conversations landed well, with genuine interest in extending the pattern beyond a single product. What started as a single-product proposal turned into an active cross-team opportunity, with conversations underway at both the design and leadership level.

Stakeholder Pitch Presentation

MVP scope

Rather than building every validated capability at once, I scoped the MVP to the highest-impact gaps confirmed across all four research sessions:

  • Active tab with step-level plan visibility (Processing / Queued / Complete) instead of a single progress bar

  • Attention tab with Approve, Reject, Request Changes, and Add Note, plus a risk/confidence indicator and expandable impact-and-evidence detail

  • Completed tab with agent/domain tagging, finished status, and basic report access (View Report, View in Chat)

  • Explicit upfront scope messaging distinguishing what the panel can act on versus what it only surfaces

Full reporting depth (cross-event audit trails, root-cause linkage) and full multi-agent orchestration polish were pushed to a later phase; research confirmed they mattered, but they weren't required to prove the core concept.

One senior design leader also pushed to narrow the MVP further, from the full multi-agent set down to a single agent. Since research had already validated visibility, approval richness, and trust across a multi-agent view, narrowing scope wasn't re-proving the concept; it was a deliberate trade-off to reduce build complexity and prove the pattern (step visibility, approval richness, completed history) inside one contained workflow before committing engineering effort to full cross-agent orchestration. The cost was clear-eyed: a single-agent MVP wouldn't demonstrate the "unified across all agents" value that had tested well, but expanding later would be scaling the validated pattern, not redesigning it.

Impact & Outcomes

User impact: Higher trust through transparent, explainable agent behavior; less manual status-tracking and re-asking in complex multi-step workflows; reduced cognitive load from offloading task tracking to the system; confidence that work continues reliably across parallel tasks and when stepping away.

Product impact: A scalable orchestration pattern any current or future agent can plug into, rather than each agent building its own tracking approach; a more predictable and consistent AI experience across the product; a reusable foundation (states, approval pattern, history model) that reduces design and engineering cost for future agent-driven features.

Business impact: Lower uncertainty as an adoption barrier for AI features; reduced time spent on manual status-checking for teams running complex multi-agent work; a foundation for premium, auditable AI orchestration offerings that enterprise and compliance-sensitive customers require; a pattern that extends beyond a single product's roadmap into at least two other teams' work.

Reflection and what's next

The MVP is now on the roadmap, and I'm continuing to drive it on several fronts rather than treating it as a finish line: extending the pattern to the two partner product teams already engaged, designing the next layer of multi-agent collaboration visibility (which agents are working on which jobs, and how they coordinate; a direct extension of the multi-agent value that already tested well), and building out a patent case around a novelty angle the design surfaced.

What I'd point to most in this project isn't the interface; it's the discipline of the loop: notice a real pattern in how people actually work, build just enough to test it, let research findings (not opinion) decide what changes, and use every one of those findings as the argument for why the next investment is worth making. That loop is what took this from a design-jam sketch to a cross-team initiative with a validated MVP and a defensible roadmap case.