Cloud operations
Making sense of operational noise
I helped design a greenfield service that correlated high volumes of infrastructure alerts into incidents operators could prioritise and resolve.
Cloud Event Management replaced fragmented alert handling with a single workflow for correlation, prioritisation, notification, investigation and resolution.
Project summary
A clear picture of IT environments
000s of alerts could fire each day, often from the same underlying problem. Cloud Event Management correlated them into incidents operators could prioritise, route and resolve. Notifications reached the right people; linked runbooks supported resolution. One resolved incident could clear hundreds of underlying events.
Research & discovery
000s of events. No clear place to start.
Operators had little help deciding what mattered. They had to correlate related events themselves, find diagnostic context in other tools and identify the right person to act. We turned those problems into 3 IBM Design Thinking Hills defining the MVP outcomes.
Interaction design
Exploring the operator workflow
Operators worked through dense alert and incident data under pressure and frequent interruption. Competitor products compounded the problem with dense layouts and weak prioritisation. I used research, sketches, workflow models and prototypes to explore a clearer incident workflow.
Clarity, prioritisation and action
I worked with IBM visual designers to reduce noise, surface context and make the next action clearer. Contextual runbooks moved into the incident record; an incident timeline exposed the history needed for investigation and reassignment.
Execution
From spec to live code
I worked with developers from specification through implementation. On staging, I inspected builds in-browser, tested the live experience with users and joined heuristic reviews. Design, development and user-feedback issues were prioritised together in GitHub.
Further examples
Across the service
Onboarding, event-source configuration, incident management and runbooks across Cloud Event Management. The core product remains part of IBM’s later Multicloud Management and AIOps products, with implementations reporting 55% lower mean time to detection and more than 70% alert-noise reduction.





