The Platform

One causal entity graph. Four AI engines. Built for enterprise engineering teams who need answers, not dashboards.

⬡ Platform overview → ◯ Four engines → ▶ How it works → → Start free trial →

What we solve

Applicare enables teams to observe, automate, and resolve incidents faster across every stack.

★ Customer stories → → Try it free → ◯ Platform overview →
Learn
Blog
Engineering insights & deep dives
Webinars
Live sessions & on-demand recordings
Case Studies
Proven results across industries
White Papers
Research, architecture guides & reports
Customer Stories
Real teams, real outcomes
Product News
What’s new in Applicare
Adopt & Grow
Applicare University
Free courses for Applicare engineers
Learning Center
Every video lesson, searchable in one place
Documentation
Guides, APIs & integration docs
About Arcturus
Our mission, story & team
Partners
SI, MSP & technology partners
Community
Slack, GitHub & developer forum
Support
Helpdesk & customer portal
Downloads
Ops Vault
Datasheets, runbooks & toolkits
Featured Case Studies
AeroMexico
Digital ticketing · MTTR 4.5h → 11min
Leading Private Bank
MTTR 3.2h → 18min · first month
Mediclinic
Audit prep 11 weeks → 18 days
NTT DATA
80% on-call page reduction
Danube Group
94% SLO compliance
ONP
0 violations at last audit
Seygen
78% downtime reduction · GxP compliance
Insurance Tech Platform
67% P1 reduction · $2.4M saved
IIS & Server Availability
100% SLA report accuracy
Health Check Offering
4.2x ROI · 48hr delivery
Bank of Muscat
99.95% core banking uptime
By Industry
Financial Services
Airlines & Transport
Healthcare
Government
Retail & E-Commerce
Pharmaceuticals
Insurance Technology
Managed Services
Enterprise IT
Professional Services
HomePlatformSolutionsArcIn AIResourcesCustomers Login Book a Demo
Case Study · Platform Engineering · Resource Monitoring

Turning a resource monitoring platform into one administrators could trust.

A system resource monitoring platform was exhibiting inconsistencies in breach detection, reporting, configuration management, and deployment reliability. While the core monitoring functionality was operational, several edge cases prevented administrators from receiving accurate, timely, and consistent information during system resource events.

Breach DetectionReportingProcess SnapshotsJava / SpringBuild Stability
Real-Time
Sustained breaches recorded while still active, not after they end
Continuous
Threshold detection preserved through Warning → Critical escalation
Zero
Breaches left indefinitely open by missing tracking data
Stable
Builds no longer corrupted by stale Maven processes
Overview

A monitoring engine, hardened end to end.

This initiative focused on improving the monitoring engine, reporting capabilities, user interface, and build process to deliver a more dependable monitoring experience — addressing issues that affected the accuracy and usability of the platform across breach detection, diagnostics, configuration, and deployment.

Objectives
  • Improve the accuracy of breach detection.
  • Capture diagnostic information at the appropriate time.
  • Ensure consistency across dashboards, reports, and exported documents.
  • Simplify administration through UI improvements.
  • Increase build stability for development and deployment.

The challenge

Several issues affected the accuracy and usability of the platform:

  • Sustained threshold breaches were not recorded until they had already ended, delaying visibility into active incidents.
  • Escalating a resource breach from Warning to Critical reset threshold detection instead of preserving the existing monitoring state.
  • Process snapshots were captured too late during active breaches, limiting their usefulness for troubleshooting.
  • Breaches with missing tracking information could remain open indefinitely.
  • Configured Cooldown periods were not functioning as intended.
  • Reports included unnecessary system processes while omitting complete command-line information for long-running processes.
  • Dashboard charts and data grids occasionally displayed inconsistent Warning and Critical counts.
  • Configuration screens contained usability issues, including incorrect values after saving and an inaccessible Save button.
  • Monitoring grids required unnecessary horizontal scrolling, reducing usability.
  • Development builds could become corrupted due to stale Maven processes interfering with WAR generation.

Escalating a resource breach from Warning to Critical reset threshold detection instead of preserving the existing monitoring state — the exact moment administrators most needed continuous tracking.

The solution

Improved breach detection

The monitoring engine was enhanced to provide more reliable event tracking. Key improvements included:

  • Recording sustained threshold breaches while they are still active instead of waiting until they end.
  • Preserving threshold detection when a breach escalates from Warning to Critical.
  • Capturing process snapshots immediately during active breaches to improve diagnostic value.
  • Automatically closing breaches that cannot be tracked due to missing tracking data.
  • Implementing fully functional Cooldown support to prevent unnecessary repeated breach notifications.

Enhanced reporting

Reporting was refined to improve clarity and diagnostic usefulness. Enhancements included:

  • Reducing the default maximum number of captured processes per snapshot from 50 to improve readability.
  • Excluding the Windows Idle system process from reports and snapshots.
  • Preserving complete process command lines instead of truncating commands longer than 2048 characters.
  • Adding the missing process command column to PDF reports.
  • Standardizing Warning and Critical counting logic across dashboards, charts, grids, and reports.

User experience improvements

Several usability issues were resolved to simplify policy management. Updates included:

  • Correcting process snapshot settings so saved values display accurately.
  • Restoring accessibility of the Save button within the Edit Policy panel.
  • Removing unnecessary horizontal scrolling from CPU, Memory, and Disk Spike grids.

Build stability

The deployment workflow was made more reliable by resolving build corruption caused by stale Maven processes interfering with WAR generation.

Results

The improvements delivered a more reliable and consistent monitoring platform by:

  • Providing immediate visibility into sustained resource breaches.
  • Preserving monitoring continuity during Warning-to-Critical escalations.
  • Capturing more useful diagnostic data during active incidents.
  • Preventing indefinitely open breaches caused by missing tracking information.
  • Producing cleaner and more informative reports.
  • Delivering consistent breach counts across dashboards and exported reports.
  • Improving policy management through UI fixes.
  • Increasing development and deployment reliability by eliminating build corruption caused by stale Maven processes.

Key takeaways

This project focused on improving the reliability of an existing monitoring platform by addressing issues across the monitoring engine, reporting pipeline, user interface, and build infrastructure. The result was a more dependable system that provides timely breach detection, consistent reporting, improved diagnostics, and a smoother administration experience while increasing overall maintainability.

At a glance
FocusResource monitoring reliability
Monitored resourcesCPU, Memory, Disk
StackJava, Spring, Maven (WAR)
ChallengeBreach detection accuracy, reporting consistency
OutcomeReliable, consistent monitoring platform
Technologies
JavaSpring FrameworkMavenWAR-based deploymentPDF reportingProcess monitoring & snapshots
Get similar results → ← All case studies
Want more reliable resource monitoring for your own platform?
See Applicare on your environment in 30 minutes.
Book a live demo →